Antisense compounds and methods for targeting CUG repeats
Patent Information
- Application Number
- JP2023579457
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-06
- Filing Date
- 2022-06-22
- Publication Date
- 2025-07-01
AI Technical Summary
Current therapeutic oligonucleotide compounds face challenges in effectively delivering intracellularly to treat diseases associated with expanded CTG·CUG repeats, such as myotonic dystrophy type 1, due to limited access to intracellular compartments when administered systemically.
Development of compounds comprising cyclic peptides and endosomal escape vehicles that facilitate intracellular localization of antisense compounds, which bind to expanded CUG repeats, allowing for effective modulation of gene activity and expression.
The described compounds enhance the delivery and efficacy of antisense oligonucleotides to target expanded CUG repeats, restoring normal splicing patterns and reducing disease symptoms by modulating the activity of proteins trapped by these repeats.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. Provisional Application Nos. 63 / 213,900 filed June 23, 2021, 63 / 239,847 filed September 1, 2021, 63 / 290,892 filed December 17, 2021, 63 / 305,071 filed January 31, 2022, 63 / 314,369 filed February 26, 2022, 63 / 316,634 filed March 4, 2022, 63 / 317,856 filed March 8, 2022, and 63 / 318,899 filed March 31, 2022. This application claims the benefit of patent application Ser. No. 63 / 326,201, filed April 4, 2022; patent application Ser. No. 63 / 327,179, filed May 6, 2022; patent application Ser. No. 63 / 339,250, filed March 31, 2022; patent application Ser. No. 63 / 362,295, filed September 1, 2021; patent application Ser. No. 63 / 239,671, filed December 17, 2021; patent application Ser. No. 63 / 290,960, filed January 11, 2022; and patent application Ser. No. 63 / 268,577, filed February 25, 2022.
[0002] The present disclosure relates to compounds, compositions, and methods for modulating the activity and / or levels of genes containing expanded nucleotide repeats, specifically expanded CTG·CUG repeats. The compounds and compositions containing them can be used to treat diseases associated with genes containing expanded nucleotide repeats, specifically expanded CTG·CUG repeats.
[0003] preface Some diseases are associated with extended nucleotide repeats, i.e., genes with a greater number of nucleotide repeats than observed in healthy phenotypes. Extended repeats can cause the aggregation and / or nucleation of extended repeat-containing transcripts, and / or the nucleation of proteins that bind to extended repeat-containing transcripts. Extended repeats can cause some proteins, such as pre-mRNA processing proteins, to become trapped on the repeats, thereby preventing these proteins from performing their normal functions, such as processing pre-mRNA transcripts of other genes that do not contain extended repeats.
[0004] Several diseases are associated with genes containing expanded CTG·CUG trinucleotide repeats (CTG refers to the DNA repeat, and CUG refers to the corresponding RNA repeat that is generated during transcription). Diseases associated with genes containing expanded CTG·CUG trinucleotide repeats include, but are not limited to, myotonic dystrophy type 1 (DM1), spinocerebellar ataxia-8 (SCA8), Huntington's disease-like 2 (HDL2), and Fuchs endothelial corneal dystrophy (FECD).
[0005] Myotonic dystrophy type 1 (DM1), the most common cause of muscular dystrophy in adults, affecting 1 in 8,500 people worldwide, is associated with genes containing expanded trinucleotide repeats (Lee and Cooper. (2009) “Pathogenic mechanisms of myotonic dystrophy,” Biochem Soc Trans. 37(06):10.1042 / BST0371281). DM1 is a disorder affecting skeletal and smooth muscles, as well as the eyes, heart, endocrine system, and central nervous system. DM1 is caused by an abnormal expansion of a CTG-trinucleotide repeat within the noncoding region of the gene encoding myotonic dystrophy protein kinase (DMPK). The CTG expansion is located within the region corresponding to the 3'-untranslated region (3'-UTR) of DMPK mRNA. While the DMPK gene in healthy individuals contains 5-40 CTG trinucleotide repeats, patients with DM1 have 50 to several thousand CTG trinucleotide repeats. CTG trinucleotide repeat expansions result in global deregulation of gene expression in affected individuals due to the nucleation of several RNA-binding regulatory proteins at the CUG-expansion within the 3' untranslated region (3'-UTR), preventing RNA-binding proteins, such as muscleblind-like proteins (MBNL1-3), from performing their normal cellular functions. Nucleated RNA-binding proteins are unable to bind to and affect the translation of other mRNA transcripts. These CUG-expanded mRNA protein aggregates form distinct nuclear foci. The activity of additional splicing factors, such as CUGBP Elav-like family member 1 (CELF1), is also disrupted, leading to the missplicing of numerous downstream gene transcripts associated with DM1 symptoms. As the number of repeats increases, disease severity increases and the age of onset decreases (Pettersson et al. (2015) “Molecular mechanisms in DM1 - a focus on foci.” Nucleic Acids Res. 43(4):2433-2441).
[0006] CUG-trinucleotide repeats within the 3' untranslated region of DMPK mRNA form imperfect, stable hairpin structures that accumulate in the nucleus in small ribonucleocomplexes or microscopically visible inclusion bodies, impairing the function of proteins involved in transcription, splicing, or RNA export. Although the DMPK gene containing the CUG repeats is transcribed into mRNA, mutant transcripts are trapped in the nucleus as aggregates (foci), resulting in reduced cytoplasmic DMPK mRNA levels. These aggregates lead to deregulation of alternative splicing of many different transcripts due to the trapping of two RNA-binding proteins: MBNL1 (myoblind-like 1) and CUGBP1 (CUG-binding protein 1), resulting in loss of MBNL1 function and upregulation of CUGBP1 (Lee and Cooper. (2009) "Pathogenic mechanisms of myotonic dystrophy," Biochem Soc Trans. 37(06):10.1042 / BST0371281). MBNL1 and CUGBP-ETR-3-like factor 1 (CELF1) are developmental regulators of splicing events during the fetal-to-adult transition, and their altered activity in DM1 leads to the expression of fetal splicing patterns in adult tissues. The downstream effects of reduced MBNL1 levels and increased CELF1 levels include disruption of alternative splicing, mRNA translation, and mRNA decay in proteins such as cardiac troponin T (cTNT), insulin receptor (INSR), muscle-specific chloride channel (CLCN1), and sarcoplasmic / endoplasmic reticulum calcium ATPase 1 (ATP2A1) transcripts, in addition to MBNL1 (Konieczny et al. (2017) “Myotonic dystrophy: candidate small molecule therapeutics.” Drug Discovery Today. 22(11):1740-174).
[0007] Potential therapeutic approaches for treating DM1 or other diseases associated with expanded CTG-CUG repeats include the use of therapeutic oligonucleotide-containing compounds. However, a major problem associated with the use of oligonucleotide compounds in therapy is their limited ability to access intracellular compartments when administered systemically. Intracellular delivery of oligonucleotide compounds can be facilitated by the use of carrier systems such as polymers, cationic liposomes, or by chemical modification of the construct, for example, by covalent attachment of cholesterol molecules. However, the efficiency of intracellular delivery of oligonucleotide compounds remains low. Improved delivery systems are still needed to increase the efficacy of these compounds.
[0008] There is an unmet need for effective compositions that deliver therapeutic oligonucleotide compounds to intracellular compartments to treat diseases caused by expanded CTG·CUG repeats, such as DM1. Summary of the Invention
[0009] Described herein are compounds, compositions, and methods for treating diseases associated with expanded CTG-CUG repeats. In some embodiments, the disclosure relates to a compound comprising an antisense compound (AC) and a cyclic peptide, e.g., a cyclic cell-penetrating peptide (cCPP). In some embodiments, the AC binds to a gene or gene transcript containing an expanded CUG repeat. In some embodiments, the cyclic peptide facilitates intracellular localization of the AC. The compound may comprise an endosomal escape vehicle (EEV). The EEV may comprise a cyclic peptide and an exocyclic peptide.
[0010] In some embodiments, provided herein are compounds comprising (a) at least one cyclic peptide and (b) an antisense compound (AC) complementary to a target nucleotide. In some embodiments, the target nucleotide comprises at least one extended CUG or CTG repeat. In some embodiments, the target nucleotide is a gene comprising at least one extended CTG repeat. In some embodiments, the target nucleotide is an RNA comprising at least one extended CUG repeat. In some embodiments, the RNA comprising at least one extended CUG repeat is a pre-mRNA sequence. In some embodiments, the extended CUG repeat corresponds to an expanded CTG repeat in the gene from which the pre-mRNA is transcribed. In some embodiments, the antisense compound binds to the expanded CTG repeat or the expanded CUG repeat. In some embodiments, the AC comprises 5 to 40 CAG repeats (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 repeats). In some embodiments, the AC comprises a sequence of nucleotides listed in Table 2, Table 10, or Table 11. In some embodiments, the AC comprises a sequence of nucleotides listed in Table 2.
[0011] In some embodiments, the AC comprises at least one modified nucleotide or nucleic acid selected from phosphorothioate (PS) nucleotides, phosphorodiamidate morpholino (PMO) nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNAs), nucleotides containing 2'-O-methyl (2'-OMe) modified backbones, 2'O-methoxy-ethyl (2'-MOE) nucleotides, 2',4' constrained ethyl (cEt) nucleotides, 2'-deoxy-2'-fluoro-beta-D-arabinonucleic acid (2'F-ANA), and combinations thereof. In some embodiments, the AC comprises PMO nucleotides.
[0012] In some embodiments, compounds are provided that include a cyclic peptide having 6-12 amino acids, wherein at least two amino acids of the cyclic peptide are charged amino acids, at least two amino acids of the cyclic peptide are aromatic hydrophobic amino acids, and at least two amino acids of the cyclic peptide are uncharged non-aromatic amino acids. In some embodiments, the antisense compound (AC) is complementary to at least a portion of an expanded CUG repeat in a target mRNA sequence. In some embodiments, the AC is a phosphorodiamidate morpholino (PMO) nucleotide.
[0013] In some embodiments, at least two charged amino acids of the cyclic peptide are arginine. In some embodiments, at least two aromatic hydrophobic amino acids of the cyclic peptide are phenylalanine, naphthylalanine (3-naphth-2-yl-alanine), or a combination thereof. In some embodiments, at least two uncharged non-aromatic amino acids of the cyclic peptide are citrulline, glycine, or a combination thereof. In some embodiments, the compound is a cyclic peptide having 6 to 12 amino acids, wherein two amino acids of the cyclic peptide are arginine, at least two amino acids are aromatic hydrophobic amino acids selected from phenylalanine, naphthylalanine, and combinations thereof, and at least two amino acids are uncharged non-aromatic amino acids selected from citrulline, glycine, and combinations thereof.
[0014] In some embodiments, the compound comprises an endosomal escape vehicle comprising a cyclic peptide and an exocyclic peptide (EP). In some embodiments, the EP is conjugated to a linker at the amino group. The linker can be a linker described herein. In some embodiments, the EP is conjugated to the cyclic peptide via a linker. In some embodiments, the EP is conjugated to the AC via a linker. In some embodiments, the EP is conjugated to a linker that conjugates the AC to the cyclic peptide.
[0015] In some embodiments, the EP comprises 2-10 amino acids. In some embodiments, the EP comprises 4-8 amino acid residues. In some embodiments, the EP comprises one or two amino acids comprising a side chain containing a guanidine group, or a protonated form thereof. In some embodiments, the EP comprises one, two, three, or four lysine residues. In some embodiments, the amino group on the side chain of each lysine residue is substituted with a trifluoroacetyl (-COCF3) group, an allyloxycarbonyl (Alloc), a 1-(4,4-dimethyl-2,6-dioxocyclohexylidene)ethyl (Dde), or a (4,4-dimethyl-2,6-dioxocyclohex-1-ylidene-3)-methylbutyl (ivDde) group. In some embodiments, the EP comprises at least two amino acid residues having hydrophobic side chains. In some embodiments, the amino acid residues having hydrophobic side chains are selected from valine, proline, alanine, leucine, isoleucine, and methionine. In embodiments, the exocyclic peptide comprises one of the following sequences: PKKKRKV, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK, KKRK, KRKK, KRRK, RKKR, RRRR, KGKK, KKGK, KKKKK, KKKRK, KBKBK, KKKRKV, PGKKRKV, PKGKRKV, PKKGRKV, PKKKGKV, PKKKRGV, or PKKKRKG. In embodiments, the exocyclic peptide consists of one of the following sequences: PKKKRKV, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK, KKRK, KRKK, KRRK, RKKR, RRRR, KGKK, KKGK, KKKKK, KKKRK, KBKBK, KKKRKV, PGKKRKV, PKGKRKV, PKKGRKV, PKKKGKV, PKKKRGV, or PKKKRKG. In embodiments, the exocyclic peptide has the following structure: Ac-PKKKRKV-.
[0016] In some embodiments, the cyclic peptide comprises 4 to 12 amino acids. In some embodiments, the cyclic peptide comprises 6 to 12 amino acids. In some embodiments, at least two amino acids of the cyclic peptide are charged amino acids, at least two amino acids of the cyclic peptide are aromatic hydrophobic amino acids, and at least two amino acids of the cyclic peptide are uncharged non-aromatic amino acids. In some embodiments, at least two charged amino acids of the cyclic peptide are arginine, at least two aromatic hydrophobic amino acids of the cyclic peptide are phenylalanine, naphthylalanine, or a combination thereof, and at least two uncharged non-aromatic amino acids are citrulline, glycine, or a combination thereof.
[0017] In some embodiments, the cyclic peptide has 4-12 amino acids, at least two of which are arginine and at least two of which comprise a hydrophobic side chain, provided that the cyclic peptide is not a cyclic peptide having a sequence of SEQ ID NO: 89-117. In some embodiments, the cyclic peptide is not a cyclic peptide having a sequence of SEQ ID NO: 89-117. [Table 21] In the formula, F is L-phenylalanine, f is D-phenylalanine, Φ is L-3-(2-naphthyl)-alanine, Φ is D-3-(2-naphthyl)-alanine, R is L-arginine, r is D-arginine, Q is L-glutamine, q is D-glutamine, C is L-cysteine, U is L-selenocysteine, W is L-tryptophan, K is L-lysine, D is L-aspartic acid, and Ω is L-norleucine.
[0018] In embodiments, the cyclic peptide has the following structure: [ka] , or having its protonated form, During the ceremony, R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4, R5, R6, and R7 are independently H or an amino acid side chain; at least one of R4, R5, R6, and R7 is a side chain of 3-guanidino-2-aminopropionic acid, 4-guanidino-2-aminobutanoic acid, arginine, homoarginine, N-methylarginine, N,N-dimethylarginine, 2,3-diaminopropionic acid, 2,4-diaminobutanoic acid, lysine, N-methyllysine, N,N-dimethyllysine, N-ethyllysine, N,N,N-trimethyllysine, 4-guanidinophenylalanine, citrulline, N,N-dimethyllysine, β-homoarginine, or 3-(1-piperidinyl)alanine; AA SC is the amino acid side chain to which the antisense compound is conjugated; q is 1, 2, 3, or 4.
[0019] In some embodiments, at least one of R4, R5, R6, and R7 is independently an uncharged, non-aromatic side chain of an amino acid, hi some embodiments, at least one of R4, R5, R6, and R7 is independently H or the side chain of citrulline.
[0020] In embodiments, the cyclic peptide has the structure of Formula I: [ka] or its protonated form, During the ceremony, R1, R2, and R3 are each independently H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R7 are independently H or an amino acid side chain; AA SC is the amino acid side chain to which the antisense compound is conjugated; q is 1, 2, 3, or 4; Each m is independently an integer of 0, 1, 2, or 3.
[0021] In embodiments, the cyclic peptide of Formula (I) has one of the following structures: [ka] or It has its protonated form.
[0022] In embodiments, the cyclic peptide of Formula (I) has one of the following structures: [ka] [ka] or It has its protonated form.
[0023] In embodiments, the cyclic peptide of Formula (I) has one of the following structures: [ka] or It has its protonated form.
[0024] In embodiments, the cyclic peptide of Formula (I) has the following structure: [ka] , or It has its protonated form.
[0025] In embodiments, the compound has the structure of Formula C: [ka] , or a protonated form or salt thereof, During the ceremony, R1, R2, and R3 are each independently H or a side chain containing an aryl or heteroaryl group, and at least one of R1, R2, and R3 is a side chain containing an aryl or heteroaryl group; R4 and R7 are independently H or an amino acid side chain; EP is an exocyclic peptide; each m is independently an integer of 0 to 3; n is an integer from 0 to 2, x' is an integer from 1 to 23, y is an integer from 1 to 5, q is an integer from 1 to 4, z' is an integer from 1 to 23, The cargo is the antisense compound.
[0026] In embodiments, the compound has one or more of the following structures: [ka] [ka] or a protonated form or salt thereof, wherein EP is an exocyclic peptide; Oligonucleotides are the antisense compounds.
[0027] In embodiments, the oligonucleotide of the compound of Formula (C-1), Formula (C-2), Formula (C-3), or Formula (C-4) comprises the following sequence: 5'-CAG CAG CAG CAG CAG CAG CAG-3'.
[0028] In embodiments, the EP of the compound of Formula (C-1), Formula (C-2), Formula (C-3), or Formula (C-4) comprises the following sequence: PKKKRKV. [Brief explanation of the drawings]
[0029] [Figure 1] FIG. 1 is a schematic diagram showing multiple strategies for targeting CUG repeats in mRNA. [Figure 2]
[0023] Figures 1-3 show modified nucleotides used in the antisense oligonucleotides described herein. Structures 1-3 (1 = phosphorothioate, 2 = (SC5-Rp)-α,β-CAN, 3 = PMO) are phosphate backbone modifications, structure 4 (2-thio-dT) is a base modification, structures 5-8 (5 = 2'-OMe-RNA, 6 = 2'O-MOE-RNA, 7 = 2'F-RNA, 8 = 2'F-ANA) are 2' sugar modifications, structures 9-11 are constrained nucleotides, and structures 12-13 are constrained nucleotides. Structures 4 (9 = LNA, 10 = (S)-cET, 11 = tcDNA, 12 = FHNA, 13 = (S)-5'-C-methyl, 14 = UNA) are additional sugar modifications, structures 15-18 (15 = E-VP, 16 = methyl phosphonate, 17 = 5' phosphorothioate, 18 = (S)-5'-C-methyl with phosphate) are 5' phosphate-stabilizing modifications, and structure 19 is a morpholino sugar. Reformatted from Khvorova, A., et al., Nat. Biotechnol. 2017 Mar;35(3):238-248. [Figure 3A] Figure 3A shows conjugation chemistries for linking ACs to cyclic cell-penetrating peptides. Figure 3A shows amide bond formation between a peptide bearing a carboxylic acid group or a peptide bearing a TFP-activated ester and a primary amine residue at the 5' end of the AC. Figure 3B shows the conjugation of a peptide-TFP ester to a secondary amine or primary amine-modified AC at the 3' end via amide bond formation. Figure 3C shows the conjugation of a peptide-azide to a 5'-cyclooctyne-modified AC via copper-free azide-alkyne cycloaddition. Figure 3D illustrates another exemplary conjugation between a 3'-modified cyclooctyne AC or a 3'-modified azide AC and a CPP containing a linker-azide or linker-alkyne / cyclooctyne moiety via copper-free azide-alkyne cycloaddition or copper-catalyzed azide-alkyne cycloaddition, respectively (click reaction). [Figure 3B] Same as above. [Figure 3C] Same as above. [Figure 3D] Same as above. [Figure 4] Conjugation chemistry for linking AC and CPP with additional linker modalities including polyethylene glycol (PEG) moieties is shown. [Figure 5A] 5A, 5B, 5C, and 5D provide structures of adenine (FIG. 5A), cytosine (FIG. 5B), guanine (FIG. 5C), and thymine (FIG. 5D) morpholino subunit monomers that can be used to synthesize phosphorodiamidate-linked morpholino oligomers (PMOs). [Figure 5B] Same as above. [Figure 5C] Same as above. [Figure 5D] Same as above. [Figure 6A] Figure 6 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of MBNL1 (exon 5, Figures 6A, 6B, and 6E) and CLASP1 (exon 19, Figures 6B, 6D, and 6F) after treatment of HeLa-48 cells with various PMO or PMO-EEV compounds at 1 μM, 3 μM, or 10 μM for 24 h (Figures 6A, 6B, and 6E) and 48 h (Figures 6C, 6D) with or without Endo-Porter transfection agents (Figures 6A, 6C, and 6E). Parental HeLa and HeLa-480 cell lines treated with (Figures 6A, 6D) or without (Figures 6E, 6F) Endo-Porter agents were included as controls. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 6D] Same as above. [Figure 6E] Same as above. [Figure 6F] Same as above. [Figure 7A]Figure 7A shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of MBNL1 (exon 5, Figure 7A) and CLASP1 (exon 19, Figure 7B) after 48 h treatment of DM1 myoblasts with 1 μM of various PMO or PMO-EEV compounds without Endo-Porter transfection reagent. Two controls, DM-04 without Endo-Porter treatment and DM-05 without Endo-Porter treatment, were included as controls. [Figure 7B] Same as above. [Figure 8A] Figure 8A shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of Atp2a1 (exon 22, Figure 8A), Nfix (exon 7, Figure 8B), Clcn1 (exon 7a, Figure 8C), and Mbnl1 (exon 5, Figure 8D) from gastrocnemius muscle tissue of HSA-LR (DM1 mouse model) mice treated with PMO, 20 mpk PMO-EEV 221-1106, or 40 mpk PMO-EEV 221-1106 for 1 week. FVB / NJ (wild-type inbred mice) and HSA-LR (untreated) mice were included as controls. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 9A] Figure 9 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of Atp2a1 (exon 22, Figure 9A), Nfix (exon 7, Figure 9B), Clcn1 (exon 7a, Figure 9C), and Mbnl1 (exon 5, Figure 9D) from quadriceps muscle tissue of HSA-LR (DM1 mouse model) mice treated with PMO, 20 mpk PMO-EEV 221-1106, or 40 mpk PMO-EEV 221-1106 for 1 week. FVB / NJ (wild-type inbred mice) and HSA-LR (untreated) mice were included as controls. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 10A]Figure 10 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of Atp2a1 (exon 22, Figure 10A), Nfix (exon 7, Figure 10B), Clcn1 (exon 7a, Figure 10C), and Mbnl1 (exon 5, Figure 10D) from tibialis anterior muscle tissue after 1-week treatment of HSA-LR (DM1 mouse model) mice with PMO, 20 mpk PMO-EEV 221-1106, or 40 mpk PMO-EEV 221-1106. FVB / NJ (wild-type inbred mice) and HSA-LR (untreated) mice were included as controls. [Figure 10B] Same as above. [Figure 10C] Same as above. [Figure 10D] Same as above. [Figure 11A] Figure 11 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) in MBNL1 (exon 5, Figure 11A), SOS1 (exon 25, Figure 11B), IR (exon 11, Figure 11C), DMD (exon 78, Figure 11D), BIN1 (exon 11, Figure 11E), and LDB3 (exon 11, Figure 11F) after treatment of muscle cells from DM1 patients with different concentrations (10 μM, 3 μM, 1 μM, and 0.3 μM) of DMPK CUG-targeting EEV-PMOs (CUGexp 197-777 and CUGexp 221-1106). Muscle cells from two groups, healthy individuals (negative control) and DM1 patients (positive control), were tested for alternative RNA splicing events as controls. All data were collected from three separate experiments (n = 3). A t-test was performed comparing treated and untreated DM1 myotubes. * p<0.05, ** p<0.01, *** p<0.001. [Figure 11B] Same as above. [Figure 11C] Same as above. [Figure 11D] Same as above. [Figure 11E] Same as above. [Figure 11F] Same as above. [Figure 12A]Figure 12 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) in MBL1 (exon 5, Figure 12A), SOS1 (exon 25, Figure 12B), INSR (exon 11, Figure 12C), DMD (exon 78, Figure 12D), BIN1 (exon 11, Figure 12E), and LDB3 (exon 11, Figure 12F) after treatment of patient-derived DM1 myoblasts and myotubes with 10 μm, 3 μm, or 1 μm of the DMPK CUG-targeting EEV-PMO 197-777. Healthy patient cells and DM1 cells served as controls. All data were collected from three separate experiments (n=3). t-tests comparing treated and untreated DM1 myotubes were performed. *p<0.05, **p<0.01, ***p<0.001. [Figure 12B] Same as above. [Figure 12C] Same as above. [Figure 12D] Same as above. [Figure 12E] Same as above. [Figure 12F] Same as above. [Figure 13A] Figure 13 shows the relative mRNA levels after treatment of HSA-LR mice with various concentrations of PMO-EEV 221-1120. Figure 13A shows the relative mRNA levels in the gastrocnemius, triceps, tibialis anterior, and diaphragm. Figure 13B shows the relative mRNA levels in the diaphragm. [Figure 13B] Same as above. [Figure 14A] Relative mRNA levels in quadriceps (FIG. 14A), gastrocnemius (FIG. 14B), triceps (FIG. 14C), and tibialis anterior (FIG. 14D) tissues after treatment of HSA-LR mice with various concentrations of PMO-EEV 221-1120 are shown. [Figure 14B] Same as above. [Figure 14C] Same as above. [Figure 14D] Same as above. [Figure 15A]Mouse DM1 splicing index (mDSI) for various genes in quadriceps (FIG. 15A), gastrocnemius (FIG. 15B), triceps (FIG. 15C), and tibialis anterior (FIG. 15D) tissues after treatment of HSA-LR mice with various concentrations of PMO-EEV 221-1120 is shown. [Figure 15B] Same as above. [Figure 15C] Same as above. [Figure 15D] Same as above. [Figure 16A] Figures 16A-B show the prevalence of RNA foci in the tibialis anterior muscle of HSA-LR mice after untreated or treatment with EEV-PMO 221-1120 (EEV-PMO-DM1-3, DM1-3). Figures 16A-B show images of tibialis anterior muscle tissue stained for RNA CUG foci (red) and nuclei (blue). Figure 16C is a plot quantitating the percentage of nuclei with CUG foci from the data associated with the images in Figures 16A-B. [Figure 16B] Same as above. [Figure 16C] Same as above. [Figure 17A] Figure 17A shows the dose-dependent response of drug levels in quadriceps (17A), triceps (17B), heart (17C), gastrocnemius (17D), tibialis anterior (17E), diaphragm (17F), brain (17H), liver (17I), and kidney (17J) tissues after treatment of HSA-LR mice with various concentrations of EEV-PMO-DM1-3. Figure 17K shows drug exposure in various tissues at a dosage level of 60 mpk. [Figure 17B] Same as above. [Figure 17C] Same as above. [Figure 17E] Same as above. [Figure 17F] Same as above. [Figure 17G] Same as above. [Figure 17H] Same as above. [Figure 17I] Same as above. [Figure 17J] Same as above. [Figure 17K] Same as above. [Figure 18]1 shows dose-dependent reduction of myotonia in HSA-LR mice after 7 days of treatment with 15, 30, 60, and 90 mpk EEV-PMO-DM1-3. [Figure 19A] 19A and 19C are plots showing the results of principal component analysis comparing gene expression in non-diseased mice (WT), DM1 mice (HSA-LR), and HSA-LR mice treated with PMO-EEV 221-1120. Figures 19A and 19C are plots showing three principal components, and Figures 19B and 19D are plots showing two principal components. [Figure 19B] Same as above. [Figure 19C] Same as above. [Figure 19D] Same as above. [Figure 20A] Figure 20 shows heat maps of differentially expressed genes between non-diseased mice (WT), DM1 mice (HSA-LR), and HSA-LR mice treated with 60 mpk of PMO-EEV 221-1120. Figure 20A shows a clustered heat map of 513 differentially expressed genes. Figure 20B shows a clustered heat map of 40 genes known to contain CTG-CUG repeats. [Figure 20B] Same as above. [Figure 21] Volcano plot showing global transcriptional changes across untreated HSA-LR mice and mice treated with PMO-EEV 221-1120. [Figure 22] 22A, 22B, 22C, 22D, and 22E are plots showing the results of principal component analysis of the Scube2 (FIG. 22A), Greb1 (FIG. 22B), Ttc7 (FIG. 22C), Txlnb(CUG)9 (FIG. 22D), and Ndrg3 (FIG. 22E) genes from non-diseased mice, HSA-LR mice, and HSA-LR mice treated with PMO-EEV 221-1120. [Figure 23A]RNA sequencing (RNAseq) data for Atp2a1 (Figure 23A, exon 22 boxed), Clcn1 (Figure 23B, exon 7a boxed), Nfix (Figure 23C, exon 7 boxed), and Mbn1 (Figure 23D, exon 5 boxed) from non-diseased mice (WT-saline), HSA-LR mice (HSA-LR saline), and HSA-LR mice treated with PMO-EEV 221-1120 are shown. Two reads are shown per treatment group. [Figure 23B] Same as above. [Figure 23C] Same as above. [Figure 23D] Same as above. [Figure 24] Shown are the percent splicing index (PSI) of individual exons of various genes of interest in non-diseased mice (WT-Saline), HSA-LR mice (HSA-LR Saline), and HSA-LR mice treated with PMO-EEV 221-1120. [Figure 25A] Drug levels in tibialis anterior (Figure 25A), gastrocnemius (Figure 25B), triceps (Figure 25C), and quadriceps (Figure 25D) tissues of HSA-LR mice after treatment with 80 mpk (60 mpk oligo, 80 mpk total drug) of EEV-PMO-DM1-3 for 1 to 4 weeks are shown. [Figure 25B] Same as above. [Figure 25C] Same as above. [Figure 25D] Same as above. [Figure 26A] Figures 26A-26D show drug levels in HSA-LR mice after treatment with a single dose of 80 mpk EEV-PMO-DM1-3. Figures 26A-26B show drug levels in the liver from 1 week to 12 weeks after treatment. Figures 26C-26D show drug levels in the kidney from 1 week to 12 weeks after treatment. [Figure 26B] Same as above. [Figure 26C] Same as above. [Figure 26D] Same as above. [Figure 27A]26A and 26B are plots showing exon inclusion levels of MBNL1 (exon 5, FIG. 26A), SOS1 (exon 25, FIG. 26B), and NFIX (exon 7, FIG. 26C) after treatment of muscle cells from DM1 patients with 30 μM EEV-PMO-DM1-3. [Figure 27B] Same as above. [Figure 27C] Same as above. [Figure 28A] Figures 28A-B show that EEV-PMO-DM1-3 reduces CUG nuclear foci (green) in nuclei (blue) in myocytes from DM1 patients. Figures 28A-B show images of myocytes from DM1 patients that were either untreated, treated with EEV-PMO-DM1-3, or untreated. Figure 28C shows quantification of the number of CUG foci per nucleus for the data associated with the image in Figure 28A. [Figure 28B] Same as above. [Figure 28C] Same as above. [Figure 29A] Raw (FIG. 29A) and normalized (FIG. 29B) data from a CELLTITER-GLO luminescent viability assay in which RPTEC cells were treated with various concentrations of PMO-DM1 or EEV-PMO-DM1-3 are shown. Melittin was used as a positive control. [Figure 29B] Same as above. [Figure 30A] Figure 30A shows images depicting RNA CUG repeat foci in cells from a DM1 patient (Figure 30A) and cells from a DM1 patient treated with EEV-PMO 221-1113 (Figure 30B). Cells were stained for nuclei (blue, Hoechst) and RNA CUG foci (green). Figure 30C is a plot of CUG RNA foci per nuclear area for data related to the images in Figures 30A-B. [Figure 30B] Same as above. [Figure 30C] Same as above. [Figure 31A]Figure 31A shows the prevalence of RNA CUG7 foci in HeLa, untreated HeLa480 cells, and HeLa480 cells treated with EEV-PMO 221-1113. Figure 31A shows images of cells stained for RNA CUG7 foci (green) and nuclei (blue). Figure 31B is a plot quantitating CUG7 foci per nuclear area for the data associated with the image in Figure 31A. [Figure 31B] Same as above. [Figure 32A] 32A and 32B are plots showing the percent inclusion of exon 5 in MBNL1 (FIG. 32A), exon 25 in SOS1 (FIG. 32B), and exon 7 in NFIX (FIG. 32C) after treatment of cells from a DM1 patient with 30 μM EEV-PMO 221-1113. [Figure 32B] Same as above. [Figure 32C] Same as above. [Figure 33A] Figure 33 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion) of MBNL1 (exon 5, Figure 33A), SOS1 (exon 25, Figure 33B), CLASP1 (exon 19, Figure 33C), NFIX (exon 7, Figure 33D), and INSR (exon 11, Figure 33E) after treatment of muscle cells from DM1 patients with various concentrations of PMO-EEV 221-1113. Significance was determined using a t-test. * p<0.05, ** p<0.01, *** p<0.001. [Figure 33B] Same as above. [Figure 33C] Same as above. [Figure 33D] Same as above. [Figure 33E] Same as above. [Figure 34A] RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion) of Atp2a1 (exon 22, FIG. 34A), Nfix (exon 7, FIG. 34B), Clcn1 (exon 7a, FIG. 34C), and Mbnl1 (exon 5, FIG. 34D) in gastrocnemius muscle tissue from mice treated with various concentrations of PMO 221 or EEV-PMO 221-1106. [Figure 34B] Same as above. [Figure 34C] Same as above. [Figure 34D] Same as above. [Figure 35A] Figure 35 shows RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion or exclusion) of Mbnl1 (exon 5, Figure 35A), Nfix (exon 7, Figure 35B), and Atp2a1 (exon 22, Figure 35C) in tibialis anterior muscle tissue from HSA-LR mice treated with either EEV 0221-1121 (21mer) or PMO-EEV 0325-1121 (24mer). [Figure 35B] Same as above. [Figure 35C] Same as above. [Figure 36A] 36A , 36B , and 36C show RT-PCR analysis of alternative RNA splicing events (e.g., exon inclusion) of Mbnl1 (exon 5, FIG. 36A ), Nfix (exon 7, FIG. 36B ), and Atp2a1 (exon 22, FIG. 36C ) in gastrocnemius muscle tissue from HSA-LR mice treated with either PMO-EEV 0221-1121 (21 mer) or PMO-EEV 0325-1121 (24 mer). [Figure 36B] Same as above. [Figure 36C] Same as above. [Figure 37] The major metabolite of PMO-EEV 220-1120 detected in vivo, PMO-0221a, is shown. [Figure 38A] Percent exon inclusion of MBNL1 (exon 5) in tibialis anterior (FIG. 38A) and gastrocnemius (FIG. 38B) muscles after treating Hela480 cells with various concentrations of EEV-PMO 221-1120 is shown. [Figure 38B] Same as above. [Figure 39A] Percent exon inclusion of NFIX (exon 7) in tibialis anterior (FIG. 39A) and gastrocnemius (FIG. 39B) muscles after treatment of Hela480 cells with various concentrations of EEV-PMO 221-1120 is shown. [Figure 39B] Same as above. [Figure 40A]Percent exon inclusion of Atp2a1 (exon 22) in tibialis anterior (FIG. 40A) and gastrocnemius (FIG. 40B) muscles after treatment of Hela480 cells with various concentrations of EEV-PMO 221-1120 is shown. [Figure 40B] Same as above. [Figure 41A] Figure 41A shows an image depicting RNA CUG repeat foci in Hela480 cells after treatment with various concentrations of EEV-PMO 221-1120. Figure 41B is a plot of RNA foci per nuclear area of data related to the image in Figure 41A. [Figure 41B] Same as above. [Figure 42A] Relative r(CUG480) repeat mRNA levels (Figure 42A), relative DMPK mRNA levels (Figure 42B), percent exon 5 inclusion of MBNL1 (Figure 42C), and percent exon 25 inclusion of SOS1 (Figure 42D) are shown in HeLa480 cells after treatment with various concentrations of EEV-PMO 221-1120. [Figure 42B] Same as above. [Figure 42C] Same as above. [Figure 42D] Same as above. [Figure 43] 1 is a bar graph showing examples of genes expressed in muscle tissue known to have CTG·CUG repeats. [Figure 44A] Figure 44 shows the reduction of phenotypic myotonia in the HSA-LR mouse model treated with 20 mpk PMO-EEV 221-1106. Figures 44A and 44C show relaxation plots. Figure 44B shows an example of a raw force trace. Figure 44D shows a representative electromyogram trace. [Figure 44B] Same as above. [Figure 44C] Same as above. [Figure 44D] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0030] compound In some embodiments, compounds are provided that modulate the level and / or activity of gene transcripts having an expanded CUG trinucleotide repeat. In some embodiments, the compounds of the present disclosure comprise at least one cyclic cell-penetrating peptide (cCPP) and a therapeutic moiety (TM). The cCPP facilitates entry of the TM into cells. In some embodiments, the compounds comprise an enosome escape vehicle (EEV) comprising the cCPP and an exocyclic peptide (EP). The cCPP or EEV can enable the TM to enter the cytosol or a cellular compartment and interact with a target transcript.
[0031] treatment part Generally, a TM is an effector moiety that elicits a response. In some embodiments, the TM induces a response by modulating the expression, activity, and / or level of a target transcript and / or target protein. In some embodiments, the target transcript comprises an expanded CUG trinucleotide repeat. In some embodiments, the TM modulates the level of a target transcript and / or target protein in a cell. In some embodiments, the TM reduces the level of a target transcript and / or target protein in a cell.
[0032] In several embodiments, a TM modulates the activity of a target transcript by decreasing the affinity between the target transcript and one or more proteins that bind to the target transcript. By decreasing the affinity between the target transcript and one or more proteins, the TM can effectively modulate the activity of one or more proteins that would otherwise be associated with the target transcript. For example, when one or more proteins are not bound to the target transcript, they are available to perform their function on other molecules. For example, if one or more proteins are involved in pre-mRNA processing, decreasing the affinity of one or more proteins for transcripts containing expanded CUG repeats may enable the one or more proteins to process pre-mRNA transcripts that do not contain expanded CUG repeats. Thus, the TM can modulate the activity, expression, and / or level of downstream genes (genes that do not contain expanded CTG repeats) that are regulated by one or more proteins whose interaction with the target transcript is disrupted by the TM.
[0033] In some embodiments, the TM comprises an oligonucleotide, a peptide, an antibody, and / or a small molecule. The class and identity of the TM depends on the mechanism used to regulate the level and / or activity of the target transcript containing the expanded CUG trinucleotide repeat.
[0034] Antisense Compounds In various embodiments, the compounds disclosed herein comprise a cell-penetrating peptide (CPP) conjugated to an antisense compound (AC).
[0035] The term "antisense compound" refers to an oligonucleotide sequence that is complementary or at least partially complementary to a target nucleotide sequence. An antisense compound is an oligonucleotide that contains natural DNA bases, modified DNA bases, natural RNA bases, modified RNA bases, natural RNA sugars, modified RNA sugars, natural DNA sugars, modified DNA sugars, natural internucleoside linkages, modified internucleoside linkages, or any combination thereof. Antisense compounds include, but are not limited to, antisense oligonucleotides, RNAi, microRNA, antagomir, aptamers, ribozymes, immunostimulatory oligonucleotides, decoy oligonucleotides, supermir, miRNA mimics, miRNA inhibitors, U1 adaptors, and combinations thereof.
[0036] In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a target transcript having an expanded CUG trinucleotide repeat. In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to an expanded CUG trinucleotide repeat in a target mRNA sequence. Several diseases, such as myotonic dystrophy type 1 (DM1), Fuchs endothelial corneal dystrophy (FECD), spinocerebellar ataxia-8 (SCA8), and Huntington's disease-like (HDL2), are associated with expanded CUG trinucleotide repeats. Table 1 provides examples of nucleotide repeat disorders and characteristics of genes with expanded nucleotide repeats associated with such disorders. The following documents describe exemplary oligonucleotides for treating tandem repeat diseases and are incorporated herein by reference in their entireties: Zain et al.Neurotherapeutics.2019;16(2):248-262, Zarouchlioti et al.Am J Hum Genet.2018;102(4):528-539, Fautsch et al.Prog Retin Eye Res.2021;81:100883. [Table 1]
[0037] In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to a nucleotide sequence in the target mRNA transcript that comprises an expanded CTG·CUG trinucleotide repeat. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to a nucleotide sequence in the target mRNA transcript that comprises an expanded CTG·CUG trinucleotide repeat.
[0038] In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a nucleotide sequence in a DMPK1 target transcript that contains an expanded CTG·CUG trinucleotide repeat. In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a nucleotide sequence in a TCF4 target transcript that contains an expanded CTG·CUG trinucleotide repeat. In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a nucleotide sequence in an ATXN8OS / ATXN8 target transcript that contains an expanded CTG·CUG trinucleotide repeat. In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a nucleotide sequence in a JPH3 target transcript that contains an expanded CTG·CUG trinucleotide repeat.
[0039] In embodiments, the AC comprises a nucleotide sequence at least partially complementary to a trinucleotide repeat in the 3' UTR of the target mRNA transcript. In embodiments, the AC comprises a nucleotide sequence at least partially complementary to an expanded CTG·CUG trinucleotide repeat in the 3' UTR of a DMPK1 target transcript. In embodiments, the AC comprises a nucleotide sequence at least partially complementary to an expanded CTG·CUG trinucleotide repeat in the 3' UTR of an ATXN8OS / ATXN8 target transcript. In embodiments, the AC comprises a nucleotide sequence at least partially complementary to an expanded CTG·CUG trinucleotide repeat in the 3' UTR of a JPH3 target transcript.
[0040] In some embodiments, the AC comprises a nucleotide sequence at least partially complementary to a trinucleotide repeat, such as a CTG·CUG repeat. In some embodiments, the target nucleotide sequence comprises at least one extended trinucleotide repeat (e.g., a CTG·CUG repeat). In some embodiments, the target nucleotide sequence comprises at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, or at least 2000 CTG·CUG trinucleotide repeats. In some embodiments, the extended trinucleotide repeat is within the 3'UTR of the target nucleotide sequence.
[0041] In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a contiguous extended trinucleotide repeat present in the target transcript, hi some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, and up to 50, up to 100, up to 150, up to 200, up to 300, up to 400, up to 500, up to 600, up to 700, up to 800, up to 900, up to 1000, or up to 2000 trinucleotide repeats in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 trinucleotide repeats in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 5 to 10 trinucleotide repeats in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 5 to 9 trinucleotide repeats in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 5 to 8 trinucleotide repeats in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 5-7 trinucleotide repeats in the target transcript, hi some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to 5-6 trinucleotide repeats in the target transcript.In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing with five trinucleotide repeats in the target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing with six trinucleotide repeats in the target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing with seven trinucleotide repeats in the target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing with eight trinucleotide repeats in the target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing with nine trinucleotide repeats in the target transcript. In embodiments, the AC comprises a nucleotide sequence.
[0042] In some embodiments, the AC may comprise a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a contiguous extended trinucleotide repeat present anywhere in the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a contiguous extended trinucleotide repeat present in the 3'UTR of the target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a contiguous extended trinucleotide repeat present in the 3'UTR of the DMPK1, SCA8, and / or HDL2 target transcript. In some embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a contiguous extended trinucleotide repeat present in an intron of the target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a consecutive extended trinucleotide repeat present in intron 3 of the TCF4 target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a consecutive extended trinucleotide repeat present in the CTG18.1 locus of the TCF4 target transcript. In embodiments, the AC comprises a nucleotide sequence that is at least partially complementary to and capable of hybridizing to at least a portion of a consecutive extended trinucleotide repeat present in an exon of the target transcript.
[0043] In some embodiments, AC is 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, or 45 or more nucleic acids in length. In some embodiments, AC is 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, or 10 or less nucleic acids in length. In some embodiments, AC is 5 to 50, 5 to 45, 5 to 40, 5 to 35, 5 to 30, 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleic acids in length. In some embodiments, AC is 10 to 50, 10 to 45, 10 to 40, 10 to 35, 10 to 30, 10 to 25, 10 to 20, or 10 to 15 nucleic acids in length. In several embodiments, the AC is 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, 15 to 25, or 15 to 20 nucleic acids in length. In several embodiments, the AC is 20 to 50, 20 to 45, 20 to 40, 20 to 35, 20 to 30, or 20 to 25 nucleic acids in length. In several embodiments, the AC is 25 to 50, 25 to 45, 25 to 40, 25 to 35, or 25 to 30 nucleic acids in length. In several embodiments, the AC is 30 to 50, 30 to 45, 30 to 40, or 30 to 35 nucleic acids in length. In several embodiments, the AC is 35 to 50, 35 to 45, or 35 to 40 nucleic acids in length. In several embodiments, the AC is 40 to 50 or 40 to 45 nucleic acids in length. In several embodiments, the AC is 45 to 50 nucleic acids in length. In embodiments, the AC is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleic acids in length.
[0044] In some embodiments, the AC has 100% complementarity to the target nucleotide sequence. In some embodiments, the AC does not have 100% complementarity to the target nucleotide sequence. As used herein, the term "percent complementarity" refers to the number of nucleobases (e.g., natural or modified nucleobases) of an AC that have nucleobase complementarity with corresponding nucleobases of an oligomeric compound or nucleic acid (e.g., a target nucleotide sequence), divided by the total length (number of nucleobases) of the AC. Those of skill in the art will recognize that mismatches are possible without eliminating activity of the antisense compound.
[0045] In some embodiments, the AC contains 20% or less, 15% or less, 10% or less, 5% or less, or 0% mismatch with the target nucleotide sequence. In some embodiments, the AC contains 5% or more, 10% or more, or 15% or more mismatch. In some embodiments, the AC contains 0-5%, 0-10%, 0-15%, or 0-20% mismatch with the target nucleotide sequence. In some embodiments, the AC contains 5%-10%, 5%-15%, or 5%-20% mismatch with the target nucleotide sequence. In some embodiments, the AC contains 10%-15% or 10%-20% mismatch with the target nucleotide sequence. In some embodiments, the AC contains 10%-20% mismatch with the target nucleotide sequence.
[0046] In some embodiments, the AC has 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementarity to the target nucleotide sequence. In some embodiments, the AC has 100% or less, 99% or less, 98% or less, 97% or less, 96% or less, 95% or less, 90% or less, or 85% or less complementarity to the target nucleotide sequence. In some embodiments, the AC has 80% to 100%, 80% to 99%, 80% to 98%, 80% to 97%, 80% to 96%, 80% to 95%, 80% to 90%, or 80% to 85% complementarity to the target nucleotide sequence. In some embodiments, the AC has 85% to 100%, 85% to 99%, 85% to 98%, 85% to 97%, 85% to 96%, 85% to 95%, or 85% to 90% complementarity to the target nucleotide sequence. In some embodiments, the AC has 90% to 100%, 90% to 99%, 90% to 98%, 90% to 97%, 90% to 96%, or 90% to 95% complementarity to the target nucleotide sequence. In some embodiments, the AC has 95% to 100%, 95% to 99%, 95% to 98%, 95% to 97%, or 95% to 96% complementarity to the target nucleotide sequence. In some embodiments, the AC has 96% to 100%, 96% to 99%, 96% to 98%, or 96% to 97% complementarity to the target nucleotide sequence. In some embodiments, the AC has 97% to 100%, 97% to 99%, or 97% to 98% complementarity to the target nucleotide sequence. In some embodiments, the AC has 98% to 100% or 98% to 99% complementarity to the target nucleotide sequence. In some embodiments, the AC has 99% to 100% complementarity to the target nucleotide sequence.
[0047] In some embodiments, the incorporation of nucleotide affinity modifications allows for more mismatches compared to unmodified compounds.Similarly, certain oligonucleotide sequences may tolerate mismatches better than other oligonucleotide sequences.Those skilled in the art can determine the appropriate number of mismatches between AC and target nucleotide sequence, for example, by determining the thermal melting temperature (Tm).Tm or ΔTm can be calculated by techniques well known to those skilled in the art.For example, the technique described in Freier et al. (Nucleic Acids Research, 1997, 25, 22: 4429-4443) allows those skilled in the art to evaluate nucleotide modifications for their ability to increase the melting temperature of RNA:DNA duplexes.
[0048] In some embodiments, the AC comprises a nucleotide sequence that is itself a trinucleotide repeat, i.e., a CAG trinucleotide repeat. The reverse complement of 5'-CAG-3' has 100% complementarity and can hybridize to a 5'-CUG-3' trinucleotide repeat. In some embodiments, the AC comprises 1 to 50 CAG repeats. In some embodiments, the CAG repeats are contiguous. In some embodiments, the CAG repeats are discontinuous. In some embodiments, the AC comprises a nucleotide sequence that includes an imperfect CAG repeat at either the 5' or 3' end. For example, in some embodiments, the AC is AG(CAG) n , G(CAG) n , (CAG) n AG, or (CAG) nA, where n is an integer between 1 and 50. In embodiments, the AC comprises a nucleotide sequence comprising 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 20 or more, 30 or more, or 50 or more CAG repeats. In embodiments, the AC comprises a nucleotide sequence comprising 50 or fewer, 40 or fewer, 30 or fewer, 20 or fewer, 10 or fewer, 9 or fewer, 8 or fewer, 7 or fewer, 6 or fewer, 5 or fewer, 4 or fewer, 3 or fewer, or 2 or fewer CAG repeats. In embodiments, the AC comprises a nucleotide sequence comprising 2 to 50, 2 to 20, 2 to 10, 4 to 10, 5 to 10, 6 to 10, 6 to 9, 6 to 8, or 6 to 7 CAG repeats. In embodiments, the AC comprises any one of the nucleotide sequences in Table 2 (SEQ ID NOs: 151-291). [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8]
[0049] In some embodiments, an AC having nucleotides containing a CAG repeat can include additional nucleotide sequences at the 5' end, the 3' end, or both of the CAG repeats. In some embodiments, the additional nucleotide sequences can have 80% to 100% or 95% to 100% complementarity to the portion of the target transcript to which they hybridize. Additional nucleotide sequences can be added to the CAG repeat nucleotide sequence to increase the selectivity of the AC to hybridize to a particular target transcript.
[0050] In some embodiments, the AC comprises a nucleotide sequence that includes 1 to 50 CAG repeats and is a gapmer. A gapmer is an oligonucleotide that is a DNA / RNA hybrid and induces RNase degradation. For example, a gapmer can have a central DNA or DNA mimetic segment flanked by RNA or RNA mimetic segments on both the 5' and 3' ends of the DNA or DNA mimetic segment. In some embodiments, the AC comprises a gapmer that includes a nucleotide sequence that hybridizes to a target nucleic acid sequence of a target transcript that is distinct from the extended CUG repeats of the target transcript.
[0051] In embodiments, the AC of the present disclosure is a gapmer oligonucleotide as disclosed in US Pat. No. 9,550,988, the disclosure of which is incorporated herein by reference.
[0052] In embodiments, the AC of the present disclosure comprises the sequence and / or structure of any one of the DMPK-targeting ACs disclosed in U.S. Patent Publication No. 2017 / 0260524, the disclosure of which is incorporated herein by reference.
[0053] In embodiments, the AC of the present disclosure can be a polymerizable compound selected from the group consisting of U.S. Patent Publication Nos. 2003 / 0235845 A1, 2006 / 0099616 A1, 2013 / 0072671 A1, 2014 / 0275212 A1, 2009 / 0312532 A1, 2010 / 0125099 A1, 2010 / 0125099 A1, 2009 / 0269755 A1, 2011 / 0294753 A1, 2012 / 0022134 A1, 2011 / 0263682 A1, 2014 / 0128592 A1, and 2015 / 0073037 A1, and the sequences and / or structures of any one of the ACs or oligonucleotides disclosed in 2012 / 0059042 A1, the contents of each of which are incorporated herein in their entirety for all purposes.
[0054] When using ACs to target and / or hybridize to expanded CTG·CUG repeats, care must be taken to avoid off-target effects, such as unintended binding of ACs to off-target transcripts containing CTG·CUG repeats (e.g., transcripts containing CTG·CUG repeats that are not expanded CTG·CUG repeats). In silico analysis of the human genome reveals that a total of 63 human genes contain CTG·CUG repeats (Uhlen, et al., Science 2015 347(6220):1260419). These 63 genes can be ranked by mRNA expression and the amount of protein expressed in total muscle (cardiac, skeletal, and smooth muscle). Expression levels can be quantified using RPM (reads per million), mRNA expression in FPKM (fragments per kilobase of transcript per million mapped fragments), and protein expression in pTPM (transcripts per million protein-coding genes), using >10 RPM as a cutoff for insignificant expression. Figure 43 shows the results of such an in silico analysis. Thirty-six genes exhibit expression levels greater than 10 RPM. Of these 36 genes, only three (excluding DMPK) had more than 10 CTG·CUG repeats. Genes with ≤10 CTG·CUG repeats represent the lowest risk of off-target binding and toxicity. Nevertheless, the number of CTG·CUG repeats (11–24) in these three genes (TCF4, CASK, and MAP3k4) is significantly lower than that found in patients with classic congenital DM1. For example, patients with late-onset DM1 have 100–600 CTG·CUG repeats in DMPK, patients with classical DM1 have 250–750 CTG·CUG repeats in DMPK, and patients with congenital DM1 have 750–1,400 CTG·CUG repeats in DMPK. The same in silico analysis can be performed on the liver and kidney. CASK is the only significant gene with more than 10 CTG·CUG repeats in the kidney. In the liver, no genes with more than 10 CTG·CUG repeats were significant.
[0055] The ACs described herein contain one or more asymmetric centers and may therefore give rise to enantiomers, diastereomers, and other stereoisomeric configurations, which may be defined in terms of absolute stereochemistry as (R) or (S), α or β, (D) or (L). The antisense compounds provided herein include all such possible isomers, as well as their racemic and optically pure forms.
[0056] The effectiveness of ACs can be assessed by evaluating the antisense activity resulting from their administration. As used herein, the term "antisense activity" refers to any detectable and / or measurable activity resulting from the hybridization of an antisense compound to its target nucleotide sequence. Such detection and / or measurement can be direct or indirect. In some embodiments, antisense activity is assessed by detecting and / or measuring the amount of protein expressed from a transcript of interest. In some embodiments, antisense activity is assessed by detecting and / or measuring the amount of a transcript of interest. In some embodiments, antisense activity is assessed by detecting and / or measuring the amount of alternatively spliced RNA and / or the amount of a protein isoform translated from a target transcript.
[0057] AC adjustment mechanism In embodiments, an AC can modulate the activity and / or level of a target transcript in a cell. Figure 1 shows an exemplary mechanism of how an AC modulates the level and / or activity of a target transcript.
[0058] In some embodiments, the AC can modulate the level of a target transcript in a cell. For example, in some embodiments, where the AC is a gapper, binding of the AC to the target transcript induces degradation of the target transcript via the RNase H pathway (Figure 1, arrows A and B). In some embodiments, the gapmer hybridizes to a portion of the target transcript that is distinct from the extended CUG trinucleotide repeat, thereby inducing degradation of the target transcript via the RNase H pathway (Figure 1, arrow A). In some embodiments, the gapmer hybridizes to at least a portion of the extended CUG repeat within the target transcript, thereby inducing degradation of the target transcript via the RNase H pathway (Figure 1, arrow B).
[0059] In some embodiments, the AC may modulate the activity of a target transcript. Modulating activity may include increasing or decreasing the ability of the target transcript to bind to a binding partner. In some embodiments, the AC may modulate the activity of a target transcript by decreasing the ability of the target transcript to bind to one or more proteins that may associate with the target transcript, specifically proteins that associate with at least a portion of the expanded CUG repeat of the target transcript ( FIG. 1 , arrow C). In some embodiments, decreasing the ability of the target transcript to bind to one or more proteins includes decreasing the affinity of the target transcript for one or more proteins. In some embodiments, decreasing the ability of the target transcript to bind to one or more proteins includes partial or complete steric blocking of the target transcript from binding to one or more proteins. For example, the AC may occupy at least a portion of a binding site that would be occupied by one or more proteins if not sterically blocked. In some embodiments, the binding site of one or more proteins to the target transcript includes at least a portion of the expanded trinucleotide repeat of the target transcript. Thus, in some embodiments, the AC may occupy at least a portion of an extended trinucleotide repeat (e.g., an extended CUG repeat) that would otherwise be occupied by one or more proteins. Partial steric blocking of the target transcript may result in a decrease in affinity between the target transcript and one or more proteins. For example, in some embodiments, the AC binds to at least a portion of the extended CUG repeat of the target transcript, thereby sterically blocking the target transcript and / or decreasing the affinity of proteins that may bind to the extended CUG repeat (Figure 1, arrow C). The following review describes additional uses for sterically blocked antisense oligonucleotides and is incorporated herein by reference in its entirety: Roberts et al. Nature Reviews Drug Discovery (2020) 19:673-694.
[0060] The CUG repeats of the extended CUG repeat can form a double-stranded hairpin structure. In pathological conditions, proteins bind to this double-stranded hairpin structure and become trapped, preventing them from performing other functions. In embodiments, an AC binds to at least a portion of the double-stranded hairpin structure, thereby sterically blocking the double-stranded hairpin structure for a protein-binding partner and / or reducing the affinity of the double-stranded hairpin structure for a protein-binding partner. In embodiments, an AC binds to at least a portion of the single-stranded extended CUG repeat, thereby inhibiting the formation of the double-stranded hairpin structure and therefore inhibiting the binding of one or more proteins to the double-stranded hairpin structure. In embodiments, hybridization of an AC to the double-stranded hairpin structure sterically blocks binding to the double-stranded hairpin structure and / or reduces the affinity for one or more proteins that bind to the double-stranded hairpin structure. In embodiments, hybridization of an AC to at least a portion of the single-stranded region of the extended trinucleotide repeat inhibits the formation of the double-stranded hairpin structure.
[0061] Reducing the ability of a target transcript to bind one or more proteins may allow the one or more proteins to perform other functions, such as, for example, regulating the splicing of downstream transcripts (transcripts that do not contain the expanded CUG repeat). Thus, in embodiments, reducing the ability of a target transcript to bind one or more proteins may increase the levels of one or more proteins in the cell that are available to provide other function(s) to other transcripts. In embodiments, reducing the ability of a target transcript to bind one or more proteins may increase the cytoplasmic levels of one or more proteins in the cell that are available to provide other function(s) to other transcripts. Thus, in embodiments, binding of an AC to a target transcript may result in modulation of the level and / or activity of one or more proteins that interact with the target transcript.
[0062] In embodiments, hybridization of AC to at least a portion of the expanded CUG repeat in the target transcript reduces the affinity of MNBL1 for the target transcript and / or sterically blocks the binding of MNBL1 to the target transcript. MNBL1 is a splicing factor that regulates the splicing of downstream gene transcripts. In the DM1 disease phenotype, MNBL1 binds to the expanded CUG repeat in the target transcript. Although bound to the target transcript, MNBL1 is trapped in the nucleus and is unable to regulate the splicing of downstream gene transcripts (transcripts that do not contain the expanded CUG repeat). In embodiments, hybridization of AC to at least a portion of the expanded CUG repeat in the target transcript sterically blocks MNBL1 from the target transcript and / or reduces the affinity of MNBL1 for the target transcript, thereby enabling it to regulate the splicing of downstream gene transcripts. In some embodiments, hybridization of the AC to at least a portion of the expanded CUG repeat in the target transcript sterically blocks the target transcript of MNBL1 and / or reduces the affinity of MNBL1 for the target transcript, thereby increasing the amount of free (e.g., not bound to a transcript having a CUG repeat). In some embodiments, hybridization of the AC to at least a portion of the expanded CUG repeat in the target transcript sterically blocks the target transcript of MNBL1 and / or reduces the affinity of MNBL1 for the target transcript, thereby reducing the amount of MBNL1 that binds to and is captured by the target transcript.
[0063] In embodiments where the target transcript is DMPK, hybridization of AC to at least a portion of the expanded CUG repeat reduces the affinity of MNBL1 for the target transcript and / or sterically blocks the binding of MNBL1 to the target transcript. MNBL1 is a splicing factor that regulates the splicing of downstream gene transcripts. In the DM1 disease phenotype, MNBL1 binds to the expanded CUG repeat of DMPK1. While bound to DMPK1, MNBL1 is trapped in the nucleus and is unable to regulate the splicing of downstream gene transcripts (transcripts that do not contain the expanded CUG repeat). In embodiments, hybridization of AC to at least a portion of the expanded CUG repeat in DMPK sterically blocks MNBL1 from the DMPK transcript and / or reduces the affinity of MNBL1 for the DMPK transcript, thereby enabling it to regulate the splicing of downstream gene transcripts. In some embodiments, hybridization of the AC to at least a portion of the expanded CUG repeat in DMPK sterically blocks the DMPK transcript of MNBL1 and / or reduces the affinity of MNBL1 for the DMPK transcript, thereby increasing the amount of free (e.g., not bound to a transcript having a CUG repeat). In some embodiments, hybridization of the AC to at least a portion of the expanded CUG repeat in DMPK sterically blocks the DMPK transcript of MNBL1 and / or reduces the affinity of MNBL1 for the DMPK transcript, thereby reducing the amount of MBNL1 that binds to and is captured by the DMPK1 transcript.
[0064] In embodiments where the target transcript is DMPK, hybridization of AC to at least a portion of the expanded CUG repeat results in a decrease in CUGBP1 levels. In DM1 pathology, free (functional) MBNL1 levels are decreased and free (functional) CUGBP1 levels are increased. Increased CUGBP1 levels are associated with the pathology. Thus, in embodiments, hybridization of AC to at least a portion of the expanded CUG repeat results in an increase in free (functional) MBNL1 levels and / or a decrease in free (functional) CUGBP1 levels.
[0065] Reducing the ability of a target transcript to bind to one or more proteins can reduce or inhibit the formation of CUG repeat foci. Transcripts containing extended nucleotide repeats (e.g., extended CUG repeats) can be transcribed and then trapped in the nucleus. Within the nucleus, the trapped transcripts can form aggregates. Proteins that bind to the transcripts can then nucleate on the sequenced transcripts and / or trap the transcript aggregates, thereby forming extended nucleotide repeat (e.g., CUG repeat) foci. The CUG repeat foci can be visible using a microscope. In some embodiments, reducing the ability of a target transcript to bind to one or more proteins can reduce or inhibit the formation of aggregates containing the target transcript. In some embodiments, reducing the ability of a target transcript to bind to one or more proteins can reduce or inhibit the nucleation of those one or more proteins on the target transcript, on the double-stranded hairpin region of the transcript, or on aggregates of the target transcript. In embodiments where the target transcript is DMPK, reducing the ability of the DMPK1 target transcript to bind to MNBL1 can reduce or inhibit MNBL1 nucleation on the DMPK1 target transcript or DMPK1 target transcript aggregates. In embodiments, hybridization of an AC to a target transcript can result in the inhibition or reduction of the formation of CUG repeat nuclear foci. In embodiments, hybridization of an AC to a target transcript can result in the inhibition or reduction of the formation of CUG repeat nuclear foci formed from DMPK, TCF4, JPH3, and / or ATXN8OS / ATXN8 target transcripts.
[0066] In some embodiments, hybridization of an AC to a target transcript can result in modulation of the level, expression, and / or activity of one or more downstream genes. For example, hybridization of an AC to a target transcript can be used to induce degradation of the target transcript or to sterically block or reduce the affinity of the target transcript for one or more proteins, thereby enabling one or more proteins captured by the target transcript to modulate the expression, level, and / or activity of downstream genes. For example, in some embodiments, the one or more proteins can include proteins involved in regulating splicing of one or more downstream transcripts (transcripts that do not contain expanded CUG repeats). In some embodiments, splicing of the downstream transcript is altered when a protein involved in splicing binds to and is captured by the target transcript. For example, altered splicing can include the elimination of one or more exons or the inclusion of one or more introns in the transcript, resulting in the expression of various protein isoforms. In some embodiments, altered splicing can result in the inclusion of exons and / or introns containing premature stop codons, resulting in truncated isoforms that may have no activity or adverse activity. Altered splicing of downstream gene transcripts can result in changes in the level, folding, and / or activity of downstream gene products that may be associated with disease phenotypes. When not bound to target transcripts containing expanded CUG repeats, proteins involved in splicing are free to regulate splicing, resulting in corrective (or rescue) splicing of the downstream gene transcript, thereby at least partially restoring the protein level, folding, and / or activity of downstream gene products associated with a healthy phenotype.
[0067] In some embodiments, hybridization of an AC to a target transcript can result in modulation of splicing of a downstream gene transcript regulated by a protein captured by the target transcript during a pathology associated with an extended nucleotide repeat (e.g., an extended trinucleotide repeat). In diseases associated with extended trinucleotide repeats, the downstream gene transcript is often misprocessed, e.g., misspliced. Missplicing of the downstream gene transcript can result in a gene product that is disrupted pre-translationally or translated into a protein with abnormal structure and / or function. For example, capture of a protein that regulates processing of a downstream gene transcript can result in inclusion of an exon and / or intron with a premature stop codon, inclusion of an intron, exclusion of an exon, and / or inclusion of an alternative exon, which can result in a transcript and / or gene product that is disrupted pre-translationally or translated into a gene product with abnormal function. Changes in the levels of downstream gene transcripts and / or gene products and / or the abnormal structure and / or function of downstream gene products are associated with expanded trinucleotide disease phenotypes. In some embodiments, hybridization of an AC to a target transcript can result in modulation of exon inclusion, exon exclusion, intron inclusion, and / or intron exclusion in downstream transcripts whose splicing is regulated by proteins trapped in the target transcript during pathology. Thus, hybridization of an AC to a target transcript can result in upregulation of downstream protein isomers and / or transcripts associated with a healthy phenotype. Similarly, hybridization of an AC to a target transcript can result in downregulation (e.g., suppression) of downstream transcripts and / or protein isomers associated with a disease phenotype.
[0068] In embodiments where the target transcript is DMPK, hybridization of an AC to the target transcript can result in modulation of splicing of downstream gene transcripts regulated by proteins captured by DMPK target transcripts during pathology. DM1 results in missplicing of several downstream gene transcripts. Misspliced genes are associated with disease phenotypes. Thus, modulation of gene splicing can include correcting (e.g., rescuing) gene splicing to result in gene products of downstream genes associated with healthy phenotypes. In embodiments where the target transcript is DMPK, hybridization of an AC to the target transcript can result in modulation of splicing of downstream gene transcripts regulated by MNBL1, a splicing regulator captured by DMPK target transcripts during pathology. In embodiments where the target transcript is DMPK, hybridization of an AC to the target transcript can result in modulation of splicing of downstream gene transcripts regulated by CUGBP1, a protein whose activity is affected by expanded CUG repeats. In embodiments in which the target transcript is DMPK, hybridization of AC to the target transcript can result in the correct processing (eg, splicing) of downstream genes regulated by MNBL1 and / or CUGBP1.In embodiments where the target transcript is DMPK, hybridization of the AC to the target transcript identifies the following: 4833439L19Rik, Abcc9, Atp2a1, Arhgef10, Arhgap28, Armcx6, Angel1, Best3, Bin1, Brd2, Cacna1s, Cacna2d1, Cpd, Cpeb3, Ccpg1, Clasp1, ClC-1, Clcn1, Clk4, Cpeb2, Camk2g, Capzb, Copz2, Coch, cTNT, Ctu2, Cyp2s1, Dctn4, Dnm1l, Eya4, Efna3, Efna2, Fbxo31, Fbxo21, Frem2, Fgd4, Fuca1, Fn1, Gogla4, Gpr37l1, Greb1, Heg1, Insr, Impdh2, IR, Itgav, Jag2, Klc1, Kcan6, Kif13a, Ldb3, Lrrfip2, Mapt, Macf1, Map3k4, Mapkap 1, Mbnl1, Mllt3, Mbnl2, Mef2c, Mpdz, Mrpl1, Mxra7, Mybpc1, Myo9a, Ncapd3, Ngfr, Ndrg3, Ndufv3, Neb, Nfix, Numa1, Opa 1, Pacsin2, Pcolce, Pdlim3, Pla2g15, Phactr4, Phka1, Phtf2, Ppp1r12b, Ppp3cc, Ppp1cc, Ramp2, Rapgef1, Rur1, Ryr1, S This may result in modulation of splicing of downstream genes, including, but not limited to, orcs2, Spsb4, Scube2, Sema6c, Sfc8a3, Slain2, Sorbsl, Spag9, Tmem28, Tacc1, Tacc2, Ttc7, Tnik, Tnfrsf22, Tnfrsf25, Trappc9, Trim55, Ttn, Txnl4a, Txlnb, Ube2d3, Vsp39, or any combination thereof.
[0069] Missplicing of many of the downstream gene transcripts mentioned above results in specific DM1 disease phenotypes. For example, MNBL1 is a splicing factor with loss of function in DM1 due to exon 5 inclusion. MNBL1 is trapped by the DMPK CUG expansion, forming RNA nuclear foci. In addition, SOS1 promotes Ras activation and positively regulates the RAS / MAPK signaling pathway. In DM1, exon 25 of SOS1 is excluded, leading to inhibition of the muscle hypertrophy pathway. In DM1, IR / INSR has exon 11 exclusion, which increases the levels of low-signaling non-muscle isoforms in DM1 and reduces the metabolic response to insulin (insulin resistance). Similarly, exon 78 exclusion of DMD is observed in DM. Exon 78 exclusion results in an out-of-frame transcript in the C-terminal domain. This mutant protein is expressed in DM1 patients and is associated with the mechanism responsible for muscle wasting in these patients. BIN1 is required for proper myotubule formation (EC coupling). Exon 11 exclusion results in an inactive isoform and is found in DM1 patients. LDB3 interacts with α-actinin at the Z-zone of striated muscle, maintaining muscle structure. Exon 11 inclusion of LDB3 detected in DM1 results in reduced affinity for protein kinase C (PKC). As a result, PKC is overactive in DM1. In embodiments, modulation of one or more downstream genes results in correction or rescue of transcript splicing associated with a healthy phenotype. Thus, in embodiments, hybridization of AC to DMPK target transcripts results in rescue of missplicing of downstream genes / transcripts, thereby reducing the levels of downstream genes / transcripts associated with a disease phenotype. Thus, in embodiments, hybridization of AC to DMPK target transcripts results in rescue of missplicing of downstream genes / transcripts, thereby increasing the levels of downstream genes / transcripts associated with a healthy phenotype.
[0070] In some embodiments, the AC inhibits expression of the target transcript. In some embodiments, the AC inhibits expression of the target transcript by blocking the pre-mRNA processing machinery and / or translation machinery from accessing and / or completing translation and / or pre-mRNA processing. In some embodiments, the AC inhibits expression of the target transcript by inducing degradation of the target transcript, for example, via the RNase H pathway.
[0071] AC structure ACs include oligonucleotides and / or oligonucleosides. Oligonucleotides and / or oligonucleosides are nucleotides or nucleosides linked via internucleoside linkages. Nucleosides comprise a pentose sugar (e.g., ribose or deoxyribose) and a nitrogenous base covalently linked to the sugar. Naturally occurring (or traditional bases) bases found in DNA and / or RNA are adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Naturally occurring sugars (or traditional sugars) found in DNA and / or RNA are deoxyribose (DNA) and ribose (RNA). The naturally occurring nucleoside bond (or traditional internucleoside bond) is a phosphodiester bond. In embodiments, ACs of the present disclosure can have all natural sugars, bases, and internucleoside linkages.
[0072] Chemically modified nucleosides are routinely incorporated into antisense compounds to enhance one or more properties, such as nuclease resistance, pharmacokinetics, or affinity for target RNA. In embodiments, the ACs of the present disclosure may have one or more modified nucleosides. In embodiments, the ACs of the present disclosure may have one or more modified sugars. In embodiments, the ACs of the present disclosure may have one or more modified bases. In embodiments, the ACs of the present disclosure may have one or more modified internucleoside linkages.
[0073] Generally, a nucleobase is any group containing one or more atoms or groups of atoms that can hydrogen bond to the base of another nucleic acid. In addition to "unmodified" nucleobases or "natural" nucleobases (A, G, T, C, and U), many modified nucleobases or nucleobase mimics are known to those skilled in the art and are suitable for the compounds described herein. Generally, modified nucleobases refer to nucleobases that are structurally very similar to their parent nucleobases, such as 7-deazapurine, 5-methylcytosine, 2-thio-dT (Figure 2), or G-clamp. Generally, nucleobase mimics are nucleobases that contain more complex structures than modified nucleobases, such as tricyclic phenoxazine nucleobase mimics. Methods for preparing the above-mentioned modified nucleobases are well known to those skilled in the art.
[0074] In embodiments, the AC may include one or more nucleosides with modified sugar moieties. In embodiments, the furanosyl sugar of a natural nucleoside may have a 2' modification, a modification to create a constrained nucleoside, etc. (See Figure 2). For example, in embodiments, the furanosyl sugar ring of a natural nucleoside can be modified in several ways, including, but not limited to, adding a substituent, bridging two non-geminal ring atoms to form a bicyclic nucleic acid (BNA) or locked nucleic acid, replacing an oxygen of the furanosyl ring with a C or N, and / or substituting an atom or group (See Figure 2). Modified sugars are well known and can be used to increase or decrease the affinity of the AC for its target nucleotide sequence. Modified sugars can also be used to increase the AC's resistance to nucleases. The sugar can also be replaced with, among other things, a sugar mimetic group. In embodiments, one or more sugars of the nucleosides of the AC are replaced with a methylenemorpholine ring, as shown at 19 in Figure 2.
[0075] In embodiments, the AC comprises one or more nucleosides containing a bicyclic modified sugar (BNA, also called a bridged nucleic acid). Examples of BNAs include, but are not limited to, LNA (4'-(CH)-O-2' bridge), 2'-thio-LNA (4'-(CH)-S-2' bridge), 2'-amino-LNA (4'-(CH)-NR-2' bridge), ENA (4'-(CH)-O-2' bridge), 4'-(CH)-2' bridged BNA, 4'-(CHCH(CH))-2' bridged BNA," cEt (4'-(CH(CH)-O-2' bridge), and cMOE BNA (4'-(CH(CHOCH)-O-2' bridge). BNAs have been prepared and are described in the patent and scientific literature (Srivastava, et al. J. Am. Chem. Soc. (2007), ACS online advance publication, 10.1021 / ja071106y; Albaek et al. al.J.Org.Chem.(2006),71,7731-7740, Fluiter,et al.Chembiochem(2005),6,1104-1109,Singh et al.,Chem.Commun.(1998),4,455-456,Koshkin et al. al., Tetrahedron (1998), 54, 3607-3630, Wahlestedt et al., Proc. Natl. Acad. Sci. USA (2000), 97, 5633-5638, Kumar et al. al.,Bioorg.Med.Chem.Lett.(1998),8,2219-2222, WO94 / 14226, WO2005 / 021570, Singh et al. al., J. Org. Chem. (1998), 63, 10035-10039, WO2007 / 090071, U.S. Patent Nos. 7,053,207, 6,268,490, 6,770,748, 6,794,499, 7,034,133, and 6,525,191, and U.S. Pre-Grant Publication Nos. 2004-0171570, 2004-0219565, 2004-0014959, 2003-0207841, 2004-0143114, and 20030082807).
[0076] In some embodiments, the AC comprises one or more nucleosides, including locked nucleic acids (LNAs), in which the 2'-hydroxyl group of the ribosyl sugar ring is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene bond to form a bicyclic sugar moiety (reviewed in Elayadi et al., Curr. Opinion Invens. Drugs (2001), 2,558-561; Braasch et al., Chem. Biol. (2001), 8, 1-7; and Orum et al., Curr. Opinion Mol. Ther. (2001), 3,239-243; see also U.S. Patent Nos. 6,268,490 and 6,670,461). This bond can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom; the term LNA is used for the bicyclic moiety; if this position is an ethylene group, the term ENA™ is used (Singh et al., Chem. Commun. (1998), 4, 455-456; ENA™; Morita et al., Bioorganic Medicinal Chemistry (2003), 11, 2211-2226). LNA and other bicyclic sugar analogs exhibit very high duplex thermal stability with complementary DNA and RNA (Tm = +3 to +10°C), stability against 3'-exonucleolysis, and good solubility. Potent, non-toxic antisense oligonucleotides containing LNA have been described (Wahlestedt et al., Proc. Natl. Acad. Sci. USA (2000), 97, 5633-5638).
[0077] A similarly studied isomer of LNA is alpha-L-LNA, which has been shown to have excellent stability against 3'-exonucleases. Alpha-L-LNA has been incorporated into antisense gapmers and chimeras that have demonstrated potent antisense activity (Frieden et al., Nucleic Acids Research (2003), 21, 6365-6372).
[0078] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine, and uracil, as well as their oligomerization and nucleic acid recognition properties, have been described (Koshkin et al., Tetrahedron, 1998, 54, 3607-3630). LNAs and their preparation are also described in WO98 / 39352 and WO99 / 14226.
[0079] Analogs of LNA, phosphorothioate-LNA, and 2'-thio-LNA have also been prepared (Kumar et al., Bioorg. Med. Chem. Lett., 1998, 8, 2219-2222). The preparation of LNA analogs containing oligodeoxyribonucleotide duplexes as substrates for nucleic acid polymerases has also been described (Wengel et al., WO99 / 14226). Furthermore, the synthesis of 2'-amino-LNA, a conformationally restricted, high-affinity oligonucleotide analog, has been described (Singh et al., J. Org. Chem. (1998), 63, 10035-10039). In addition, 2'-amino-LNA and 2'-methylamino-LNA have been prepared, and the thermal stability of their duplexes with complementary RNA and DNA strands has previously been reported.
[0080] Methods for preparing modified sugars are well known to those of skill in the art. Some representative patents and publications that teach the preparation of such modified sugars include U.S. Patent Nos. 4,981,957, 5,118,800, 5,319,080, 5,359,044, 5,393,878, 5,446,137, 5,466,786, 5,514,785, 5,519,134, 5,567,811, and 5,576,427. , 5,591,722, 5,597,909, 5,610,300, 5,627,053, 5,639,873, 5,646,265, 5,658,873, 5,670,633, 5,792,747, 5,700,920, and 6,600,032, and WO2005 / 121371.
[0081] internucleoside bond Described herein are internucleoside linking groups that link nucleosides or otherwise modified nucleoside monomer units together to thereby form oligonucleotides and / or oligonucleotides that include AC, which can include naturally occurring internucleoside linkages, non-natural internucleoside linkages, or both.
[0082] In naturally occurring DNA and RNA, the internucleoside linking group is phosphodiester, which covalently links adjacent nucleosides to one another to form a linear polymeric compound. In naturally occurring DNA and RNA, the phosphodiester is linked to the 2', 3', or 5' hydroxyl moiety of the sugar. Within oligonucleotides, the phosphate groups are commonly referred to as forming the internucleoside backbone of the oligonucleotide. In naturally occurring DNA and RNA, the linkage between RNA and DNA or their backbones is a 3' to 5' phosphodiester linkage. In some embodiments, the internucleoside linking group of AC is phosphodiester. In some embodiments, the internucleoside linking group of AC is a 3' to 5' phosphodiester linkage.
[0083] Two major classes of non-natural internucleoside linking groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus containing internucleoside linkages include, but are not limited to, phosphotriesters, methylphosphonates, phosphoramidates, and phosphorothioates. Representative non-phosphorus containing internucleoside linking groups include, but are not limited to, methylenemethylimino (-CH2-N(CH3)-O-CH2-), thiodiesters (-OC(O)-S-), thionocarbamate (-OC(O)(NH)-S-), siloxanes (-O-Si(H2-O-), and N,N'-dimethylhydrazine (-CH2-N(CH3)-N(CH3)-). ACs with phosphorus internucleoside linking groups are referred to as oligonucleotides. Antisense compounds with non-phosphorus internucleoside linking groups include, but are not limited to, methylenemethylimino (-CH2-N(CH3)-O-CH2-), thiodiesters (-OC(O)-S-), thionocarbamate (-OC(O)(NH)-S-), siloxanes (-O-Si(H2-O-), and N,N'-dimethylhydrazine (-CH2-N(CH3)-N(CH3)-). , referred to as oligonucleosides. Modified internucleoside linkages can be used to alter, typically increase, the nuclease resistance of antisense compounds compared to natural phosphodiester linkages. Internucleoside linkages containing chiral atoms can be prepared as racemic, chiral, or mixtures. Representative chiral internucleoside linkages include, but are not limited to, alkylphosphonates and phosphorothioates. Methods for preparing phosphorus-containing and non-phosphorus-containing linkages are well known to those skilled in the art.
[0084] In some embodiments, two or more nucleosides having modified sugars and / or modified nucleobases can be linked using phosphoramidates, and in some embodiments, two or more nucleosides having a methylenemorpholine ring can be linked via phosphoramidate internucleoside linkages.
[0085] Antisense compounds comprising nucleobases having methylenemorpholine rings linked via phosphoramidate internucleoside linkages can be referred to as phosphoramidate morpholino oligomers (PMOs).
[0086] Conjugate Group In some embodiments, the AC is modified by the covalent attachment of one or more conjugate groups. Generally, the conjugate group modifies one or more properties of the attached AC, including, but not limited to, pharmacodynamics, pharmacokinetics, binding, absorption, cellular distribution, cellular uptake, charge, and clearance. Conjugate groups are routinely used in chemistry and are either directly linked to a parent compound, such as the AC, or linked via an optional linking moiety or group. Conjugate groups include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, thioethers, polyethers, cholesterol, thiocholesterol, cholic acid moieties, folic acid, lipids, phospholipids, biotin, phenazines, phenanthridines, anthraquinones, adamantanes, acridines, fluoresceins, rhodamines, coumarins, and dyes. In some embodiments, the conjugate group is polyethylene glycol (PEG), and the PEG is conjugated to either the AC or a CPP (as discussed elsewhere herein).
[0087] In some embodiments, the conjugate group includes a lipid moiety, e.g., a cholesterol moiety (Letsinger et al., Proc. Natl. Acad. Sci. USA (1989), 86, 6553), cholic acid (Manoharan et al., Bioorg. Med. Chem. Lett. (1994), 4, 1053), a thioether, e.g., hexyl-S-tritylthiol (Manoharan et al., Ann. NY Acad. Sci. (1992), 660, 306; Manoharan et al., Bioorg. Med. Chem. Let. (1993), 3, 2765), a thiocholesterol (Oberhauser et al., Nucl. Acids Res. (1992), 20, 533), an aliphatic chain, e.g., a dodecanediol or undecyl residue (Saison-Behmoaras et al., EMBO J. (1991), 10, 111; Kabanov et al., FEBS Lett. (1990), 259, 327; Svinarchuk et al., Biochimie (1993), 75, 49), phospholipids, such as di-hexadecyl-rac-glycerol or triethylammonium-1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett. (1995), 36, 3651; Shea et al., Nucl. Acids Res. (1990), 18, 3777), polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides (1995), 14, 969), adamantane acetic acid (Manoharan et al., Tetrahedron Lett. (1995), 36, 3651), palmityl moiety (Mishra et al., Biochim. Biophys. Acta. (1995), 1264, 229), or octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther. (1996), 277, 923).
[0088] Types of antisense compounds Various types of ACs may be used, including, for example, antisense oligonucleotides, siRNAs, microRNAs, antagomirs, aptamers, ribozymes, supermirs, miRNA mimics, miRNA inhibitors, or combinations thereof.
[0089] antisense oligonucleotides In various embodiments, an antisense compound (AC) is an antisense oligonucleotide (ASO) complementary to a target nucleotide sequence. The term "antisense oligonucleotide (ASO)" or simply "antisense" is intended to include an oligonucleotide complementary to a target nucleotide sequence. This term also encompasses ASOs that may not be perfectly complementary to a desired target nucleotide sequence. An ASO comprises a single strand of DNA and / or RNA complementary to a selected target nucleotide sequence or target gene. An ASO may contain one or more modified DNA and / or RNA bases, modified sugars, and / or non-natural internucleoside linkages. In some embodiments, an ASO may contain one or more phosphoramidate internucleoside linkages. In some embodiments, an ASO is a phosphoramidate morpholino oligomer (PMO). An ASO may have any properties, be any length, bind to any target nucleotide sequence and / or sequence element, and provide any of the mechanisms described for an AC.
[0090] Antisense oligonucleotides have been demonstrated to be effective as targeted inhibitors of protein synthesis, and as a result, can be used to specifically inhibit protein synthesis by target genes. The effectiveness of ASOs for inhibiting protein synthesis has been well established. To date, these compounds have shown promise in several in vitro and in vivo models, including models of inflammatory disease, cancer, and HIV (Agrawal, Trends in Biotech. (1996), 14:376-387). Antisense can also affect cellular activity by specifically hybridizing with chromosomal DNA.
[0091] Methods for producing antisense oligonucleotides are known in the art and can be easily adapted to generate antisense oligonucleotides targeting any polynucleotide sequence. Selection of an antisense oligonucleotide sequence specific for a given target sequence is based on analysis of the selected target sequence and determination of secondary structure, Tm, binding energy, and relative stability. Antisense oligonucleotides can be selected based on their relative inability to form dimers, hairpins, or other secondary structures that would reduce or prohibit specific binding to the target mRNA in a host cell. Target regions of mRNA include regions at or near the AUG translation initiation codon and sequences substantially complementary to the 5' region of the mRNA. These secondary structure analysis and target site selection considerations can be performed using, for example, OLIGO primer analysis software version 4 (Molecular Biology Insights) and / or BLASTN 2.0.5 algorithm software (Altschul et al., Nucleic Acids Res. 1997, 25(17):3389-402).
[0092] RNA interference In some embodiments, the AC comprises a molecule that mediates RNA interference (RNAi). As used herein, the phrase "mediating RNAi" refers to the ability to silence a target transcript in a sequence-specific manner. Without wishing to be bound by theory, it is believed that silencing utilizes the RNAi machinery or process and a guide RNA, e.g., an siRNA compound of about 21 to about 23 nucleotides. In some embodiments, the AC targets the target transcript for degradation. Thus, in some embodiments, RNAi molecules can be used to disrupt expression of a gene or polynucleotide of interest. In some embodiments, RNAi molecules are used to induce degradation of a target transcript, such as a pre-mRNA or mature mRNA.
[0093] In embodiments, the AC comprises a small interfering RNA (siRNA) that elicits an RNAi response.
[0094] Small interfering RNAs (siRNAs) are nucleic acid duplexes, typically about 16 to 30 nucleotides long, that can associate with a cytoplasmic multiprotein complex known as the RNAi-induced silencing complex (RISC). RISC loaded with siRNA mediates the degradation of homologous transcripts; therefore, siRNAs can be designed to knock down protein expression with high specificity. Unlike other antisense technologies, siRNAs function through a natural mechanism that evolved to control gene expression via non-coding RNA. Various RNAi reagents, including siRNAs targeting clinically relevant targets, are currently in pharmaceutical development, as described, for example, in de Fougerolles, A. et al., Nature Reviews (2007) 6:443-453.
[0095] The first RNAi molecules described were RNA:RNA hybrids containing both RNA sense and RNA antisense strands, but it has been demonstrated that DNA sense:RNA antisense hybrids, RNA sense:DNA antisense hybrids, and DNA:DNA hybrids can mediate RNAi (Lamberton, JS and Christian, AT, Molecular Biotechnology (2003), 24:111-119). In some embodiments, RNAi molecules containing any of these different types of double-stranded molecules are used. In addition, it is understood that RNAi molecules can be used and introduced into cells in various forms. Thus, as used herein, RNAi molecules encompass any and all molecules capable of mediating RNAi in cells, including, but not limited to, double-stranded oligonucleotides comprising two separate strands, i.e., a sense strand and an antisense strand, e.g., small interfering RNAs (siRNAs), double-stranded oligonucleotides comprising two separate strands linked together by a non-nucleotidyl linker, oligonucleotides comprising a hairpin loop of complementary sequences that form a double-stranded region, e.g., shRNAi molecules, and expression vectors that express one or more polynucleotides that can form a double-stranded polynucleotide, either alone or in combination with another polynucleotide.
[0096] " single-stranded siRNA compound " as used herein refers to the siRNA compound that is composed of a single molecule.It can comprise the double-stranded region that is formed by intrastrand pairing, for example, it can be or comprise hairpin structure or panhandle structure.Single-stranded siRNA compound can be antisense with respect to target molecule.
[0097] Single-stranded siRNA compound can be long enough so that it can enter RISC and participate in the RISC-mediated cleavage of target mRNA.Single-stranded siRNA compound is at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, or at most about 50 nucleotides long.In certain embodiments, single-stranded siRNA is less than about 200, about 100, or about 60 nucleotides long.
[0098] Hairpin siRNA compounds can have a double-stranded region of at least about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25 nucleotide pairs, or an equivalent number thereof. The double-stranded region can be about 200, about 100, or about 50 nucleotide pairs in length, or an equivalent number thereof. In certain embodiments, the double-stranded region ranges from about 15 to about 30, about 17 to about 23, about 19 to about 23, and about 19 to about 21 nucleotide pairs in length. The hairpin can have a single-stranded overhang or terminal unpaired region. In certain embodiments, the overhang is about 2 to about 3 nucleotides in length. In some embodiments, the overhangs are on the same side of the hairpin, and in some embodiments, on the antisense side of the hairpin.
[0099] As used herein, a "double-stranded siRNA compound" is an siRNA compound that contains two or more strands, and in some cases two strands, capable of interstrand hybridization to form a region of double-stranded structure.
[0100] The antisense strand of a double-stranded siRNA compound can be at least about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 40, or about 60 nucleotides in length, or a length equal thereto. It can be up to about 200, about 100, or about 50 nucleotides in length. Ranges can be about 17 to about 25, about 19 to about 23, and about 19 to about 21 nucleotides in length. As used herein, the term "antisense strand" refers to the strand of an siRNA compound that is sufficiently complementary to the target nucleotide sequence of a target molecule, e.g., a target transcript.
[0101] The sense strand of a double-stranded siRNA compound can be at least about 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, or 60 nucleotides in length, or a length equal to or less than about 200, 100, or 50 nucleotides in length. Ranges can be about 17 to about 25, about 19 to about 23, and about 19 to about 21 nucleotides in length.
[0102] The double-stranded portion of a double-stranded siRNA compound can be at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 60 nucleotide pairs in length, or a length equal thereto. It can be up to about 200, 100, or 50 nucleotide pairs in length. Ranges can be about 15 to about 30, about 17 to about 23, about 19 to about 23, and about 19 to about 21 nucleotide pairs in length.
[0103] In embodiments, the siRNA compound is large enough so that it can be cleaved by an endogenous molecule, eg, Dicer, to produce smaller siRNA compounds, eg, siRNA agents.
[0104] The sense and antisense strands can be selected so that the double-stranded siRNA compound contains a single-stranded or unpaired region at one or both ends of the molecule. Thus, the double-stranded siRNA compound can contain a paired sense and antisense strand to contain an overhang, for example, one or two 5' or 3' overhangs, or a 3' overhang of 1 to 3 nucleotides. The overhang can be the result of one strand being longer than the other, or the result of two strands of the same length being staggered. Some embodiments have at least one 3' overhang. In some embodiments, both ends of the siRNA molecule have a 3' overhang. In some embodiments, the overhang is 2 nucleotides.
[0105] In some embodiments, the length of the double-stranded region is, for example, about 15 to about 30, or about 18, about 19, about 20, about 21, about 22, or about 23 nucleotides, within the range of ssiRNA (siRNA with sticky overhangs) compounds described above. ssiRNA compounds can be similar in length and construct natural Dicer processed products derived from longer dsiRNAs. Also included are embodiments in which the two strands of the ssiRNA compound are linked, e.g., covalently linked. Some embodiments include a hairpin or other single-stranded structure that provides a double-stranded region and a 3' overhang.
[0106] The siRNA compounds described herein, including double-stranded siRNA compounds and single-stranded siRNA compounds, can mediate the silencing of target RNA, such as mRNA, for example, the transcript of the gene encoding protein.For convenience, this mRNA is also referred to herein as the mRNA to be silenced.This gene is also referred to as target gene.Generally, the RNA to be silenced is endogenous gene.
[0107] In some embodiments, the siRNA compound is "sufficiently complementary" to the target transcript such that the siRNA compound silences production of the protein encoded by the target mRNA. In some embodiments, the siRNA compound is "sufficiently complementary" to at least a portion of the target transcript such that the siRNA compound silences production of the gene product encoded by the target transcript. In another embodiment, the siRNA compound is "fully complementary" to the target nucleotide sequence (e.g., a portion of the target transcript) such that the target nucleotide sequence and the siRNA compound anneal to form a hybrid made up of only Watson-Crick base pairs within the exact region of complementarity. "Sufficiently complementary" to the target nucleotide sequence can include an internal region (e.g., at least about 10 nucleotides) that is exactly complementary to the target nucleotide sequence. Furthermore, in certain embodiments, the siRNA compound specifically discriminates between single-base differences. In this case, the siRNA compound mediates RNAi only when exact complementarity is found in the region of a single-base difference (e.g., within 7 nucleotides).
[0108] The therapeutic applications of RNAi are extremely broad, as siRNA and miRNA constructs can be synthesized with any nucleotide sequence directed against a target protein. To date, siRNA constructs have demonstrated the ability to specifically downregulate target proteins in both in vitro and in vivo models, as well as in clinical trials.
[0109] microRNA In some embodiments, the AC comprises a microRNA molecule. MicroRNAs (miRNAs) are a highly conserved class of small RNA molecules that are transcribed from DNA in the genomes of plants and animals but are not translated into proteins. Processed miRNAs are single-stranded 17-25 nucleotide RNA molecules that are incorporated into RNA-induced silencing complexes (RISCs) and have been identified as important regulators of development, cell proliferation, apoptosis, and differentiation. They are thought to be involved in regulating gene expression by binding to the 3'-untranslated regions of specific mRNAs. RISCs mediate downregulation of gene expression through translational inhibition, transcriptional cleavage, or both. RISCs are also involved in transcriptional silencing in the nuclei of a wide range of eukaryotic organisms.
[0110] Antagomir In some embodiments, the AC is an antagomir. An antagomir is an RNA-like oligonucleotide with various modifications for RNAse protection and pharmacological properties, such as enhanced tissue and cellular uptake. They differ from normal RNA in, for example, sugars, phosphorothioate backbones, and complete 2'-0-methylation of cholesterol moieties, for example, at the 3' end. An antagomir can be used to efficiently silence endogenous miRNAs by forming a duplex containing an antagomir and endogenous miRNA, thereby preventing miRNA-induced gene silencing. An example of antagomir-mediated miRNA silencing is the silencing of miR-122 described in Krutzfeldt et al., Nature (2005), 438:685-689, the entire contents of which are expressly incorporated herein by reference. Antagomir RNA can be synthesized using standard solid phase oligonucleotide synthesis protocols (U.S. Patent Application Nos. 11 / 502,158 and 11 / 657,341, the disclosures of each of which are incorporated herein by reference).
[0111] Antagomir can include ligand-binding monomer subunits and monomers for oligonucleotide synthesis. Monomers are described in U.S. Patent Application No. 10 / 916,185. Antagomir can have a ZXY structure, for example, as described in PCT Application No. PCT / US2004 / 07070. Antagomir can be conjugated with an amphiphilic moiety. Amphiphilic moieties for use in combination with oligonucleotide agents are described in PCT Application No. PCT / US2004 / 07070.
[0112] Aptamers In some embodiments, the AC comprises an aptamer. Aptamers are nucleic acid or peptide molecules that bind to specific molecules of interest with high affinity and specificity (Tuerk and Gold, Science 249:505 (1990); Ellington and Szostak, Nature 346:818 (1990)). DNA or RNA aptamers have been successfully produced that bind to many different entities, from large proteins to small organic molecules (Eaton, Curr. Opin. Chem. Biol. (1997), 1:10-16; Famulok, Curr. Opin. Struct. Biol. (1999), 9:324-9; and Hermann and Patel, Science (2000), 287:820-5). Aptamers can be RNA- or DNA-based and can include riboswitches. A riboswitch is a part of an mRNA molecule that can directly bind to a small target molecule, and target binding affects the activity of the gene. Thus, the mRNA containing the riboswitch is directly involved in regulating its own activity depending on the presence or absence of its target molecule. Generally, aptamers are engineered to bind to various molecular targets, such as small molecules, proteins, nucleic acids, and even cells, tissues, and organisms, through repeated rounds of in vitro selection, i.e., SELEX (Systematic Evolution of Ligands by Exponential Enrichment). Aptamers may be prepared by any known method, including synthetic, recombinant, and purified methods, and may be used alone or in combination with other aptamers specific to the same target. Furthermore, the term "aptamer" also includes "secondary aptamers," which contain consensus sequences obtained by comparing two or more known aptamers with a given target. In several embodiments, the aptamer is an "intracellular aptamer" or "intramer" that specifically recognizes an intracellular target (Famulok et al., Chem Biol. (2001), 8(10):931-939; Yoon and Rossi, Adv. Drug Deliv. Rev. (2018), 134:22-35, each incorporated herein by reference).
[0113] Ribozymes In some embodiments, the AC is a ribozyme, which is an RNA molecular complex that has a specific catalytic domain with endonuclease activity (Kim and Cech, Proc. Natl. Acad. Sci. USA (1987), 84(24):8788-92; Forster and Symons, Cell (1987) 24, 49(2):211-20). For example, many ribozymes accelerate phosphate transfer reactions with high specificity, often cleaving only one of several phosphoesters in an oligonucleotide substrate (Cech et al., Cell (1981), 27(3 Pt 2):487-96; Michel and Westhof, J. Mol. Biol. (1990), 5, 216(3):585-610; Reinhold-Hurek and Shub, Nature (1992), 14, 357(6374):173-6). This specificity results from the requirement that the substrate bind to the ribozyme's internal guide sequence (IGS) through specific base-pairing interactions prior to chemical reaction.
[0114] Currently, at least six basic types of naturally occurring enzymatic RNAs are known. Each is capable of catalyzing the hydrolysis of RNA phosphodiester bonds in trans under physiological conditions (and thus capable of cleaving other RNA molecules). Generally, enzymatic nucleic acids act by first binding to a target RNA. Such binding occurs via the target binding portion of the enzymatic nucleic acid, which is held in close proximity to the enzymatic portion of the molecule that acts to cleave the target RNA. Thus, the enzymatic nucleic acid first recognizes and then binds to the target RNA through complementary base pairing, and once bound to the correct site, acts enzymatically to cleave the target RNA. Such strategic cleavage of the target RNA destroys its ability to direct synthesis of the encoded protein. After an enzymatic nucleic acid has bound and cleaved its RNA target, it is released from that RNA to seek another target and can repeatedly bind and cleave new targets.
[0115] The enzymatic nucleic acid molecule can be formed, for example, in a hammerhead motif, a hairpin motif, a hepatitis delta virus motif, a group I intron motif, an RNase P RNA motif (in conjunction with an RNA guide sequence), or a Neurospora VS RNA motif. Specific examples of hammerhead motifs are described by Rossi et al., Nucleic Acids Res. (1992), 20(17):4559-65. Examples of hairpin motifs are described by Hampel et al. (European Patent Application Publication No. EP 0 360 257), Hampel and Tritz, Biochemistry (1989), 28(12):4929-33, Hampel et al., Nucleic Acids Res. (1990), 18(2):299-304, and U.S. Patent No. 5,631,359. An example of a hepatitis virus motif is described by Perrotta and Been, Biochemistry (1992), 31(47):11843-52; an example of an RNase P motif is described by Guerrier-Takada et al., Cell (1983), 35(3 Pt 2):849-57; a Neurospora VS RNA ribozyme motif is described by Collins (Saville and Collins, Cell (1990), 61(4):685-96; Saville and Collins, Proc. Natl. Acad. Sci. USA (1991), 88(19):8826-30; Collins and Olive, Biochemistry (1993), 32(11):2795-9); and an example of a group I intron is described in U.S. Pat. No. 4,987,071. In some embodiments, the enzymatic nucleic acid molecule has a specific substrate binding site complementary to one or more of the target gene DNA or RNA regions, and has nucleotide sequences within or surrounding that substrate binding site that confer RNA cleavage activity to the enzymatic nucleic acid molecule. Thus, ribozyme constructs are not necessarily limited to the particular motifs referenced herein.
[0116] Ribozymes can be designed as described in International Patent Applications WO 93 / 23569 and WO 94 / 02595, each expressly incorporated herein by reference, and synthesized and tested in vitro and in vivo as described therein. In some embodiments, the ribozyme targets a target nucleotide sequence of a target transcript.
[0117] Ribozyme activity can be increased by altering the length of the ribozyme binding arms, or by chemically synthesizing ribozymes with modifications that prevent their degradation by serum ribonucleases (see, e.g., WO 92 / 07065, WO 93 / 15187, WO 91 / 03162, EP 92110298.4, U.S. Pat. No. 5,334,711, and WO 94 / 13688, which describe various chemical modifications that can be made to the sugar portion of enzymatic RNA molecules), modifications that enhance their efficacy within cells, and removal of the stem Π base to shorten RNA synthesis time and reduce chemical requirements.
[0118] Super Mill In some embodiments, the AC is a supermil. Supermil refers to a single-stranded, double-stranded, or partially double-stranded oligomer, or a polymer of RNA, DNA, or both, or modifications thereof, that has a nucleotide sequence that is substantially identical to an miRNA and antisense with respect to its target. This term includes oligonucleotides composed of naturally occurring nucleobases, sugars, and covalent internucleoside (backbone) linkages, and that contain at least one non-naturally occurring moiety that functions similarly. Such modified or substituted oligonucleotides have desirable properties, such as enhanced cellular uptake, enhanced affinity for nucleic acid targets, and increased stability in the presence of nucleases. In some embodiments, supermil does not contain a sense strand, and in other embodiments, supermil does not self-hybridize to a significant extent. While supermil can have secondary structure, it is substantially single-stranded under physiological conditions. A substantially single-stranded supermil is single-stranded to the extent that less than about 50% (e.g., about 40%, about 30%, about 20%, about 10%, or less than about 5%) of the supermil is double-stranded with itself. The supermil can include a hairpin segment, and a sequence, for example, a sequence at the 3' end, can self-hybridize to form a double-stranded region, for example, a double-stranded region of at least about 1, about 2, about 3, or about 4 nucleotides, or about 8, about 7, about 6, or about 5 nucleotides, or less than about 5 nucleotides. The double-stranded region can be linked by a linker, for example, a nucleotide linker, for example, about 3, about 4, about 5, or about 6 dTs, for example, modified dTs. In another embodiment, the supermill is duplexed with a shorter oligo, for example, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides in length, at one or both of the 3' and 5' ends, and at a non-terminal or intermediate end of the supermill.
[0119] miRNA mimics In several embodiments, the AC is an miRNA mimic. miRNA mimics represent a class of molecules that can be used to mimic the gene silencing capabilities of one or more miRNAs. Thus, the term "microRNA mimic" refers to a synthetic non-coding RNA that can enter the RNAi pathway and regulate gene expression (i.e., miRNAs are not obtained by purification from endogenous miRNA sources). miRNA mimics can be designed as mature molecules (e.g., single-stranded) or mimic precursors (e.g., pri-miRNA or pre-miRNA). miRNA mimics can comprise nucleic acids (modified or modified nucleic acids), such as oligonucleotides, including, but not limited to, RNA, modified RNA, DNA, modified DNA, locked nucleic acids, or 2'-0,4'-C-ethylene-bridged nucleic acids (ENA), or any combination of the above (including DNA-RNA hybrids). In addition, miRNA mimics can include conjugates that can affect delivery, intracellular compartmentalization, stability, specificity, functionality, strand usage, and / or efficacy. In one design, miRNA mimics are double-stranded molecules (e.g., having a double-stranded region about 16 to about 31 nucleotides in length) that contain one or more sequences that share identity with the mature strand of a given miRNA. Modifications can include 2' modifications (including 2'-0 methyl and 2'F modifications) on one or both strands of the molecule, as well as internucleoside modifications (e.g., phosphorothioate modifications) that enhance nucleic acid stability and / or specificity. In addition, miRNA mimics can contain overhangs. The overhangs can comprise about 1 to about 6 nucleotides at either the 3' or 5' end of either strand and can be modified to enhance stability or functionality. In embodiments, the miRNA mimic comprises a double-stranded region of about 16 to about 31 nucleotides and one or more of the following chemical modification patterns: the sense strand includes 2'-0-methyl modifications of nucleotides 1 and 2 (counting from the 5' end of the sense oligonucleotide) and all Cs and Us; the antisense strand modifications include 2'F modifications of all Cs and Us, phosphorylation of the 5' end of the oligonucleotide, and stabilized internucleoside linkages associated with the two-nucleotide 3' overhang.
[0120] miRNA inhibitors In some embodiments, the AC is an miRNA inhibitor. The terms "antimir," "microRNA inhibitor," "miR inhibitor," or "miRNA inhibitor" are used interchangeably and refer to an oligonucleotide or modified oligonucleotide that interferes with the activity of a specific miRNA. Generally, these inhibitors are natural or modified nucleic acids, including oligonucleotides containing RNA, modified RNA, DNA, modified DNA, locked nucleic acids (LNA), or any combination of the above.
[0121] Modifications include 2'-modifications (including 2'-0 alkyl and 2'F modifications) and internucleoside modifications (e.g., phosphorothioate modifications), which can affect delivery, stability, specificity, intracellular compartmentalization, or efficacy. Additionally, miRNA inhibitors can include conjugates, which can affect delivery, intracellular compartmentalization, stability, and / or efficacy. Inhibitors can adopt a variety of configurations, including single-stranded, double-stranded (RNA / RNA duplexes or RNA / DNA duplexes), and hairpin designs. Generally, microRNA inhibitors include one or more sequences or portions of sequences that are complementary or partially complementary to the mature strand (or strands) of the targeted miRNA. Additionally, miRNA inhibitors can also include additional sequences located 5' and 3' of a sequence that is the reverse complement of the mature miRNA. The additional sequences can be the reverse complement of the sequences adjacent to the mature miRNA in the pri-miRNA from which the mature miRNA is derived, or the additional sequences can be any sequence (having a mixture of A, G, C, or U). In some embodiments, one or both of the additional sequences is any sequence capable of forming a hairpin. Thus, in some embodiments, a sequence that is the reverse complement of the miRNA is flanked on the 5' and 3' sides by a hairpin structure. When double-stranded, the microRNA inhibitor may contain mismatches between nucleotides on opposite strands. Furthermore, the microRNA inhibitor may be linked to a conjugate moiety to facilitate cellular uptake of the inhibitor. For example, the microRNA inhibitor may be linked to cholesteryl 5-(bis(4-methoxyphenyl)(phenyl)methoxy)-3 hydroxypentylcarbamate), which enables passive uptake of the microRNA inhibitor into cells. MicroRNA inhibitors, including hairpin miRNA inhibitors, are described in detail in Vermeulen et al., RNA 13:723-730 (2007), and WO2007 / 095387 and WO2008 / 036825, each of which is incorporated herein by reference in its entirety. One skilled in the art can select the sequence of a desired miRNA from a database and design an inhibitor useful in the methods disclosed herein.
[0122] Linking groups or bifunctional linking moieties, such as those known in the art, are suitable for the compounds provided herein. Linking groups are useful for attaching chemical functional groups, conjugate groups, reporter groups, and other groups to selective sites in a parent compound, such as an AC. Generally, bifunctional linking moieties include a hydrocarbyl moiety with two functional groups. One of these functional groups is selected to attach to a parent molecule or compound of interest, and the other is selected to attach to essentially any selected group, such as a chemical functional group or conjugate group. Any of the linkers described herein can be used. In embodiments, the linker includes a chain structure or oligomer of repeating units, such as ethylene glycol or amino acid units. Examples of functional groups routinely used in bifunctional linking moieties include, but are not limited to, electrophiles for reacting with nucleophilic groups and nucleophiles for reacting with electrophilic groups. In embodiments, bifunctional linking moieties include amino, hydroxyl, carboxylic acid, thiol, and unsaturated groups (e.g., double or triple bonds). Some non-limiting examples of bifunctional linking moieties include 8-amino-3,6-dioxaoctanoic acid (ADO), succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (SMCC), and 6-aminohexanoic acid (AHEX or AHA). Other linking groups include, but are not limited to, substituted C1-C10 alkyl, substituted or unsubstituted C2-C10 alkenyl, or substituted or unsubstituted C2-C10 alkynyl, where a non-limiting list of substituents includes hydroxyl, amino, alkoxy, carboxy, benzyl, phenyl, nitro, thiol, thioalkoxy, halogen, alkyl, aryl, alkenyl, and alkynyl.
[0123] In some embodiments, the AC comprises a nucleotide modification designed to not support RNase H activity. Nucleotide modifications of antisense compounds that do not support RNase H activity are known, and these modifications include, but are not limited to, 2'-O-methoxyethyl / phosphorothioate (MOE) modifications. Advantageously, ACs with MOE modifications have increased affinity for target RNA and increased nuclease stability.
[0124] Immunostimulatory oligonucleotides In some embodiments, the therapeutic moiety is an immunostimulatory oligonucleotide. Immunostimulatory oligonucleotides (ISS, single-stranded or double-stranded) can induce an immune response when administered to a patient, which may be a mammal, or other subject. ISSs include, for example, certain palindromes that result in hairpin secondary structures (see Yamamoto S., et al. (1992) J. Immunol. 148:4072-4076), or CpG motifs, as well as other known ISS features (e.g., multi-G domains, see WO96 / 11266).
[0125] The immune response can be an innate immune response or an adaptive immune response. The immune system is divided into a more innate immune system and an acquired adaptive immune system in vertebrates, the latter being further divided into humoral and cellular components. In certain embodiments, the immune response can be mucosal.
[0126] An immunostimulatory nucleic acid is considered non-sequence-specific if it does not need to specifically bind to and reduce the expression of a target polynucleotide in order to elicit an immune response. Thus, a particular immunostimulatory nucleic acid may contain a sequence that corresponds to a region of a naturally occurring gene or mRNA and still be considered a non-sequence-specific immunostimulatory nucleic acid.
[0127] In several embodiments, the immunostimulatory nucleic acid or oligonucleotide comprises at least one CpG dinucleotide. The oligonucleotide or CpG dinucleotide may be unmethylated or methylated. In another embodiment, the immunostimulatory nucleic acid comprises at least one CpG dinucleotide having a methylated cytosine. In several embodiments, the nucleic acid comprises a single CpG dinucleotide, wherein the cytosine in the CpG dinucleotide is methylated. In a particular embodiment, the nucleic acid comprises the sequence 5'TAACGTTGAGGG'CAT 3' (SEQ ID NO: 369). In an alternative embodiment, the nucleic acid comprises at least two CpG dinucleotides, wherein at least one cytosine in the CpG dinucleotide is methylated. In a further embodiment, each cytosine in the CpG dinucleotide present in the sequence is methylated. In another embodiment, the nucleic acid comprises a plurality of CpG dinucleotides, wherein at least one of the CpG dinucleotides comprises a methylated cytosine.
[0128] Further specific nucleic acid sequences of oligonucleotides (ODN) suitable for use in compositions and methods are described in Raney et al., Journal of Pharmacology and Experimental Therapeutics, 298:1185-1192 (2001).In certain embodiments, the ODN used in compositions and methods has a phosphodiester ("PO") backbone or a phosphorothioate ("PS") backbone, and / or at least one methylated cytosine residue in CpG motif.
[0129] Decoy oligonucleotides In several embodiments, the therapeutic moiety is a decoy oligonucleotide. Because transcription factors recognize these relatively short binding sequences even in the absence of surrounding genomic DNA, short oligonucleotides containing the consensus binding sequence of a specific transcription factor can be used as a tool to manipulate gene expression in living cells. This strategy involves the intracellular delivery of such a "decoy oligonucleotide," which is then recognized and bound by the target factor. Occupation of the DNA binding site of the transcription factor by the decoy prevents the transcription factor from subsequently binding to the promoter region of the target gene. The decoy can be used either as a therapeutic agent to inhibit the expression of genes activated by the transcription factor or to upregulate genes repressed by the binding of the transcription factor. Examples of the use of decoy oligonucleotides can be found in Mann et al., J. Clin. Invest, 2000, 106:1071-1075, the entire contents of which are expressly incorporated herein by reference.
[0130] U1 Adapter In some embodiments, the therapeutic moiety is a U1 adaptor. A U1 adaptor is a bifunctional oligonucleotide having a targeting domain that interrupts the polyA site and is complementary to a site in the terminal exon of a target gene, and a "U1 domain" that binds to the U1 small nuclear RNA component of the U1 snRNP (Goraczniak, et al., 2008, Nature Biotechnology, 27(3), 257-263, expressly incorporated herein by reference in its entirety). The U1 snRNP is a ribonucleoprotein complex that functions primarily to direct the initial step of spliceosome formation by binding to pre-mRNA exon-intron boundaries (Brown and Simpson, 1998, Annu Rev Plant Physiol Plant Mol Biol 49:77-95). Nucleotides 2-11 at the 5' end of the U1 snRNA base pair to bind to the 5' ss of the pre-mRNA. In one embodiment, the oligonucleotide is a U1 adaptor. In one embodiment, the U1 adaptor can be administered in combination with at least one other iRNA agent.
[0131] (CRISPR) gene editing mechanism In some embodiments, the compounds disclosed herein comprise one or more CPPs (or cCPPs) conjugated to CRISPR gene editing machinery. As used herein, "CRISPR gene editing machinery" refers to a protein, nucleic acid, or combination thereof that can be used to edit genomes. Non-limiting examples of gene editing machinery include gRNA, nuclease, nuclease inhibitor, and combinations and complexes thereof. The following patent documents describe CRISPR gene editing mechanisms: U.S. Patent No. 8,697,359, U.S. Patent No. 8,771,945, U.S. Patent No. 8,795,965, U.S. Patent No. 8,865,406, U.S. Patent No. 8,871,445, U.S. Patent No. 8,889,356, U.S. Patent No. 8,895,308, U.S. Patent No. 8,906,616, U.S. Patent No. 8,932,814, U.S. Patent No. 8,945,839, U.S. Patent No. 8,993,233, U.S. Patent No. 8,999,641, U.S. Patent No. 14 / 704,551, and U.S. Patent Application No. 13 / 842,859. Each of the foregoing patent documents is incorporated herein by reference in its entirety.
[0132] In embodiments, a linker conjugates the cCPP to the CRISPR gene editing machinery. Any linker described in this disclosure or known to one of skill in the art may be utilized.
[0133] gRNA In embodiments, the compound comprises a CPP (or cCPP) conjugated to a gRNA, which targets a genomic locus in a prokaryotic or eukaryotic cell.
[0134] In some embodiments, the gRNA is a single-molecule guide RNA (sgRNA). The sgRNA comprises a spacer sequence and a scaffold sequence. The spacer sequence is a short nucleic acid sequence used to target a nuclease (e.g., a Cas9 nuclease) to a specific nucleotide region of interest (e.g., a genomic DNA sequence to be cleaved). In some embodiments, the spacer can be about 17-24 bases in length, e.g., about 20 bases in length. In some embodiments, the spacer can be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 bases in length. In embodiments, the spacer can be at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 bases in length. In embodiments, the spacer can be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 bases in length. In embodiments, the spacer sequence has a GC content of about 40% to about 80%.
[0135] In some embodiments, the spacer targets the site immediately preceding the 5' protospacer adjacent motif (PAM). The PAM sequence can be selected based on the desired nuclease. For example, the PAM sequence can be any one of the PAM sequences shown in Table 3 below, where N refers to any nucleic acid, R refers to A or G, Y refers to C or T, W refers to A or T, and V refers to A, C, or G. [Table 3]
[0136] In embodiments, the spacer may target a sequence of a mammalian gene, such as a human gene. In embodiments, the spacer may target a mutated gene. In embodiments, the spacer may target a coding sequence. In embodiments, the spacer may target an exon sequence. In embodiments, the spacer may target a polyadenylation site (PS). In embodiments, the spacer may target a sequence element of a PS. In embodiments, the spacer may target a polyadenylation signal (PAS), an intervening sequence (IS), a cleavage site (CS), a downstream element (DES), or portions or combinations thereof. In embodiments, the spacer may target a splicing element (SE) or a cis-splicing regulatory element (SRE).
[0137] The scaffold sequence is a sequence within the sgRNA that is involved in nuclease (e.g., Cas9) binding. The scaffold sequence does not include a spacer / target sequence. In embodiments, the scaffold can be about 1 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, or about 120 to about 130 nucleotides in length. In embodiments, the scaffold comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, About 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67 7, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84 , about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100 pieces, about 1 about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115, about 116, about 117, about 118, about 119, about 120, about 121, about 122, about 123, about 124, or about 125 nucleotides in length. In embodiments, the scaffold can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or at least 125 nucleotides in length.
[0138] In some embodiments, the gRNA is a dual-molecule guide RNA, e.g., a crRNA and a tracrRNA. In some embodiments, the gRNA may further comprise a poly(A) tail.
[0139] In some embodiments, the compound comprising a CPP is conjugated to a nucleic acid comprising a gRNA. In some embodiments, the nucleic acid comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 gRNAs. In some embodiments, the gRNAs recognize the same target. In some embodiments, the gRNAs recognize different targets. In some embodiments, the nucleic acid comprising the gRNA comprises a sequence encoding a promoter, which drives expression of the gRNA.
[0140] nuclease In some embodiments, the compound comprises a cell membrane-permeable peptide conjugated to a nuclease. In some embodiments, the nuclease is a type II, type VA, type VB, type VC, type VU, or type VI-B nuclease. In some embodiments, the nuclease is a transcription activator-like effector nuclease (TALEN), meganuclease, or zinc finger nuclease. In some embodiments, the nuclease is a Cas9, Cas12a (Cpf1), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease. For example, in some embodiments, the nuclease is a Cas9 nuclease or a Cpf1 nuclease.
[0141] In some embodiments, the nuclease is a modified form or variant of a Cas9, Cas12a (Cpf1), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease. In some embodiments, the nuclease is a modified form or variant of a TAL nuclease, meganuclease, or zinc finger nuclease. A "modified" or "variant" nuclease may be, for example, truncated, fused to another protein (such as another nuclease), catalytically inactivated, etc. In embodiments, the nuclease can have at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or about 100% sequence identity to a naturally occurring Cas9, Cas12a (Cpf1), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease, or a TALEN, meganuclease, or zinc finger nuclease. In embodiments, the nuclease is a Cas9 nuclease from S. pyogenes (SpCas9). In some embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cas9 nuclease from S. pyogenes (SpCas9). In some embodiments, the nuclease is Cas9 from S. aureus (SaCas9). In some embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cas9 from S. aureus (SaCas9). In some embodiments, the Cpfl is the Cpfl enzyme from Acidaminococcus (species BV3L6, UniProt accession number U2UMQ6).In embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cpf1 enzyme from Acidaminococcus (species BV3L6, UniProt accession number U2UMQ6).
[0142] In some embodiments, the Cpfl is the Cpfl enzyme from Lachnospiraceae (species ND2006, UniProt accession number A0A182DWE3). In some embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cpfl enzyme from Lachnospiraceae. In some embodiments, the nuclease-encoding sequence is codon optimized for expression in mammalian cells. In some embodiments, the nuclease-encoding sequence is codon optimized for expression in human or mouse cells.
[0143] In embodiments, the compound comprising a CPP is conjugated to a nuclease, hi embodiments, the nuclease is a soluble protein.
[0144] In some embodiments, the compound comprising a CPP is conjugated to a nucleic acid encoding a nuclease, hi some embodiments, the nucleic acid encoding the nuclease comprises a sequence encoding a promoter, which drives expression of the nuclease.
[0145] gRNA and nuclease combinations In some embodiments, the compound comprises one or more CPPs (or cCPPs) conjugated to a gRNA and a nuclease. In some embodiments, the one or more CPPs (or cCPPs) are conjugated to a nucleic acid encoding a gRNA and / or a nuclease. In some embodiments, the nucleic acid encoding the nuclease and gRNA comprises a sequence encoding a promoter, which drives expression of the nuclease and gRNA. In some embodiments, the nucleic acid encoding the nuclease and gRNA comprises two promoters, a first promoter controlling expression of the nuclease and a second promoter controlling expression of the gRNA. In some embodiments, the nucleic acid encoding the gRNA and nuclease encodes about 1 to about 20 gRNAs, or about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, or about 19, and up to about 20 gRNAs. In some embodiments, the gRNAs recognize different targets. In some embodiments, the gRNAs recognize the same target.
[0146] In embodiments, the compound comprises a cell membrane-permeable peptide (or cCPP) conjugated to a ribonucleoprotein (RNP) comprising a gRNA and a nuclease.
[0147] In some embodiments, a composition comprising (a) a CPP conjugated to a gRNA and (b) a nuclease is delivered to a cell. In some embodiments, a composition comprising (a) a CPP conjugated to a nuclease and (b) a gRNA is delivered to a cell.
[0148] In some embodiments, a composition comprising (a) a first CPP conjugated to a gRNA and (b) a second CPP conjugated to a nuclease is delivered to a cell. In some embodiments, the first CPP and the second CPP are the same. In some embodiments, the first CPP and the second CPP are different.
[0149] genetic element of interest In embodiments, the compounds disclosed herein comprise a cell membrane-permeable peptide conjugated to a genetic element of interest. In embodiments, the genetic element of interest replaces a genomic DNA sequence cleaved by a nuclease. Non-limiting examples of genetic elements of interest include a gene, a single nucleotide polymorphism, a promoter, or a terminator.
[0150] Nuclease inhibitors In some embodiments, the compounds disclosed herein comprise a cell membrane-permeable peptide conjugated to a nuclease inhibitor (e.g., Cas9). A limitation of gene editing is potential off-target editing. Delivery of the nuclease inhibitor limits off-target editing. In some embodiments, the nuclease inhibitor is a polypeptide, polynucleotide, or small molecule. Exemplary nuclease inhibitors are described in U.S. Patent Application Publication No. 2020 / 087354, International Patent Application Publication No. 2018 / 085288, U.S. Patent Application Publication No. 2018 / 0382741, International Patent Application Publication No. 2019 / 089761, International Patent Application Publication No. 2020 / 068304, International Patent Application Publication No. 2020 / 041384, and International Patent Application Publication No. 2019 / 076651, each of which is incorporated herein by reference in its entirety.
[0151] Therapeutic Polypeptides In some embodiments, the therapeutic moiety comprises a polypeptide. In some embodiments, the therapeutic moiety comprises a protein or fragment thereof. In some embodiments, the therapeutic moiety comprises an RNA binding protein or an RNA binding fragment thereof. In some embodiments, the therapeutic moiety comprises an enzyme. In some embodiments, the therapeutic moiety comprises an RNA cleaving enzyme or an active fragment thereof.
[0152] Conjugate Group In some embodiments, the AC is modified by the covalent attachment of one or more conjugate groups. Generally, the conjugate group modifies one or more properties of the attached AC, including, but not limited to, pharmacodynamics, pharmacokinetics, binding, absorption, cellular distribution, cellular uptake, charge, and clearance. Conjugate groups are routinely used in chemistry and are either directly linked to a parent compound, such as the AC, or linked via an optional linking moiety or group. Conjugate groups include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, thioethers, polyethers, cholesterol, thiocholesterol, cholic acid moieties, folic acid, lipids, phospholipids, biotin, phenazines, phenanthridines, anthraquinones, adamantanes, acridines, fluoresceins, rhodamines, coumarins, and dyes. In some embodiments, the conjugate group is polyethylene glycol (PEG), and the PEG is conjugated to either the AC or the CPP.
[0153] Conjugate groups include lipid moieties, such as cholesterol moieties (Letsinger et al., Proc. Natl. Acad. Sci. USA, 1989, 86, 6553), cholic acid (Manoharan et al., Bioorg. Med. Chem. Lett., 1994, 4, 1053), thioethers, such as hexyl-S-tritylthiol (Manoharan et al., Ann. NY Acad. Sci., 1992, 660, 306; Manoharan et al., Bioorg. Med. Chem. Lett., 1993, 3, 2765), thiocholesterol (Oberhauser et al., Nucl. Acids Res., 1992, 20, 533), aliphatic chains, such as dodecanediol or undecyl residues (Saison-Behmoaras et al., EMBO J., 1991, 10, 111; Kabanov et al., FEBS Lett., 1990, 259, 327; Svinarchuk et al., Biochimie, 1993, 75, 49), phospholipids, e.g., di-hexadecyl-rac-glycerol or triethylammonium-1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651; Shea et al., Nucl. Acids Res., 1990, 18, 3777), polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides, 1995, 14, 969), adamantane acetic acid (Manoharan et al., Tetrahedron Lett., 1995, 36, 3651), palmityl moiety (Mishra et al., Biochim. Biophys. Acta, 1995, 1264, 229), or octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther., 1996, 277, 923).
[0154] Linking groups or bifunctional linking moieties, such as those known in the art, are suitable for the compounds provided herein. Linking groups are useful for attaching chemical functional groups, conjugate groups, reporter groups, and other groups to selective sites in a parent compound, such as an AC. Generally, bifunctional linking moieties include a hydrocarbyl moiety with two functional groups. One of these functional groups is selected to attach to a parent molecule or compound of interest, and the other is selected to attach to essentially any selected group, such as a chemical functional group or conjugate group. Any of the linkers described herein can be used. In embodiments, the linker includes a chain structure or oligomer of repeating units, such as ethylene glycol or amino acid units. Examples of functional groups routinely used in bifunctional linking moieties include, but are not limited to, electrophiles for reacting with nucleophilic groups and nucleophiles for reacting with electrophilic groups. In embodiments, bifunctional linking moieties include amino, hydroxyl, carboxylic acid, thiol, and unsaturated groups (e.g., double or triple bonds). Some non-limiting examples of bifunctional linking moieties include 8-amino-3,6-dioxaoctanoic acid (ADO), succinimidyl 4-(N-maleimidomethyl)cyclohexane-1-carboxylate (SMCC), and 6-aminohexanoic acid (AHEX or AHA). Other linking groups include, but are not limited to, substituted C1-C10 alkyl, substituted or unsubstituted C2-C10 alkenyl, or substituted or unsubstituted C2-C10 alkynyl, where a non-limiting list of substituents includes hydroxyl, amino, alkoxy, carboxy, benzyl, phenyl, nitro, thiol, thioalkoxy, halogen, alkyl, aryl, alkenyl, and alkynyl.
[0155] In some embodiments, the AC can be linked to 10 arginine-serine dipeptide repeats. AC linked to 10 arginine-serine dipeptide repeats for artificial recruitment of splicing enhancer factors has been applied in vitro to induce the inclusion of mutant BRCA1 and SMN2 exons that would otherwise be skipped. See Cartegni and Krainer 2003, incorporated herein by reference.
[0156] Endosomal escape vehicles (EEVs) Endosomal escape vehicles (EEVs) can be used to transport cargo across cell membranes, for example, to deliver cargo to the cytosol or nucleus of a cell. The cargo can include a therapeutic moiety (TM). The EEV can include a cell membrane-penetrating peptide (CPP), for example, a cyclic cell membrane-penetrating peptide (cCPP). In some embodiments, the EEV includes a cCPP conjugated to an exocyclic peptide (EP). The EP can be interchangeably referred to as a regulatory peptide (MP). The EP can include a nuclear localization signal (NLS) sequence. The EP can be coupled to the cargo. The EP can be coupled to the cCPP. The EP can be coupled to the cargo and the cCPP. The coupling between the EP, cargo, cCPP, or a combination thereof can be non-covalent or covalent. The EP can be attached to the N-terminus of the cCPP via a peptide bond. The EP can be attached to the C-terminus of the cCPP via a peptide bond. The EP can be bound to the cCPP via the side chain of an amino acid in the cCPP. The EP can be bound to the cCPP via the side chain of a lysine, which can be conjugated to the side chain of a glutamine in the cCPP. The EP can be conjugated to the 5'-end or 3'-end of the oligonucleotide cargo. The EP can be coupled to a linker. The exocyclic peptide can be conjugated to the amino group of the linker. The EP can be coupled to the linker via a side chain on the cCPP and / or the EP, or via the C-terminus of the EP and the cCPP. For example, the EP may contain a terminal lysine and then be coupled to a cCPP containing glutamine via an amide bond. If the EP contains a terminal lysine and can be bound to the cCPP using the side chain of the lysine, the C-terminus or N-terminus can be bound to the linker on the cargo.
[0157] exocyclic peptides The exocyclic peptide (EP) can contain 2 to 10 amino acid residues, for example, 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residues, including all ranges and values therebetween. The EP can contain 6 to 9 amino acid residues. The EP can contain 4 to 8 amino acid residues.
[0158] Each amino acid in the exocyclic peptide can be a natural amino acid or an unnatural amino acid. The term "unnatural amino acid" refers to an organic compound that is a homolog of a natural amino acid in that it has a structure similar to a natural amino acid so as to mimic the structure and reactivity of the natural amino acid. An unnatural amino acid can be a modified amino acid and / or an amino acid analog that is neither one of the 20 common naturally occurring amino acids nor the rare natural amino acids selenocysteine or pyrrolysine. An unnatural amino acid can also be a D-isomer of a natural amino acid. Examples of suitable amino acids include, but are not limited to, alanine, allosoleucine, arginine, citrulline, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, naphthylalanine, phenylalanine, proline, pyroglutamic acid, serine, threonine, tryptophan, tyrosine, valine, derivatives thereof, or combinations thereof. These and other amino acids, along with their abbreviations used herein, are listed in Table 4. For example, an amino acid can be A, G, P, K, R, V, F, H, NaI, or citrulline.
[0159] The EP may contain at least one positively charged amino acid residue, such as at least one lysine residue and / or at least one amino acid residue containing a side chain containing a guanidine group, or a protonated form thereof. The EP may contain one or two amino acid residues containing a side chain containing a guanidine group, or a protonated form thereof. The amino acid residue containing a side chain containing a guanidine group may be an arginine residue. The protonated form may refer to a salt thereof throughout this disclosure.
[0160] The EP can contain at least two, at least three, or at least four or more lysine residues. The EP can contain two, three, or four lysine residues. The amino group on the side chain of each lysine residue can be substituted with a protecting group, including, for example, trifluoroacetyl (-COCF), allyloxycarbonyl (Alloc), 1-(4,4-dimethyl-2,6-dioxocyclohexylidene)ethyl (Dde), or (4,4-dimethyl-2,6-dioxocyclohex-1-ylidene-3)-methylbutyl (ivDde) groups. The amino group on the side chain of each lysine residue can be substituted with a trifluoroacetyl (-COCF) group. The protecting group can be included to enable amide conjugation. The protecting group can be removed after the EP is conjugated to the cCPP.
[0161] The EP may comprise at least two amino acid residues having a hydrophobic side chain. The amino acid residues having a hydrophobic side chain may be selected from valine, proline, alanine, leucine, isoleucine, and methionine. The amino acid residue having a hydrophobic side chain may be valine or proline.
[0162] The EP can contain at least one positively charged amino acid residue, for example, at least one lysine residue and / or at least one arginine residue. The EP can contain at least two, at least three, or at least four or more lysine and / or arginine residues.
[0163] EP is KK, KR, RR, HH, HK, HR, RH, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKH, KHK, HKK, HRR, HRH, HHR, HBH, HHH, HHHH (SEQ ID NO: 1), KHKK (SEQ ID NO: 2), KKHK (SEQ ID NO: 3), KKKH (SEQ ID NO: 4), KHKH (SEQ ID NO: 5), HKHK (SEQ ID NO: 6), KKKK (SEQ ID NO: 7), KKRK (SEQ ID NO: 8), KRKK (SEQ ID NO: 9), KR RK (SEQ ID NO: 10), RKKR (SEQ ID NO: 11), RRRR (SEQ ID NO: 12), KGKK (SEQ ID NO: 13), KKGK (SEQ ID NO: 14), HBHBH (SEQ ID NO: 15), HBKBH (SEQ ID NO: 16), RRRRR (SEQ ID NO: 17), KKKKK (SEQ ID NO: 18), KKKRK (SEQ ID NO: 19), RKKKK (SEQ ID NO: 20), KRKKK (SEQ ID NO: 21), KRKK (SEQ ID NO: 22), KKKKR (SEQ ID NO: 23), KBKBK (SEQ ID NO: 24), RKKKKG (SEQ ID NO: 25), KRKKKG (SEQ ID NO: 26), KRKRKKG (SEQ ID NO: 27), KKKKRG (SEQ ID NO: 28), RKKKKB (SEQ ID NO: 29), KRKKKB (SEQ ID NO: 30), KKRKKB (SEQ ID NO: 31), KKKKRB (SEQ ID NO: 32), KKKRKV (SEQ ID NO: 33), RRRRRR (SEQ ID NO: 34), HHHHHH (SEQ ID NO: 35), RHRHRH (SEQ ID NO: 36), HRHRHR (SEQ ID NO: 37), KRKRKR (SEQ ID NO:38), RKRKRK (SEQ ID NO:39), RBRBRB (SEQ ID NO:40), KBKBKB (SEQ ID NO:41), PKKKRKV (SEQ ID NO:42), PGKKRKV (SEQ ID NO:43), PKGKRKV (SEQ ID NO:44), PKKGRKV (SEQ ID NO:45), PKKKGKV (SEQ ID NO:46), PKKKRGV (SEQ ID NO:47), or PKKKRKG (SEQ ID NO:48), where B is beta-alanine. The amino acids in the EP can have D or L stereochemistry.
[0164] The EP can include KK, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK (SEQ ID NO:7), KKRK (SEQ ID NO:8), KRKK (SEQ ID NO:9), KRRK (SEQ ID NO:10), RKKR (SEQ ID NO:11), RRRR (SEQ ID NO:12), KGKK (SEQ ID NO:13), KKGK (SEQ ID NO:14), KKKKK (SEQ ID NO:18), KKKRK (SEQ ID NO:19), KBKBK (SEQ ID NO:24), KKKRKV (SEQ ID NO:33), PKKKRKV (SEQ ID NO:42), PGKKRKV (SEQ ID NO:43), PKGKRKV (SEQ ID NO:44), PKKGRKV (SEQ ID NO:45), PKKKGKV (SEQ ID NO:46), PKKKRGV (SEQ ID NO:47), or PKKKRKG (SEQ ID NO:48). The EP can include PKKKRKV (SEQ ID NO: 42), RR, RRR, RHR, RBR, RBRBR (SEQ ID NO: 49), RBHBR (SEQ ID NO: 50), or HBRBH (SEQ ID NO: 51), where B is beta-alanine. The amino acids in the EP can have D or L stereochemistry.
[0165] The EP can consist of KK, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK (SEQ ID NO:7), KKRK (SEQ ID NO:8), KRKK (SEQ ID NO:9), KRRK (SEQ ID NO:10), RKKR (SEQ ID NO:11), RRRR (SEQ ID NO:12), KGKK (SEQ ID NO:13), KKGK (SEQ ID NO:14), KKKKK (SEQ ID NO:18), KKKRK (SEQ ID NO:19), KBKBK (SEQ ID NO:24), KKKRKV (SEQ ID NO:33), PKKKRKV (SEQ ID NO:42), PGKKRKV (SEQ ID NO:Z43), PKGKRKV (SEQ ID NO:Z44), PKKGRKV (SEQ ID NO:Z45), PKKKGKV (SEQ ID NO:46), PKKKRGV (SEQ ID NO:47), or PKKKRKG (SEQ ID NO:48). The EP can consist of PKKKRKV (SEQ ID NO: 42), RR, RRR, RHR, RBR, RBRBR (SEQ ID NO: 49), RBHBR (SEQ ID NO: 50), or HBRBH (SEQ ID NO: 51), where B is beta-alanine. The amino acids in the EP can have D or L stereochemistry.
[0166] The EP can comprise an amino acid sequence identified in the art as a nuclear localization sequence (NLS). The EP can consist of an amino acid sequence identified in the art as a nuclear localization sequence (NLS). The EP can comprise an NLS comprising the amino acid sequence PKKKRKV (SEQ ID NO: 42). The EP can consist of an NLS comprising the amino acid sequence PKKKRKV (SEQ ID NO: 42). The EP can include an NLS comprising an amino acid sequence selected from NLSKRPAAIKKAGQAKKKK (SEQ ID NO:52), PAAKRVKLD (SEQ ID NO:53), RQRRNELKRSF (SEQ ID NO:54), RMRKFKNKGKDTAELRRRRVEVSVELR (SEQ ID NO:Z55), KAKKDEQILKRRNV (SEQ ID NO:56), VSRKRPRP (SEQ ID NO:57), PPKKARED (SEQ ID NO:58), PQPKKKPL (SEQ ID NO:59), SALIKKKKKMAP (SEQ ID NO:60), DRLRR (SEQ ID NO:61), PKQKKRK (SEQ ID NO:62), RKLKKKIKKL (SEQ ID NO:63), REKKKFLKRR (SEQ ID NO:64), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:65), and RKCLQAGMNLEARKTKK (SEQ ID NO:66). The EP can consist of an NLS comprising an amino acid sequence selected from NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 52), PAAKRVKLD (SEQ ID NO: 53), RQRRNELKRSF (SEQ ID NO: 54), RMRKFKNKGKDTAELRRRRVEVSVELR (SEQ ID NO: 55), KAKKDEQILKRRNV (SEQ ID NO: 56), VSRKRPRP (SEQ ID NO: 57), PPKKARED (SEQ ID NO: 58), PQPKKKPL (SEQ ID NO: 59), SALIKKKKKMAP (SEQ ID NO: 60), DRLRR (SEQ ID NO: 61), PKQKKRK (SEQ ID NO: 62), RKLKKKIKKL (SEQ ID NO: 63), REKKKFLKRR (SEQ ID NO: 64), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 65), and RKCLQAGMNLEARKTKK (SEQ ID NO: 66).
[0167] All exocyclic sequences can also include an N-terminal acetyl group. Thus, for example, an EP can have the following structure: Ac-PKKKRKV (SEQ ID NO: 42).
[0168] Cell-penetrating peptides (CPPs) A cell membrane-penetrating peptide (CPP) can contain 6 to 20 amino acid residues. The cell membrane-penetrating peptide can be a cyclic cell membrane-penetrating peptide (cCPP). The cCPP can permeate the cell membrane. An exocyclic peptide (EP) can be conjugated to the cCPP, and the resulting construct can be referred to as an endosomal escape vehicle (EEV). The cCPP can direct a cargo (e.g., a therapeutic moiety (TM), such as an oligonucleotide, peptide, or small molecule) to permeate the cell membrane. The cCPP can deliver the cargo to the cytosol of a cell. The cCPP can deliver the cargo to a cellular location where a target (e.g., pre-mRNA) is located. To conjugate the cCPP to a cargo (e.g., a peptide, oligonucleotide, or small molecule), at least one bond or lone pair of electrons on the cCPP can be replaced.
[0169] The total number of amino acid residues in a cCPP can range from 6 to 20 amino acid residues, for example, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid residues, including all ranges and subranges therebetween. A cCPP can contain 6 to 13 amino acid residues. A cCPP disclosed herein can contain 6 to 10 amino acids. In one example, a cCPP containing 6 to 10 amino acid residues can be represented by Formulas IA to IE: [ka] wherein AA1, AA2, AA3, AA4, AA5, AA6, AA7, AA8, AA9, and AA 10 are amino acid residues.
[0170] The cCPP can contain 6 to 8 amino acids. The cCPP can contain 8 amino acids.
[0171] Each amino acid in a cCPP can be a natural amino acid or an unnatural amino acid. The term "unnatural amino acid" refers to an organic compound that is a homolog of a natural amino acid in that it has a structure similar to a natural amino acid so as to mimic the structure and reactivity of the natural amino acid. An unnatural amino acid can be a modified amino acid and / or amino acid analog that is neither one of the 20 common naturally occurring amino acids nor the rare natural amino acids selenocysteine or pyrrolysine. An unnatural amino acid can also be a D-isomer of a natural amino acid. Examples of suitable amino acids include, but are not limited to, alanine, allosoleucine, arginine, citrulline, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, naphthylalanine, phenylalanine, proline, pyroglutamic acid, serine, threonine, tryptophan, tyrosine, valine, derivatives thereof, or combinations thereof. These and other amino acids, along with their abbreviations used herein, are listed in Table 5. [Table 4]
[0172] As used herein, "polyethylene glycol" and "PEG" are used interchangeably. m " has the formula HO(CO)-(CH2) n -(OCH2CH2) mIn some embodiments, the compound is a molecule of —NH2, wherein n is any integer from 1 to 5, and m is any integer from 1 to 23. In some embodiments, n is 1 or 2. In some embodiments, n is 1. In some embodiments, n is 2. In some embodiments, n is 1 and m is 2. In some embodiments, n is 2 and m is 2. In some embodiments, n is 1 and m is 4. In some embodiments, n is 2 and m is 4. In some embodiments, n is 1 and m is 12. In some embodiments, n is 2 and m is 12.
[0173] As used herein, "miniPEGm" or "miniPEG m " has the formula HO(CO)-(CH2) n -(OCH2CH2) m -NH2, where n is 1 and m is any integer from 1 to 23. For example, "miniPEG2" or "miniPEG2" is or is derived from (2-[2-[2-aminoethoxy]ethoxy]acetic acid), and "miniPEG4" or "miniPEG4" is HO(CO)-(CH2). n -(OCH2CH2) m -NH2, where n is 1 and m is 4.
[0174] A cCPP can contain 4 to 20 amino acids, in which (i) at least one amino acid has a side chain containing a guanidine group or its protonated form, and (ii) at least one amino acid has no side chain, or [ka] or their protonated forms, and (iii) at least two amino acids independently have side chains that include an aromatic group or a heteroaromatic group.
[0175] At least two amino acids have no side chains, or [ka] or have a side chain including their protonated form. As used herein, when no side chain is present, an amino acid has two hydrogen atoms on the carbon atom connecting the amine and carboxylic acid (e.g., -CH-).
[0176] The amino acid having no side chain can be glycine or β-alanine.
[0177] The cCPP can include 6 to 20 amino acid residues that form the cCPP, where (i) at least one amino acid can be a glycine, β-alanine, or 4-aminobutyric acid residue, (ii) at least one amino acid can have a side chain that includes an aryl or heteroaryl group, (iii) at least one amino acid can have a guanidine group, [ka] or have side chains containing their protonated forms.
[0178] The cCPP can include 6 to 20 amino acid residues that form the cCPP, where (i) at least two amino acids can be independently glycine, β-alanine, or 4-aminobutyric acid residues, (ii) at least one amino acid can have a side chain that includes an aryl or heteroaryl group, and (iii) at least one amino acid can have a guanidine group, [ka] or have side chains containing their protonated forms.
[0179] The cCPP can include 6 to 20 amino acid residues that form the cCPP, where (i) at least three amino acids can be independently glycine, β-alanine, or 4-aminobutyric acid residues, (ii) at least one amino acid can have a side chain that includes an aromatic or heteroaromatic group, and (iii) at least one amino acid can have a side chain that includes a guanidine group, [ka] or may have side chains containing their protonated forms.
[0180] Glycine and related amino acid residues A cCPP can contain (i) 1, 2, 3, 4, 5, or 6 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 2 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 3 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 4 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 5 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 6 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. A cCPP can contain (i) 3, 4, or 5 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) three or four glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof.
[0181] The cCPP can comprise (i) 1, 2, 3, 4, 5, or 6 glycine residues. The cCPP can comprise (i) 2 glycine residues. The cCPP can comprise (i) 3 glycine residues. The cCPP can comprise (i) 4 glycine residues. The cCPP can comprise (i) 5 glycine residues. The cCPP can comprise (i) 6 glycine residues. The cCPP can comprise (i) 3, 4, or 5 glycine residues. The cCPP can comprise (i) 3 or 4 glycine residues. The cCPP can comprise (i) 2 or 3 glycine residues. The cCPP can comprise (i) 1 or 2 glycine residues.
[0182] The cCPP can include (i) three, four, five, or six glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) three glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) four glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) five glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) six glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) three, four, or five glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP can include (i) three or four glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof.
[0183] The cCPP can comprise at least three glycine residues. The cCPP can comprise (i) 3, 4, 5, or 6 glycine residues. The cCPP can comprise (i) 3 glycine residues. The cCPP can comprise (i) 4 glycine residues. The cCPP can comprise (i) 5 glycine residues. The cCPP can comprise (i) 6 glycine residues. The cCPP can comprise (i) 3, 4, or 5 glycine residues. The cCPP can comprise (i) 3 or 4 glycine residues.
[0184] In embodiments, none of the glycine, β-alanine, or 4-aminobutyric acid residues in the cCPP are contiguous. Two or three glycine, β-alanine, 4- or 4-aminobutyric acid residues can be contiguous. Two glycine, β-alanine, or 4-aminobutyric acid residues can be contiguous.
[0185] In embodiments, none of the glycine residues in a cCPP are contiguous. Each glycine residue in a cCPP can be separated by an amino acid residue that cannot be glycine. Two or three glycine residues can be contiguous. Two glycine residues can be contiguous.
[0186] Amino acid side chains containing aromatic or heteroaromatic groups A cCPP can comprise 2, 3, 4, 5, or 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 2 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 5 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 2, 3, or 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can comprise 2 or 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group.
[0187] A cCPP can comprise 2, 3, 4, 5, or 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 2 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 5 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 2, 3, or 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group. A cCPP can comprise 2 or 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group.
[0188] The aromatic group can be a 6- to 14-membered aryl. The aryl can be phenyl, naphthyl, or anthracenyl, each of which is optionally substituted. The aryl can be phenyl or naphthyl, each of which is optionally substituted. The heteroaromatic group can be a 6- to 14-membered heteroaryl having 1, 2, or 3 heteroatoms selected from N, O, and S. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl.
[0189] The amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be bis(homonaphthylalanine), homonaphthylalanine, naphthylalanine, phenylglycine, bis(homophenylalanine), homophenylalanine, phenylalanine, tryptophan, 3-(3-benzothienyl)-alanine, 3-(2-quinolyl)-alanine, O-benzylserine, 3-(4-(benzyloxy)phenyl)-alanine, S-(4-methylbenzyl)cysteine, N-(naphthalen-2-yl)glutamine, 3-(1,1'-biphenyl-4-yl)-alanine, 3-(3-benzothienyl)-alanine, or tyrosine, each optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be: [ka] wherein H on the N-terminus and / or H on the C-terminus is replaced by a peptide bond.
[0190] The amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be a phenylalanine, naphthylalanine, phenylglycine, homophenylalanine, homonaphthylalanine, bis(homophenylalanine), bis-(homonaphthylalanine), tryptophan, or tyrosine residue, each optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group can each independently be tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butylphenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, or 3-(9-anthryl)alanine. The amino acid residues having a side chain containing an aromatic group can each independently be phenylalanine, naphthylalanine, phenylglycine, homophenylalanine, or homonaphthylalanine, each optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group can each independently be a phenylalanine, naphthylalanine, homophenylalanine, homonaphthylalanine, bis(homonaphthylalanine), or bis(homonaphthylalanine) residue, each optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group can each independently be a phenylalanine or naphthylalanine residue, each optionally substituted with one or more substituents. At least one amino acid residue having a side chain containing an aromatic group can be a phenylalanine residue. At least two amino acid residues having side chains containing an aromatic group can be phenylalanine residues. The amino acid residues having side chains containing an aromatic group can each be a phenylalanine residue.
[0191] In embodiments, none of the amino acids having a side chain containing an aromatic or heteroaromatic group are contiguous. Two amino acids having a side chain containing an aromatic or heteroaromatic group can be contiguous. Two consecutive amino acids can have opposite stereochemistry. Two consecutive amino acids can have the same stereochemistry. Three amino acids having a side chain containing an aromatic or heteroaromatic group can be contiguous. Three consecutive amino acids can have the same stereochemistry. Three consecutive amino acids can have alternating stereochemistry.
[0192] The amino acid residues containing aromatic or heteroaromatic groups can be L-amino acids. The amino acid residues containing aromatic or heteroaromatic groups can be D-amino acids. The amino acid residues containing aromatic or heteroaromatic groups can be a mixture of D- and L-amino acids.
[0193] The optional substituent can be, for example, any atom or group that does not significantly (e.g., by more than 50%) reduce the cytoplasmic delivery efficiency of the cCPP compared to an otherwise identical sequence without the substituent. The optional substituent can be a hydrophobic or hydrophilic substituent. The optional substituent can be a hydrophobic substituent. The substituent can increase the solvent-accessible surface area (as defined herein) of the hydrophobic amino acid. The substituent can be a halogen, alkyl, alkenyl, alkynyl, cycloalkyl, cycloalkenyl, cycloalkynyl, heterocyclyl, aryl, heteroaryl, alkoxy, aryloxy, acyl, alkylcarbamoyl, alkylcarboxamidyl, alkoxycarbonyl, alkylthio, or arylthio. The substituent can be a halogen.
[0194] Without wishing to be bound by theory, it is believed that amino acids having aromatic or heteroaromatic groups with higher hydrophobicity values (i.e., amino acids having side chains containing aromatic or heteroaromatic groups) can improve the cytoplasmic delivery efficiency of cCPPs compared to amino acids with lower hydrophobicity values. Each hydrophobic amino acid can independently have a hydrophobicity value greater than that of glycine. Each hydrophobic amino acid can independently have a hydrophobicity value greater than that of alanine. Each hydrophobic amino acid can independently have a hydrophobicity value equal to or greater than that of phenylalanine. Hydrophobicity can be measured using hydrophobicity scales known in the art. Table 5 lists the hydrophobicity values of various amino acids reported by Eisenberg and Weiss (Proc. Natl. Acad. Sci. USA 1984; 81(1): 140-144), Engleman, et al. (Ann. Rev. of Biophys. Biophys. Chem. 1986; 1986(15): 321-53), Kyte and Doolittle (J. Mol. Biol. 1982; 157(1): 105-132), Hoop and Woods (Proc. Natl. Acad. Sci. USA 1981; 78(6): 3824-3828), and Janin (Nature. 1979; 277(5696): 491-492), each of which is incorporated herein by reference in its entirety. Hydrophobicity can be measured using the hydrophobicity scale reported in Engleman, et al. [Table 5]
[0195] The size of the aromatic or heteroaromatic group can be selected to improve the cytoplasmic delivery efficiency of cCPPs. Without wishing to be bound by theory, it is believed that larger aromatic or heteroaromatic groups on the side chain of an amino acid can improve cytoplasmic delivery efficiency compared to an otherwise identical sequence having a smaller hydrophobic amino acid. The size of a hydrophobic amino acid can be measured in terms of the molecular weight of the hydrophobic amino acid, the steric effect of the hydrophobic amino acid, the solvent-accessible surface area (SASA) of the side chain, or a combination thereof. The size of a hydrophobic amino acid can be measured in terms of the molecular weight of the hydrophobic amino acid, with larger hydrophobic amino acids having side chains with molecular weights of at least about 90 g / mol, or at least about 130 g / mol, or at least about 141 g / mol. The size of an amino acid can be measured in terms of the SASA of the hydrophobic side chain. A hydrophobic amino acid can have a side chain with a SASA greater than or equal to that of alanine or glycine. A larger hydrophobic amino acid can have a side chain with a SASA greater than that of alanine or glycine. The hydrophobic amino acid can have an aromatic or heteroaromatic group with a SASA of about piperidine-2-carboxylic acid or greater, about tryptophan or greater, about phenylalanine or greater, or about naphthylalanine or greater. H1 ) is at least about 200 Å 2 , at least about 210 Å 2 , at least about 220 Å 2 , at least about 240 Å 2 , at least about 250 Å 2 , at least about 260 Å 2 , at least about 270 Å 2 , at least about 280 Å 2 , at least about 290 Å 2 , at least about 300 Å 2 , at least about 310 Å 2 , at least about 320 Å 2 , or at least about 330 Å 2 The second hydrophobic amino acid (AA H2 ) is at least about 200 Å 2 , at least about 210 Å2 , at least about 220 Å 2 , at least about 240 Å 2 , at least about 250 Å 2 , at least about 260 Å 2 , at least about 270 Å 2 , at least about 280 Å 2 , at least about 290 Å 2 , at least about 300 Å 2 , at least about 310 Å 2 , at least about 320 Å 2 , or at least about 330 Å 2 The side chain may have a SASA of AA H1 The side chain of AA H2 The side chains of 2 , at least 360 Å 2 , at least 370 Å 2 , at least 380 Å 2 , at least 390 Å 2 , at least 400 Å 2 , at least 410 Å 2 , at least 420 Å 2 , at least 430 Å 2 , at least 440 Å 2 , at least 450 Å 2 , at least 460 Å 2 , at least 470 Å 2 , at least 480 Å 2 , at least 490 Å 2 , about 500Å 2 Over, at least about 510 Å 2 , at least about 520 Å 2 , at least about 530 Å 2 , at least about 540 Å 2 , at least about 550 Å 2 , at least about 560 Å 2 , at least about 570 Å 2 , at least about 580 Å 2 , at least about 590 Å 2 , at least about 600 Å 2 , at least about 610 Å 2 , at least about 620 Å 2, at least about 630 Å 2 , at least about 640 Å 2 , approximately 650 Å 2 Over, at least about 660 Å 2 , at least about 670 Å 2 , at least about 680 Å 2 , at least about 690 Å 2 , or at least about 700 Å 2 It can have a SASA combination of AA H2 AA H1 The hydrophobic side chain may be a hydrophobic amino acid residue having a side chain with a SASA less than or equal to the SASA of the hydrophobic side chain of
[0043] By way of example and not limitation, a cCPP having a NaI-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a Phe-Arg motif, a cCPP having a Phe-NaI-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a NaI-Phe-Arg motif, and a Phe-NaI-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a NaI-Phe-Arg motif.
[0196] As used herein, "hydrophobic surface area" or "SASA" refers to the solvent-accessible surface area of an amino acid side chain (in square angstroms, Å 2 SASA refers to the surface area of a molecule (reported as a function of the surface area). SASA can be calculated using the "rolling ball" algorithm developed by Shrake & Rupley (J Mol Biol. 79(2):351-71), which is incorporated herein by reference in its entirety for all purposes. This algorithm uses a solvent "sphere" of a specific radius to probe the surface of the molecule. A typical value for the sphere is 1.4 Å, which approximates the radius of a water molecule.
[0197] The SASA values for certain side chains are shown below in Table 6. The SASA values described herein are based on the theoretical values listed in Table 6, as reported by Tien, et al. (PLOS ONE 8(11):e80635, available at doi.org / 10.1371 / journal.pone.0080635), which is incorporated herein by reference in its entirety for all purposes. [Table 6]
[0198] Amino acid residues having a side chain containing a guanidine group, a guanidine substituent, or their protonated forms As used herein, guanidine refers to the following structure: [ka] .
[0199] As used herein, the protonated form of guanidine refers to the following structure: [ka] .
[0200] A guanidine substituent refers to a functional group on the side chain of an amino acid that is positively charged at or above physiological pH or that is capable of replicating the hydrogen bond donating and accepting activity of a guanidinium group.
[0201] The guanidine substituents facilitate cell penetration and delivery of therapeutic agents while reducing toxicity associated with the guanidine group or its protonated form. The cCPP can include at least one amino acid having a side chain containing a guanidine or guanidinium substituent. The cCPP can include at least two amino acids having side chains containing a guanidine or guanidinium substituent. The cCPP can include at least three amino acids having side chains containing a guanidine or guanidinium substituent.
[0202] The guanidine or guanidinium group may be an isostere of guanidine or guanidinium. The guanidine or guanidinium substituent may be less basic than guanidine.
[0203] As used herein, a guanidine substituent is: [ka] or their protonated forms.
[0204] The present disclosure relates to cCPPs comprising 4 to 20 amino acid residues, wherein (i) at least one amino acid has a side chain comprising a guanidine group or its protonated form, and (ii) at least one amino acid residue has no side chain, or [ka] or their protonated forms, and (iii) at least two amino acid residues independently have side chains that include an aromatic group or a heteroaromatic group.
[0205] At least two amino acid residues have no side chains, or [ka] or may have side chains including their protonated forms. As used herein, when no side chains are present, an amino acid residue has two hydrogen atoms on the carbon atom connecting the amine and carboxylic acid (e.g., -CH-).
[0206] cCPP is one of the following parts: [ka] or at least one amino acid having a side chain including its protonated form.
[0207] Each cCPP is independently one of the following moieties: [ka] or a protonated form thereof. The at least two amino acids may be [ka] or their protonated forms. At least one amino acid may have a side chain comprising the same moiety selected from: [ka] or a side chain containing the protonated form thereof. [ka] or the side chains containing the protonated forms thereof. [ka] or its protonated form. [ka] or the side chains containing the protonated forms thereof. [ka] or may have side chains containing the protonated form thereof. [ka] Or their protonated forms can be attached to the termini of amino acid side chains. [ka] can be attached to the terminus of an amino acid side chain.
[0208] A cCPP can include 2, 3, 4, 5, or 6 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. A cCPP can include 2 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. A cCPP can include 3 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. A cCPP can include 4 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. A cCPP can include 5 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. A cCPP can include 6 amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP can comprise two, three, four, or five amino acid residues that independently have (iii) a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP can comprise two, three, or four amino acid residues that independently have (iii) a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP can comprise two or three amino acid residues that independently have (iii) a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP can comprise at least one amino acid residue that has (iii) a side chain comprising a guanidine group or a protonated form thereof. The cCPP can comprise two amino acid residues that have (iii) a side chain comprising a guanidine group or a protonated form thereof. The cCPP can comprise three amino acid residues that have (iii) a side chain comprising a guanidine group or a protonated form thereof.
[0209] The amino acid residues can independently have side chains containing non-contiguous guanidine groups, guanidine substituents, or their protonated forms. Two amino acid residues can independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms, which may be contiguous. Three amino acid residues can independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms, which may be contiguous. Four amino acid residues can independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms, which may be contiguous. Contiguous amino acid residues can have the same stereochemistry. Contiguous amino acids can have alternating stereochemistry.
[0210] The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or their protonated forms can be L-amino acids. The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or their protonated forms can be D-amino acids. The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or their protonated forms can be a mixture of D-amino acids or L-amino acids.
[0211] The amino acid residues having a side chain containing a guanidine group or its protonated form can each independently be an arginine residue, a homoarginine residue, a 2-amino-3-propionic acid residue, a 2-amino-4-guanidinobutyric acid residue, or a protonated form thereof. The amino acid residues having a side chain containing a guanidine group or its protonated form can each independently be an arginine residue or a protonated form thereof.
[0212] The amino acids having a side chain containing a guanidine substituent or its protonated form are each independently [ka] or their protonated forms.
[0213] Without being bound by theory, it is hypothesized that the guanidine substituent has reduced basicity compared to arginine and, in some cases, is uncharged (e.g., -N(H)C(O)) at physiological pH, allowing it to maintain bidentate hydrogen-bonding interactions with phospholipids on cell membranes, which is believed to facilitate effective membrane association and subsequent internalization. Removal of the positive charge is also believed to reduce the toxicity of cCPPs.
[0214] Those skilled in the art will understand that the N- and / or C-termini of the above non-natural aromatic hydrophobic amino acids will form an amide bond when incorporated into the peptides disclosed herein.
[0215] A cCPP can include a first amino acid having a side chain comprising an aromatic or heteroaromatic group and a second amino acid having a side chain comprising an aromatic or heteroaromatic group, where the N-terminus of the first glycine forms a peptide bond with the first amino acid having a side chain comprising an aromatic or heteroaromatic group and the C-terminus of the first glycine forms a peptide bond with the second amino acid having a side chain comprising an aromatic or heteroaromatic group. By convention, the term "first amino acid" often refers to the N-terminal amino acid of a peptide sequence; however, as used herein, the term "first amino acid" is used to distinguish the reference amino acid from another amino acid (e.g., a second amino acid) in the cCPP, such that the term "first amino acid" can refer or may refer to the amino acid located at the N-terminus of a peptide sequence.
[0216] The cCPP can include a second glycine N-terminus that forms a peptide bond with an amino acid having a side chain containing an aromatic or heteroaromatic group, and a second glycine C-terminus that forms a peptide bond with an amino acid having a side chain containing a guanidine group or its protonated form.
[0217] The cCPP can include a first amino acid having a side chain comprising a guanidine group or its protonated form, and a second amino acid having a side chain comprising a guanidine group or its protonated form, wherein the N-terminus of the third glycine forms a peptide bond with the first amino acid having a side chain comprising a guanidine group or its protonated form, and the C-terminus of the third glycine forms a peptide bond with the second amino acid having a side chain comprising a guanidine group or its protonated form.
[0218] The cCPP may comprise an asparagine, aspartic acid, glutamine, glutamic acid, or homoglutamine residue. The cCPP may comprise an asparagine residue. The cCPP may comprise a glutamine residue.
[0219] cCPPs can contain tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butyl-phenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, and 3-(9-anthryl)-alanine residues.
[0220] Without wishing to be bound by theory, it is believed that the chirality of amino acids in a cCPP can affect cytoplasmic uptake efficiency. A cCPP can contain at least one D amino acid. A cCPP can contain 1 to 15 D amino acids. A cCPP can contain 1 to 10 D amino acids. A cCPP can contain 1, 2, 3, or 4 D amino acids. A cCPP can contain 2, 3, 4, 5, 6, 7, or 8 consecutive amino acids with alternating D and L chirality. A cCPP can contain three consecutive amino acids with the same chirality. A cCPP can contain two consecutive amino acids with the same chirality. At least two of these amino acids can have opposite chirality. At least two amino acids with opposite chirality can be adjacent to each other. At least three amino acids can have alternating stereochemistry relative to each other. At least three amino acids with alternating chirality relative to each other can be adjacent to each other. At least four amino acids have alternating stereochemistry relative to each other. At least four amino acids of alternating chirality relative to one another can be adjacent to one another. At least two of these amino acids can have the same chirality. At least two amino acids of the same chirality can be adjacent to one another. At least two amino acids have the same chirality and at least two amino acids have opposite chirality. At least two amino acids of opposite chirality can be adjacent to at least two amino acids of the same chirality. Thus, adjacent amino acids in a cCPP can have any of the following sequences: DL, LD, DLLD, LDDL, LDLLD, DLDDL, DLLDL, or LDDLD. All amino acid residues forming a cCPP can be L-amino acids. All amino acid residues forming a cCPP can be D-amino acids.
[0221] At least two of these amino acids can have different chiralities. At least two amino acids with different chiralities can be adjacent to each other. At least three amino acids can have different chiralities relative to adjacent amino acids. At least four amino acids can have different chiralities relative to adjacent amino acids. At least two amino acids have the same chirality and at least two amino acids have different chiralities. One or more amino acid residues forming the cCPP can be achiral. The cCPP can include a motif of 3, 4, or 5 amino acids, where two amino acids with the same chirality can be separated by an achiral amino acid. The cCPP can include the following sequences: DXD, DXDX, DXDXD, LXL, LXLX, or LXLXL, where X is an achiral amino acid. The achiral amino acid can be glycine.
[0222] [ka] Or, amino acids having side chains containing their protonated forms can be adjacent to amino acids having side chains containing aromatic or heteroaromatic groups. [ka] or their protonated form can be adjacent to at least one amino acid having a side chain comprising guanidine or its protonated form. An amino acid having a side chain comprising guanidine or its protonated form can be adjacent to an amino acid having a side chain comprising an aromatic group or a heteroaromatic group. [ka] Two amino acids having a side chain containing guanidine or its protonated form can be adjacent to each other. Two amino acids having a side chain containing guanidine or its protonated form are adjacent to each other. A cCPP comprises at least two consecutive amino acids having side chains that can include an aromatic or heteroaromatic group, and [ka] or their protonated forms. A cCPP can comprise at least two non-adjacent amino acids having side chains containing an aromatic or heteroaromatic group, and [ka] or at least two non-adjacent amino acids having side chains containing the protonated form thereof. Adjacent amino acids can have the same chirality. Adjacent amino acids can have opposite chiralities. Other combinations of amino acids can have any arrangement of D and L amino acids, for example, any of the sequences described in the previous paragraph.
[0223] [ka] Or at least two amino acids having side chains containing the protonated form thereof alternate with at least two amino acids having side chains containing a guanidine group or the protonated form thereof.
[0224] cCPP has the structure of formula (A): [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4, R5, R6, and R7 are independently H or an amino acid side chain; at least one of R4, R5, R6, and R7 is a side chain of 3-guanidino-2-aminopropionic acid, 4-guanidino-2-aminobutanoic acid, arginine, homoarginine, N-methylarginine, N,N-dimethylarginine, 2,3-diaminopropionic acid, 2,4-diaminobutanoic acid, lysine, N-methyllysine, N,N-dimethyllysine, N-ethyllysine, N,N,N-trimethyllysine, 4-guanidinophenylalanine, citrulline, N,N-dimethyllysine, β-homoarginine, or 3-(1-piperidinyl)alanine; AA SC is the amino acid side chain, q is 1, 2, 3, or 4.
[0225] In some embodiments, at least one of R4, R5, R6, and R7 is independently an uncharged, non-aromatic side chain of an amino acid, hi some embodiments, at least one of R4, R5, R6, and R7 is independently H or the side chain of citrulline.
[0226] In some embodiments, compounds are provided that comprise a cyclic peptide having 6 to 12 amino acids, wherein at least two amino acids of the cyclic peptide are charged amino acids, at least two amino acids of the cyclic peptide are aromatic hydrophobic amino acids, and at least two amino acids of the cyclic peptide are uncharged non-aromatic amino acids. In some embodiments, at least two charged amino acids of the cyclic peptide are arginine. In some embodiments, at least two aromatic hydrophobic amino acids of the cyclic peptide are phenylalanine or naphthylalanine. In some embodiments, at least two uncharged non-aromatic amino acids of the cyclic peptide are citrulline or glycine.
[0227] In embodiments, the cyclic peptide of Formula (A) is not selected from cyclic peptides having the sequences of SEQ ID NOs: 89-117.
[0228] In embodiments, the cyclic peptide of Formula (A) is selected from cyclic peptides having the sequences of SEQ ID NOs: 89-117. [Table 22]
[0229] cCPPs have the structure of formula (I): [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R7 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; Each m is independently an integer of 0, 1, 2, or 3.
[0230] R1, R2, and R3 can each independently be H, -alkylene-aryl, or -alkylene-heteroaryl. R1, R2, and R3 can each independently be H, -C 1-3 Alkylene-aryl, or -C 1-3 R1, R2, and R3 can each independently be H or -alkylene-aryl. R1, R2, and R3 can each independently be H or -C 1-3 It can be alkylene-aryl. 1-3The alkylene can be methylene. The aryl can be a 6- to 14-membered aryl. The heteroaryl can be a 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, and isoquinolyl. R1, R2, and R3 are each independently selected from H, -C, 1-3 Alkylene-Ph or -C 1-3 R1, R2, and R3 can each independently be H, -CH2Ph, or -CH2naphthyl. R1, R2, and R3 can each independently be H or -CH2Ph.
[0231] R1, R2, and R3 can each independently be the side chain of tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butyl-phenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, 3-(9-anthryl)-alanine.
[0232] R1 can be the side chain of tyrosine. R1 can be the side chain of phenylalanine. R1 can be the side chain of 1-naphthylalanine. R1 can be the side chain of 2-naphthylalanine. R1 can be the side chain of tryptophan. R1 can be the side chain of 3-benzothienylalanine. R1 can be the side chain of 4-phenylphenylalanine. R1 can be the side chain of 3,4-difluorophenylalanine. R1 can be the side chain of 4-trifluoromethylphenylalanine. R1 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R1 can be the side chain of homophenylalanine. R1 can be the side chain of β-homophenylalanine. R1 can be the side chain of 4-tert-butyl-phenylalanine. R1 can be the side chain of 4-pyridinylalanine. R1 can be the side chain of 3-pyridinylalanine. R1 can be the side chain of 4-methylphenylalanine. R1 can be the side chain of 4-fluorophenylalanine. R1 can be the side chain of 4-chlorophenylalanine. R1 can be the side chain of 3-(9-anthryl)-alanine.
[0233] R2 can be the side chain of tyrosine. R2 can be the side chain of phenylalanine. R2 can be the side chain of 1-naphthylalanine. R1 can be the side chain of 2-naphthylalanine. R2 can be the side chain of tryptophan. R2 can be the side chain of 3-benzothienylalanine. R2 can be the side chain of 4-phenylphenylalanine. R2 can be the side chain of 3,4-difluorophenylalanine. R2 can be the side chain of 4-trifluoromethylphenylalanine. R2 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R2 can be the side chain of homophenylalanine. R2 can be the side chain of β-homophenylalanine. R2 can be the side chain of 4-tert-butyl-phenylalanine. R2 can be the side chain of 4-pyridinylalanine. R2 can be the side chain of 3-pyridinylalanine. R2 can be the side chain of 4-methylphenylalanine. R2 can be the side chain of 4-fluorophenylalanine. R2 can be the side chain of 4-chlorophenylalanine. R2 can be the side chain of 3-(9-anthryl)-alanine.
[0234] R3 can be the side chain of tyrosine. R3 can be the side chain of phenylalanine. R3 can be the side chain of 1-naphthylalanine. R3 can be the side chain of 2-naphthylalanine. R3 can be the side chain of tryptophan. R3 can be the side chain of 3-benzothienylalanine. R3 can be the side chain of 4-phenylphenylalanine. R3 can be the side chain of 3,4-difluorophenylalanine. R3 can be the side chain of 4-trifluoromethylphenylalanine. R3 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R3 can be the side chain of homophenylalanine. R3 can be the side chain of β-homophenylalanine. R3 can be the side chain of 4-tert-butyl-phenylalanine. R3 can be the side chain of 4-pyridinylalanine. R3 can be the side chain of 3-pyridinylalanine. R3 can be the side chain of 4-methylphenylalanine. R3 can be the side chain of 4-fluorophenylalanine. R3 can be the side chain of 4-chlorophenylalanine. R3 can be the side chain of 3-(9-anthryl)-alanine.
[0235] R4 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1-3 Alkylene-aryl, or -C 1-3 R4 can be H or -alkylene-aryl. R4 can be H or -C 1-3 It can be alkylene-aryl. 1-3The alkylene can be methylene. The aryl can be a 6-14 membered aryl. The heteroaryl can be a 6-14 membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, and isoquinolyl. R4 can be H, -C 1-3 Alkylene-Ph or -C 1-3 R4 can be alkylene-naphthyl. R4 can be H or the side chain of an amino acid in Table 4 or Table 6. R4 can be H or an amino acid residue having a side chain containing an aromatic group. R4 can be H, -CH2Ph, or -CH2naphthyl. R4 can be H or -CH2Ph.
[0236] R5 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1-3 Alkylene-aryl, or -C 1-3 R5 can be H or -alkylene-aryl. R5 can be H or -C 1-3 It can be alkylene-aryl. 1-3 The alkylene can be methylene. The aryl can be a 6-14 membered aryl. The heteroaryl can be a 6-14 membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, and isoquinolyl. R5 can be H, -C 1-3 Alkylene-Ph or -C 1-3R5 can be H or the side chain of an amino acid in Table 4 or Table 6. R4 can be H or an amino acid residue having a side chain containing an aromatic group. R5 can be H, -CH2Ph, or -CH2naphthyl. R4 can be H or -CH2Ph.
[0237] R6 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1-3 Alkylene-aryl, or -C 1-3 R6 can be H or -alkylene-aryl. R6 can be H or -C 1-3 It can be alkylene-aryl. 1-3 The alkylene can be methylene. The aryl can be a 6-14 membered aryl. The heteroaryl can be a 6-14 membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, and isoquinolyl. R6 can be H, -C 1-3 Alkylene-Ph or -C 1-3 R6 can be alkylene-naphthyl. R6 can be H or the side chain of an amino acid in Table 4 or Table 6. R6 can be H or an amino acid residue having a side chain containing an aromatic group. R6 can be H, -CH2Ph, or -CH2naphthyl. R6 can be H or -CH2Ph.
[0238] R7 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1-3 Alkylene-aryl, or -C 1-3 R7 can be H or -alkylene-aryl. R7 can be H or -C1-3 It can be alkylene-aryl. 1-3 The alkylene can be methylene. The aryl can be a 6-14 membered aryl. The heteroaryl can be a 6-14 membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, and isoquinolyl. R7 can be H, -C 1-3 Alkylene-Ph or -C 1-3 R7 can be alkylene-naphthyl. R7 can be H or the side chain of an amino acid in Table 4 or Table 6. R7 can be H or an amino acid residue having a side chain containing an aromatic group. R7 can be H, -CH2Ph, or -CH2naphthyl. R7 can be H or -CH2Ph.
[0239] One, two, or three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. One of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Two of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. At least one of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Up to four of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph.
[0240] One, two, or three of R1, R2, R3, and R4 are -CH2Ph. One of R1, R2, R3, and R4 is -CH2Ph. Two of R1, R2, R3, and R4 are -CH2Ph. Three of R1, R2, R3, and R4 are -CH2Ph. At least one of R1, R2, R3, and R4 is -CH2Ph.
[0241] One, two, or three of R1, R2, R3, R4, R5, R6, and R7 can be H. One of R1, R2, R3, R4, R5, R6, and R7 can be H. Two of R1, R2, R3, R4, R5, R6, and R7 are H. Three of R1, R2, R3, R5, R6, and R7 can be H. At least one of R1, R2, R3, R4, R5, R6, and R7 can be H. Up to three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph.
[0242] One, two, or three of R1, R2, R3, and R4 are H. One of R1, R2, R3, and R4 is H. Two of R1, R2, R3, and R4 are H. Three of R1, R2, R3, and R4 are H. At least one of R1, R2, R3, and R4 is H.
[0243] At least one of R4, R5, R6, and R7 can be the side chain of 3-guanidino-2-aminopropionic acid. At least one of R4, R5, R6, and R7 can be the side chain of 4-guanidino-2-aminobutanoic acid. At least one of R4, R5, R6, and R7 can be the side chain of arginine. At least one of R4, R5, R6, and R7 can be the side chain of homoarginine. At least one of R4, R5, R6, and R7 can be the side chain of N-methylarginine. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethylarginine. At least one of R4, R5, R6, and R7 can be the side chain of 2,3-diaminopropionic acid. At least one of R4, R5, R6, and R7 can be the side chain of 2,4-diaminobutanoic acid, lysine. At least one of R4, R5, R6, and R7 can be the side chain of N-methyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N-ethyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N,N,N-trimethyllysine, 4-guanidinophenylalanine. At least one of R4, R5, R6, and R7 can be the side chain of citrulline. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethyllysine, β-homoarginine. At least one of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.
[0244] At least two of R4, R5, R6, and R7 can be side chains of 3-guanidino-2-aminopropionic acid. At least two of R4, R5, R6, and R7 can be side chains of 4-guanidino-2-aminobutanoic acid. At least two of R4, R5, R6, and R7 can be side chains of arginine. At least two of R4, R5, R6, and R7 can be side chains of homoarginine. At least two of R4, R5, R6, and R7 can be side chains of N-methylarginine. At least two of R4, R5, R6, and R7 can be side chains of N,N-dimethylarginine. At least two of R4, R5, R6, and R7 can be side chains of 2,3-diaminopropionic acid. At least two of R4, R5, R6, and R7 can be the side chains of 2,4-diaminobutanoic acid, lysine. At least two of R4, R5, R6, and R7 can be the side chains of N-methyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N-ethyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N,N,N-trimethyllysine, 4-guanidinophenylalanine. At least two of R4, R5, R6, and R7 can be the side chains of citrulline. At least two of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine, β-homoarginine. At least two of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.
[0245] At least three of R4, R5, R6, and R7 can be the side chains of 3-guanidino-2-aminopropionic acid. At least three of R4, R5, R6, and R7 can be the side chains of 4-guanidino-2-aminobutanoic acid. At least three of R4, R5, R6, and R7 can be the side chains of arginine. At least three of R4, R5, R6, and R7 can be the side chains of homoarginine. At least three of R4, R5, R6, and R7 can be the side chains of N-methylarginine. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethylarginine. At least three of R4, R5, R6, and R7 can be the side chains of 2,3-diaminopropionic acid. At least three of R4, R5, R6, and R7 can be the side chains of 2,4-diaminobutanoic acid or lysine. At least three of R4, R5, R6, and R7 can be the side chains of N-methyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N-ethyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N,N,N-trimethyllysine or 4-guanidinophenylalanine. At least three of R4, R5, R6, and R7 can be the side chains of citrulline. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine or β-homoarginine. At least three of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.
[0246] AA SC can be the side chain of an asparagine, glutamine, or homoglutamine residue. SC can be the side chain of a glutamine residue. SCFor example, a cCPP may further comprise a linker conjugated to an asparagine, glutamine, or homoglutamine residue. Thus, a cCPP may further comprise a linker conjugated to an asparagine, glutamine, or homoglutamine residue. A cCPP may further comprise a linker conjugated to a glutamine residue.
[0247] q can be 1, 2, or 3. q can be 1 or 2. q can be 1. q can be 2. q can be 3. q can be 4.
[0248] m can be 1 to 3. m can be 1 or 2. m can be 0. m can be 1. m can be 2. m can be 3.
[0249] The cCPP of formula (A) has the structure of formula (I): [ka] , or a protonated form thereof, wherein AA SC , R1, R2, R3, R 4- , R7, m, and q are as defined herein.
[0250] The cCPP of formula (A) has the structure of formula (Ia) or formula (Ib): [ka] or a protonated form thereof, wherein AA SC , R1, R2, R3, R 4- , and m is as defined herein.
[0251] The cCPP of formula (A) has the structure of formula (I-1), formula (I-2), formula (I-3), or formula (I-4): [ka] [ka] or a protonated form thereof, wherein AA SC and m is as defined herein.
[0252] The cCPP of formula (A) has the structure of formula (I-5) or formula (I-6): [ka] or a protonated form thereof, wherein AA SC- is as defined herein.
[0253] The cCPP of formula (A) has the structure of formula (I-1): [ka] , or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.
[0254] The cCPP of formula (A) has the structure of formula (I-2): [ka] , or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.
[0255] The cCPP of formula (A) has the structure of formula (I-3): [ka] , or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.
[0256] The cCPP of formula (A) has the structure of formula (I-4): [ka] , or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.
[0257] The cCPP of formula (A) has the structure of formula (I-5): [ka] , or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.
[0258] The cCPP of formula (A) has the structure of formula (I-6): [ka] , or a protonated form thereof, wherein AA SC- and m is as defined herein.
[0259] The cCPP can comprise one of the following sequences: FGFGRGR (SEQ ID NO: 68), GfFGrGr (SEQ ID NO: 69), FfΦGRGR (SEQ ID NO: 70), FfFGRGR (SEQ ID NO: 71), or FfΦGrGr (SEQ ID NO: 72). The cCPP can have one of the following sequences: FGFΦ (SEQ ID NO: 73), GfFGrGrQ (SEQ ID NO: 74), FfΦGRGRQ (SEQ ID NO: 75), FfFGRGRQ (SEQ ID NO: 76), or FfΦGrGrQ (SEQ ID NO: 77).
[0260] The present disclosure also relates to a cCPP having the structure of formula (II): [ka] During the ceremony, AA SC is the amino acid side chain, R 1a , R 1b , and R 1c are each independently a 6- to 14-membered aryl or a 6- to 14-membered heteroaryl, R 2a , R 2b , R 2c , and R 2d are independently amino acid side chains, R 2a , R 2b , R 2c , and R 2d At least one of the [ka] or their protonated forms, R 2a , R 2b , R 2c , and R 2d at least one of is guanidine or its protonated form; each n" is independently an integer of 0, 1, 2, 3, 4, or 5; each n' is independently an integer of 0, 1, 2, or 3; If n' is 0, then R 2a , R 2b , R 2b , or R 2d is absent.
[0261] R 2a , R 2b , R 2c , and R 2d At least two of the [ka] or their protonated forms. 2a , R2 b , R2 c , and R 2d Two or three of them are [ka] or their protonated forms. 2a , R 2b , R 2c , and R 2d One of them is [ka] or their protonated forms. 2a , R 2b , R 2c , and R 2d At least one of the [ka] or its protonated form, R 2a , R 2b , R 2c , and R 2d The remainder of R can be guanidine or its protonated form. 2a , R 2b , R 2c , and R 2d At least two of the [ka] or its protonated form, R 2a , R 2b , R 2c , and R 2d The remainder can be guanidine or its protonated form.
[0262] R 2a , R 2b , R 2c , and R 2d are all [ka] or their protonated forms. 2a , R 2b , R 2c , and R 2d At least some of [ka] or its protonated form, R 2a , R 2b , R 2c , and R 2d The remainder of R can be guanide or its protonated form. 2a , R 2b , R 2c , and R 2d teeth, [ka] or its protonated form, R 2a , R 2b , R 2c , and R 2d The remainder is guanidine or its protonated form.
[0263] R 2a , R 2b , R 2c , and R 2d can each independently be the side chain of 2,3-diaminopropionic acid, 2,4-diaminobutyric acid, ornithine, lysine, methyllysine, dimethyllysine, trimethyllysine, homolysine, serine, homoserine, threonine, allotreonine, histidine, 1-methylhistidine, 2-aminobutanedioic acid, aspartic acid, glutamic acid, or homoglutamic acid.
[0264] AA SC teeth, [ka] where t is an integer from 0 to 5. SC teeth, [ka] wherein t is an integer of 0 to 5. t is 1 to 5. t is 2 or 3. t is 2. t is 3.
[0265] R 1a , R 1b , and R 1c Each R can independently be a 6- to 14-membered aryl. 1a , R 1b , and R 1c Each R can independently be a 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, or S. 1a , R 1b , and R 1c Each R can be independently selected from phenyl, naphthyl, anthracenyl, pyridyl, quinolyl, or isoquinolyl. 1a , R 1b , and R 1c Each R can be independently selected from phenyl, naphthyl, or anthracenyl. 1a , R 1b , and R 1c Each R can independently be phenyl or naphthyl. 1a , R 1b , and R 1c may each independently be selected from pyridyl, quinolyl, or isoquinolyl.
[0266] Each n' can independently be 1 or 2. Each n' can be 1. Each n' can be 2. At least one n' can be 0. At least one n' can be 1. At least one n' can be 2. At least one n' can be 3. At least one n' can be 4. At least one n' can be 5.
[0267] Each n" can be independently an integer from 1 to 3. Each n" can be independently 2 or 3. Each n" can be 2. Each n" can be 3. At least one n" can be 0. At least one n" can be 1. At least one n" can be 2. At least one n" can be 3.
[0268] Each n" can independently be 1 or 2, and each n' can independently be 2 or 3. Each n" can be 1, and each n' can independently be 2 or 3. Each n" can be 1, and each n' can be 2. Each n" is 1, and each n' is 3.
[0269] The cCPP of formula (II) can have the structure of formula (II-1): [ka] In the formula, R 1a , R 1b , R 1c , R 2a , R 2b , R 2c , R 2d , A.A. SC- , n', and n" are as defined herein.
[0270] The cCPP of formula (II) can have the structure of formula (IIa): [ka] In the formula, R 1a , R 1b , R 1c , R 2a , R 2b , R 2c , R 2d , A.A. SC- , and n' are as defined herein.
[0271] The cCPP of formula (II) can have the structure of formula (IIb): [ka] In the formula, R 2a , R 2b , A.A. SC- , and n' are as defined herein.
[0272] cCPP has the structure of formula (IIc): [ka] , or its protonated form, During the ceremony, AA SC and n' is as defined herein.
[0273] The cCPP of formula (IIa) has the following structure: [ka] [ka] wherein AA SC and n is as defined herein.
[0274] The cCPP of formula (IIa) has the following structure: [ka] [ka] wherein AA SC and n is as defined herein.
[0275] The cCPP of formula (IIa) has the following structure: [ka] [ka] wherein AA SC and n is as defined herein.
[0276] The cCPP of formula (II) can have the following structure: [ka] .
[0277] The cCPP of formula (II) has the following structure: [ka] It can have:
[0278] The cCPP can have the structure of formula (III): [ka] During the ceremony, AA SC is the amino acid side chain, R 1a , R 1b , and R 1c are each independently a 6- to 14-membered aryl or a 6- to 14-membered heteroaryl, R 2a and R 2c are each independently H, [ka] or their protonated forms, R 2b and R 2d are each independently guanidine or its protonated form; n" is independently an integer from 1 to 3, n' is independently an integer from 1 to 5, Each p' is independently an integer of 0 to 5.
[0279] The cCPP of formula (III) can have the structure of formula (III-1): [ka] During the ceremony, AA SC , R 1a , R 1b , R 1c , R 2a , R 2c , R 2b , R 2d , n', n'', and p' are as defined herein.
[0280] The cCPP of formula (III) can have the structure of formula (IIIa): [ka] During the ceremony, AA SC , R 2a , R 2c , R 2b , R 2d , n', n'', and p' are as defined herein.
[0281] In formula (III), formula (III-1), and formula (IIIa), R a and R c can be H. R a and R c can be H and R b and R d Each R can independently be guanidine or its protonated form. a can be H. R b can be H. p' can be 0. R a and R c can be H, and each of p' can be 0.
[0282] In formula (III), formula (III-1), and formula (IIIa), R a and R c can be H and R b and R dcan each independently be guanidine or its protonated form, n" can be 2 or 3, and each p' can be 0.
[0283] p' can be 0. p' can be 1. p' can be 2. p' can be 3. p' can be 4. p' can be 5.
[0284] The cCPP can have the following structure: [ka] .
[0285] The cCPP of formula (A) can be selected from: [Table 23]
[0286] The cCPP of formula (A) can be selected from: [Table 24]
[0287] In some embodiments, the cCPP is selected from the following: [Table 25]
[0288] In some embodiments, the cCPP is not selected from the following: [Table 26]
[0289] cCPP has the structure of formula (D): [ka] , or a protonated form thereof, wherein: R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; Y is, [ka] and AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; Each n is independently an integer of 0, 1, 2, or 3.
[0290] The cCPP of formula (D) has the structure of formula (DI): [ka] or its protonated form, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; Y is, [ka] is.
[0291] The cCPP of formula (D) has the structure of formula (D-II): [ka] or its protonated form, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; each m is independently an integer of 0, 1, 2, or 3; Y is, [ka] is.
[0292] The cCPP of formula (D) has the structure of formula (D-III): [ka] or its protonated form, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; n is independently an integer of 0, 1, 2, or 3; Y is, [ka] is.
[0293] The cCPP of formula (D) has the structure of formula (D-IV): [ka] or its protonated form, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; Y is, [ka] is.
[0294] The cCPP of formula (D) has the structure of formula (DV): [ka] or its protonated form, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R6 are independently H or an amino acid side chain; AA SC is the amino acid side chain, q is 1, 2, 3, or 4; each m is independently an integer of 0, 1, 2, or 3; Y is, [ka] is.
[0295] AA SC can be conjugated to a linker.
[0296] Linker The cCPP of the present disclosure can be conjugated to a linker. The linker can link a cargo to the cCPP. The linker can be attached to the side chain of an amino acid of the cCPP, and the cargo can be attached at a suitable position on the linker.
[0297] The linker can be any suitable moiety capable of conjugating the cCPP to one or more additional moieties, such as an exocyclic peptide (EP) and / or cargo. Prior to conjugation to the cCPP and one or more additional moieties, the linker has two or more functional groups, each of which can independently form a covalent bond to the cCPP and one or more additional moieties. When the cargo is an oligonucleotide, the linker can be covalently attached to the 5'-end of the cargo or the 3'-end of the cargo. The linker can be covalently attached to the 5'-end of the cargo. The linker can be covalently attached to the 3'-end of the cargo. When the cargo is a peptide, the linker can be covalently attached to the N-terminus or C-terminus of the cargo. The linker can be covalently attached to the backbone of an oligonucleotide or peptide cargo. The linker can be any suitable moiety capable of conjugating the cCPP described herein to a cargo, such as an oligonucleotide, peptide, or small molecule.
[0298] The linker can include a hydrocarbon linker.
[0299] The linker may comprise a cleavage site, which may be a disulfide or a caspase cleavage site (e.g., Val-Cit-PABC).
[0300] The linker may be selected from the group consisting of (i) one or more D or L amino acids (each optionally substituted), (ii) an optionally substituted alkylene, (iii) an optionally substituted alkenylene, (iv) an optionally substituted alkynylene, (v) an optionally substituted carbocyclyl, (vi) an optionally substituted heterocyclyl, (vii) one or more -(R 1- JR 2 )z″-subunit (wherein R 1 and R 2 are each independently selected at each occurrence from alkylene, alkenylene, alkynylene, carbocyclyl, and heterocyclyl; J is each independently selected from C, NR 3 , -NR 3 C(O)—, S, and O, where R 3 is independently selected from H, alkyl, alkenyl, alkynyl, carbocyclyl, and heterocyclyl (each optionally substituted), and z″ is an integer from 1 to 50; (viii)—(R 1- J)z”-or-(JR 1 )z”-(wherein, R 1 is independently at each occurrence alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl; J is independently at each occurrence C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3 is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl (each optionally substituted); and z″ is an integer from 1 to 50; or (ix) the linker can comprise one or more of (i) through (x).
[0301] The linker may comprise one or more D or L amino acids and / or -(R 1- JR 2 )z”-(wherein, R1 and R 2 is independently at each occurrence alkylene; and J is independently at each occurrence C, NR 3 , -NR 3 C(O)—, S, and O, where R 4 are independently selected from H and alkyl, and z″ is an integer from 1 to 50), or a combination thereof.
[0302] The linker may be (for example, as a spacer) -(OCH2CH2) z’ - (wherein z' is an integer from 1 to 23, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23). "-(OCH2CH2)z'" can also be referred to as polyethylene glycol (PEG).
[0303] The linker can comprise one or more amino acids. The linker can comprise a peptide. The linker can comprise -(OCH2CH2) z’ - (wherein z' is an integer from 1 to 23) and a peptide. The peptide can include 2 to 10 amino acids. The linker can further include a functional group (FG) that can react via click chemistry. FG can be an azide or an alkyne, which forms a triazole when the cargo is conjugated to the linker.
[0304] The linker consists of (i) β-alanine and lysine residues, (ii) -(JR 1 )z”, or (iii) a combination thereof. 1 each independently can be alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl; and each J independently can be C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl (each optionally substituted), and z″ can be an integer from 1 to 50. R 1 Each R can be alkylene and each J can be O.
[0305] The linker may comprise (i) β-alanine, glycine, lysine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid residues, or a combination thereof, and (ii) -(R 1- J)z”-or-(JR 1 )z”. R 1 each independently can be alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl; and each J independently can be C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3 is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl (each optionally substituted), and z″ can be an integer from 1 to 50. R 1 Each R can be alkylene and each J can be O. The linker can include glycine, beta-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid, or a combination thereof.
[0306] The linker can be a trivalent linker. The linker has the structure: [ka] wherein A1, B1, and C1 are independently a hydrocarbon linker (e.g., NRH-(CH2) n -COOH), PEG linker (e.g., NRH-(CHO) n -COOH (wherein R is H, methyl, or ethyl), or one or more amino acid residues, and Z is independently a protecting group. The linker can be a disulfide [NH-(CHO) n -SS-(CH2O) n-COOH] or a cleavage site containing a caspase cleavage site (Val-Cit-PABC) can also be incorporated.
[0307] The carbohydrate can be a glycine or beta-alanine residue.
[0308] The linker can be bivalent and can link the cCPP to the cargo. The linker can be bivalent and can link the cCPP to the exocyclic peptide (EP).
[0309] The linker can be trivalent and can connect the cCPP to the cargo and the EP.
[0310] The linker may be a divalent or trivalent C1-C 50 and alkylene, wherein 1 to 25 methylene groups are optionally and independently replaced by -N(H)-, -N(C-C alkyl)-, -N(cycloalkyl)-, -O-, -C(O)-, -C(O)O-, -S-, -S(O)-, -S(O)-, -S(O)N(C-C alkyl)-, -S(O)N(cycloalkyl)-, -N(H)C(O)-, -N(C-C alkyl)C(O)-, -N(cycloalkyl)C(O)-, -C(O)N(H)-, -C(O)N(C-C alkyl), -C(O)N(cycloalkyl), aryl, heterocyclyl, heteroaryl, cycloalkyl, or cycloalkenyl. 50 alkylene, in which 1 to 25 methylene groups are optionally and independently replaced by -N(H)-, -O-, -C(O)N(H)-, or combinations thereof.
[0311] The linker has the following structure: [ka] wherein each AA is independently an amino acid residue and * is AA SC AA SCis a side chain of an amino acid residue of cCPP, x is an integer of 1 to 10, y is an integer of 1 to 5, and z is an integer of 1 to 10. x can be an integer of 1 to 5. x can be an integer of 1 to 3. x can be 1. y can be an integer of 2 to 4. y can be 4. z can be an integer of 1 to 5. z can be an integer of 1 to 3. z can be 1. Each AA can be independently selected from glycine, β-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, and 6-aminohexanoic acid.
[0312] The cCPP can be linked to the cargo via a linker ("L"), which can be conjugated to the cargo via a linking group ("M").
[0313] The linker has the following structure: [ka] wherein x is an integer from 1 to 10, y is an integer from 1 to 5, z is an integer from 1 to 10, each AA is independently an amino acid residue, * is AA SC AA SC is the side chain of an amino acid residue of the cCPP, and M is a linking group as defined herein.
[0314] The linker can have the following structure: [ka] In the formula, x' is an integer of 1 to 23, y is an integer of 1 to 5, z' is an integer of 1 to 23, and * is AA SC AA SC is the side chain of an amino acid residue of the cCPP, and M is a linking group as defined herein.
[0315] The linker can have the following structure: [ka] In the formula, x' is an integer of 1 to 23, y is an integer of 1 to 5, z' is an integer of 1 to 23, and * is AA SC AA SC are the side chains of the amino acid residues of cCPP.
[0316] x can be an integer between 1 and 10, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, including all ranges and subranges therebetween.
[0317] x' can be an integer from 1 to 23, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23, including all ranges and subranges therebetween. x' can be an integer from 5 to 15. x' can be an integer from 9 to 13. x' can be an integer from 1 to 5. x' can be 1.
[0318] y can be an integer from 1 to 5, for example, 1, 2, 3, 4, or 5, including all ranges and subranges therebetween. y can be an integer from 2 to 5. y can be an integer from 3 to 5. y can be 3 or 4. y can be 4 or 5. y can be 3. y can be 4. y can be 5.
[0319] z can be an integer between 1 and 10, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, including all ranges and subranges therebetween.
[0320] z' can be an integer from 1 to 23, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23, including all ranges and subranges therebetween. z' can be an integer from 5 to 15. z' can be an integer from 9 to 13. z' can be 11.
[0321] As mentioned above, the linker or M (where M is part of the linker) can be covalently attached to the cargo at any suitable position on the cargo. The linker or M (where M is part of the linker) can be covalently attached to the 3'-end of the oligonucleotide cargo or the 5'-end of the oligonucleotide cargo. The linker or M (where M is part of the linker) can be covalently attached to the N-terminus or C-terminus of the peptide cargo. The linker or M (where M is part of the linker) can be covalently attached to the backbone of the oligonucleotide or peptide cargo.
[0322] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the cCPP, or to a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the cCPP.
[0323] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the peptide cargo, or to a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the peptide cargo.
[0324] The linker can have the following structure: [ka] During the ceremony, M is a group that conjugates L to a cargo, e.g., an oligonucleotide; AA s is the side chain or terminus of an amino acid on the cCPP, AA x are each independently an amino acid residue, o is an integer from 0 to 10, p is an integer of 0 to 5.
[0325] The linker can have the following structure: [ka] During the ceremony, M is a group that conjugates L to a cargo, e.g., an oligonucleotide; AA s is the side chain or terminus of an amino acid on the cCPP, AA x are each independently an amino acid residue, o is an integer from 0 to 10, p is an integer of 0 to 5.
[0326] M can include alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl, each of which is optionally substituted. [ka] wherein R is alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl.
[0327] M is [ka] wherein R 10 is alkylene, cycloalkyl, or [ka] In the formula, a is 0 to 10.
[0328] M is [ka] and R 10 teeth, [ka] where a is 0 to 10. M can be [ka] It can be said that:
[0329] M is a heterobifunctional cross-linker, e.g., [ka] which is disclosed in Williams et al. Curr. Protoc Nucleic Acid Chem. 2010, 42, 4.41.1-4.41.20, which is incorporated herein by reference in its entirety.
[0330] M can be —C(O)—.
[0331] AA s can be the side chain or terminus of an amino acid on the cCPP. s Non-limiting examples of AA include aspartic acid, glutamic acid, glutamine, asparagine, or lysine, or a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). s is AA as defined herein SC It can be said that:
[0332] AA x are each independently a natural amino acid or an unnatural amino acid. x can be a natural amino acid. x can be an unnatural amino acid. x can be a β-amino acid. The β-amino acid can be β-alanine.
[0333] o can be an integer from 0 to 10, for example, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. o can be 0, 1, 2, or 3. o can be 0. o can be 1. o can be 2. o can be 3.
[0334] p can be 0 to 5, for example, 0, 1, 2, 3, 4, or 5. p can be 0. p can be 1. p can be 2. p can be 3. p can be 4. p can be 5.
[0335] The linker has the following structure: [ka] and In the formula, M, AA s , each -(R 1- JR 2 ) z″-, o, and z″ are as defined herein, and r can be 0 or 1.
[0336] r can be 0. r can be 1.
[0337] The linker has the following structure: [ka] and In the formula, M, AA s , o, p, q, r, and z″ can each be as defined herein.
[0338] z" can be an integer from 1 to 50, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50, including all ranges and subranges therebetween. z" can be an integer from 5 to 20. z" can be an integer from 10 to 15.
[0339] The linker can have the following structure: [ka] During the ceremony, M, A.A. s , and o are as defined herein.
[0340] Other non-limiting examples of suitable linkers include: [ka] [ka]
[0341] In the formula, M and AA s is as defined herein.
[0342] A compound comprising a cCPP and an AC complementary to a target in a pre-mRNA sequence further comprising an L, wherein a linker is conjugated to the AC via a linking group (M), and M is [ka] Provided herein are compounds wherein:
[0343] A compound comprising a cCPP and a cargo comprising an antisense compound (AC), e.g., an antisense oligonucleotide, complementary to a target in a pre-mRNA sequence, the compound further comprising L, wherein a linker is conjugated to AC via a linking group (M), wherein M is [ka] wherein R 1 is alkylene, cycloalkyl, or [ka] wherein t' is 0 to 10, each R is independently alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl, and R 1 but, [ka] and t' is 2.
[0344] The linker can have the following structure: [ka] During the ceremony, A.A. s is as defined herein, and m' is 0-10.
[0345] The linker can be of the formula: [ka] .
[0346] The linker has the formula: [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.
[0347] The linker has the formula: [ka] where "base" corresponds to the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.
[0348] The linker has the formula: [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.
[0349] The linker has the formula: [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.
[0350] The linker can be of the formula: [ka] .
[0351] The linker can be covalently attached to the cargo at any suitable position on the cargo. The linker is covalently attached to the 3' end of the cargo or the 5' end of an oligonucleotide cargo. The linker can be covalently attached to the backbone of the cargo.
[0352] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the cCPP, or to a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the cCPP.
[0353] cCPP-linker conjugates The cCPP can be conjugated to a linker as defined herein. The linker can be a linker between the AA SC can be conjugated to
[0354] The linker may be (for example, as a spacer) -(OCH2CH2) z’ -subunits (wherein z' is an integer from 1 to 23, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23). z’ " is also referred to as PEG. The cCPP-linker conjugate can have a structure selected from Table 7. [Table 7]
[0355] The linker is -(OCH2CH2) z’ The cCPP-linker conjugate can have a structure selected from Table 8. [Table 8]
[0356] An EEV is provided that includes a cyclic cell-penetrating peptide (cCPP), a linker, and an exocyclic peptide (EP). The EEV has the structure of formula (B): [ka] , or a protonated form thereof, During the ceremony, R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; R4 and R7 are independently H or an amino acid side chain; EP is an exocyclic peptide as defined herein; each m is independently an integer of 0 to 3; n is an integer from 0 to 2, x' is an integer from 1 to 20, y is an integer from 1 to 5, q is 1 to 4; z' is an integer from 1 to 23.
[0357] R1, R2, R3, R4, R7, EP, m, q, y, x', z' are as described herein.
[0358] n can be 0. n can be 1. n can be 2.
[0359] EEV has the structure of formula (Ba) or formula (Bb): [ka] or protonated forms thereof, wherein EP, R 1 , R 2 , R 3 , R 4 , m, and z' are as defined in formula (B) above.
[0360] EEV has the structure of formula (Bc): [ka] , or protonated forms thereof, wherein EP, R 1 , R 2 , R 3 , R 4 and m are as defined in formula (B) above, AA is an amino acid as defined herein, M is as defined herein, n is an integer of 0 to 2, x is an integer of 1 to 10, y is an integer of 1 to 5, and z is an integer of 1 to 10.
[0361] EEV has a structure of formula (B-1), formula (B-2), formula (B-3), or formula (B-4): [ka] [ka] or a protonated form thereof, where EP is as defined in formula (B) above.
[0362] The EEV can comprise formula (B) and has the following structure: Ac-PKKKRKVAEEA-K(cyclo[FGFGRGRQ])-PEG 12 —OH(Ac-SEQ ID NO:132-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) or Ac-PK-KKR-KV-AEEA-K(cyclo[GfFGrGrQ])-PEG 12 —OH(Ac-SEQ ID NO:133-K(cyclo[SEQ ID NO:83])-PEG 12 -OH).
[0363] The EEV can include a cCPP of the following formula: [ka]
[0364] The EEV can comprise the formula: Ac-PKKKRKV-miniPEG2-Lys(cyclo(FfFGRGRQ)-PEG2-K(N3)(Ac-SEQ ID NO:42-miniPEG2-Lys(cyclo(SEQ ID NO:81)-PEG2-K(N3)).
[0365] The EEV can be: [ka]
[0366] The EEV can be: [ka] .
[0367] EEV was prepared by the procedure described in Example 1. EEV was prepared by the procedure described in Example 1. Ac-PK(Tfa)-K(Tfa)-K(Tfa)-RK(Tfa)-V-miniPEG2-K(cyclo(Ff-Nal-GrGrQ)-PEG 12 —OH(Ac-SEQ ID NO: 134)-miniPEG2-K(cyclo(SEQ ID NO: 135)-PEG 12 -OH).
[0368] The EEV can be: [ka] .
[0369] EEV was synthesized using Ac-PKKKRKV-miniPEG2-K(cyclo(Ff-Nal-GrGrQ)-PEG 12 —OH(Ac-SEQ ID NO: 42)-miniPEG2-K(cyclo(SEQ ID NO: 135)-PEG 12 -OH).
[0370] The EEV can be: [ka]
[0371] The EEV can be: [ka] The EEV can be: [ka]
[0372] The EEV can be: [ka]
[0373] The EEV can be: [ka] .
[0374] The EEV can be: [ka]
[0375] The EEV can be: [ka]
[0376] The EEV can be: [ka]
[0377] The EEV can be: [ka] .
[0378] The EEV can be: [ka]
[0379] The EEV can be: [ka] .
[0380] The EEV can be selected from the following: [Table 27-1] [Table 27-2] [Table 27-3]
[0381] EEV is Ac-PKKKRKV-Lys(cyclo[FfΦGrGrQ])-PEG 12 -K(N3)-NH2 (Ac-SEQ ID NO:42-Lys(cyclo[SEQ ID NO:80])-PEG 12 -K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FfΦGrGrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 80])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRGRQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 82])-miniPEG2-K(N3)-NH2) Ac-KR-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-KR-PEG2-K(cyclo[SEQ ID NO: 82])-PEG2-K(N3)-NH2) Ac-PKKKGKV-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:46-PEG2-K(cyclo[SEQ ID NO:82])-PEG2-K(N3)-NH2) Ac-PKKKRKG-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:48-PEG2-K(cyclo[SEQ ID NO:82])-PEG2-K(N3)-NH2) Ac-KKKRK-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:19-PEG-K(cyclo[SEQ ID NO:82])-PEG-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FFΦGRGRQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 80])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[βhFfΦGrGrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 142])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FfΦSrSrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 143])-miniPEG2-K(N3)-NH2).
[0382] EEV is Ac-PKKKRKV-miniPEG2-Lys(cyclo(GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo(SEQ ID NO: 133))-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFKRKRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 144])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFRGRGQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 145])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 146])-PEG 12-OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRrRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 147])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 84])-PEG 12 -OH), and Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 85])-PEG 12 -OH).
[0383] EEV is Ac-KKKRKG-miniPEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 148-miniPEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-KKKRK-miniPEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 19-miniPEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-KKRKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:22-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) Ac-KRKKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:21-PEG4-K(cyclo[SEQ ID NO:82])-PEG12 -OH) Ac-KKKKR-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:23-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) Ac-RKKKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:20-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH), and Ac-KKKRK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 19-PEG4-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH).
[0384] EEV is Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:42-PEG-K(cyclo[SEQ ID NO:82])-PEG-K(N3)-NH2) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:42-PEG-K(cyclo[SEQ ID NO:133])-PEG-K(N3)-NH2), and Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH).
[0385] The cargo can be AC and the EEV can be Ac-PKKKRKV-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfΦGrGrQ])-PEG12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfF-GRGRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH), wherein b is beta-alanine and the exocyclic sequence can be of D or L stereochemistry.
[0386] cargo A cell membrane penetrating peptide (CPP), such as a cyclic cell membrane penetrating peptide (e.g., cCPP), can be conjugated to a cargo. As used herein, "cargo" refers to a compound or moiety that is desired to be delivered to a cell. The cargo can be conjugated to the terminal carbonyl group of the linker. At least one atom of the cyclic peptide can be replaced by the cargo, or at least one lone pair can form a bond to the cargo. The cargo can be conjugated to the cCPP via a linker. The cargo can be linked to the AA SC At least one atom of the cCPP can be replaced by a cargo, or at least one lone pair of the cCPP forms a bond to a cargo. A hydroxyl group on an amino acid side chain of the cCPP can be replaced by a bond to a cargo. A hydroxyl group on a glutamine side chain of the cCPP can be replaced by a bond to a cargo. The cargo can be conjugated to the cCPP via a linker. The cargo can be linked to an AA SC can be conjugated to
[0387] In some embodiments, the amino acid side chain comprises a chemically reactive group to which a linker or cargo is conjugated. The chemically reactive group can comprise an amine group, a carboxylic acid group, an amide group, a hydroxyl group, a sulfhydryl group, a guanidinyl group, a phenol group, a thioether group, an imidazolyl group, or an indolyl group. In some embodiments, the amino acid of the cCPP to which the cargo is conjugated comprises lysine, arginine, aspartic acid, glutamic acid, asparagine, glutamine, homoglutamine, serine, threonine, tyrosine, cysteine, arginine, tyrosine, methionine, histidine, or tryptophan.
[0388] The cargo can comprise one or more detectable moieties, one or more therapeutic moieties (TM), one or more targeting moieties, or any combination thereof. In some embodiments, the cargo comprises a TM. In some embodiments, the cargo comprises an AC.
[0389] Cyclic cell-penetrating peptides (cCPPs) conjugated to cargo moieties Cyclic cell-penetrating peptides (cCPPs) can be conjugated to cargo moieties.
[0390] The cargo moiety is conjugated to the linker at the terminal carbonyl group to form the following structure: [ka] wherein EP is an exocyclic peptide, M, AA SC , cargo, x', y, and z' are as defined above, and * is AA SC. x' can be 1. y can be 4. z' can be 11. -(OCH2CH-2) x’ -and / or-(OCH2CH-2) z’ can be independently replaced with one or more amino acids, including, for example, glycine, beta-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid, or combinations thereof.
[0391] The endosomal escape vehicle (EEV) can comprise a cyclic cell-penetrating peptide (cCPP), an exocyclic peptide (EP), and a linker, conjugated to a cargo and having the structure of formula (C): [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 can each independently be H or an amino acid residue having a side chain containing an aromatic group; R4 is H or an amino acid side chain; EP is an exocyclic peptide as defined herein; cargo is a moiety as defined herein; each m is independently an integer of 0 to 3; n is an integer from 0 to 2, x' is an integer from 2 to 20, y is an integer from 1 to 5, q is an integer from 1 to 4, z' is an integer of 2 to 20.
[0392] R1, R2, R 3、 R4, EP, cargo, m, n, x', y, q, and z' are as defined herein.
[0393] The EEV can be conjugated to a cargo, the EEV-conjugate having the structure of formula (Ca) or formula (Cb): [ka] or a protonated form thereof, where EP, m, and z are as defined in formula (C) above.
[0394] The EEV can be conjugated to a cargo, the EEV-conjugate having the structure of formula (Cc): [ka] , or protonated forms thereof, wherein EP, R 1 , R 2 , R 3 , R 4 and m are as defined in formula (III) above, AA can be an amino acid as defined herein, n can be an integer from 0 to 2, x can be an integer from 1 to 10, y can be an integer from 1 to 5, and z can be an integer from 1 to 10.
[0395] The EEV can be conjugated to an oligonucleotide cargo, and the EEV-oligonucleotide conjugate can comprise a structure of formula (C-1), formula (C-2), formula (C-3), or formula (C-4). [ka] [ka]
[0396] The EEV can be conjugated to an oligonucleotide cargo, and the EEV-conjugate can comprise the following structure: [ka] .
[0397] Cytoplasmic delivery efficiency Modifications to cyclic cell-penetrating peptides (cCPPs) can improve cytoplasmic delivery efficiency. Improved cytoplasmic uptake efficiency can be measured by comparing the cytoplasmic delivery efficiency of a cCPP having a modified sequence with a control sequence. The control sequence is identical except that it does not contain specific substituted amino acid residues (including, but not limited to, arginine, phenylalanine, and / or glycine) in the modified sequence.
[0398] As used herein, cytoplasmic delivery efficiency refers to the ability of a cCPP to cross the cell membrane and enter the cytoplasm of a cell. The cytoplasmic delivery efficiency of a cCPP does not necessarily depend on the receptor or cell type. Cytoplasmic delivery efficiency can refer to absolute cytoplasmic delivery efficiency or relative cytoplasmic delivery efficiency.
[0399] Absolute cytoplasmic delivery efficiency is the ratio of the cytoplasmic concentration of cCPP (or cCPP-cargo conjugate) to the concentration of cCPP (or cCPP-cargo conjugate) in the growth medium. Relative cytoplasmic delivery efficiency refers to the concentration of cCPP in the cytosol compared to the concentration of a control cCPP in the cytosol. Quantitation can be achieved by fluorescently labeling the cCPP (e.g., with FITC dye) and measuring the fluorescence intensity using techniques well known in the art.
[0400] The relative cytoplasmic delivery efficiency is determined by comparing (i) the amount of the cCPP of the present invention internalized by a cell type (e.g., HeLa cells) with (ii) the amount of a control cCPP internalized by the same cell type. To measure the relative cytoplasmic delivery efficiency, the cell type may be incubated in the presence of the cCPP for a specific period of time (e.g., 30 minutes, 1 hour, 2 hours, etc.), and then the amount of cCPP internalized by the cells is quantified using methods known in the art, such as fluorescence microscopy. Separately, the same concentration of the control cCPP is incubated in the presence of the cell type for the same period of time, and the amount of the control cCPP internalized by the cells is quantified.
[0401] Relative cytoplasmic delivery efficiency is the IC of cCPPs with modified sequences for intracellular targets 50 and measuring the IC of cCPP having the modified sequence. 50 can be determined by comparing with a control sequence (described herein).
[0402] The relative cytoplasmic delivery efficiency of cCPPs compared to cyclo(FfΦRrRrQ, SEQ ID NO: 150) can range from about 50% to about 450%, for example, about 60%, about 70%, about 80%, about 90%, about 100%, about 110%, about 120%, about 130%, about 140%, about 150%, about 160%, about 170%, about 180%, about 190%, about 200%, about 210%, about 220%, about 230%, about 240%, about 250%, about 260%, about 270%, about 280%, about 290%, about 300%, about 310%, about 320%, about 330%, about 340%, about 350%, about 360%, about 370%, about 380%, about 390%, about 400%, about 410%, about 420%, about 430%, about 440%, about 450%, about 460%, about 470%, about 480%, about 490%, about 500%, about 510%, about 520%, about 530%, about 540%, about 550%, about 560%, about 570%, about 580%, about 590%, about 600%, about 610%, about 620%, about 630%, about 640%, about 650%, about 660%, about 670%, about 680%, about 690%, about 700%, about 710%, about 720%, about 730%, about 740%, about 750%, about 760%, about 770%, about 780%, about 790%, about 800%, about 810%, about 8 The relative cytoplasmic delivery efficiency of a cCPP can be improved by more than about 600% compared to a cyclic peptide containing cyclo(FfΦRrRrQ, SEQ ID NO: 150).
[0403] Absolute cytoplasmic delivery efficacy of about 40% to about 100%, for example, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, including all values and subranges therebetween.
[0404] The cCPPs of the present disclosure may increase cytoplasmic delivery efficiency by about 1.1 fold to about 30 fold compared to an otherwise identical sequence, for example, about 1.2, about 1.3, about 1.4, about 1.5, about 1.6, about 1.7, about 1.8, about 1.9, about 2.0, about 2.5, about 3.0, about 3.5, about 4.0, about 4.5, about 5.0, about 5.5, about 6.0, about 6.5, about 7.0, about 7.5, about 8.0, about 8.5, about 9.0, about 10, about 10.5, about 11.0, about 11.5, about 12.0, about 12.5, including all values and subranges therebetween. 5, about 13.0, about 13.5, about 14.0, about 14.5, about 15.0, about 15.5, about 16.0, about 16.5, about 17.0, about 17.5, about 18.0, about 18.5, about 19.0, about 19.5, about 20, about 20.5, about 21.0, about 21.5, about 22.0, about 22.5, about 23.0, about 23.5, about 24.0, about 24.5, about 25.0, about 25.5, about 26.0, about 26.5, about 27.0, about 27.5, about 28.0, about 28.5, about 29.0, or about 29.5 times improvement.
[0405] Detectable Part In some embodiments, the compounds disclosed herein comprise a detectable moiety. In some embodiments, the detectable moiety is attached to the cell membrane-penetrating peptide at the amino group, carboxylate group, or side chain of any of the amino acids of the cell membrane-penetrating peptide moiety (e.g., the amino group, carboxylate group, or side chain of any of the amino acids in the CPP). In some embodiments, the therapeutic moiety comprises a detectable moiety. The detectable moiety can comprise any detectable label. Examples of suitable detectable labels include, but are not limited to, a UV-Vis label, a near-infrared label, a luminescent group, a fluorescent group, a magnetic spin resonance label, a photosensitizer, a photocleavable moiety, a chelating center, a heavy atom, a radioisotope, an isotopically detectable spin resonance label, a paramagnetic moiety, a chromophore, or any combination thereof. In some embodiments, the label is detectable without the addition of additional reagents.
[0406] In embodiments, the detectable moiety is a biocompatible detectable moiety, such that the compound may be suitable for use in a variety of biological applications. As used herein, "biocompatible" and "biologically compatible" generally refer to compounds that, along with any metabolites or breakdown products thereof, are generally non-toxic to cells and tissues and do not cause any significant adverse effects to cells and tissues when they are incubated (e.g., cultured) in their presence.
[0407] The detectable moiety can comprise a luminophore, such as a fluorescent label or a near-infrared label. Examples of suitable luminophores include, but are not limited to, metalloporphyrins, benzoporphyrins, azabenzoporphyrins, naphthoporphyrins, phthalocyanines, polycyclic aromatic hydrocarbons such as perylenedimine and pyrene, azo dyes, xanthene dyes, boron dipyromethenes, aza-boron dipyromethenes, cyanine dyes, metal-ligand complexes such as bipyridines, bipyridyls, phenanthrolines, coumarins, and ruthenium and iridium acetylacetonates, acridines, oxazine derivatives such as benzophenoxazines, aza-annulenes, squaraines, 8-hydroxyquinolines, polymethines, luminescence-generating nanoparticles such as quantum dots and nanocrystals, carbostyrils, terbium complexes, inorganic phosphorus, ionophores such as crown ether-related dyes or derivatized dyes, or combinations thereof. Specific examples of suitable luminophores include Pd(II) octaethylporphyrin; Pt(II)-octaethylporphyrin; Pd(II) tetraphenylporphyrin; Pt(II) tetraphenylporphyrin; Pd(II) meso-tetraphenylporphyrin tetrabenzoporphine; Pt(II) meso-tetraphenylmethylbenzoporphyrin; Pd(II) octaethylporphyrin ketone; Pt(II) octaethylporphyrin ketone; Pd(II) meso-tetra(pentafluorophenyl)porphyrin; Pt(II) meso-tetra(pentafluorophenyl)porphyrin; Ru(II) tris(4,7-diphenyl-1,10-phenanthroline) (Ru(dpp)3); Ru(II) tris(1,10-phenanthroline) (Ru(phen)3 ), tris(2,2'-bipyridine)ruthenium(II) chloride hexahydrate (Ru(bpy)3); erythrosin B; fluorescein; fluorescein isothiocyanate (FITC); eosin; iridium(III) ((N-methyl-benzimidazol-2-yl)-7-(diethylamino)-coumarin)); XXX benzothiazole ((benzothiazol-2-yl)-7-(diethylamino)-coumarin))-2-(acetylacetonate); Lumogen dye; Macroflex fluorescent red;Macrolex Fluorescent Yellow; Texas Red; Rhodamine B; Rhodamine 6G; Sulfur Rhodamine; m-Cresol; Thymol Blue; Xylenol Blue; Cresol Red; Chlorophenol Blue; Bromocresol Green; Bromocresol Red; Bromothymol Blue; Cy2; Cy3; Cy5; Cy5.5; Cy7; 4-Nitrophenol; Alizarin; Phenolphthalein; o-Cresolphthalein; Chlorophenol Red; Calmagite; Bromo-Xylenol; Phenol Red; Neutral Red; Nitrazine; 3,4,5,6-Tetrabromophenolphthalein; Congo Red; Fluor'sc'in; Eosin; 2',7'-Dichlorofluoro Examples of suitable fluorescent dyes include, but are not limited to, fluorescein; 5(6)-carboxy-fluorescein; carboxynaphthofluorescein; 8-hydroxypyrene-1,3,6-trisulfonic acid; semi-naphthorhodafluor; semi-naphthofluorescein; tris(4,7-diphenyl-1,10-phenanthroline)ruthenium(II) dichloride; (4,7-diphenyl-1,10-phenanthroline)ruthenium(II) tetraphenylboron; platinum(II) octaethylporphyrin; dialkylcarbocyanines; dioctadecylcycloxacarbocyanine; fluorenylmethyloxycarbonyl chloride; 7-amino-4-methylcoumarin (Amc); green fluorescent protein (GFP); and derivatives or combinations thereof.
[0408] In some examples, the detectable moiety can include rhodamine B (Rho), fluorescein isothiocyanate (FITC), 7-amino-4-methylcoumarin (Amc), green fluorescent protein (GFP), or derivatives or combinations thereof.
[0409] Production method The compounds described herein can be prepared by various organic synthesis methods known to those skilled in the art or by modifications thereto that will be appreciated by those skilled in the art. The compounds described herein can be prepared from readily available starting materials. Optimum reaction conditions may vary with the particular reactants or solvents used, but such conditions can be determined by one skilled in the art.
[0410] Changes made to the compounds described herein include the addition, subtraction, or movement of various components described for each compound. Similarly, if one or more chiral centers are present in a molecule, the chirality of the molecule can be altered. Furthermore, the synthesis of a compound can involve the protection and deprotection of various chemical groups. The use of protection and deprotection, and the selection of appropriate protecting groups, can be determined by those skilled in the art. The chemistry of protecting groups can be found, for example, in Wuts and Greene, Protective Groups in Organic Synthesis, 4th Ed., Wiley & Sons, 2006, which is incorporated herein by reference in its entirety.
[0411] Starting materials and reagents used in preparing the disclosed compounds and compositions may be purchased from Aldrich Chemical Co. (Milwaukee, WI), Acros Organics (Morris Plains, NJ), Fisher Scientific (Pittsburgh, PA), Sigma (St. Louis, MO), Pfizer (New York, NY), GlaxoSmithKline (Raleigh, NC), Merck (Whitehouse Station, NJ), Johnson & Johnson (New Brunswick, NJ), Aventis (Bridgewater, NJ), AstraZeneca (Wilmington, DE), Novartis (Basel, Switzerland), Wyeth (Madison, NJ), Bristol-Myers-Squibb (New York, NY), Roche (Basel, Switzerland), Lilly (Indianapolis, IN), Abbott (Abbott Park, IL), Schering They are either available from commercial suppliers such as Plough (Kenilworth, NJ), or Boehringer Ingelheim (Ingelheim, Germany), or prepared by methods known to those skilled in the art following procedures described in references such as Fieser and Fieser's Reagents for Organic Synthesis, Volumes 1-17 (John Wiley and Sons, 1991), Rodd's Chemistry of Carbon Compounds, Volumes 1-5 and Supplementals (Elsevier Science Publishers, 1989), Organic Reactions, Volumes 1-40 (John Wiley and Sons, 1991), March's Advanced Organic Chemistry, (John Wiley and Sons, 4th Edition), and Larock's Comprehensive Organic Transformations (VCH Publishers Inc., 1989).Other materials, such as the pharmaceutical carriers disclosed herein, can be obtained from commercial sources.
[0412] The reactions to produce the compounds described herein can be carried out in a solvent that can be selected by one skilled in the art of organic synthesis. The solvent can be substantially non-reactive with the starting materials (reactants), intermediates, or products under the conditions, i.e., temperature and pressure, at which the reaction is carried out. The reaction can be carried out in one solvent or a mixture of two or more solvents. The formation of the product or intermediate can be monitored according to any suitable method known in the art. For example, the formation of the product can be monitored by nuclear magnetic resonance spectroscopy (e.g., 1 H or 13 C), can be monitored by spectroscopic means such as infrared spectroscopy, spectrophotometry (e.g., UV-visible), or mass spectrometry, or by chromatography such as high performance liquid chromatography (HPLC) or thin layer chromatography.
[0413] The disclosed compounds can be prepared by solid-phase peptide synthesis, in which the α-N-terminal amino acid is protected with an acid or base protecting group. Such protecting groups should have the properties of being stable to the conditions of peptide bond formation while being easily removable without disrupting the growing peptide chain or racemization of any of the chiral centers contained therein. Suitable protecting groups include 9-fluorenylmethyloxycarbonyl (Fmoc), t-butyloxycarbonyl (Boc), benzyloxycarbonyl (Cbz), biphenylisopropyloxycarbonyl, t-amyloxycarbonyl, isobornyloxycarbonyl, α,α-dimethyl-3,5-dimethoxybenzyloxycarbonyl, o-nitrophenylsulfenyl, 2-cyano-t-butyloxycarbonyl, and the like. The 9-fluorenylmethyloxycarbonyl (Fmoc) protecting group is particularly preferred for synthesizing the disclosed compounds. Other preferred side chain protecting groups are 2,2,5,7,8-pentamethylchroman-6-sulfonyl (pmc), nitro, p-toluenesulfonyl, 4-methoxybenzenesulfonyl, Cbz, Boc, and adamantyloxycarbonyl for side chain amino groups such as lysine and arginine; benzyl, o-bromobenzyloxycarbonyl, 2,6-dichlorobenzyl, isopropyl, t-butyl (t-Bu), cyclohexyl, cyclopenyl, and acetyl (Ac) for tyrosine; t-butyl, benzyl, and tetrahydropyranyl for serine; trityl, benzyl, Cbz, p-toluenesulfonyl, and 2,4-dinitrophenyl for histidine; formyl for tryptophan; benzyl and t-butyl for aspartic acid and glutamic acid; and triphenylmethyl (trityl) for cysteine.
[0414] In solid-phase peptide synthesis, the α-C-terminal amino acid is attached to a suitable solid support or resin. Suitable solid supports useful for the above synthesis are those materials that are inert to the reagents and reaction conditions of the stepwise condensation-deprotection reactions and are insoluble in the media used. Solid supports for the synthesis of α-C-terminal carboxypeptides are 4-hydroxymethylphenoxymethyl-copoly(styrene-1% divinylbenzene) or 4-(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxyacetamidoethyl resin, available from Applied Biosystems (Foster City, Calif.). The α-C-terminal amino acid is coupled to the resin by N,N'-dicyclohexylcarbodiimide (DCC), N,N'-diisopropylcarbodiimide (DIC), or O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU) with or without 4-dimethylaminopyridine (DMAP), 1-hydroxybenzotriazole (HOBT), benzotriazol-1-yloxy-tris(dimethylamino)phosphonium hexafluorophosphate (BOP), or bis(2-oxo-3-oxazolidinyl)phosphine chloride (BOPCl) mediated coupling in a solvent such as dichloromethane or DMF at a temperature of 10°C to 50°C for about 1 to about 24 hours. When the solid support is a 4-(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxy-acetamidoethyl resin, the Fmoc group is cleaved with a secondary amine, preferably piperidine, before coupling to the α-C-terminal amino acid described above. One method for coupling to the deprotected 4(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxy-acetamidoethyl resin is with O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU, 1 equivalent) and 1-hydroxybenzotriazole (HOBT, 1 equivalent) in DMF. The coupling of successive protected amino acids can be carried out in an automated polypeptide synthesizer. In one example, the α-N-terminus of the amino acid in the growing peptide chain is protected with Fmoc.Removal of the Fmoc protecting group from the α-N-terminus of the growing peptide is accomplished by treatment with a secondary amine, preferably piperidine. The protected amino acids are then introduced in approximately a three-fold molar excess, and coupling is preferably carried out in DMF. The coupling agents can be O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU, 1 equivalent) and 1-hydroxybenzotriazole (HOBT, 1 equivalent). At the end of solid-phase synthesis, the polypeptide is removed from the resin and deprotected, either sequentially or in a single operation. Removal and deprotection of the polypeptide can be accomplished in a single operation by treating the resin-bound polypeptide with a cleavage reagent containing thianisole, water, ethanedithiol, and trifluoroacetic acid. If the α-C-terminus of the polypeptide is an alkylamide, the resin is cleaved by aminolysis with an alkylamine. Alternatively, the peptide can be removed by transesterification, e.g., with methanol, followed by aminolysis or direct transamidation. The protected peptide can be purified at this point or can proceed directly to the next step. Removal of side chain protecting groups can be achieved using the cleavage cocktail described above. The fully deprotected peptide can be purified by a series of chromatographic steps using any or all of the following types: ion exchange on a weakly basic resin (acetate form), hydrophobic adsorption chromatography on undifferentiated polystyrene-divinylbenzene (e.g., Amberlite XAD), silica gel adsorption chromatography, ion exchange chromatography on carboxymethylcellulose, partition chromatography, e.g., Sephadex G-25, LH-20, or partition chromatography on countercurrent distribution, high performance liquid chromatography (HPLC), particularly reverse-phase HPLC on octyl-silica or octadecylsilyl-silica bonded phase columns.
[0415] The above-mentioned polymers, such as PEG groups, can be attached to oligonucleotides, such as AC, under any suitable conditions. Any means known in the art can be used, including acylation, reductive alkylation, Michael addition, thiol alkylation, or other chemoselective conjugation / ligation methods via reactive groups on the PEG moiety (e.g., aldehyde, amino, ester, thiol, α-haloacetyl, maleimide, or hydrazino groups) to reactive groups on AC (e.g., aldehyde, amino, ester, thiol, α-haloacetyl, maleimide, or hydrazino groups). Activated groups that can be used to link water-soluble polymers to one or more proteins include, but are not limited to, sulfone, maleimide, sulfhydryl, thiol, triflate, tresylate, azidiline, oxirane, 5-pyridyl, and alpha-halogenated acyl groups (e.g., α-iodoacetic acid, α-bromoacetic acid, α-chloroacetic acid). When conjugated to AC by reductive alkylation, the selected polymer should have a single reactive aldehyde so that the degree of polymerization can be controlled. See, for example, Kinstler et al., Adv. Drug. Delivery Rev. (2002), 54:477-485; Roberts et al., Adv. Drug Delivery Rev. (2002), 54:459-476; and Zalipsky et al., Adv. Drug Delivery Rev. (1995), 16:157-182.
[0416] To direct the covalent attachment of the AC or linker to the CPP, appropriate amino acid residues of the CPP may be reacted with an organic derivatizing agent capable of reacting with selected side chains of amino acids or the N- or C-terminus. Reactive groups on the peptide or conjugate moiety include, for example, aldehyde, amino, ester, thiol, α-haloacetyl, maleimide, or hydrazino groups. Derivatizing agents include, for example, maleimidobenzoyl sulfosuccinimide ester (conjugation via cysteine residues), N-hydroxysuccinimide (via lysine residues), glutaraldehyde, succinic anhydride, or other agents known in the art.
[0417] Methods for making ACs and conjugating ACs to linear CPPs are generally described in U.S. Publication No. 2018 / 0298383, which is incorporated herein by reference for all purposes. These methods can be applied to the cyclic CPPs disclosed herein.
[0418] The synthetic scheme is provided in FIGS. 3A-3D and 4.
[0419] Non-limiting examples of compounds containing a CPP and a reactive group useful for conjugation to an AC are shown in Table 9. Exemplary linker groups are also shown. Examples of reactive groups include tetrafluorophenyl ester (TFP), free carboxylic acid (COOH), and azide (N3). In Table 9, n is an integer between 0 and 20, Pipa6 is AcRXRRBRRXRYQFLIRXRBRXRB (where B is β-alanine and X is aminohexanoic acid), Dap is 2,3-diaminopropionic acid, NLS is a nuclear localization sequence, βA is beta-alanine, -ss- is a disulfide, PABC is poly(A) binding protein C-terminal domain, and C x (where x is a number) is an alkyl chain of length x, and BCN is bicyclo[6.1.0]nonyne. [Table 9-1] [Table 9-2]
[0420] In embodiments, the CPP has a free carboxylic acid group available for conjugation to an AC. In embodiments, the EEV has a free carboxylic acid group available for conjugation to an AC.
[0421] The following structure: [ka] is a 3'cyclooctyne-modified PMO used in click reactions with azide-containing compounds.
[0422] An exemplary scheme for conjugation of a CPP and a linker to the 3' end of an AC via an amide bond is shown below. [ka]
[0423] An exemplary scheme for the conjugation of a CPP and linker to a 3'-cyclooctyne-modified PMO via strain-promoted azide-alkyne cycloaddition is shown below. [ka]
[0424] Examples of conjugation chemistries used to attach ACs and CPPs to additional linkers containing polyethylene glycol moieties are shown below. [ka]
[0425] An example of the conjugation of a CPP linker to a 5′-cyclooctyne-modified PMO via strain-promoted azide-alkyne cycloaddition (click chemistry) is shown below. [ka]
[0426] Methods for synthesizing oligomeric antisense compounds are known in the art. The present disclosure is not limited by the method for synthesizing the AC. In several embodiments, compounds having reactive phosphorus groups useful for forming internucleoside linkages, including, for example, phosphodiester and phosphorothioate internucleoside linkages, are provided herein. The methods for preparing and / or purifying precursors or antisense compounds are not intended to limit the compositions or methods provided herein. Methods for synthesizing and purifying DNA, RNA, and antisense compounds are well known to those skilled in the art.
[0427] Oligomerization of modified and unmodified nucleosides can be routinely carried out according to reference procedures for DNA (Protocols for Oligonucleotides and Analogs, Ed. Agrawal (1993), Humana Press) and / or RNA (Scaringe, Methods (2001), 23, 206-217; Gait et al., Applications of Chemically Synthesized RNA in RNA: Protein Interactions, Ed. Smith (1998), 1-36; Gallo et al., Tetrahedron (2001), 57, 5707-5713).
[0428] The antisense compounds provided herein can be conveniently and routinely produced by the well-known technique of solid-phase synthesis. Equipment for such synthesis is sold by several vendors, including, for example, Applied Biosystems (Foster City, CA). Any other means for such synthesis known in the art can additionally or alternatively be used. It is well known to use similar techniques to prepare oligonucleotides such as phosphorothioates and alkylated derivatives. The present invention is not limited by the method of synthesis of antisense compounds.
[0429] Methods for purifying and analyzing oligonucleotides are known to those skilled in the art. Analytical methods include capillary electrophoresis (CE) and electrospray mass spectrometry. Such synthesis and analysis methods can be performed in multi-well plates. The method of the present invention is not limited by the method for purifying the oligomers.
[0430] In the compounds disclosed herein, the AC is coupled to a CPP (e.g., a cyclic peptide). As used herein, "coupled" can refer to a covalent or non-covalent association between the CPP and the AC, including fusion of the CPP to the AC and chemical conjugation of the CPP (e.g., a cyclic peptide) to the AC. A non-limiting example of a means of non-covalently linking a CPP to an AC is via a streptavidin / biotin interaction, for example, conjugating biotin to the CPP and fusing the AC to streptavidin. In the resulting compound, the CPP is coupled to the AC via a non-covalent association between biotin and streptavidin.
[0431] In some embodiments, a CPP (e.g., a cyclic peptide) is directly or indirectly conjugated to an AC, thereby forming a CPP-AC conjugate. Conjugation of the AC to the CPP can occur at any suitable site on these moieties. For example, in some embodiments, the 5' or 3' end of the AC can be conjugated to the C-terminus, N-terminus, or side chain of an amino acid in the CPP.
[0432] In some embodiments, the AC is covalently linked to the CPP (e.g., a cyclic peptide). As used herein, covalent linkage refers to a construct in which the CPP moiety is covalently linked to the 5' and / or 3' end of the AC moiety. Alternatively, such a conjugate may be said to have a CPP moiety (e.g., a cyclic peptide moiety) and an oligonucleotide moiety. Covalently linked AC-CPP or CPP-AC conjugates according to certain embodiments comprise an AC moiety and a CPP moiety associated with each other by a linker as described herein.
[0433] In embodiments, the AC can be conjugated to the CPP (e.g., a cyclic peptide) via the side chain of an amino acid on the CPP. Any amino acid side chain on the CPP that can form a covalent bond or can be modified to do so can be used to link the AC to the CPP. The amino acid on the CPP can be a natural amino acid or an unnatural amino acid. In embodiments, the amino acid on the CPP used to conjugate the AC is aspartic acid, glutamic acid, glutamine, asparagine, lysine, ornithine, 2,3-diaminopropionic acid, or an analog thereof, where the side chain is replaced with a bond to the AC or a linker. In embodiments, the amino acid is lysine or an analog thereof. In embodiments, the amino acid is glutamic acid or an analog thereof. In embodiments, the amino acid is aspartic acid or an analog thereof.
[0434] In some embodiments, the CPP is cyclic. There are many possible configurations of the compounds disclosed herein. In some embodiments, the compounds disclosed herein include compounds in which AC is conjugated to the side chain of an amino acid in a cyclic peptide. In some embodiments, the compounds disclosed herein have a structure according to formula IA (i.e., exocyclic), [ka] , wherein the linker is covalently attached to the side chain of an amino acid on the CPP and to the 5' end of AC, the backbone of AC, or the 3' end of AC.
[0435] Diseases and target genes In embodiments, compounds and methods are provided for treating a disease or disorder associated with one or more genes having an expanded nucleotide repeat (e.g., an expanded trinucleotide repeat, such as an extended trinucleotide repeat). In embodiments, compounds and methods are provided for treating a disease or disorder associated with one or more genes having an expanded CTG·CUG trinucleotide repeat. In embodiments, compounds and methods are provided for treating a disease or disorder associated with one or more genes having an expanded CTG·CUG trinucleotide repeat in the 3′-UTR of the gene. In embodiments, compounds and methods are provided for treating a disease or disorder associated with a gene having an expanded CTG·CUG trinucleotide repeat in the 3′-UTR, such as DMPK, ATXN8OS, and / or JPH3. In embodiments, compounds and methods are provided for treating a disease or disorder associated with one or more genes having an expanded CTG·CUG trinucleotide repeat in an intron of the gene. In embodiments, compounds and methods are provided for treating a disease or disorder associated with an expanded CTG·CUG trinucleotide repeat in the intron of TCF4. In several embodiments, compounds and methods are provided for treating myotonic dystrophy type 1 (DM1), Fuchs endothelial corneal dystrophy (FECD), spinocerebellar ataxia-8 (SCA8), and / or Huntington's disease-like (HDL2).
[0436] Myotonic dystrophy type 1 (DM1) In several embodiments, compounds, compositions, and methods for treating myotonic dystrophy (DM1 or Steinert's disease) are provided. DM1 is a multisystem disorder often characterized by muscle degeneration and delayed muscle stiffness or relaxation due to repetitive activity in muscle fibers. Myotonic dystrophy type 1 (DM1) is the most common form of muscular dystrophy, affecting 1 in 8,000 people. DM1 is a paradigm genetic disorder caused by a CTG·CUG expansion. DM1 is a neuromuscular disorder caused by a CTG·CUG repeat expansion in the 3'-untranslated region (UTR) of the myotonic dystrophy protein kinase (DMPK) gene. At the RNA level, DMPK transcripts (e.g., expanded CUG repeats) capture splicing regulator proteins, such as muscleblind-like (MBNL) protein, leading to the incorrect splicing of some downstream pre-mRNAs (pre-mRNAs that do not contain the expanded CUG repeats) regulated by MBNL1. This gain of function is responsible for DM1.
[0437] The excess number of CUG repeats confers toxic activity, termed gain of toxicity, leading to the misprocessing of multiple critical proteins, which contributes to the multisystemic nature of the disease, including generalized limb weakness, respiratory muscle disorders, cardiac abnormalities, fatigue, gastrointestinal complications, cataracts, incontinence, and excessive daytime sleepiness.
[0438] DM1 patients with a CTG-CUG expansion within the 3'-untranslated region of the DMPK gene are at increased risk for FECD, forming CUGexp-MBNL1 foci in the corneal endothelium (Mootha et al., Investigative Ophthalmology & Visual Science, 2017;58,4579-4585). Association of MBNL1 with mutant RNA affects the cellular pool of free MBNL1 and induces missplicing of several MBNL1 target genes (e.g., regulated by MBNL1) in affected brain, muscle, and heart tissue (Jiang et al., Hum Mol Genet. 2004;13:3079-3088). Gattey et al. (Cornea. 2014;33:96-98) reported FECD in four DM1 subjects, including a mother-daughter pair. Therefore, an association between DM1 and FECD is likely (FECD is described in more detail elsewhere herein).
[0439] Without wishing to be bound by theory, there are at least two hypotheses proposed to explain the pathogenesis of DM1. One is that the expanded CTG-CUG repeat inhibits DMPK mRNA or protein production, resulting in DMPK haploinsufficiency. This is supported by studies showing decreased expression of DMPK mRNA and protein in DM1 muscle (Fu, Y., et al. (1993) Decreased expression of myotonin-protein kinase messenger RNA and protein in adult forms of myotonic dystrophy. Science 260, 235-238). In some embodiments, the compounds and methods described herein ameliorate DMPK haploinsufficiency. Another RNA gain-of-function hypothesis proposes that mutant RNA transcribed from the expanded allele is sufficient to induce disease symptoms. This was suggested by the following observations: (i) expanded CTG repeats are transcribed into CUG repeats that accumulate in discrete nuclear foci, and (ii) expression of only the DMPK 3'-UTR, which contains 200 CTG repeats, is sufficient to inhibit myogenesis (Davis, BM, et al. (1997) Expansion of a CUG trinucleotide repeat in the 31 untranslated region of myotonic dystrophy protein kinase transcripts results in nuclear retention of transcripts. Proc. Natl. Acad. Sci. USA 94, 7388-7393; Amack, JD et al., (1999) Cis and trans effects of the myotonic dystrophy (DM) mutation in a cell culture model. Hum. Mol. Genet. 8, 1975-1984). In some embodiments, the compounds and methods described herein reduce transcription of mutant RNA associated with the expanded allele.
[0440] An expanded CTG-CUG trinucleotide repeat within the 3' untranslated region of DMPK mRNA forms imperfect, stable hairpin structures that accumulate in the nucleus in small ribonucleocomplexes or microscopically visible inclusion bodies, impairing the function of proteins involved in transcription, splicing, or RNA export. Although the DMPK gene containing the CUG repeat is transcribed into mRNA, mutant transcripts are trapped in the nucleus as aggregates (foci), resulting in reduced cytoplasmic DMPK mRNA levels. These aggregates lead to deregulation of alternative splicing of many different transcripts due to the trapping of two RNA-binding proteins: MBNL1 (myoblind-like 1) and CUGBP1 (CUG-binding protein 1), resulting in loss of MBNL1 function and upregulation of CUGBP1 (Lee and Cooper. (2009) "Pathogenic mechanisms of myotonic dystrophy," Biochem Soc Trans. 37(06):1281-1286).
[0441] In DM1, the RNA-binding protein MBNL1 is trapped in a double-stranded hairpin structure formed by CUG repeats, causing its depletion from the nucleoplasm. The MBNL1-bound CUG repeats then stimulate protein kinase C (PKC) activation through an unknown mechanism, leading to CUGBP1 hyperphosphorylation and stabilization. Downstream effects include disruption of alternative splicing of downstream genes, mRNA translation, and mRNA decay. A key molecular hallmark of DM1 is misregulation of alternative splicing due to MBNL1 trapping in the CUG repeats with the double-stranded hairpin structure. Among the more than 24 splicing events misregulated in DM1, aberrant splicing of the skeletal muscle-specific ClC-1 (chloride channel 1) is known to be one of the causes of myotonia. Increased inclusion of exons containing premature stop codons leads to downregulation of ClC-1 mRNA and protein sufficient to cause myotonia (Charlet-B et al. (2002) Loss of the muscle-specific chloride channel in type 1 myotonic dystrophy due to misregulated alternative splicing. Mol. Cell 10, 45-53; Mankodi, A. et al. (2002) Expanded CUG repeats trigger aberrant splicing of ClC-1 chloride channel pre-mRNA and hyperexcitability of skeletal muscle in myotonic dystrophy. Mol. Cell 10, 35-44). In embodiments, the compounds and methods described herein improve downstream effects, including alternative splicing of downstream genes, mRNA translation, and mRNA decay. In embodiments, the compounds and methods described herein reduce the number of misregulated splicing events in DM1 compared to subjects with DM1 who are not treated with the compounds or methods disclosed herein.For example, in some embodiments, the compounds and methods described herein may be used in combination with 4833439L19Rik, Abcc9, Atp2a1, Arhgef10, Arhgap28, Armcx6, Angel1, Best3, Bin1, Brd2, Cacna1s, Cacna2d1, Cpd, Cpeb3, Ccpg1, Clasp1, Clcn1, Clk4, Cpeb2, Camk2g, Capzb, Copz2, Coch, cT NT, Ctu2, Cyp2s1, Dctn4, Dnm1l, Eya4, Efna3, Efna2, Fbxo31, Fbxo21, Frem2, Fgd4, Fuca1, Fn1, Gogla4, Gpr37l1, G reb1, Heg1, Insr, Impdh2, IR, Itgav, Jag2, Klc1, Kcan6, Kif13a, Ldb3, Lrrfip2, Mapt, Macf1, Map3k4, Mapkap1, Mb nl1, Mllt3, Mbnl2, Mef2c, Mpdz, Mrpl1, Mxra7, Mybpc1, Myo9a, Ncapd3, Ngfr, Ndrg3, Ndufv3, Neb, Nfix, Numa1, Opa 1, Pacsin2, Pcolce, Pdlim3, Pla2g15, Phactr4, Phka1, Phtf2, Ppp1r12b, Ppp3cc, Ppp1cc, Ramp2, Rapgef1, Rur1, R The compounds and methods described herein may reduce the number of misregulated splicing events in one or more downstream genes, such as yr1, Sorcs2, Spsb4, Scube2, Sema6c, Sfc8a3, Slain2, Sorbsl, Spag9, Tmem28, Tacc1, Tacc2, Ttc7, Tnik, Tnfrsf22, Tnfrsf25, Trappc9, Trim55, Ttn, Txnl4a, Txlnb, Ube2d3, or Vsp39. In embodiments, the compounds and methods described herein reduce the number of exons containing premature stop codons that result in downregulation of ClC-1 mRNA, compared to subjects with DM1 who are not treated with the disclosed compounds or methods.
[0442] Nuclear levels of MBNL1 and CUGBP1 control a subset of developmentally regulated splicing events that are reversed in DM1. During embryonic development, MBNL1 nuclear levels are low and CUGBP1 levels are high. During development, MBNL1 nuclear levels increase while CUGBP1 levels decrease, inducing the embryonic-to-adult transition of downstream splicing targets (including IR exon 11, the ClC-1 exon containing the stop codon, and cTNT exon 5). However, in DM1, MBNL1 becomes trapped by CUG repeats, resulting in a decrease in functional MBNL1 while CUGBP1 levels increase due to phosphorylation and stabilization. This simulates an embryonic state and enhances the expression of embryonic isoforms in adults, resulting in multiple disease symptoms (Lee and Cooper, 2009). In several embodiments, the compounds and methods described herein reduce the amount of trapped MBNL1, increase the amount of functional MBNL1, and reduce CUGBP1 levels compared to subjects with DM1 who are not treated with the compounds or methods disclosed herein.
[0443] MBNL1 and CELF1 (also referred to as "CUGBP1") are developmental regulators of splicing events during the fetal-to-adult transition, and modification of their activity in DM1 results in the expression of fetal splicing patterns in adult tissues. Downstream effects of low MBNL1 and high CELF1 include disruption of alternative splicing, mRNA translation, and mRNA decay in proteins such as cardiac troponin T (cTNT), insulin receptor (INSR), muscle-specific chloride channel (CLCN1), and sarcoplasmic / endoplasmic reticulum calcium ATPase 1 (ATP2A1) transcripts, in addition to MBNL1. Konieczny et al. (2017) "Myotonic dystrophy: candidate small molecule therapeutics," Drug Discovery Today. 22(11):1740-1748.
[0444] Compounds and methods for treating myotonic dystrophy using antisense oligomers targeting polyCUG repeats within the 3'-UTR of the DMPK gene are described in US10106796B2, US10111962B2, and US20150080311A1, each of which is incorporated by reference in its entirety for all purposes. However, such PMOs or PPMOs targeting CUG repeats for treating DM1 may have limitations in oligo delivery to the affected tissue, muscle.
[0445] In embodiments, the present disclosure teaches the use of various cell-penetrating peptides (CPPs) to deliver ACs (e.g., PMOs or ASOs) and degradation sequences described herein, e.g., in Tables 2 and 10, to the cytosol of a cell. In embodiments, a CPP or EEV conjugated to an AC delivers the AC of interest to the cellular location where the target sequence on the pre-mRNA is located.
[0446] In some embodiments, the disease is a form of myotonic dystrophy (e.g., myotonic dystrophy type 1 or myotonic dystrophy type 2). In some embodiments, the target gene is the DMPK gene, which encodes myotonic protein kinase. In some embodiments, the compounds provided herein comprise an AC (e.g., an ASO) that targets DMPK (e.g., the 3'-untranslated region / polyadenylation of the DMPK gene) to degrade the DMPK gene. Exemplary oligonucleotides that target DMPK for degradation are provided in Table 10. Degradation sequences may be used in combination with AC sequences containing 10 to 40 CAG repeats, including, but not limited to, the ACs provided in Table 2. [Table 10]
[0447] Spinocerebellar ataxia-8 (SCA8) In several embodiments, compounds, compositions, and methods are provided for treating spinocerebellar ataxia-8 (SCA8). SCA8 is an inherited neurodegenerative condition characterized by slowly progressive ataxia. Symptoms usually appear between the ages of 30 and 50. Symptoms include eye movement abnormalities, sensory neuropathy, dysphagia, cerebellar ataxia, and cognitive impairment.
[0448] SCA8 is associated with heterozygous expanded CTG·CUG repeats within the 3'UTR of two overlapping genes, ATXN8OS and ATXN8. Healthy individuals typically have 15-50 CTG·CUG repeats in the ATXN8OS and ATXN8 genes. Patients with SCA8 have more than 50 CTG·CUG repeats in the ATXN8OS and ATXN8 genes and may have as many as 240 CTG·CUG repeats.
[0449] Huntington's disease-like-2 (HDL2) In several embodiments, compounds, compositions, and methods are provided for treating Huntington's disease-like 2 (HDL2) disease. HDL2 is an autosomal dominant neurodegenerative disorder phenotypically related to Huntington's disease. HDL2 is characterized by symptoms including chorea, dystonia, rigidity, bradykinesia, and psychiatric symptoms such as dementia. HDL2 symptoms typically occur in middle age and can lead to premature death by approximately 10 to 15 years.
[0450] HDL2 is associated with an expansion of CTG·CUG within the 3'UTR of the junctophilin 3 (JPH3) gene (16q24.3). Healthy individuals typically have 6 to 27 CTG·CUG repeats in the JPH3 gene. Patients with HDL2 have more than 40 CTG·CUG repeats in the JPH3 gene, and may have as many as 60 or more CTG·CUG repeats.
[0451] Fuchs endothelial corneal dystrophy In some embodiments, compounds, compositions, and methods for treating Fuchs' endothelial corneal dystrophy (FECD) are provided. FECD (MIM 136800) is an age-related degenerative disorder of the corneal endothelium. FECD is characterized by the progressive loss of corneal endothelial cells, thickening of Descemet's membrane, and deposition of extracellular matrix in droplet-like forms. When the number of endothelial cells becomes too low, the cornea swells, causing blindness (Elhalis et al. Ocul Surf. 2010;8(4):173-184).
[0452] FECD may be inherited as an autosomal dominant trait with genetic heterogeneity. Rare heterozygous mutations in the collagen type VIII alpha 2 gene (COL8A2, MIM 120252) can cause early-onset corneal endothelial dystrophy. Other genes, such as solute carrier family 4, sodium borate transporter, member 11 (SLC4A11, MIM 610206), transcription factor 8 (TCF8, MIM 189909), lipoxygenase homology domain 1 (LOXHD1, MIM 613267), and ATP / GTP-binding protein-like 1 (AGBL1, MIM 615523), are collectively associated with a small percentage of adult-onset FECD cases. Genome-wide association studies of adult-onset FECD have identified transcription factor 4 (TCF4, MIM 602272), and more recently KN motif and ankyrin repeat domain-containing protein 4 (KANK4, MIM 614612), laminin gamma-1 (LAMC1, MIM150290), Na + / K + It has been suggested that transport ATPase, as well as beta-1 polypeptide (ATP1B1, MIM 182330), together with the TCF4 locus, have a dominant influence on FECD (Mootha et al., Investigative ophthalmology & visual science, 2017;58,4579-4585).
[0453] An expanded trinucleotide repeat at the CTG18.1 locus in intron 2 of TCF4 has been associated with FECD (Wieben et al., PLoS One. 2012;7(11):e49083). Each copy of the CTG18.1 allele, which contains more than 40 CTG-CUG trinucleotide repeats, confers a significant risk for developing FECD (Mootha et al., Invest Ophthalmol Vis Sci. 2014;55:33-42). RNA nuclear foci, a hallmark of toxic gain-of-function RNAs, have been reported in neurodegenerative disorders caused by simple repeat expansions. Expanded CUG repeat RNA accumulates as nuclear foci in the corneal endothelium of FECD subjects with a CTG18.1 triplet repeat expansion, but is absent in control samples lacking the triplet expansion (Mootha et al., Invest Ophthalmol Vis Sci. 2015;56(3):2003-2011). Expanded CUG repeat RNA colocalizes with the mRNA splicing factor myoblind-like 1 (MBNL1) in nuclear foci in endothelial tissue as a molecular signature. Triplet repeat expansions in the CTG18.1 locus may mediate endothelial dysfunction through aberrant gene splicing as a result of mutant CUG RNA transcripts that trap MBNL1 (Du et al., J Biol Chem. 2015;290:5979-5990). Thus, two distinct triplet repeats converge on the RNA lesion and FECD, and the lesion may play a causal role in FECD.
[0454] In some embodiments, compounds and methods useful for treating Fuchs' endothelial corneal dystrophy (FECD) using antisense oligonucleotides to reduce expanded CUG repeat RNA are described in WO2018165541A1 and US10760076B2, each of which is incorporated herein by reference in its entirety for all purposes. However, such phosphorodiamidate morpholino oligomers (PMOs) or peptide-conjugated PMOs (PPMOs) targeting CUG repeats, such as those described in the prior art for treating DM1, may have limitations in oligo delivery to affected target tissues (e.g., the corneal endothelium).
[0455] In embodiments, the present disclosure teaches the use of various cell-penetrating peptides (CPPs) or endosomal escape vehicles (EEVs) to deliver an AC (e.g., a PMO or ASO) described herein, e.g., in Table 6, to the cytosol of a cell. In embodiments, the CPP or EEV conjugated to the AC delivers the AC of interest to the cellular location where the target sequence on the pre-mRNA is located.
[0456] In some embodiments, the disease is Fuchs' endothelial corneal dystrophy (FECD). In some embodiments, the target gene is TCF4, which encodes transcription factor 4 (TCF-4), also known as immunoglobulin transcription factor 2 (ITF-2). In some embodiments, the compounds provided herein include antisense oligonucleotides that target TCF4. Exemplary oligonucleotides that can be used to target TCF4 are provided in Tables 2 and 11. [Table 11]
[0457] Mootha et al. (2017) reported that DM1 and FECD are not identical diseases, but that they result from a non-coding CTG expansion. DMPK expansion in DM1 results in a multisystem disease involving various tissues of the eye, including the lens, retina, and corneal endothelium. In contrast, TCF4 repeat expansion appears to affect the corneal endothelium without causing clinically apparent sequelae in other ocular tissues or body organs. Mutational expansions in DMPK and TCF4 share important similarities, including (i) nuclear foci containing the expanded CUG repeat, (ii) association of the foci with the MBNL1 protein, and (iii) the ability to cause FECD. It has been suggested that triplet expansions in both DMPK and TCF4 may cause the same corneal endothelial tissue phenotype of FECD through a shared molecular mechanism.
[0458] See U.S. Patent No. 10,760,076 B2, International Application Publication No. 2018165541 A1, U.S. Patent Application Publication No. 2016 / 0355796, and U.S. Patent Application Publication No. 2018 / 0344817, each of which is incorporated herein by reference and discloses that diseases and corresponding genes tend to form and / or expand tandem nucleotide repeats.
[0459] Compositions and Methods of Administration The compounds of the present disclosure can be formulated into compositions suitable for in vivo use, and the compounds and / or compositions can be administered to patients having or suspected of having a disease associated with an e...
Claims
1. An endosomal escape vehicle comprising a cyclic peptide and an extra-cyclic peptide, wherein the cyclic peptide comprises 6 to 12 amino acids and the extra-cyclic peptide comprises 2 to 10 amino acids, and an endosomal escape vehicle, An antisense compound (AC) complementary to at least a portion of an extended CUG repeat in a target mRNA sequence, wherein the AC comprises a phosphorodiamidate morpholino (PMO) oligonucleotide, and an antisense compound A compound comprising, wherein the cyclic peptide has the following structure 【Chemical 207】 Or one of its protonated forms (wherein, R1, R2, and R3 are each independently H or an amino acid residue having a side chain containing an aromatic group, At least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid, R4 is H or an amino acid side chain, AASC is an amino acid side chain, Each m is independently an integer of 0, 1, 2, or 3), a compound.
2. The compound according to claim 1, wherein at least two amino acids of the cyclic peptide are aromatic hydrophobic amino acids and at least two amino acids of the cyclic peptide are uncharged non-aromatic amino acids.
3. At least two of R1, R2, and R3 are independently aromatic or heteroaromatic, and optionally, R1, R2, and R3 are each independently tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butyl-phenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, 3-(9-anthryl)-alanine side chains, the compound according to claim 1 or 2.
4. The compound according to claim 1, wherein at least two of R1, R2, and R3 are independently phenylalanine or naphthylalanine. Compound according to claim 1, wherein two of R1, R2, R3 and R4 are -CH2Ph. Compound according to claim 1, wherein two of R1, R2, R3 and R4 are H. Compound according to claim 1, wherein R4 is H or -CH2Ph.
8. The cyclic peptide has the structure of formula (I-1) 【Chemical 208】 Or a protonated form thereof, the compound according to claim 1.
9. (i) The exocyclic peptide contains 2, 3, or 4 lysine residues, and / or (ii) The exocyclic peptide contains at least 2 amino acid residues having hydrophobic side chains, The compound according to claim 1.
10. The exocyclic peptide contains the sequence PKKKRKV, the compound according to claim 1, 8, or 9.
11. The exocyclic peptide has the structure Ac-P-K-K-K-R-K-V, the compound according to claim 9.
12. The cyclic peptide is conjugated to a linker, the linker conjugates the cyclic peptide to the antisense compound and the exocyclic peptide, and the linker has the structure 【Chemical 209】 (wherein x' is an integer from 1 to 23, y is an integer from 1 to 5, z' is an integer from 1 to 23, * is the binding point to AASC, and M is a linking group) having, The compound according to claim 1.
13. The compound according to claim 12, wherein z' is 11 and / or x' is 1.
14. The exocyclic peptide is conjugated to the linker at the amino terminus of the linker, the antisense compound is conjugated to M, and / or M is -C(O)-, the compound according to claim 12 or 13.
15. The antisense compound (AC) contains the sequence AG(CAG)n, G(CAG)n, (CAG)nAG, or (CAG)nA, wherein n is an integer from 1 to 50, and / or the antisense compound (AC) contains 5 to 10 CAG repeats, the compound according to claim 1.
16. The antisense compound (AC) is (i) an antisense nucleotide which is 5'-CAG CAG CAG CAG CAG CAG CAG CAG CAG CAG-3'; (ii)an antisense nucleotide that is 5'-CAG CAG CAG CAG CAG CAG CAG CAG-3'; or (iii)an antisense nucleotide that is 5'-CAG CAG CAG CAG CAG CAG CAG-3' The compound according to claim 1, comprising the same.
17. The cyclic peptide has a structure of formula (I-1) 【Chemical 210】 or a protonated form thereof (wherein m is 2), the antisense compound and the exocyclic peptide include a linker that conjugates to the AAS C, the AAS C is the side chain of a glutamic acid residue, and the linker has a structure 【Chemical 211】 (wherein x' is 1, y is an integer from 1 to 5, z' is 11, * is the bonding point to the side chain of the glutamic acid residue, and M is -C(O)-), and the exocyclic peptide includes PKKKRKV. The compound according to claim 1.
18. The compound according to claim 17, wherein the antisense compound comprises the nucleotide sequence 5'-CAG CAG CAG CAG CAG CAG CAG-3'.
19. The compound according to claim 17 or 18, wherein y is 4.
20. The compound according to claim 17, wherein the exocyclic peptide has the structure Ac-P-K-K-K-R-K-V-.
21. The compound according to claim 17, wherein the exocyclic peptide is conjugated to the amino group of the linker.
22. The compound according to claim 17, wherein M is covalently bonded to the antisense compound.
23. The PMO has the sequence 5'-CAG CAG CAG CAG CAG CAG CAG-3', y is 4, the exocyclic peptide is Ac-P-K-K-K-R-K-V-, the exocyclic peptide is conjugated to the amino group of the linker, M is covalently bonded to the PMO, The compound according to claim 17. **Claim 24**: The compound according to claim 17, wherein the compound is EEV-PMO 221-1120 (where PMO 221 = 5'-CAG CAG CAG CAG CAG CAG CAG-3' (SEQ ID NO: 154, all PMO monomers), EEV 1120 = Ac-PKKKRKV-AEEA-Lys(cyclo[FGFRGRGQ]-PEG12-OH (Ac-(SEQ ID NO: 42)-AEEA-Lys(SEQ ID NO: 82)-PEG12-OH))), and the PMO and EEV are conjugated using amide chemistry). **Claim 25**: The formula 【Chemical 212】 (wherein the cargo is a phosphorodiamidate morpholino (PMO) nucleotide containing the sequence 5'-CAG CAG CAG CAG CAG CAG CAG-3', the EP is an exocyclic peptide containing PKKKRKV,[[]] R1 is the side chain of phenylalanine,[[]] R2 is H,[[]] R3 is the side chain of phenylalanine,[[]] R4 is H,[[]] R6 is H,[[]] each m is 2,[[]] q is 1,[[]] n is 1,[[]] y is 4,[[]] x' is 1,[[]] z' is 11), the compound according to claim 1. **Claim 26**: The formula 【Chemical 213】 (wherein the cargo is a phosphorodiamidate morpholino (PMO) nucleotide containing the sequence 5'-CAG CAG CAG CAG CAG CAG CAG-3', the EP is an exocyclic peptide containing PKKKRKV), the compound according to claim 1. **Claim 27**: A pharmaceutical composition comprising the compound according to claim 1 in combination with a pharmaceutically acceptable carrier. **Claim 28**: A composition comprising the compound according to claim 1, or the composition according to claim 27, for use in a method of treating myotonic dystrophy (DM). **Claim 29**: The composition according to claim 28, wherein administration of the composition results in an increase in the expression of the wild-type protein in muscle tissue, the wild-type protein is a protein expressed from a gene that does not have an extended CUG repeat, optionally, the muscle tissue is diaphragm tissue, quadriceps muscle tissue, and / or heart tissue, and / or the administration prevents or reduces lesion formation.