Compositions and methods for modulating mRNA splicing

JP2024518476A5Pending Publication Date: 2025-05-13ENTRADA THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023569638
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-03-31
Filing Date
2022-05-09
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing antisense compounds face challenges in gaining access to subcellular compartments, achieving broad or targeted tissue distribution, and ensuring specificity to modulate splicing of target transcripts effectively, leading to limited efficacy in treating diseases caused by aberrant gene transcription and splicing.

Method used

The development of compounds comprising a therapeutic moiety and a cell-penetrating peptide (CPP) that binds to splice elements or splice regulatory elements of target transcripts, enhancing intracellular delivery and modulating splicing by inducing exon skipping and frameshifts, resulting in nonsense-mediated decay.

Benefits of technology

The compounds enhance the specificity and efficacy of splicing modulation, leading to reduced expression of target proteins and potential therapeutic benefits for genetic diseases by inducing exon skipping and premature stop codons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The compound comprises at least one cyclic cell penetrating peptide (cCPP) conjugated to an antisense compound (AC). The AC regulates splicing of RNA transcripts. For example, the AC induces exon skipping. Exon skipping can result in downregulation of protein expression or activity. Exon skipping can cause a frameshift in the resulting mRNA. The frameshift can result in a premature stop codon. The frameshift can result in nonsense-mediated decay.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application Nos. 63 / 186,664, filed May 10, 2021, 63 / 210,882, filed June 15, 2021, 63 / 321,921, filed March 21, 2022, 63 / 362,295, filed March 31, 2022, 63 / 239,671, filed September 1, 2021, 63 / 210,866, filed June 15, 2021, 63 / 298,587, filed January 11, 2022, and 63 / 318,201, filed March 9, 2022, each of which is incorporated by reference in its entirety to the extent not inconsistent with the disclosures presented herein.

[0002] Provided herein are compositions and methods for modulating mRNA splicing, particularly for modulating the expression or activity of a protein of interest by inducing exon skipping, for example, by introducing a frameshift into an RNA transcript that can result in nonsense-mediated decay of the RNA transcript. [Background technology]

[0003] A gene is a deoxyribonucleic acid (DNA) sequence that encodes a functional gene product, such as a protein. The process of converting a gene's code into a functional gene product involves transcribing RNA (transcript) from the gene's DNA and translating the RNA into protein. RNA is initially transcribed from DNA as an immature "pre-mRNA" that undergoes processing to become mature messenger RNA (mRNA) that can be translated into protein. In eukaryotes, processing steps include the addition of a single-nucleotide modified guanine (G) nucleotide cap to the 5' end of the RNA, the addition of a polyadenosine sequence to the 3' end of the RNA (poly-A tail), and RNA splicing.

[0004] Splicing refers to the process by which introns (intervening sequences) are removed from pre-mRNA and exons (coding sequences) are joined together to form the mature mRNA.

[0005] Many mammalian genes are alternatively spliced, with different exons in the pre-mRNA sequence being included or excluded in the mature mRNA transcript, so that a single gene can produce different mRNA messages that are translated into proteins (isoforms) with different sizes and / or functions.

[0006] Alternative splicing can involve cryptic splice sites within the exon and / or intron regions of a transcript. Cryptic splice sites are splice sites that are typically unused but can be used when the normal splice site is blocked or unavailable, or when a mutation makes a normally dormant site an active splice site. In cryptic splicing, the splicing machinery recognizes the cryptic splice site rather than the canonical splice site. In many cases, cryptic splicing results in the inclusion or exclusion of part or all of an intron or exon sequence in the mRNA.

[0007] Antisense modulation of pre-mRNA splicing has been used to restore cryptic splicing, to alter the levels of alternatively spliced ​​genes (isoform switching), and for exon skipping, e.g., to restore a disrupted reading frame or knock down the function of an undesirable gene (Aartsma-Rus and Ommen, RNA (2007), 13:1609-1624).

[0008] The major problems with the use of antisense compounds in therapy include their limited ability to gain access to intracellular compartments when administered systemically, their limited ability to achieve widespread or specifically targeted tissue distribution, and the challenge of achieving sufficient specificity for targeted RNA to minimize off-target effects. Intracellular delivery of antisense compounds can be facilitated by the use of carrier systems such as polymers, cationic liposomes, or by chemical modification of the construct, for example, by covalent attachment of cholesterol molecules. However, intracellular delivery efficiency can be low and tissue distribution can be narrow. Furthermore, existing technologies remain hampered by off-target interactions. As a result, improved delivery systems are still needed to increase the effectiveness of these antisense approaches, and there remains an unmet need for effective compositions for delivering antisense compounds to intracellular compartments and widely to all affected tissue types, for example, to specifically target a given gene product to treat diseases caused by abnormal gene transcription, splicing, and / or translation. Summary of the Invention

[0009] The present disclosure generally relates to compounds, compositions, and methods for modulating splicing of target transcripts (e.g., pre-mRNA) of genes, such as genes associated with disease. In embodiments, the disclosure relates to compounds and compositions comprising a therapeutic moiety (TM) and a cell penetrating peptide (CPP). The TM can be an antisense compound (AC) that binds to the target transcript and modulates splicing of the target transcript. In embodiments, the AC binds to at least a portion of a splice element (SE) or cis-acting splice regulatory element (SRE) of the target transcript, or in proximity to a splice element or cis-acting splice regulatory element of the target transcript, thereby modulating splicing of the target transcript. In embodiments, binding of the AC to the transcript results in downregulation of expression or activity of a protein expressed from the target transcript.

[0010] In embodiments, binding of the AC to the target transcript results in exon skipping. In embodiments, exon skipping results in frameshifting. In embodiments, frameshifting results in a premature stop codon. In embodiments, frameshifting results in nonsense-mediated decay. In embodiments, frameshifting results in a premature stop codon and nonsense-mediated decay.

[0011] Described herein are methods in which a compound or composition described herein is used to treat a disease. In embodiments, the disease is a genetic disease. In embodiments, the compound or composition is used to treat the genetic disease by modulating splicing of a gene associated with the disease. In embodiments, the compound or composition treats the genetic disease by modulating splicing of a gene transcript associated with the disease. In embodiments, the method comprises administering a compound or composition described herein to a subject in need thereof. In embodiments, the subject in need thereof is a patient having or at risk of having a genetic disease. In embodiments, the method comprises administering a therapeutically effective amount of a compound or composition described herein to a subject in need thereof. In embodiments, the genetic disease is a disease associated with abnormal expression of IRF-5, DUX4, or GYS1, or a genetic variant thereof.

[0012] The CPP can enhance the intracellular delivery of the AC, thereby enhancing the effectiveness of the AC in modulating the splicing of target transcripts. The CPP can be a cyclic CPP (cCPP).

[0013] The compounds described herein may comprise an Endosomal Escape Vehicle (EEV) configured such that the compound, or portions thereof, internalized in a cell within an endosome, can escape the endosome and enter the cytosol or a cellular compartment, allowing an AC to act on a target transcript to modulate splicing. In embodiments, the EEV comprises a CPP, such as a cCPP.

[0014] In embodiments, the cCPP is of formula (A):

[0015] [ka] or a protonated form thereof, wherein: R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4, R5, R6, R7 are independently H or an amino acid side chain; at least one of R4, R5, R6, and R7 is a side chain of 3-guanidino-2-aminopropionic acid, 4-guanidino-2-aminobutanoic acid, arginine, homoarginine, N-methylarginine, N,N-dimethylarginine, 2,3-diaminopropionic acid, 2,4-diaminobutanoic acid, lysine, N-methyllysine, N,N-dimethyllysine, N-ethyllysine, N,N,N-trimethyllysine, 4-guanidinophenylalanine, citrulline, N,N-dimethyllysine, β-homoarginine, or 3-(1-piperidyl)alanine; AA SC is an amino acid side chain, q is 1, 2, 3 or 4.

[0016] In an embodiment, the cCPP of formula (A) is of formula (I):

[0017] [ka] or a protonated form or salt thereof, having a structure selected from: Each m is independently an integer of 0 to 3.

[0018] In embodiments, the cCPP of formula (A) has the formula (I-1):

[0019] [ka] or a protonated form or salt thereof.

[0020] In embodiments, the cCPP of formula (A) has the formula (I-2):

[0021] [ka] or a protonated form or salt thereof.

[0022] In embodiments, the cCPP of formula (A) has the formula (I-3):

[0023] [ka] or a protonated form or salt thereof.

[0024] In embodiments, the cCPP of formula (A) has the formula (I-4):

[0025] [ka] or a protonated form or salt thereof.

[0026] In embodiments, the cCPP of formula (A) has the formula (I-5):

[0027] [ka] or a protonated form or salt thereof.

[0028] In embodiments, the cCPP of formula (A) has the formula (I-6):

[0029] [ka] or a protonated form or salt thereof.

[0030] In an embodiment, the cCPP is of formula (II):

[0031] [ka] During the ceremony, AA SC is an amino acid side chain, R 1a, R 1b , and R 1c are each independently a 6- to 14-membered aryl or a 6- to 14-membered heteroaryl; R 2a , R 2b , R 2c and R 2d are independently amino acid side chains, R 2a , R 2b , R 2c and R 2d At least one of the

[0032] [ka] or a protonated form or salt thereof, R 2a , R 2b , R 2c and R 2d at least one of is guanidine or a protonated form or salt thereof; each n'' is independently an integer from 0 to 5; each n' is independently an integer from 0 to 3; If n' is 0, then R 2a , R 2b , R 2b or R 2d does not exist.

[0033] In embodiments, the cCPP of formula (II) has the formula (II-1):

[0034] [ka] It is of the type.

[0035] In embodiments, the cCPP of formula (II) has the formula (IIa):

[0036] [ka] It is of the type.

[0037] In embodiments, the cCPP of formula (II) has the formula (IIb):

[0038] [ka] It is of the type.

[0039] In embodiments, the cCPP of formula (II) has the formula (IIc):

[0040] [ka] or a protonated form or salt thereof.

[0041] In embodiments, the cCPP has the structure:

[0042] [ka] or a protonated form or salt thereof, wherein at least one atom of the amino acid side chain is replaced by a therapeutic moiety or a linker, or at least one lone pair of electrons forms a bond to a therapeutic moiety or a linker.

[0043] In embodiments, the cCPP has the structure:

[0044] [ka] or a protonated form or salt thereof, wherein at least one atom of the amino acid side chain is replaced by a therapeutic moiety or a linker, or at least one lone pair of electrons forms a bond to a therapeutic moiety or a linker.

[0045] In embodiments, the compound comprises an exocyclic peptide (EP). In ,EP is, Contains one of the following sequences::KK、KR、RR、HH、HK、HR、RH、KKK、KGK、KBK、KBR、KRK、KRR、RKK、RRR、KKH、KHK、HKK、HRR、HRH、HHR、HBH、HHH、HHHH (SEQ ID NO: 1) 、KHKK (SEQ ID NO: 2) 、KKHK (SEQ ID NO: 3) 、KKKH (SEQ ID NO: 4) 、KHKH (SEQ ID NO: 5) 、HKHK (SEQ ID NO: 6) 、KKKK (SEQ ID NO: 7) 、KKRK (SEQ ID NO: 8) 、KRKK (SEQ ID NO: 9) 、KRRK (SEQ ID NO: 10) 、RKKR (SEQ ID NO: 11) 、RRRR (SEQ ID NO: 12) 、KGKK (SEQ ID NO: 13) 、KKGK (SEQ ID NO: 14) 、HBHBH (SEQ ID NO: 15) 、HBKBH (SEQ ID NO: 16) 、RRRRRR (SEQ ID NO: 17) 、KKKKK (SEQ ID NO: 18) 、KKKRK (SEQ ID NO: 19) 、RKKKK (SEQ ID NO: 20) 、KRKKK (SEQ ID NO: 21) 、KKRKK (SEQ ID NO: 22) 、KKKKR (SEQ ID NO: 23) 、KBKBK (SEQ ID NO: 24) 、RKKKKG (SEQ ID NO: 25) 、KRKKKG (SEQ ID NO: 26) 、KKRKKG (SEQ ID NO: 27) 、KKKKRG (SEQ ID NO: 28) 、RKKKKB (SEQ ID NO: 29) 、KRKKKB (SEQ ID NO: 30) 、KKRKKB (SEQ ID NO: 31) 、KKKKRB (SEQ ID NO: 32) 、KKKRKV (SEQ ID NO: 33) 、RRRRRR (SEQ ID NO: 34) 、HHHHHH (SEQ ID NO: 35) 、RHRHRH (SEQ ID NO: 36) 、HRHRHR (SEQ ID NO: 37) 、KRKRKR (SEQ ID NO: 38) 、RKRKRK (SEQ ID NO: 39) 、RBRBRB (SEQ ID NO: 40) 、KBKBKB (SEQ ID NO: 41) 、PKKKRKV (SEQ ID NO:42) 、PGKKRKV (SEQ ID NO: 43) 、PKGKRKV (SEQ ID NO: 44) 、PKKGRKV (SEQ ID NO: 45) 、PKKKGKV (SEQ ID NO: 46) 、PKKKRGV (SEQ ID NO: 47), or PKKKRKG (SEQ ID NO: 48) . B is β-alanine.

[0046] In embodiments, the compound has formula (C): :

[0047] [ka] or a protonated form or salt thereof. R1, R2, and R3 are each independently H or a side chain comprising an aryl or heteroaryl group, wherein at least one of R1, R2, and R3 is a side chain comprising an aryl or heteroaryl group; R4 and R7 are independently H or an amino acid side chain; EP is an exocyclic peptide, each m is independently an integer from 0 to 3; n is an integer from 0 to 2, x' is an integer from 1 to 23, y is an integer from 1 to 5, q is an integer from 1 to 4, z' is an integer from 1 to 23, Cargo is AC.

[0048] In embodiments, the compound comprises the structure of formula (C-1), (C-2), (C-3), or (C-4):

[0049] [ka]

[0050] [ka] or a protonated form or salt thereof, where EP is an exocyclic peptide and the oligonucleotide is AC. [Brief explanation of the drawings]

[0051] [Figure 1A]Schematic diagram showing a splicing regulatory element containing a splice site (A) and a general splicing reaction (two transesterification reactions) (B). [Figure 1B] Schematic diagram showing a splicing regulatory element containing a splice site (A) and a general splicing reaction (two transesterification reactions) (B). [Figure 2] FIG. 1 is a schematic diagram showing antisense compound-mediated exon skipping to create a premature stop codon that ultimately leads to nonsense-mediated decay of the target transcript. [Figure 3] 1 depicts modified nucleotides used in the antisense oligonucleotides described herein. Structures 1–3 (1 = phosphorothioate; 2 = (SC5-Rp)-α,β-CAN; 3 = PMO) are phosphate backbone modifications; 4 (2-thio-dT) is a base modification; 5–8 (5 = 2'-OMe-RNA; 6 = 2'O-MOE-RNA; 7 = 2'F-RNA; 8 = 2'F-ANA) are 2' sugar modifications; 9–11 are constrained nucleotides; 12–14 (9 = LNA; 10 = (S)-cEt; 11 = tcDNA; 12 = FHNA; 13 = (S)5'-C-methyl; 14 = UNA) are additional sugar modifications; 15–18 (15 = E-VP; 16 = methyl phosphonate; 17 = 5'-phosphorothioate; 18 = (S)-5'-C-methyl with phosphate) are 5' phosphate-stabilizing modifications; and 19 is a morpholino sugar. Reformatted from Khvorova, A., et al., Nat. Biotechnol. (2017) March;35(3):238-248. [Figure 4A] 1 provides the structures of adenine (A), cytosine (B), guanine (C), and thymine (D) morpholino subunit monomers used in synthesizing phosphorodiamidate-linked morpholino oligomers (PMOs). [Figure 4B]1 provides the structures of adenine (A), cytosine (B), guanine (C), and thymine (D) morpholino subunit monomers used in synthesizing phosphorodiamidate-linked morpholino oligomers (PMOs). [Figure 4C] 1 provides the structures of adenine (A), cytosine (B), guanine (C), and thymine (D) morpholino subunit monomers used in synthesizing phosphorodiamidate-linked morpholino oligomers (PMOs). [Figure 4D] 1 provides the structures of adenine (A), cytosine (B), guanine (C), and thymine (D) morpholino subunit monomers used in synthesizing phosphorodiamidate-linked morpholino oligomers (PMOs). [Figure 5A] Figure 1 shows the conjugation chemistry for connecting antisense compounds (ACs) to peptides such as cyclic cell penetrating peptides (cCPPs). Reagents for amide bond formation between peptides with N-hydroxysuccinimide activated esters (top) or free carboxylic acids (bottom) and the primary amine at the 5' end of ACs are shown. [Figure 5B] This figure shows the conjugation chemistry for connecting an antisense compound (AC) to a peptide such as a cyclic cell penetrating peptide (cCPP). The figure also shows reagents for the amide bond formation reaction between the primary or secondary amine at the 3' end of the AC and a peptide bearing a tetrafluorophenyl (TFP)-activated ester. [Figure 5C]Conjugation chemistry for connecting antisense compounds (ACs) to peptides such as cyclic cell penetrating peptides (cCPPs) is shown. Reagents for peptide-azide conjugation to 5' cyclooctyne-modified ACs via copper-free azide-alkyne cycloaddition are shown. [Figure 5D] Figure 1 shows conjugation chemistries for connecting antisense compounds (ACs) to peptides such as cyclic cell penetrating peptides (cCPPs). Other exemplary reagents are shown for conjugation between a 3'-modified cyclooctyne AC or a 3'-modified azide AC and a peptide such as a cCPP containing a linker-azide or linker-alkyne / cyclooctyne moiety via copper-free or copper-catalyzed azide-alkyne cycloaddition (click reaction), respectively. [Figure 6] Conjugation chemistry is shown for connecting ACs and CPPs with additional linker modalities containing polyethylene glycol (PEG) moieties using the conjugation chemistry shown in Figure 5. Purification methods are shown. [Figure 7A] The levels of GYS1 protein (A and C) and GYS1 mRNA (B and D) in the diaphragm (A and B) and heart (C and D) of untreated mice, mice treated with PMO, and mice treated with various concentrations of EEV-PMO in a GAA knockout mouse model are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 7B] The levels of GYS1 protein (A and C) and GYS1 mRNA (B and D) in the diaphragm (A and B) and heart (C and D) of untreated mice, mice treated with PMO, and mice treated with various concentrations of EEV-PMO in a GAA knockout mouse model are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 7C] The levels of GYS1 protein (A and C) and GYS1 mRNA (B and D) in the diaphragm (A and B) and heart (C and D) of untreated mice, mice treated with PMO, and mice treated with various concentrations of EEV-PMO in a GAA knockout mouse model are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 7D] The levels of GYS1 protein (A and C) and GYS1 mRNA (B and D) in the diaphragm (A and B) and heart (C and D) of untreated mice, mice treated with PMO, and mice treated with various concentrations of EEV-PMO in a GAA knockout mouse model are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 8A] Plots of GYS1 mRNA levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 8B] Plots of GYS1 mRNA levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 8C] Plots of GYS1 mRNA levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 8D]Plots of GYS1 mRNA levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 9A] Plots of GYS1 protein levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 9B] Plots of GYS1 protein levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 9C] Plots of GYS1 protein levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 9D] Plots of GYS1 protein levels in the heart (A), diaphragm (B), quadriceps (C), and triceps (D) of untreated, PMO-treated, and EEV-PMO-treated mice at various time points after treatment are shown (P > 0.05 = NS; P ≤ 0.05 = *; P ≤ 0.01 = **; P ≤ 0.001 = ***). [Figure 10A] Plots showing the levels of IRF5 mRNA expression in the liver (A), small intestine (B), and tibialis anterior muscle (C) of mice treated with various concentrations of EEV-PMO (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). MPK (mpk)=mg / kg. [Figure 10B] Plots showing the levels of IRF5 mRNA expression in the liver (A), small intestine (B), and tibialis anterior muscle (C) of mice treated with various concentrations of EEV-PMO (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). MPK (mpk)=mg / kg. [Figure 10C] Plots showing the levels of IRF5 mRNA expression in the liver (A), small intestine (B), and tibialis anterior muscle (C) of mice treated with various concentrations of EEV-PMO (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). MPK (mpk)=mg / kg. [Figure 11A] 1 is a plot showing the level of IRF5 protein expression in an in vitro experiment in which mouse macrophage cells were treated with various concentrations of EEV#1-PMO, EEV#2-PMO, EEV#3-PMO, and EEV#4-PMO (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 11B] 1 is a plot showing the level of IRF5 protein expression in an in vitro experiment in which mouse macrophage cells were treated with various concentrations of EEV#1-PMO, EEV#2-PMO, EEV#3-PMO, and EEV#4-PMO (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 12] 1 is a plot showing knockdown of GYS1 mRNA levels in the wild-type mouse myoblast cell line C2C12 after treatment with various concentrations of PMO220 or EEV-PMO220-814. N=3, *p<0.05, **p<0.01 vs. 0 (no treatment) by Student's t-test. [Figure 13A] Plots showing knockdown of GYS1 mRNA levels in mouse myoblasts (A) and mouse fibroblasts (B) after treatment with various concentrations of PMO220. N=2, *p<0.05 vs. NT (no treatment) by Student's t-test. [Figure 13B]Plots showing knockdown of GYS1 mRNA levels in mouse myoblasts (A) and mouse fibroblasts (B) after treatment with various concentrations of PMO220. N=2, *p<0.05 vs. NT (no treatment) by Student's t-test. [Figure 14A] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-814. MPK (mpk) = mg / kg. [Figure 14B] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-814. MPK (mpk) = mg / kg. [Figure 14C] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-814. MPK (mpk) = mg / kg. [Figure 14D] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-814. MPK (mpk) = mg / kg. [Figure 15] 1 is a plot showing GYS2 mRNA levels in the liver of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-814. MPK (mpk) = mg / kg. [Figure 16A] 1A-D are plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-1055. [Figure 16B]1A-D are plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-1055. [Figure 16C] 1A-D are plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-1055. [Figure 16D] 1A-D are plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) of GAA knockout mice after treatment with PMO220 or various concentrations of PMO-EEV220-1055. [Figure 17A] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 17B] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 17C] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 17D] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 18A]Plots showing drug exposure levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO220 or 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 18B] Plots showing drug exposure levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO220 or 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 18C] Plots showing drug exposure levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO220 or 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 18D] Plots showing drug exposure levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) at various time points after treatment of GAA knockout mice with 20 mpk of PMO220 or 20 mpk of PMO-EEV220-1055. MPK (mpk) = mg / kg. [Figure 19A] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 19B] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 19C]Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 19D] Plots showing GYS1 mRNA levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 20A] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 20B] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 20C] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 20D] Plots showing GYS1 protein levels in the heart (A), diaphragm (B), triceps (C), and quadriceps (D) for wild-type, GAA knockout, and GAA knockout mice treated with various concentrations of EEV-PMO220-1120. MPK (mpk) = mg / kg. [Figure 21A]1A-C are plots showing GYS1 protein levels in the heart (A), diaphragm (B), and quadriceps (C) for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 21B] 1A-C are plots showing GYS1 protein levels in the heart (A), diaphragm (B), and quadriceps (C) for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 21C] 1A-C are plots showing GYS1 protein levels in the heart (A), diaphragm (B), and quadriceps (C) for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 21D] 1A-C are plots showing GYS1 protein levels in the heart (A), diaphragm (B), and quadriceps (C) for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 22A] 1 is a plot showing GYS1 (A) and GYS2 (B) levels in the liver for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 22B] 1 is a plot showing GYS1 (A) and GYS2 (B) levels in the liver for wild-type mice, GAA knockout mice, and GAA knockout mice treated with multiple doses of EEV-PMO220-1055. [Figure 23A] The expression levels of IRF-5 in mouse TIa tissue (A), liver tissue (B), and small intestine tissue (C) after treating mice with two doses of PMO or EEV-PMO278-1120 are shown. MPK (mpk) = mg / kg. [Figure 23B]The expression levels of IRF-5 in mouse TIa tissue (A), liver tissue (B), and small intestine tissue (C) after treating mice with two doses of PMO or EEV-PMO278-1120 are shown. MPK (mpk) = mg / kg. [Figure 23C] The expression levels of IRF-5 in mouse TIa tissue (A), liver tissue (B), and small intestine tissue (C) after treating mice with two doses of PMO or EEV-PMO278-1120 are shown. MPK (mpk) = mg / kg. [Figure 24A] IRF-5 expression levels in mouse liver (A), kidney (B), and tibialis anterior muscle (C) tissues after treatment with one dose of PMO278 or PMO-EEV278-1120 are shown (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 24B] IRF-5 expression levels in mouse liver (A), kidney (B), and tibialis anterior muscle (C) tissues after treatment with one dose of PMO278 or PMO-EEV278-1120 are shown (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 24C] IRF-5 expression levels in mouse liver (A), kidney (B), and tibialis anterior muscle (C) tissues after treatment with one dose of PMO278 or PMO-EEV278-1120 are shown (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 25A] GYS1 protein levels in quadriceps (A) and triceps (B) muscles are shown using a GYS antibody not specific for GYS1 after mice were treated with various concentrations of EEV-PMO construct 220-814. [Figure 25B] GYS1 protein levels in quadriceps (A) and triceps (B) muscles are shown using a GYS antibody not specific for GYS1 after mice were treated with various concentrations of EEV-PMO construct 220-814. [Figure 26A]GYS1 protein levels in the diaphragm (A), heart (B), and triceps (C) using a GYS1-specific antibody are shown after mice were treated with various concentrations of the EEV-PMO construct 220-814. [Figure 26B] GYS1 protein levels in the diaphragm (A), heart (B), and triceps (C) using a GYS1-specific antibody are shown after mice were treated with various concentrations of the EEV-PMO construct 220-814. [Figure 26C] GYS1 protein levels in the diaphragm (A), heart (B), and triceps (C) using a GYS1-specific antibody are shown after mice were treated with various concentrations of the EEV-PMO construct 220-814. [Figure 27A] Figure 1 shows the IRF-5 expression levels in RAW264.7 monocyte / macrophage cells after treatment with various concentrations of PMO-EEV277-1120 and 278-1120 (P>0.05=NS; P≦0.05=*; P≦0.01=**; P≦0.001=***). [Figure 27B] 1 is a bar graph of the exon skipping percentage at various time points after treatment of RAW264.7 monocyte / macrophage cells with EEV-PMO278-1120. NT = no treatment. [Figure 28A] 1A-B are bar graphs showing the level of IRF-5 expression (A) and exon 4 skipping percentage (B) in RAW264.7 monocyte / macrophage cells after treatment with various EEV-PMOs at various concentrations followed by stimulation with R848. [Figure 28B] 1A-B are bar graphs showing the level of IRF-5 expression (A) and exon 4 skipping percentage (B) in RAW264.7 monocyte / macrophage cells after treatment with various EEV-PMOs at various concentrations followed by stimulation with R848. [Figure 29A] 1 is a plot showing IRF-5 exon 4 and exon 5 skipping levels in human THP1 cells after treatment with various EEV-PMOs at different concentrations. [Figure 29B]1 is a plot showing IRF-5 exon 4 and exon 5 skipping levels in human THP1 cells after treatment with various EEV-PMOs at different concentrations. DETAILED DESCRIPTION OF THE INVENTION

[0052] Splicing Pre-mRNA molecules are produced in the nucleus and processed before or during transport to the cytoplasm for translation. Pre-mRNA processing involves the addition of a 5'-methylated guanine cap and a poly(A) tail of approximately 200-250 bases to the 3' end of the transcript. Pre-mRNA processing also includes splicing, which occurs in approximately 90% to 95% of mature mammalian mRNAs. Introns (or intervening sequences) are regions of the primary transcript (or the DNA encoding it) that are not included in the coding sequence of the mature mRNA. Exons are regions of the primary transcript that remain in the mature mRNA upon reaching the cytoplasm. A transcript may have multiple introns and exons. Exons are spliced ​​together to form the mature mRNA sequence. Splice junctions are also called splice sites; the 5' side of the junction is often referred to as the "5' splice site" or "splice donor site," and the 3' side is referred to as the "3' splice site" or "splice acceptor site." In splicing, the 3' end of the upstream exon is linked to the 5' end of the downstream exon. Thus, a transcript (e.g., pre-mRNA) has an exon / intron junction at the 5' end of the intron and an intron / exon junction at the 3' end of the intron. After the intron is removed, the exons are adjacent in the mature mRNA at what are sometimes called exon / exon junctions or boundaries. Cryptic splice sites are sites that are used less frequently but can be used when the regular splice sites are blocked or unavailable. Alternative splicing, defined as the splicing together of different combinations of exons, often results in multiple mRNA transcripts from a single gene.

[0053] The removal of introns from pre-mRNA is catalyzed by the spliceosome, a ribonucleoprotein (RNP) complex containing five small nuclear ribonucleoproteins (snRNPs), and numerous other proteins (Will and Luhrmann, Cold Spring Harb. Perspect. Biol. (2011), 3(7):a003707; Havens, et al., Wiley Interdiscip. RNA (2014), 4(3), 247-266. doi:10.1002 / wrna.1158). Splicing is governed in part by splice elements (SEs). As used herein, "splice elements" are sequence elements found in pre-mRNA that are necessary for splicing, e.g., canonical splicing, to occur (Figure 1A). SEs include a 5' splice site (5'ss) and a 3' splice site (3'ss). The 5'ss, also called the donor splice site, contains a largely invariant "GU" dinucleotide sequence with less conserved downstream residues. The 5'splice site also contains the exon / intron junction. As used herein, the exon / intron junction is the nucleotide sequence 10 nucleotides upstream and 10 nucleotides (+10 and -10) from the G of the GU sequence in the 5'ss. The 3'ss, or acceptor splice site, contains three conserved elements: a branch splice point (BSP), sometimes called a branch point, a polypyrimidine or Py tract, and a terminal "AG." The BSP is typically an adenosine located about 18 to about 40 nucleotides from the 3'ss. The Py tract typically contains about 15 to about 20 pyrimidine residues, particularly uracil (U) (represented by X in Figure 1A). n(Indicated as ). However, atypical branch points exist. They are located further away from the 3' splice site and / or utilize non-adenosine bases (Montes et al., Trends Genet. (2019), 35(1):68-87). The 3'ss also includes intron / exon junctions. As used herein, an intron / exon junction is the nucleotide sequence 10 nucleotides upstream and 10 nucleotides (+10 and -10) from the G of the AG sequence of the 3'ss.

[0054] In most splicing reactions, exons are recognized by specific base-pairing interactions with small nuclear RNA (snRNA) components of five small ribonucleoproteins (snRNPs) (U1, U2, U4, U5, and U6 (Havens et al., (2014) Wiley Interdiscip. RNA. 2013, 4(3), 247-266. doi:10.1002 / wrna.1158; Wahl MC et al., Cell (2009), 136:701-718)). Each snRNP contains a small nuclear RNA that is configured to recognize a specific nucleotide sequence and one or more proteins. Exon splicing involves two sequential spliceosome-catalyzed transesterification reactions (Figure 1B). Generally, the splicing reaction begins with U1 binding to the 5' ss, followed by U2 binding to the branch splice point (BPS), and finally U4, U5, and U6 binding near the 5' and 3' splice sites. U1 and U4 are then displaced, followed by a first transesterification reaction in which the 2'-OH of the branch point nucleotide within the intron (A, as shown in Figure 1B) performs a nucleophilic attack on the first nucleotide of the intron at the 5' splice site (G, as shown in Figure 1B), forming a lariat intermediate. In the second reaction, the 3'-OH of the released 5' exon performs a nucleophilic attack on the last nucleotide of the intron at the 3' splice site (G, as shown in Figure 1B), joining the exons and releasing the intron lariat. U4, U5, and U6 are also released.

[0055] In addition to SEs, splicing is regulated in part by splicing regulatory elements (SREs). SREs include cis-regulatory elements and trans-acting splicing factors. Cis-regulatory elements and trans-acting splicing factors can promote canonical splicing, alternative splicing, or cryptic splicing.

[0056] Cis-regulatory elements are nucleotide sequences within a transcript that repress or enhance splicing. Trans-acting splicing factors are proteins and / or oligonucleotides that are not located within the transcript and act to enhance or repress splicing. Cis-regulatory elements generally function to recruit trans-acting splicing factors that activate or repress splicing. Trans-acting splice factors regulate splicing by associating with cis-regulatory elements. Trans-acting splice factors include serine / arginine-rich (SR-rich) proteins and heterogeneous nuclear ribonucleoproteins (hnRNPs).

[0057] Splicing cis-regulatory elements include exon splicing enhancer (ESE) sequences, exon splicing silencer (ESS) sequences, intron splicing enhancer (ISE) sequences, and intron splicing silencer (ISS) sequences (Figure 1A). ESE sequences promote the inclusion of the exon in which they reside into mRNA. ESS sequences inhibit the inclusion of the exon in which they reside into mRNA. ISE sequences enhance the use of alternative splice sites from their position within the intron. ISS sequences inhibit the use of alternative splice sites from their position within the intron. ISS sequences are typically 8–16 nucleotides in length and are less conserved than splice sites at exon-intron junctions.

[0058] Pre-mRNA splicing can also be regulated by the formation of secondary structures within the transcript, such as terminal stem loops (TSLs), which can affect the binding of spliceosomes or other regulatory proteins. Terminal stem loop sequences can be SREs and are typically about 12 to about 24 nucleotides in length, forming secondary loop structures due to complementarity (and thus binding) within the 12-24 nucleotide sequence.

[0059] Each SE and / or cis-acting SRE is separated from adjacent cis-acting SREs and / or SEs by an intervening sequence (IS).

[0060] Exon skipping Most eukaryotic pre-mRNAs can be differentially spliced, often by skipping exons, to produce distinct mature mRNA isoforms in a process called alternative splicing. The term "alternative splicing" refers to the joining of exons in different combinations (e.g., different 5' and 3' splice sites are joined). Alternative splicing can insert or remove amino acids, shift the reading frame, and / or introduce stop codons, which contribute to the complexity, flexibility, and abundance of genes and proteins expressed from them. Alternative splicing can also affect gene expression by removing or inserting regulatory elements, controlling translation, mRNA stability, and / or localization. Splicing-disrupting mutations are estimated to account for up to one-third of all disease-causing mutations (Havens et al. (2014) Wiley Interdiscip. RNA. 2013, 4(3), 247-266. doi:10.1002 / wrna.1158; Lim KH et al., Proc. Natl. Acad. Sci. USA (2011), 108:11093-11098; Faustino and Cooper, Genes & Dev. (2003), 17:419-437; and Sterne-Weiler T. et al., Genome Res. (2011), 21:1563-1571).

[0061] Mutations that affect the splicing process can occur in many different ways (Havens et al., (2014) Wiley Interdiscip. RNA. 2013, 4(3), 247-266. doi:10.1002 / wrna.1158). For example, intronic mutations can disrupt core splice sites (5'ss or 3'ss, sequences within the Py tract or BPS), resulting in skipping of exon(s) upstream or downstream from the mutated splice site (5'ss and / or 3ss) or retention of the intron. Often, when a splice site is mutated, a false splice site is activated within an adjacent exon or intron, which, after splicing, generates an alternative transcript. Mutations within introns can also disrupt or create de novo splicing silencers and / or enhancers and / or create de novo cryptic splice sites. Intron splice site mutations may account for approximately 10-15% of disease mutations (Havens et al. (2014) Wiley Interdiscip. RNA. 2013, 4(3), 247-266. doi:10.1002 / wrna.1158; Stenson PD et al., The Human Gene Mutation Database: 2008 update. Genome Med 2009, 1:13). Mutations occurring within coding exons (exonic mutations) can result in the creation of de novo cryptic splice sites, disruption of regulatory RNA secondary structures, and / or disruption of splicing silencers or enhancers, rendering the splice site unrecognizable by sequence-specific RNA-binding proteins required for splicing. Analysis of exonic mutations predicts that as many as 25% of mutations within exons can alter splicing (ibid.; Proc. Natl. Acad. Sci. USA (2011), 108:11093-11098).Cryptic splicing is caused by sequences in pre-mRNA that are not normally used as splice sites but are activated by mutations that either inactivate canonical splice sites or create splice sites where none previously existed (Arechavala-Gomeza et al., The Application of Clinical Genetics (2014), 4(7), 245-252; Roca X. et al., Genes Dev. (2013); 27(2):129-144). Furthermore, alternative splicing that contributes to different proteins produced from pre-mRNA can cause disease by shifting expression from one isoform to a different isoform associated with the disease (ibid.).

[0062] Targeting splice elements (e.g., SEs and / or SREs) involved in the splicing reaction or splicing to induce aberrant splicing can be used to disrupt gene expression of proteins involved in disease pathogenesis. For example, splicing can be targeted to cause exon skipping, thereby introducing frameshift or stop codons that result in non-functional or truncated proteins or degradation of RNA transcripts (Stenson PD et al., Genome Med. 2008;1(13)). Splicing-induced reading frame correction, reframing, and / or nonsense-mediated decay of target transcripts offers opportunities to treat many diseases and disorders.

[0063] compound Disclosed herein are compounds that modulate the expression and / or activity of a gene of interest. In embodiments, the compound modulates splicing of a target transcript of a target gene. In embodiments, the compound comprises at least one cell penetrating peptide (CPP) and at least one therapeutic moiety (TM) that binds to a target nucleotide sequence. In embodiments, the TM is an antisense compound (AC). In embodiments, the target nucleotide sequence comprises a nucleotide sequence adjacent to or including at least a portion of a cis-acting splicing regulatory element (SRE) and / or a nucleotide sequence adjacent to or including at least a portion of a splicing element (SE).

[0064] As used herein, "modulating splicing" and "modulating splicing" refer to altering the processing of a pre-mRNA transcript so that the spliced ​​mRNA molecule contains either a different combination of exons as a result of exon skipping or exon inclusion, a deletion in one or more exons, or the deletion or addition of sequences not normally found in the spliced ​​mRNA (e.g., intron sequences). Modulating splicing can include interfering with or promoting one or more steps in the splicing process. As used herein, the term "splicing process" encompasses all steps of the splicing reaction, including, for example, the binding of various snRNPs (e.g., U1, U2, U3, U4, and U5) to splicing elements and / or cis-acting splicing regulatory elements, the binding of various proteins and / or oligonucleotides to cis-regulatory elements, and the two sequential transesterification reactions, e.g., as shown in Figure 1B.

[0065] treatment part In embodiments, the present disclosure describes compounds comprising one or more Therapeutic Moieties (TM) capable of modulating the splicing of a transcript of interest from a gene of interest. In embodiments, the gene of interest may be a disease-causing gene.

[0066] The TM binds to (e.g., hybridizes with) a target nucleotide sequence. The target nucleotide sequence is generally contained within the target transcript of a gene of interest. For example, a TM targeting a gene of interest can bind to a target nucleotide sequence (e.g., a splicing element) within the target transcript.

[0067] The TM can be an antisense compound (AC), one or more elements associated with the Clustered Regularly Interspaced Short Palindromic Repeat (CRISPR) gene editing machinery, a polypeptide, or a combination thereof.

[0068] Antisense Compounds (AC) In embodiments, the therapeutic moiety comprises an antisense compound (AC) capable of modulating splicing of a target transcript of a target gene. The AC is an oligonucleotide comprising DNA bases, modified DNA bases, RNA bases, modified RNA bases, modified internucleoside linkages, conventional internucleoside linkages, conventional DNA sugars, modified DNA sugars, conventional RNA sugars, modified RNA sugars, or combinations thereof. In embodiments, the AC comprises a nucleotide sequence complementary to a target nucleotide sequence found within the target transcript. In embodiments, the AC comprises a nucleotide sequence complementary to a target nucleotide sequence adjacent to or comprising at least a portion of a splicing element and / or splicing regulatory element within the target transcript.

[0069] The ACs described herein may contain one or more asymmetric centers and thus give rise to enantiomers, diastereomers, and other stereoisomeric configurations which may be defined, with respect to absolute stereochemistry, as (R) or (S), or (D) or (L). The antisense compounds provided herein include all such possible isomers, as well as their racemic and optically pure forms.

[0070] In embodiments, the AC induces alternative splicing that results in the addition or deletion of nucleotides in the target transcript. In some embodiments, the AC induces alternative splicing that results in the addition or deletion of nucleotides within a single exon of the target transcript. In embodiments, the AC induces alternative splicing that results in the deletion of nucleotides within a single exon of the target transcript. In embodiments, the deletion of nucleotides within a single exon results in the translation of a truncated protein. In embodiments, the truncated protein is less toxic to cells than the non-truncated protein.

[0071] In embodiments, ACs are designed to skip exons (sometimes referred to as exon skipping), resulting in increased or decreased expression or activity of a target protein and / or downstream proteins regulated by the target gene. In embodiments, ACs are provided that generate mRNAs encoding truncated and / or non-functional proteins. In embodiments, ACs are provided that generate mRNAs encoding truncated and / or non-functional proteins by alternative splicing. In embodiments, ACs are provided that induce degradation of target transcripts, for example, via nonsense-mediated decay. In embodiments, antisense compounds (ACs) are provided that generate alternative mRNA isoforms with beneficial properties.

[0072] Antisense compounds (ACs) can be used to regulate splicing in any suitable manner. In embodiments, ACs can be designed to sterically block access to a splice site or at least a portion of a splicing element (SE) and / or a cis-acting splicing regulatory element (SRE), thereby redirecting splicing to a cryptic or de novo splice site. In embodiments, ACs can be targeted to splicing enhancer sequences (e.g., ESEs and / or ISEs) or splicing silencer sequences (e.g., ESSs and / or ISSs) to prevent binding of trans-acting regulatory splicing factors at the target site, effectively blocking or promoting splicing. In embodiments, ACs can be designed to base-pair across the bases of a splicing regulatory stem-loop to reinforce the stem-loop structure.

[0073] In embodiments, AC induces the addition or deletion of one or more nucleotides in the resulting processed transcript, e.g., mRNA. If the number of nucleotides added to or removed from the open reading frame is divisible by three to produce an integer, the resulting transcript can be translated into a functional or non-functional protein that has more or fewer amino acids than the corresponding protein expressed from the transcript, but has the same amino acid sequence, other than the added or deleted amino acids, as the protein expressed from a transcript in which no nucleotides were added or removed. If the number of nucleotides added to or removed from the open reading frame is not divisible by three to produce an integer, the open reading frame of the resulting processed transcript (e.g., mRNA) will be shifted. For example, the number of nucleotides added or deleted to induce such a "frameshift" modification can be 1, 2, 4, 5, 7, 8, 10, 11, 13, 14, 16, 17, 19, 20, 22, 23, etc. Due to the triplet nature of the genetic code, the addition or deletion of a number of nucleotides not divisible by 3 shifts the reading frame of the resulting processed transcript (e.g., mRNA) downstream of the frameshift. The shifted reading frame can result in nonsense-mediated decay, can result in a premature stop codon within the nonsense downstream of the frameshift, and / or can result in the expression of a protein with an entirely different sequence of amino acids downstream of the frameshift.

[0074] In embodiments, the AC induces the introduction of a premature termination codon (PTC) into an open reading frame. As used herein, a "premature termination codon" is a termination codon that is synchronous with a translation initiation codon and located upstream of a physiological termination codon synchronous with the translation initiation codon. Target transcripts with PTCs can be destabilized and degraded through various mechanisms, including nonsense-mediated decay.

[0075] Nonsense-mediated decay is a surveillance mechanism that recognizes the initiation of exonucleolytic and endonucleolytic degradation pathways to remove mRNA transcripts with PTCs to prevent the expression of truncated proteins that may have deleterious effects on cells. Several nonsense-mediated decay pathways have been proposed and reviewed (Lejeune et al., Biomedicines (2020), 10(1):141; Brogna et al., Nature Structural and Molecular Biology (2009), 16, 108-113; Karousis et al., Wiley Interdiscip. Rev. RNA (2016), 7(5):661-682). In embodiments where a target gene is overexpressed in a disease, inducing nonsense-mediated decay can be used to reduce the concentration of the target protein and thus treat the disease.

[0076] In embodiments, AC induces exon skipping, resulting in nonsense-mediated decay of target transcripts, in contrast to traditional exon skipping, which aims to skip exons to induce expression of specific protein isoforms, correct mis-splicing, alternative splicing, and / or avoid deleterious mutations in specific exons.

[0077] In embodiments, AC induces exon skipping of an exon in a target transcript, the exon having a number of nucleotides not divisible by 3. In embodiments, AC induces exon skipping of an exon having a number of nucleotides not divisible by 3, resulting in a PCT in the target transcript. In embodiments, AC induces exon skipping of an exon having a number of nucleotides not divisible by 3, resulting in a PCT in the target transcript, which results in nonsense-mediated decay of the target transcript. In embodiments, inducing nonsense-mediated decay of a target transcript results in a decrease in the concentration of the target transcript. In embodiments, inducing nonsense-mediated decay of a target transcript results in a decrease in the concentration of a target protein encoded by the target transcript. In embodiments, inducing nonsense-mediated decay of a target transcript results in an increase and / or decrease in the level of a protein of a downstream gene regulated by the target gene.

[0078] Figure 2 shows an example of AC-induced exon skipping, which results in nonsense-mediated decay of a target transcript or premature termination of protein translation. An AC binds to pre-mRNA. In an exemplary embodiment, an AC binds at the intron / exon junction of exon 3. In other embodiments, an AC can bind to a target transcript at various other locations to induce exon skipping and result in nonsense-mediated decay of the target transcript (discussed elsewhere). The number of nucleotides in exon 3 is not divisible by 3 (e.g., 52, 106, 232, 365, etc.). Binding of an AC to an intron / exon junction induces exon skipping of exon 3 through various possible mechanisms. For example, binding of an AC to an intron / exon junction prevents the splicing machinery from accessing splicing elements. Additionally or alternatively, binding of an AC to an intron / exon junction prevents completion of one or both of the transesterification reactions required to complete the splicing process. As a result of AC binding to the target transcript, exon 3 is skipped and the resulting transcript contains exon 2 linked to exon 4. As a result of AC binding to the target transcript and skipping of exon 3, the reading frame of the resulting transcript shifts in exon 4. The reading frame shift in the illustrated embodiment introduces a PTC into the resulting transcript. As a result of AC binding to the target transcript, skipping exon 3 and exon 4 with the PTC, the resulting transcript is targeted and undergoes nonsense-mediated decay.

[0079] Determining the target sequence and designing an antisense compound (AC) for inducing exon skipping can be achieved using a variety of different methods, including, for example, those disclosed by Aartsma-Rus, A. et al., Molecular Therapy (2008), 17(3) 548-553 and Aartsma-Rus, A. et al., RNA (2007), 13(10) 1609-1624. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of a splice element (SE) of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising the entire SE of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SEs of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SEs of the target transcript and intervening sequences between the SEs.

[0080] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the SRE of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising the entire SRE of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of the target transcript and intervening sequences between the SREs.

[0081] In embodiments, the target nucleotide sequence comprises the entire SE and / or SRE and one or more flanking sequences upstream and / or downstream of the SE and / or SRE of the target transcript. In embodiments, the target nucleotide sequence comprises a portion, but not all, of the SE and / or SRE and one or more flanking sequences upstream and / or downstream of the SE and / or SRE of the target transcript.

[0082] In embodiments, the flanking sequence comprises 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, or 20 or more bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 25 or less, 20 or less, 15 or less, 10 or less, 5 or less, 4 or less, 3 or less, or 2 or less bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 1 to 4, 1 to 3, or 1 to 2 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 2 to 25, 2 to 20, 2 to 15, 2 to 10, 2 to 5, 2 to 4, or 2 to 3 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 3 to 25, 3 to 20, 3 to 15, 3 to 10, 3 to 5, or 3 to 4 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 4 to 25, 4 to 20, 4 to 15, 4 to 10, or 4 to 5 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 5 to 25, 5 to 20, 5 to 15, or 5 to 10 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 10 to 25, 10 to 20, or 10 to 15 bases on either or both sides of the SE and / or SRE. In aspects, the flanking sequence comprises 15 to 25 or 15 to 20 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequence comprises 20 to 25 bases on either or both sides of the SE and / or SRE. In embodiments, the flanking sequences include intervening sequences or portions thereof.

[0083] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the 5' ss of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an exon / intron junction of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the 3' ss of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of a Py tract, BPS, terminal "AG", and / or intron / exon junction of the target transcript.

[0084] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of a splice regulatory element (SRE) of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising all of the SREs of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of the target transcript and intervening sequences between the SREs of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ESE of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ISE. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ESS of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ISS of the target transcript.

[0085] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the terminal stem-loop (TLS) of the target transcript(s).

[0086] In embodiments, the AC hybridizes with at least a portion of an aberrant SE and / or SRE of the target transcript, wherein the aberrant SE and / or SRE arises from a mutation in the target gene.

[0087] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an SE and / or SRE, an exon / intron junction, or an intron / exon junction of the target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising an aberrant fusion junction resulting from a rearrangement or deletion of the target transcript. In embodiments, the AC hybridizes to a specific exon in an alternatively spliced ​​mRNA of the target transcript.

[0088] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of a splice element (SE) of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising the entire SE of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SEs of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SEs of a target transcript and intervening sequences between the SEs of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to at least a portion of an SE of an IRF-5, GYS1, and / or DUX4 target transcript and one or more flanking sequences of the SE.

[0089] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the 5' ss of an IRF-5 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an exon / intron junction of an IRF-5 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the 3' ss of an IRF-5 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the Py tract, BPS, terminal "AG", and / or intron / exon junction of an IRF-5 target transcript.

[0090] In embodiments, the AC binds to a target nucleotide sequence that does not contain at least a portion of an SE or at least a portion of an SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence that is sufficiently close to an SE and / or SRE to modulate splicing of the target transcript. In embodiments, an AC that binds to a target nucleotide sequence that does not contain at least a portion of an SE or at least a portion of an SRE of the target transcript and modulates splicing of the target transcript can bind to the target transcript and sterically block binding of a translation factor or trans-acting regulatory factor to the SE or SRE.

[0091] In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, or 20 or more nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 25 or less, 20 or less, 15 or less, 10 or less, 5 or less, 4 or less, 3 or less, or 2 or less nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 1 to 4, 1 to 3, or 1 to 2 nucleotides that form the 5' end and / or the 3' end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or a 5'-end that is 2 to 25, 2 to 20, 2 to 15, 2 to 10, 2 to 5, 2 to 4, or 2 to 3 nucleotides, forming the 5'-end and / or the 3'-end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or a 5'-end that is 3 to 25, 3 to 20, 3 to 15, 3 to 10, 3 to 5, or 3 to 4 nucleotides, forming the 5'-end and / or the 3'-end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or a 5'-end that is 4 to 25, 4 to 20, 4 to 15, 4 to 10, or 4 to 5 nucleotides, forming the 5'-end and / or the 3'-end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' and / or 5' end that is 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleotides from the 5' and / or 3' end of the SE and / or SRE of the target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' and / or 5' end that is 10 to 25 or 10 to 20 nucleotides from the 5' and / or 3' end of the SE and / or SRE of the target transcript.In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 20-25 nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of the target transcript.

[0092] In embodiments, the AC hybridizes to a target nucleotide sequence that is about 5 to about 50 nucleic acids in length. In embodiments, the AC is the same length as the target nucleotide sequence. In embodiments, the AC is a different length from the target nucleotide sequence. In embodiments, the AC is longer than the target nucleic acid sequence.

[0093] In embodiments, AC is 5 or more, 10 or more, 15 or more, 20 or more, 25 or more, 30 or more, 35 or more, 40 or more, or 45 or more nucleic acids in length. In embodiments, AC is 50 or less, 45 or less, 40 or less, 35 or less, 30 or less, 25 or less, 20 or less, 15 or less, or 10 or less nucleic acids in length. In embodiments, AC is 5 to 50, 5 to 45, 5 to 40, 5 to 35, 5 to 30, 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleic acids in length. In embodiments, AC is 10 to 50, 10 to 45, 10 to 40, 10 to 35, 10 to 30, 10 to 25, 10 to 20, or 10 to 15 nucleic acids in length. In embodiments, the AC is 15 to 50, 15 to 45, 15 to 40, 15 to 35, 15 to 30, 15 to 25, or 15 to 20 nucleic acids in length. In embodiments, the AC is 20 to 50, 20 to 45, 20 to 40, 20 to 35, 20 to 30, or 20 to 25 nucleic acids in length. In embodiments, the AC is 25 to 50, 25 to 45, 25 to 40, 25 to 35, or 25 to 30 nucleic acids in length. In embodiments, the AC is 30 to 50, 30 to 45, 30 to 40, or 30 to 35 nucleic acids in length. In embodiments, the AC is 35 to 50, 35 to 45, or 35 to 40 nucleic acids in length. In embodiments, the AC is 40 to 50 or 40 to 45 nucleic acids in length. In embodiments, the AC is 45 to 50 nucleic acids in length. In embodiments, the AC is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 nucleic acids in length.

[0094] In embodiments, the AC has 100% complementarity to the target nucleotide sequence. In embodiments, the AC does not have 100% complementarity to the target nucleotide sequence. As used herein, the term "percent complementarity" refers to the number of nucleobases of the AC that have nucleobase complementarity with corresponding nucleobases of an oligomeric compound or nucleic acid (e.g., a target nucleotide sequence), divided by the total length (number of nucleobases) of the AC. Those skilled in the art will understand that the inclusion of mismatches is possible without eliminating the activity of the antisense compound.

[0095] In embodiments, the AC contains 20% or less, 15% or less, 10% or less, 5% or less, or zero mismatches to the target nucleotide sequence. In some embodiments, the AC contains 5% or more, 10% or more, or 15% or more mismatches to the target nucleotide sequence. In embodiments, the AC contains 0-5%, 0-10%, 0-15%, or 0-20% mismatches to the target nucleotide sequence. In embodiments, the AC contains 5%-10%, 5%-15%, or 5%-20% mismatches to the target nucleotide sequence. In embodiments, the AC contains 10%-15% or 10%-20% mismatches to the target nucleotide sequence. In embodiments, the AC contains 10%-20% mismatches to the target nucleotide sequence.

[0096] In embodiments, the AC has 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementarity to the target nucleotide sequence. In embodiments, the AC has 100% or less, 99% or less, 98% or less, 97% or less, 96% or less, 95% or less, 90% or less, or 85% or less complementarity to the target nucleotide sequence. In embodiments, the AC has 80% to 100%, 80% to 99%, 80% to 98%, 80% to 97%, 80% to 96%, 80% to 95%, 80% to 90%, or 80% to 85% complementarity to the target nucleotide sequence. In embodiments, the AC has 85% to 100%, 85% to 99%, 85% to 98%, 85% to 97%, 85% to 96%, 85% to 95%, or 85% to 90% complementarity to the target nucleotide sequence. In embodiments, the AC has 90% to 100%, 90% to 99%, 90% to 98%, 90% to 97%, 90% to 96%, or 90% to 95% complementarity to the target nucleotide sequence. In embodiments, the AC has 95% to 100%, 95% to 99%, 95% to 98%, 95% to 97%, or 95% to 96% complementarity to the target nucleotide sequence. In embodiments, the AC has 96% to 100%, 96% to 99%, 96% to 98%, or 96% to 97% complementarity to the target nucleotide sequence. In embodiments, the AC has 97% to 100%, 97% to 99%, or 97% to 98% complementarity to the target nucleotide sequence. In embodiments, the AC has 98% to 100% or 98% to 99% complementarity to the target nucleotide sequence. In embodiments, the AC has 99% to 100% complementarity to the target nucleotide sequence. The percent complementarity of an oligonucleotide is calculated by dividing the number of complementary nucleobases by the total number of nucleobases in the oligonucleotide.

[0097] In embodiments, the AC contains 1, 2, 3, 4, or 5 mismatches with the target nucleic acid sequence to which it hybridizes. In embodiments, the AC contains 1 or 2 mismatches with the target nucleic acid sequence to which it hybridizes. In embodiments, the AC contains no mismatches with the target nucleic acid sequence to which it hybridizes.

[0098] The incorporation of nucleotide affinity modifications can allow a greater number of mismatches compared to unmodified compounds.Similarly, certain oligonucleotide sequences may be more tolerant of mismatches than other oligonucleotide sequences.Those skilled in the art can determine the appropriate number of mismatches between AC and target nucleotide sequence, for example, by determining the thermal melting temperature (Tm).Tm or ΔTm can be calculated by techniques well known to those skilled in the art.For example, the technique described in Freier et al. (Nucleic Acids Research, 1997, 25, 22: 4429-4443) allows those skilled in the art to evaluate nucleotide modifications for their ability to increase the melting temperature of RNA:DNA duplexes.

[0099] In embodiments, the AC comprises a sequence that hybridizes to the target transcript under stringent conditions and a sequence that does not hybridize to the target transcript under stringent conditions. In embodiments, the AC comprises a first sequence that does not hybridize to the target sequence under stringent conditions, a second sequence that does not hybridize to the target sequence under stringent conditions, and a third sequence that hybridizes to the target sequence under stringent conditions, the third sequence being located between the first and second sequences.

[0100] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of a splice regulatory element (SRE) of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising the entire SRE of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising multiple SREs of an IRF-5, GYS1, and / or DUX4 target transcript and intervening sequences between the SREs. In embodiments, the AC hybridizes to at least a portion of an SE of an IRF-5, GYS1, and / or DUX4 target transcript and one or more flanking sequences of the SE.

[0101] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ESE of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ISE of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ESS of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an ISS of an IRF-5, GYS1, and / or DUX4 target transcript.

[0102] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of the terminal stem loop (TLS) of an IRF-5, GYS1, and / or DUX4 target transcript.

[0103] In an embodiment, the AC hybridizes with at least a portion of an aberrant SE and / or SRE of an IRF-5, GYS1, and / or DUX4 target transcript, wherein the aberrant SE and / or SRE arise from a mutation in the IRF-5, GYS1, and / or DUX4 target transcript.

[0104] In embodiments, the AC hybridizes to a target nucleotide sequence comprising at least a portion of an exon-exon junction, an intron-exon junction, and / or an exon-intron junction of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a target nucleotide sequence comprising an aberrant fusion junction resulting from rearrangement or deletion of a portion of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC hybridizes to a specific exon in an alternatively spliced ​​mRNA in an IRF-5, GYS1, and / or DUX4 target transcript.

[0105] In embodiments, the AC binds to a target nucleotide sequence that does not include at least a portion of the SE or at least a portion of the SRE of an ISS of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence that is sufficiently proximal to the SE and / or SRE to modulate splicing of the ISS of an IRF-5, GYS1, and / or DUX4 target transcript.

[0106] In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, 10 or more, 15 or more, or 20 or more nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 25 or less, 20 or less, 15 or less, 10 or less, 5 or less, 4 or less, 3 or less, or 2 or less nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 1 to 25, 1 to 20, 1 to 15, 1 to 10, 1 to 5, 1 to 4, 1 to 3, or 1 to 2 nucleotides from the 5' end and / or 3' end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 2 to 25, 2 to 20, 2 to 15, 2 to 10, 2 to 5, 2 to 4, or 2 to 3 nucleotides from the 5' end and / or 3' end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or a 5'-end that is 3 to 25, 3 to 20, 3 to 15, 3 to 10, 3 to 5, or 3 to 4 nucleotides forming the 5'-end and / or the 3'-end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or a 5'-end that is 4 to 25, 4 to 20, 4 to 15, 4 to 10, or 4 to 5 nucleotides forming the 5'-end and / or the 3'-end of the SE and / or SRE of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3' end and / or a 5' end that is 5 to 25, 5 to 20, 5 to 15, or 5 to 10 nucleotides from the 5' end and / or the 3' end of the SE and / or SRE of an IRF-5, GYS1, and / or DUX4 target transcript.In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or 5'-end that is 10-25 or 10-20 nucleotides from the 5'-end and / or 3'-end of the SE and / or SRE of an IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC binds to a target nucleotide sequence having a 3'-end and / or 5'-end that is 20-25 nucleotides from the 5'-end and / or 3'-end of the SE and / or SRE of an IRF-5, GYS1, and / or DUX4 target transcript.

[0107] In embodiments, the AC hybridizes to a target nucleotide sequence of an IRF-5, GYS1, and / or DUX4 target transcript that is about 5 to about 50 nucleic acids in length. In embodiments, the AC is the same length as the target nucleotide sequence. In embodiments, the AC is a different length than the target nucleotide sequence. In embodiments, the AC is longer than the target nucleic acid sequence.

[0108] In embodiments, the AC has 100% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC does not have 100% complementarity to the target nucleotide sequence. As used herein, the term "percent complementarity" refers to the number of nucleobases of the AC that have nucleobase complementarity with the corresponding nucleobases of an oligomeric compound or nucleic acid (e.g., a target nucleotide sequence), divided by the total length (number of nucleobases) of the AC. Those skilled in the art will understand that the inclusion of mismatches is possible without eliminating the activity of the antisense compound.

[0109] In embodiments, the AC contains 20% or less, 15% or less, 10% or less, 5% or less, or zero mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In some embodiments, the AC contains 5% or more, 10% or more, or 15% or more mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC contains 0-5%, 0-10%, 0-15%, or 0-20% mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC contains 5%-10%, 5%-15%, or 5%-20% mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC contains 10% to 15% or 10% to 20% mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC contains 10% to 20% mismatches to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript.

[0110] In embodiments, the AC has 80% or more, 85% or more, 90% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 100% or less, 99% or less, 98% or less, 97% or less, 96% or less, 95% or less, 90% or less, or 85% or less complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 80% to 100%, 80% to 99%, 80% to 98%, 80% to 97%, 80% to 96%, 80% to 95%, 80% to 90%, or 80% to 85% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 85% to 100%, 85% to 99%, 85% to 98%, 85% to 97%, 85% to 96%, 85% to 95%, or 85% to 90% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 90% to 100%, 90% to 99%, 90% to 98%, 90% to 97%, 90% to 96%, or 90% to 95% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 95% to 100%, 95% to 99%, 95% to 98%, 95% to 97%, or 95% to 96% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 96% to 100%, 96% to 99%, 96% to 98%, or 96% to 97% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 97% to 100%, 97% to 99%, or 97% to 98% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript. In embodiments, the AC has 98% to 100% or 98% to 99% complementarity to the target nucleotide sequence. In embodiments, the AC has 99% to 100% complementarity to the target nucleotide sequence of the IRF-5, GYS1, and / or DUX4 target transcript.The percent complementarity of an oligonucleotide is calculated by dividing the number of complementary nucleobases by the total number of nucleobases in the oligonucleotide.

[0111] Antisense mechanism In embodiments, the AC regulates one or more aspects of protein transcription, translation, and expression. In embodiments, hybridization of the AC to a target nucleotide sequence of a target transcript regulates one or more aspects of pre-mRNA splicing. In embodiments, hybridization of the AC to a target nucleotide sequence of a target transcript restores native splicing to a mutant transcript sequence. In embodiments, hybridization of the AC to a target nucleotide sequence of a target transcript results in alternative splicing of the target transcript.

[0112] In embodiments, AC hybridization results in exon inclusion or exon skipping of one or more exons. In embodiments, exon skipping increases the activity of a protein expressed from the resulting mRNA. In embodiments, exon skipping decreases the activity of a protein expressed from the resulting mRNA. In embodiments, skipping one or more exons induces a frameshift in the mRNA transcript. In embodiments, the frameshift results in an mRNA encoding a protein with reduced activity. In embodiments, the frameshift results in a truncated or non-functional protein. In embodiments, skipping one or more exons results in the introduction of a premature stop codon in the mRNA. In embodiments, skipping one or more exons results in degradation of the mRNA transcript by nonsense-mediated decay. In embodiments, the skipped exon sequence comprises a nucleic acid deletion, substitution, or insertion. In embodiments, the skipped exon does not comprise a sequence mutation. In embodiments, antisense oligonucleotide hybridization to a target nucleotide sequence within a target pre-mRNA transcript results in the expression of a distinct protein isoform.

[0113] In embodiments, AC hybridization of the target transcript to the target nucleotide sequence prevents the inclusion of intron sequences in the mature mRNA molecule. In embodiments, AC hybridization of the target transcript to the target nucleotide sequence results in increased expression of a protein isoform. In embodiments, AC hybridization of the target transcript to the target nucleotide sequence results in decreased expression of a protein isoform. In embodiments, AC hybridization of the target transcript to the target nucleotide sequence results in the expression of a re-spliced ​​protein, including an inactive fragment of the protein.

[0114] In embodiments, the AC comprises DNA, and hybridization of the AC to the target transcript results in RNase H-mediated transcript degradation. In embodiments, the AC comprises a nucleotide modification designed to not support RNase H activity. Nucleotide modifications of antisense compounds that do not support RNase H activity are known, including, but not limited to, 2'-O-methoxyethyl / phosphorothioate (MOE) modifications. Advantageously, ACs with MOE modifications have increased affinity for target RNA and increased nuclease stability.

[0115] In embodiments, ACs can regulate transcription, translation, or protein expression through steric blocking. The following review articles describe the mechanisms and applications of steric blocking: Roberts et al., Nature Reviews Drug Discovery (2020) 19:673-694, which are incorporated herein by reference in their entirety.

[0116] The effectiveness of ACs can be evaluated by assessing the antisense activity resulting from their administration. As used herein, the term "antisense activity" refers to any detectable and / or measurable activity resulting from hybridization of an antisense compound to its target nucleotide sequence. Such detection and / or measurement may be direct or indirect. In embodiments, antisense activity is assessed by detecting and / or measuring the amount of a protein expressed from a transcript of interest. In embodiments, antisense activity is assessed by detecting and / or measuring the amount of a transcript of interest. In embodiments, antisense activity is assessed by detecting and / or measuring the amount of alternatively spliced ​​RNA and / or the amount of a protein isoform translated from a target transcript. In embodiments, antisense activity is assessed by detecting and / or measuring the amount of a downstream transcript and / or protein regulated by a gene of interest.

[0117] Antisense compound design The design of the AC depends on the target gene. Targeting an AC to a specific target nucleotide sequence can be a multi-step process. This process usually begins with identifying the gene of interest. The transcript of the gene of interest is analyzed to identify the target nucleotide sequence. In an embodiment, the target nucleotide sequence includes at least a portion of a splice element and / or a splice regulatory element. In an embodiment, the target gene is IRF-5. In an embodiment, the target gene is GYS1. In an embodiment, the target gene is DUX4.

[0118] Those skilled in the art will be able to design, synthesize, and screen ACs of different nucleobase sequences to identify sequences that confer antisense activity. For example, antisense compounds can be designed to inhibit the expression of target genes. Methods for designing, synthesizing, and screening ACs for antisense activity against preselected target nucleic acids and / or target genes can be found, for example, in "Antisense Drug Technology, Principles, Strategies, and Applications," edited by Stanley T. Crooke, CRC Press, Boca Raton, Florida, which is incorporated by reference in its entirety for any purpose.

[0119] AC structure ACs include oligonucleotides and / or oligonucleosides. Oligonucleotides and / or oligonucleosides are nucleosides linked via internucleoside linkages. Nucleosides comprise a pentose sugar (e.g., ribose or deoxyribose) and a nitrogenous base covalently linked to the sugar. Naturally occurring (traditional) bases found in DNA and / or RNA are adenine (A), guanine (G), thymine (T), cytosine (C), and uracil (U). Naturally occurring (traditional) sugars found in DNA and / or RNA are deoxyribose (DNA) and ribose (RNA). The naturally occurring (traditional) nucleoside bond is a phosphodiester bond. In embodiments, ACs of the present disclosure can have all natural sugars, bases, and internucleoside linkages.

[0120] Chemically modified nucleosides are routinely incorporated into antisense compounds to enhance one or more properties, such as nuclease resistance, pharmacokinetics, or affinity for target RNA. In embodiments, the ACs of the present disclosure may have one or more modified nucleosides. In embodiments, the ACs of the present disclosure may have one or more modified sugars. In embodiments, the ACs of the present disclosure may have one or more modified bases. In embodiments, the ACs of the present disclosure may have one or more modified internucleoside linkages.

[0121] Generally, a nucleobase is any group containing one or more atoms or groups of atoms that can hydrogen bond to the base of another nucleic acid. In addition to "unmodified" or "natural" nucleobases (A, G, T, C, and U), many modified nucleobases or nucleobase mimics are known to those skilled in the art and are suitable for the compounds described herein. Generally, modified nucleobases refer to nucleobases that are structurally very similar to their parent nucleobases, such as 7-deazapurine, 5-methylcytosine, 2-thio-dT (Figure 3), or G-clamp. Generally, nucleobase mimics are nucleobases that contain more complex structures than modified nucleobases (e.g., tricyclic phenoxazine nucleobase mimics, etc.). Methods for preparing the above modified nucleobases are well known to those skilled in the art.

[0122] In embodiments, an AC may include one or more nucleosides with modified sugar moieties. In embodiments, the furanosyl sugar of a natural nucleoside may have 2' modifications, modifications to create constrained nucleosides, and so forth (see Figure 3). For example, in embodiments, the furanosyl sugar ring of a natural nucleoside may be modified in several ways, including, but not limited to, adding a substituent, bridging two non-geminal ring atoms to form a bicyclic nucleic acid (BNA) or locked nucleic acid, replacing an oxygen of the furanosyl ring with a C or N, and / or substituting such atoms or groups (see Figure 3). Modified sugars are well known and can be used to increase or decrease the affinity of an AC for its target nucleotide sequence. Modified sugars can also be used to increase the AC's resistance to nucleases. Sugars can also be substituted with, among other things, sugar mimetic groups. In embodiments, one or more sugars of the nucleosides of an AC are substituted with a methylenemorpholine ring, shown as 19 in Figure 3.

[0123] In embodiments, the AC may include one or more nucleosides having modified bicyclic sugars (BNAs, sometimes called bridged nucleic acids). Examples of BNAs include, but are not limited to, LNA (4'-(CH)-O-2' bridge), 2'-thio-LNA (4'-(CH)-S-2' bridge), 2'-amino-LNA (4'-(CH)-NR-2' bridge), ENA (4'-(CH)-O-2' bridge), 4'-(CH)-2' bridged BNA, 4'-(CHCH(CH))-2' bridged BNA, cEt (4'-(CH(CH)-O-2' bridge), and cMOE BNA (4'-(CH(CHOCH)-O-2' bridge). Some examples are shown in Figure 3. BNAs have been prepared and are disclosed in the patent and scientific literature (e.g., Srivastava et al., J. Am. Chem. Soc. (2007), ACS Advanced Online). publication, 10.1021 / ja071106y; Albaek et al., J. Org. Chem. (2006), 71, 7731-7740; Fluiter et al., Chembiochem (2005), 6, 1104-1109; Singh et al., Chem. Commun. (1998), 4, 455-456; Koshkin et al., Tetrahedron (1998), 54, 3607-3630; Wahlestedt et al., Proc. Natl. Acad. Sci. USA (2000), 97, 5633-5638; Kumar et al., Bioorg. Med. Chem. Lett. (1998), 8, 2219-2222; WO 94 / 14226, WO 2005 / 021570, Singh et al. al., J. Org. Chem. (1998), 63, 10035-10039; International Publication No. 2007 / 090071, U.S. Patent Nos. 7,053,207; 6,268,490, 6,770,748, 6,794,499, 7,034,133, and 6,525,191, and U.S. Patent Publication Nos. 2004-0171570, 2004-0219565, 2004-0014959, 2003-0207841, 2004-0143114, and 20030082807.

[0124] In embodiments, the AC comprises one or more nucleosides, including locked nucleic acids (LNAs), in which the 2'-hydroxyl group of the ribosyl sugar ring is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage to form a bicyclic sugar moiety (see also Eayadi et al., Curr. Opinion Invens. Drugs (2001), 2,558-561; Braasch et al., Chem. Biol. (2001), 8,1-7; and Orum et al., Curr. Opinion Mol. Ther. (2001), 3,239-243; U.S. Patent Nos. 6,268,490 and 6,670,461). The bond can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom; the term LNA is used for the bicyclic moiety, and the term ENA™ is used for an ethylene group at this position (Singh et al., Chem. Commun. (1998), 4, 455-456; ENA™: Morita et al., Bioorganic Medicinal Chemistry (2003), 11, 2211-2226). LNA and other bicyclic sugar analogs exhibit very high duplex thermal stability (Tm = +3 to +10°C) for complementary DNA and RNA, stability against 3'-exonuclease degradation, and good solubility properties. Potent, non-toxic antisense oligonucleotides containing LNA have been described (Wahlestedt et al. Natl. Acad. Sci. USA (2000), 97, 5633-5638).

[0125] A similarly studied LNA isomer is alpha-L-LNA, which has been shown to have excellent stability against 3'-exonucleases. Alpha-L-LNA has been incorporated into antisense gapmers and chimeras that exhibit potent antisense activity (Frieden et al., Nucleic Acids Research (2003), 21, 6365-6372).

[0126] The synthesis and preparation of LNA monomers adenine, cytosine, guanine, 5-methyl-cytosine, thymine and uracil, as well as their oligomerization and nucleic acid recognition properties, have been described (Koshkin et al., Tetrahedron (1998), 54, 3607-3630). LNAs and their preparation are also described in WO 98 / 39352 and WO 99 / 14226.

[0127] Analogs of LNA, such as phosphorothioate-LNA and 2'-thio-LNA, have also been prepared (Kumar et al., Bioorg. Med. Chem. Lett., 1998, 8, 2219-2222). The preparation of LNA analogs containing oligodeoxyribonucleotide duplexes as substrates for nucleic acid polymerases has also been described (WO 99 / 14226). Furthermore, the synthesis of 2'-amino-LNA, a conformationally restricted, high-affinity oligonucleotide analog, has been described (Singh et al., J. Org. Chem. (1998), 63, 10035-10039). Furthermore, 2'-amino and 2'-methylamino-LNA have been prepared, and the thermal stability of their duplexes with complementary RNA and DNA strands has previously been reported.

[0128] In embodiments, the antisense compound is "tricyclo-DNA (tc-DNA)," which refers to a class of constrained DNA analogs in which each nucleotide is modified by the introduction of a cyclopropane ring to restrict the conformational flexibility of the backbone and enhance the backbone geometry at a torsion angle γ. Homobasic adenine- and thymine-containing tc-DNA forms highly stable AT base pairs with complementary RNA.

[0129] Methods for preparing modified sugars are well known to those of skill in the art. Some representative patents and publications that teach the preparation of such modified sugars include U.S. Patent Nos. 4,981,957, 5,118,800, 5,319,080, 5,359,044, 5,393,878, 5,446,137, 5,466,786, 5,514,785, 5,519,134, 5,567,811, 5,576,427, Nos. 5,591,722, 5,597,909, 5,610,300, 5,627,053, 5,639,873, 5,646,265, 5,658,873, 5,670,633, 5,792,747, 5,700,920, and 6,600,032, and WO 2005 / 121371.

[0130] internucleoside bond Described herein are internucleoside linking groups that link nucleosides or otherwise modified monomer units together to thereby form oligonucleotides and / or oligonucleotides containing AC, which can include naturally occurring internucleoside linkages, non-natural internucleoside linkages, or both.

[0131] In naturally occurring DNA and RNA, the internucleoside linking group is phosphodiester, which covalently bonds adjacent nucleosides to one another to form a linear polymeric compound. In naturally occurring DNA and RNA, the phosphodiester is linked to the 2', 3', or 5' hydroxyl moiety of the sugar. Within oligonucleotides, the phosphate groups are generally considered to form the internucleoside backbone of the oligonucleotide. In naturally occurring DNA and RNA, the linkage or backbone of RNA and DNA is a 3' to 5' phosphodiester bond. In embodiments, the internucleoside linking group of AC is phosphodiester. In embodiments, the internucleoside linking group of AC is a 3' to 5' phosphodiester bond.

[0132] Two major classes of non-natural internucleoside linking groups are defined by the presence or absence of a phosphorus atom. Representative phosphorus-containing internucleoside linkages include, but are not limited to, phosphotriesters, methylphosphonates, phosphoramidates, and phosphorothioates. Representative non-phosphorus-containing internucleoside linking groups include, but are not limited to, methylenemethylimino (-CH2-N(CH3)-O-CH2-), thiodiesters (-OC(O)-S-), thionocarbamate (-OC(O)(NH)-S-), siloxanes (-O-Si(H2-O-), and N,N'-dimethylhydrazine (-CH2-N(CH3-N(CH3)-). ACs with phosphorus internucleoside linking groups are referred to as oligonucleotides. Antisense compounds with non-phosphorus internucleoside linking groups are Such compounds are called oligonucleosides. Modified internucleoside linkages can be used to alter (typically increase) the nuclease resistance of antisense compounds compared to natural phosphodiester linkages. Internucleoside linkages containing chiral atoms can be prepared as racemic, chiral, or mixtures. Representative chiral internucleoside linkages include, but are not limited to, alkylphosphonates and phosphorothioates. Methods for preparing phosphorus-containing and non-phosphorus-containing linkages are well known to those skilled in the art.

[0133] In some embodiments, two or more nucleosides with modified sugars and / or modified nucleobases can be linked using phosphoramidate. In some embodiments, two or more nucleosides with methylene morpholine rings can be connected via phosphoramidate internucleoside linkages, as shown in Figure 3, where B1 and B2 are modified nucleobases or natural nucleobases, as shown in 20. Antisense compounds that include nucleobases with methylene morpholine rings linked via phosphoramidate internucleoside linkages can be referred to as phosphoramidate morpholino oligomers (PMOs).

[0134] Conjugated groups In embodiments, ACs can be modified by the covalent attachment of one or more conjugate groups. Generally, conjugate groups modify one or more properties of the attached AC, including, but not limited to, pharmacodynamics, pharmacokinetics, binding, absorption, cellular distribution, cellular uptake, charge, and clearance. Conjugate groups are routinely used in chemistry and are linked to parent compounds, such as ACs, either directly or via an optional linking moiety or group. Conjugate groups include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, thioethers, polyethers, cholesterol, thiocholesterol, cholic acid moieties, folic acid, lipids, phospholipids, biotin, phenazine, phenanthridine, anthraquinone, adamantane, acridine, fluorescein, rhodamine, coumarin, and dyes. In embodiments, the conjugate group is polyethylene glycol (PEG), and the PEG is conjugated to either the AC or a CPP (CPPs are discussed elsewhere herein).

[0135] In embodiments, the conjugate group is a lipid moiety, e.g., a cholesterol moiety (Letsinger et al., Proc. Natl. Acad. Sci. USA (1989), 86, 6553); cholic acid (Manoharan et al., Bioorg. Med. Chem. Lett. (1994), 4, 1053); thioethers, e.g., hexyl-S-tritylthiol (Manoharan et al., Ann. NY Acad. Sci. (1992), 660, 306; Manoharan et al., Bioorg. Med. Chem. Let. (1993), 3, 2765); thiocholesterol (Oberhauser et al., Nucl. Acids Res. (1992), 20, 533); aliphatic chains, e.g., dodecanediol or undecyl residues (Saison-Behmoaras et al., EMBO J., 1991, 10, 111; Kabanov et al., FEBS Lett. (1990), 259, 327; Svinarchuk et al., Biochimie (1993), 75, 49; phospholipids, such as di-hexadecyl-rac-glycerol or triethylammonium-1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate (Manoharan et al., Tetrahedron Lett. (1995), 36, 3651; Shea et al., Nucl. Acids Res. (1990), 18, 3777); polyamines or polyethylene glycol chains (Manoharan et al., Nucleosides & Nucleotides (1995), 14, 969); adamantane acetic acid (Manoharan et al., Tetrahedron Lett. (1995), 36, 3651); palmityl moiety (Mishra et al., Biochim. Biophys. Acta. (1995), 1264, 229); or octadecylamine or hexylamino-carbonyl-oxycholesterol moiety (Crooke et al., J. Pharmacol. Exp. Ther. (1996), 277, 923).

[0136] Types of antisense compounds Various types of ACs may be used, including, for example, antisense oligonucleotides, siRNAs, microRNAs, antagomirs, aptamers, ribozymes, supermirs, miRNA mimics, miRNA inhibitors, or combinations thereof.

[0137] antisense oligonucleotides In various embodiments, the antisense compound (AC) is an antisense oligonucleotide (ASO) complementary to a target nucleotide sequence. The term "antisense oligonucleotide (ASO)" or simply "antisense" is meant to include an oligonucleotide complementary to a target nucleotide sequence. This term also encompasses ASOs that may not be perfectly complementary to a desired target nucleotide sequence. ASOs comprise a single strand of DNA and / or RNA complementary to a selected target nucleotide sequence or target gene. ASOs may contain one or more modified DNA and / or RNA bases, modified sugars, and / or non-natural internucleoside linkages. In embodiments, ASOs may contain one or more phosphoramidate internucleoside linkages. In embodiments, ASOs are phosphoramidate morpholino oligomers (PMOs). ASOs may have any characteristics, be any length, bind to any splice element, and effect any of the mechanisms described for ACs. In embodiments, the ASO induces exon skipping to introduce a premature stop codon, ultimately leading to nonsense-mediated decay of the target transcript. In embodiments, the ASO is a PMO and induces exon skipping to introduce a premature stop codon, ultimately leading to nonsense-mediated decay of the target transcript.

[0138] Antisense oligonucleotides have been demonstrated to be effective and targeted inhibitors of protein synthesis, and therefore can be used to specifically inhibit the protein synthesis of target genes.The effectiveness of ASOs for inhibiting protein synthesis has been well proven.To date, these compounds have shown promise in several in vitro and in vivo models, such as inflammatory disease, cancer, and HIV models (Agrawal, Trends in Biotech. (1996), 14:376-387).Antisense can also affect cellular activity by specifically hybridizing with chromosomal DNA.

[0139] Methods for producing ASOs are known in the art and can be easily adapted to produce ASOs that bind to the target nucleotide sequences of the present disclosure. Selection of an ASO sequence specific for a given target nucleotide sequence is based on analysis of the selected target nucleotide sequence and determination of its secondary structure, Tm, binding energy, and relative stability. Antisense oligonucleotides can be selected based on their relative inability to form dimers, hairpins, or other secondary structures that reduce or prevent specific binding to the target nucleotide sequence in host cells. These secondary structure analyses and target site selection studies can be performed, for example, using OLIGO primer analysis software version 4 (Molecular Biology Insights) and / or BLASTN 2.0.5 algorithm software (Altschul et al., Nucleic Acids Res. 1997, 25(17):3389-402).

[0140] RNA interference In embodiments, the AC comprises a molecule that mediates RNA interference (RNAi). As used herein, the phrase "mediating RNAi" refers to the ability to silence a target transcript in a sequence-specific manner. Without wishing to be bound by theory, it is believed that the silencing uses the RNAi mechanism or process and guide RNAs, such as siRNAs and / or miRNA compounds of about 21 to about 23 nucleotides. In embodiments, the AC targets the target transcript for degradation. Thus, in embodiments, RNAi molecules can be used to inhibit the expression of a gene or polynucleotide of interest. In embodiments, RNAi molecules are used to induce degradation of a target transcript, such as a pre-mRNA or mature mRNA.

[0141] In embodiments, the AC comprises a small interfering RNA (siRNA) that induces an RNAi response. In embodiments, the AC comprises a microRNA (miRNA) that induces an RNAi response.

[0142] Small interfering RNAs (siRNAs) are nucleic acid duplexes, typically about 16 to about 30 nucleotides long, that can associate with a cytoplasmic multiprotein complex known as the RNAi-induced silencing complex (RISC). Because RISC loaded with siRNA mediates the degradation of homologous transcripts, siRNAs can be designed to knock down protein expression with high specificity. Unlike other antisense technologies, siRNAs function through a natural mechanism that evolved to control gene expression via non-coding RNA. Various RNAi reagents, including siRNAs targeting clinically relevant targets, are currently in pharmaceutical development, as described, for example, in de Fougerolles, A. et al., Nature Reviews (2007) 6:443-453.

[0143] Because siRNA and miRNA constructs can be synthesized using any nucleotide sequence directed against a target protein, the therapeutic applications of RNAi are extremely broad. To date, siRNA constructs have demonstrated the ability to specifically downregulate target proteins in both in vitro and in vivo models, as well as in clinical studies.

[0144] The first RNAi molecules described were RNA:RNA hybrids containing both RNA sense and RNA antisense strands, but it has now been demonstrated that DNA sense:RNA antisense hybrids, RNA sense:DNA antisense hybrids, and DNA:DNA hybrids can mediate RNAi (Lamberton, JS and Christian, AT, (2003) J. Molecular Biotechnology 24:111-119). For RNA hybrids containing both RNA sense and RNA antisense strands, DNA sense:RNA antisense hybrids, RNA sense:DNA antisense hybrids, and DNA:DNA hybrids have been shown to be able to mediate RNAi (Lamberton, JS and Christian, AT, Molecular Biotechnology (2003), 24:111-119). In embodiments, RNAi molecules containing any of these different types of double-stranded molecules are used. Furthermore, it is understood that RNAi molecules can be used and introduced into cells in various forms. Thus, as described herein, RNAi molecules include any and all molecules capable of mediating RNAi in cells, including, but not limited to, double-stranded oligonucleotides comprising two separate strands (i.e., a sense strand and an antisense strand) (e.g., small interfering RNA (siRNA)); double-stranded oligonucleotides comprising two separate strands linked to each other by a non-nucleotidyl linker; oligonucleotides comprising a hairpin loop of complementary sequences that form a double-stranded region (e.g., shRNAi molecules), and expression vectors that express one or more polynucleotides that can form a double-stranded polynucleotide, either alone or in combination with another polynucleotide.

[0145] As used herein, "single-stranded siRNA compound" refers to a siRNA compound that is composed of a single molecule.It can comprise a double-stranded region formed by intrastrand pairing, for example, it can be or comprise a hairpin or panhandle structure.Single-stranded siRNA compound can be antisense with respect to target molecule.

[0146] Single-stranded siRNA compound can be long enough to enter RISC and participate in the RISC-mediated cleavage of target mRNA.Single-stranded siRNA compound is at least about 14, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, or at most about 50 nucleotides long.In certain embodiments, single-stranded siRNA is less than about 200, less than about 100, or less than about 60 nucleotides long.

[0147] Hairpin siRNA compounds can have a duplex region of 17, 18, 19, 20, 21, 22, 23, 24, or 25, or at least about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, or about 25, nucleotide pairs. The duplex region can be about 200 or less, about 100 or less, or about 50 or less nucleotide pairs in length. In certain embodiments, the duplex region ranges from about 15 to about 30, about 17 to about 23, about 19 to about 23, and about 19 to about 21 nucleotide pairs in length. The hairpin can have a single-stranded overhang or terminal unpaired region. In certain embodiments, the overhang is about 2 to about 3 nucleotides in length. In embodiments, the overhangs are on the same side of the hairpin, and in certain embodiments, on the antisense side of the hairpin.

[0148] A "double-stranded siRNA compound," as used herein, is an siRNA compound that contains two or more (and in some cases two) strands capable of interstrand hybridization to form a region of double-stranded structure.

[0149] The antisense strand of a double-stranded siRNA compound can be 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, or 60 nucleotides in length, or at least about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 40, or about 60 nucleotides in length. It can be about 200 or less, about 100 or less, or about 50 or less nucleotides in length. The range can be about 17 to about 25, about 19 to about 23, and about 19 to about 21 nucleotides in length. As used herein, the term "antisense strand" refers to the strand of a siRNA compound that is sufficiently complementary to a target molecule (e.g., a target nucleotide sequence of a target transcript).

[0150] The sense strand of a double-stranded siRNA compound can be at least about 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, or 60 nucleotides in length. It can be no more than about 200, no more than about 100, or no more than about 50 nucleotides in length. Ranges can be from about 17 to about 25, from about 19 to about 23, and from about 19 to about 21 nucleotides in length.

[0151] The double-stranded portion of a double-stranded siRNA compound can be 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 60, or at least about 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 40, or 60, nucleotide pairs long, and can be up to about 200, up to about 100, or about 50 nucleotide pairs long. Ranges can be from about 15 to about 30, from about 17 to about 23, from about 19 to about 23, and from about 19 to about 21 nucleotide pairs long.

[0152] In embodiments, the siRNA compound is large enough that it can be cleaved by an endogenous molecule (eg, by Dicer) to produce smaller siRNA compounds (eg, siRNA agents).

[0153] The sense strand and the antisense strand can be selected so that the double-stranded siRNA compound comprises a single-stranded or unpaired region at one or both ends of the molecule.Therefore, the double-stranded siRNA compound can contain the sense strand and the antisense strand that are paired to contain an overhang, for example, one or two 5' or 3' overhangs, or a 3' overhang of 1 to 3 nucleotides.The overhang can be the result of one strand being longer than the other strand, or the result of two strands of the same length being offset.In some embodiments, it can have at least one 3' overhang.In embodiments, both ends of the siRNA molecule can have a 3' overhang.In embodiments, the overhang is 2 nucleotides.

[0154] In embodiments, the length of the duplexed region ranges from about 15 to about 30, or about 18, about 19, about 20, about 21, about 22, or about 23 nucleotides in length, for example, the ssiRNA (siRNA with sticky overhang) compounds described above. ssiRNA compounds can be similar in length and can be constructed as natural Dicer-processed products from long dsiRNAs. Also included are embodiments in which the two strands of the ssiRNA compound are linked (e.g., covalently linked). In embodiments, a hairpin or other single-stranded structure is included to provide a duplex region and a 3' overhang.

[0155] The siRNA compound described herein, for example, double-stranded siRNA compound or single-stranded siRNA compound, can mediate the silencing of target RNA, for example, mRNA, for example, the transcript of the gene encoding protein.For convenience, in this specification, this mRNA is also referred to as the mRNA to be silenced.This gene is also referred to as target gene.Generally, the RNA to be silenced is endogenous gene.

[0156] In embodiments, the siRNA compound is "sufficiently complementary" to the target transcript such that the siRNA compound silences the production of the protein encoded by the target mRNA. In embodiments, the siRNA compound is "sufficiently complementary" to at least a portion of the target transcript such that the siRNA compound silences the production of the gene product encoded by the target transcript. In another embodiment, the siRNA compound is "exactly complementary" to the target nucleotide sequence (e.g., a portion of the target transcript) such that the target nucleotide sequence and the siRNA compound anneal to form a hybrid consisting solely of Watson-Crick base pairs in the exact complementary region. "Sufficiently complementary" to the target nucleotide sequence may include an internal region (e.g., at least about 10 nucleotides) that is exactly complementary to the target nucleotide sequence. Furthermore, in certain embodiments, the siRNA compound specifically discriminates between single nucleotide differences. In this case, the siRNA compound mediates RNAi only when exact complementarity is found in the region of the single nucleotide difference (e.g., within 7 nucleotides of the region).

[0157] Because siRNA and miRNA constructs can be synthesized using any nucleotide sequence directed against a target protein, the therapeutic applications of RNAi are extremely broad. To date, siRNA constructs have demonstrated the ability to specifically downregulate target proteins in both in vitro and in vivo models, as well as in clinical studies.

[0158] microRNA In embodiments, the AC comprises a microRNA molecule. MicroRNAs (miRNAs) are a highly conserved class of small RNA molecules that are transcribed from DNA in the genomes of plants and animals but are not translated into proteins. Processed miRNAs are single-stranded 17-25 nucleotide RNA molecules that are incorporated into RNA-induced silencing complexes (RISCs) and have been identified as important regulators of development, cell proliferation, apoptosis, and differentiation. They are thought to play a role in regulating gene expression by binding to the 3' untranslated regions of specific mRNAs. RISCs mediate downregulation of gene expression through translational inhibition, transcriptional cleavage, or both. RISCs are also involved in transcriptional silencing in the nuclei of various eukaryotic organisms.

[0159] Antagomir In some embodiments, the AC is an antagomir. Antagomirs are RNA-like oligonucleotides with various modifications for RNase protection and pharmacological properties (e.g., enhanced tissue and cellular uptake). They differ from normal RNA, for example, by complete 2'-0-methylation of sugars, phosphorothioate backbones, and, for example, cholesterol moieties at the 3' end. Antagomirs can be used to efficiently silence endogenous miRNAs by forming duplexes containing the antagomir and endogenous miRNA, thereby preventing miRNA-induced gene silencing. An example of antagomir-mediated miRNA silencing is the silencing of miR-122, as described in Krutzfeldt et al., Nature (2005), 438:685-689 (expressly incorporated herein by reference in its entirety). Antagomir RNAs can be synthesized using standard solid-phase oligonucleotide synthesis protocols (US Patent Application Nos. 11 / 502,158 and 11 / 657,341; the disclosures of each of which are incorporated herein by reference).

[0160] Antagomirs may include ligand-binding monomer subunits and monomers for oligonucleotide synthesis. Monomers are described in U.S. Patent Application No. 10 / 916,185. Antagomirs may have a ZXY structure as described in PCT Application No. PCT / US2004 / 07070. Antagomirs may be conjugated with an amphiphilic moiety. Amphiphilic moieties for use with oligonucleotide agents are described in PCT Application No. PCT / US2004 / 07070.

[0161] Aptamers In embodiments, the AC comprises an aptamer. Aptamers are nucleic acid or peptide molecules that bind to specific molecules of interest with high affinity and specificity (Tuerk and Gold, Science 249:505 (1990); Ellington and Szostak, Nature 346:818 (1990)). DNA or RNA aptamers have been successfully produced that bind to many different entities, from large proteins to small organic molecules (see Eaton, Curr. Opin. Chem. Biol. (1997), 1:10-16; Famulok, Curr. Opin. Struct. Biol. (1999), 9:324-9; and Hermann and Patel, Science (2000), 287:820-5). Aptamers may be RNA- or DNA-based and may include riboswitches. A riboswitch is a part of an mRNA molecule that can directly bind to a small target molecule, and the binding of that target affects the activity of the gene. Thus, the mRNA containing the riboswitch is directly involved in regulating its own activity depending on the presence or absence of the target molecule. Generally, aptamers are engineered through repeated rounds of in vitro selection, i.e., SELEX (Systematic Evolution of Ligands by Exponential Enrichment), to bind to various molecular targets, such as small molecules, proteins, nucleic acids, and even cells, tissues, and organisms. Aptamers can be prepared by any known method, including synthetic, recombinant, and purified methods, and can be used alone or in combination with other aptamers specific to the same target. Furthermore, the term "aptamer" also includes "secondary aptamers," which contain consensus sequences derived from comparing two or more known aptamers to a given target. In embodiments, the aptamer is an "intracellular aptamer" or "intramer" that specifically recognizes an intracellular target (Famulok et al., Chem Biol. (2001), 8(10):931-939; Yoon and Rossi, Adv. Drug Deliv. Rev. (2018), 134:22-35; each incorporated herein by reference).

[0162] Ribozymes In one embodiment, AC is a ribozyme. Ribozyme is an RNA molecular complex that has a specific catalytic domain with endonuclease activity (Kim and Cech, Proc. Natl. Acad. Sci. USA (1987), 84 (24): 8788-92; Forster and Symons, Cell (1987) 24, 49 (2): 211-20). For example, many ribozymes accelerate phosphoester transfer reactions with high specificity, and often only cleave one of several phosphates in oligonucleotide substrates (Cech et al., Cell (1981), 27 (3 Pt 2): 487-96; Michel and Westhof, J. Mol. Biol. (1990), 5, 216 (3): 585-610; Reinhold-Hurek and Shub, Nature (1992), 14, 357 (6374): 173-6). This specificity results from the requirement that the substrate bind to the ribozyme's Internal Guide Sequence (IGS) through specific base-pairing interactions prior to chemical reaction.

[0163] At least six basic varieties of naturally occurring enzymatic RNAs are currently known. Each is capable of catalyzing the hydrolysis of RNA phosphodiester bonds in trans (and thus cleaving other RNA molecules) under physiological conditions. Generally, enzymatic nucleic acids act by first binding to a target RNA. Such binding occurs via the target-binding portion of the enzymatic nucleic acid, which is held in close proximity to the enzymatic portion of the molecule that acts to cleave the target RNA. Thus, the enzymatic nucleic acid first recognizes the target RNA, then binds to it through complementary base pairing, and once bound to the correct site, acts enzymatically to cleave the target RNA. This strategic cleavage of the target RNA destroys its ability to direct synthesis of the encoded protein. After an enzymatic nucleic acid binds and cleaves its RNA target, it is released from that RNA to search for another target and can repeatedly bind and cleave new targets.

[0164] Enzymatic nucleic acid molecules can be formed, for example, in hammerhead, hairpin, hepatitis delta virus, group I intron, or RNase P RNA (associated with an RNA guide sequence), or Neurospora VS RNA motifs. Specific examples of hammerhead motifs are described in Rossi et al., Nucleic Acids Res. (1992), 20(17):4559-65. Examples of hairpin motifs are described in Hampel et al. (European Patent Application Publication No. 0360257), Hampel and Tritz, Biochemistry (1989), 28(12):4929-33, Hampel et al., Nucleic Acids Res. (1990), 18(2):299-304, and U.S. Patent No. 5,631,359. Examples of hepatitis virus motifs are described in Perrotta and Been, Biochemistry (1992), 31(47):11843-52, examples of RNase P motifs are described in Guerrier-Takada et al., Cell (1983), 35(3 Pt 2):849-57, examples of Neurospora VS RNA ribozyme motifs are described in Collins (Saville and Collins, Cell (1990), 61(4):685-96; Saville and Collins, Proc. Natl. Acad. Sci. USA (1991), 88(19):8826-30; Collins and Olive, Biochemistry (1993), 32(11):2795-9), and examples of group I introns are described in U.S. Pat. No. 4,987,071. In embodiments, the enzymatic nucleic acid molecule has a specific substrate binding site complementary to one or more regions of the target gene DNA or RNA, and has nucleotide sequences within or surrounding the substrate binding site that confer RNA cleavage activity to the molecule. Thus, ribozyme constructs need not be limited to the particular motifs referenced herein.

[0165] Ribozymes can be designed as described in WO 93 / 23569 and WO 94 / 02595, each of which is specifically incorporated herein by reference, and can be synthesized and tested in vitro and in vivo as described therein. In embodiments, the ribozyme is targeted to a target nucleotide sequence of a target transcript.

[0166] Ribozyme activity can be increased by altering the length of the ribozyme binding arms or by chemically synthesizing ribozymes with modifications that prevent their degradation by serum ribonucleases (see, e.g., WO 92 / 07065, WO 93 / 15187, WO 91 / 03162, EP 92110298.4, U.S. Pat. No. 5,334,711, and U.S. Pat. No. 94 / 13688, which describe various chemical modifications that can be made to the sugar portion of enzymatic RNA molecules, modifications that enhance their effectiveness in cells, and removal of stem Π bases to shorten RNA synthesis time and reduce chemical requirements).

[0167] Supermir In embodiments, the AC is a supermir. A supermir refers to a single-stranded, double-stranded, or partially double-stranded oligomer or polymer of RNA or DNA, or both, or modifications thereof, that has a nucleotide sequence substantially identical to an miRNA and antisense to its target. This term includes oligonucleotides composed of naturally occurring nucleobases, sugars, and covalent internucleoside (backbone) linkages and that contain at least one non-naturally occurring moiety that functions similarly. Such modified or substituted oligonucleotides have desirable properties, such as enhanced cellular uptake, increased affinity for nucleic acid targets, and increased stability in the presence of nucleases. In embodiments, a supermir does not contain a sense strand, and in other embodiments, a supermir does not self-hybridize to a significant extent. While a supermir may have secondary structure, it is substantially single-stranded under physiological conditions. A substantially single-stranded supermir is single-stranded to the extent that less than about 50% (e.g., less than about 40%, less than about 30%, less than about 20%, less than about 10%, or less than about 5%) of the supermir is double-stranded with itself. A supermir can include a hairpin segment (e.g., a sequence) that can self-hybridize, for example, at the 3' end, to form a duplex region (e.g., a duplex region of at least about 1, about 2, about 3, or about 4 nucleotides, or of less than about 8, less than about 7, less than about 6, or less than about 5 nucleotides, or of about 5 nucleotides). The duplexed regions can be linked by a linker, e.g., a nucleotide linker, e.g., about 3, about 4, about 5, or about 6 dTs, e.g., modified dTs. In another embodiment, the supermir is duplexed with a shorter oligo, e.g., about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides in length, e.g., at one or both of the 3' and 5' ends, or at one end and a non-end or in the middle of the supermir.

[0168] miRNA mimics In some embodiments, the AC is an miRNA mimic. miRNA mimics represent a class of molecules that can be used to mimic the gene silencing capabilities of one or more miRNAs. Thus, the term "microRNA mimic" refers to a synthetic non-coding RNA (e.g., a miRNA not obtained by purification from an endogenous miRNA source) that can enter the RNAi pathway and regulate gene expression. miRNA mimics can be designed as mature molecules (e.g., single-stranded) or mimic precursors (e.g., pre-miRNAs). miRNA mimics can include nucleic acids (modified or modified nucleic acids) such as, but not limited to, oligonucleotides containing RNA, modified RNA, DNA, modified DNA, locked nucleic acids, or 2'-0,4'-C-ethylene-bridged nucleic acids (ENA), or any combination of the above (including DNA-RNA hybrids). Additionally, miRNA mimics can include conjugates that can affect delivery, intracellular compartmentalization, stability, specificity, functionality, strand usage, and / or efficacy. In one design, miRNA mimics are double-stranded molecules (e.g., having a double-stranded region about 16 to about 31 nucleotides in length) that contain one or more sequences that share identity with the mature strand of a given miRNA. Modifications can include 2' modifications (including 2'-0 methyl and 2'F modifications) on one or both strands of the molecule, as well as internucleoside modifications (e.g., phosphorothioate modifications) that enhance nucleic acid stability and / or specificity. In addition, miRNA mimics can include overhangs. The overhangs can comprise about 1 to about 6 nucleotides at either the 3' or 5' end of either strand and can be modified to enhance stability or functionality. In embodiments, miRNA mimics include a double-stranded region of about 16 to about 31 nucleotides and include one or more of the following chemical modification patterns: That is, the sense strand contains 2'-0-methyl modifications of nucleotides 1 and 2 (counting from the 5' end of the sense oligonucleotide) and all of the Cs and Us, while the antisense strand modifications may include 2'F modifications of all of the Cs and Us, phosphorylation of the 5' end of the oligonucleotide, and stabilized internucleoside linkages associated with the two-nucleotide 3' overhangs.

[0169] miRNA inhibitors In some embodiments, the AC is an miRNA inhibitor. The terms "anti-mir," "microRNA inhibitor," "miR inhibitor," or "miRNA inhibitor" are synonymous and refer to an oligonucleotide or modified oligonucleotide that interferes with the activity of a specific miRNA. Generally, the inhibitor is a natural or modified nucleic acid, such as an oligonucleotide of RNA, modified RNA, DNA, modified DNA, locked nucleic acid (LNA), or any combination thereof.

[0170] Modifications include 2'-modifications (including 2'-0 alkyl and 2'F modifications) and internucleoside modifications (e.g., phosphorothioate modifications), which can affect delivery, stability, specificity, intracellular compartmentalization, or efficacy. Additionally, miRNA inhibitors can include conjugates, which can affect delivery, intracellular compartmentalization, stability, and / or efficacy. Inhibitors can adopt a variety of configurations, including single-stranded, double-stranded (RNA / RNA or RNA / DNA duplexes), and hairpin designs. Generally, microRNA inhibitors include one or more sequences or portions of sequences that are complementary or partially complementary to the mature strand(s) of the targeted miRNA. Furthermore, miRNA inhibitors can also include additional sequences located 5' and 3' relative to the sequence that is the reverse complement of the mature miRNA. The additional sequences may be the reverse complement of the sequence adjacent to the mature miRNA in the pri-miRNA from which the mature miRNA is derived, or the additional sequences may be any sequence (having a mixture of A, G, C, or U). In embodiments, one or both of the additional sequences may be any sequence capable of forming a hairpin. Thus, in embodiments, a sequence that is the reverse complement of the miRNA flanks the hairpin structure on the 5' and 3' sides. When the microRNA inhibitor is double-stranded, it may contain mismatches between nucleotides on opposite strands. Furthermore, the microRNA inhibitor may be linked to a conjugate moiety to facilitate the uptake of the inhibitor into cells. For example, the microRNA inhibitor may be linked to cholesteryl 5-(bis(4-methoxyphenyl)(phenyl)methoxy)-3-hydroxypentylcarbamate), which allows passive uptake of the microRNA inhibitor into cells. MicroRNA inhibitors, including hairpin miRNA inhibitors, are described in detail in Vermeulen et al., "Double-Stranded Regions Are Essential Design Components Of Potent Inhibitors of RISC Function," RNA 13:723-730 (2007), and in International Publication Nos. WO 2007 / 095387 and WO 2008 / 036825, each of which is incorporated herein by reference in its entirety.One of skill in the art can select a sequence from a database for a desired miRNA and design an inhibitor useful in the methods disclosed herein.

[0171] CRISPR gene editing mechanism In embodiments, the therapeutic moiety comprises one or more elements of CRISPR gene editing mechanism.As used herein, "CRISPR gene editing mechanism" refers to a protein, nucleic acid, or combination thereof that can be used to edit genome.Non-limiting examples of gene editing mechanism include guide RNA (gRNA), nuclease, nuclease inhibitor, and combinations and complexes thereof. CRISPR gene editing mechanisms are described in the following patent documents: U.S. Patent No. 8,697,359, U.S. Patent No. 8,771,945, U.S. Patent No. 8,795,965, U.S. Patent No. 8,865,406, U.S. Patent No. 8,871,445, U.S. Patent No. 8,889,356, U.S. Patent No. 8,895,308, U.S. Patent No. 8,906,616, U.S. Patent No. 8,932,814, U.S. Patent No. 8,945,839, U.S. Patent No. 8,993,233, U.S. Patent No. 8,999,641, U.S. Patent Application No. 14 / 704,551, and U.S. Patent Application No. 13 / 842,859. Each of the foregoing patent documents is incorporated herein by reference in its entirety.

[0172] gRNA In embodiments, the TM comprises a gRNA that targets a genomic locus in a prokaryotic or eukaryotic cell.

[0173] In embodiments, the gRNA is a single-molecule guide RNA (sgRNA). The sgRNA comprises a spacer sequence and a scaffold sequence. The spacer sequence is a short nucleic acid sequence used to target a nuclease (e.g., Cas9 nuclease) to a specific nucleotide region of interest (e.g., a genomic DNA sequence to be cleaved). In embodiments, the spacer can be about 17 to 24 bases in length, e.g., about 20 bases in length. In embodiments, the spacer can be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 bases in length. In embodiments, the spacer can be at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, or at least 30 bases in length. In embodiments, the spacer can be about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, or about 30 bases in length. In embodiments, the spacer sequence has a GC content of about 40% to about 80%.

[0174] In embodiments, the spacer binds to the target nucleotide sequence immediately before the 5' protospacer adjacent motif (PAM). The PAM sequence can be selected based on the desired nuclease. For example, the PAM sequence can be any one of the PAM sequences shown in Table 13 below, where N refers to any nucleic acid, R refers to A or G, Y refers to C or T, W refers to A or T, and V refers to A, C, or G.

[0175] [Table 1]

[0176] In embodiments, the spacer binds to a target nucleotide sequence of a mammalian target transcript of a target gene, such as a human gene. In embodiments, the spacer may bind to a target nucleotide sequence of the target transcript. In embodiments, the spacer may bind to a target nucleotide sequence that includes at least a portion of a splice element (SE) and / or splice regulatory element (SRE) of the target transcript, or is sufficiently proximal to the SE and / or SRE of the target transcript to regulate splicing.

[0177] The scaffold sequence is a sequence within the sgRNA that is involved in nuclease (e.g., Cas9) binding. The scaffold sequence does not include a spacer / targeting sequence. In embodiments, the scaffold can be about 1 to about 10, about 10 to about 20, about 20 to about 30, about 30 to about 40, about 40 to about 50, about 50 to about 60, about 60 to about 70, about 70 to about 80, about 80 to about 90, about 90 to about 100, about 100 to about 110, about 110 to about 120, or about 120 to about 130 nucleotides in length. In embodiments, the scaffold comprises about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102 4, about 35, about 36, about 37, about 38, about 39, about 40, about 41, about 42, about 43, about 44, about 45, about 46, about 47, about 48, about 49, about 50, about 51, about 52, about 53, about 54, about 55, about 56, about 57, about 58, about 59, about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67 , about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 1 about 101, about 102, about 103, about 104, about 105, about 106, about 107, about 108, about 109, about 110, about 111, about 112, about 113, about 114, about 115, about 116, about 117, about 118, about 119, about 120, about 121, about 122, about 123, about 124, or about 125 nucleotides in length. In embodiments, the scaffold can be at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, or at least 125 nucleotides in length.

[0178] In embodiments, the gRNA is a dual-molecule guide RNA, e.g., a crRNA and a tracrRNA. In embodiments, the gRNA may further comprise a poly(A) tail.

[0179] In embodiments, multiple gRNAs can be used as a TM in a single compound. In embodiments, a TM can include about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, or about 20 gRNAs. In embodiments, the gRNAs recognize the same target. In embodiments, the gRNAs recognize different targets. In embodiments, the nucleic acid comprising the gRNA includes a sequence encoding a promoter, and the promoter drives expression of the gRNA.

[0180] nuclease In embodiments, the TM comprises a nuclease. In embodiments, the nuclease is a type II, type VA, type VB, type VC, type VU, or type VI-B nuclease. In embodiments, the nuclease is a transcription activator-like effector nuclease (TALEN), meganuclease, or zinc finger nuclease. In embodiments, the nuclease is a Cas9, Cas12a (CF3), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease. For example, in embodiments, the nuclease is a Cas9 nuclease or a Cpf1 nuclease.

[0181] In embodiments, the nuclease is a modified form or mutant of a Cas9, Cas12a (Cpf1), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease. In embodiments, the nuclease is a modified form or mutant of a TAL nuclease, meganuclease, or zinc finger nuclease. "Modified" or "mutant" nucleases can be, for example, truncated, fused to another protein (such as another nuclease), catalytically inactivated, etc. In embodiments, the nuclease can have at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99%, or about 100% sequence identity to a naturally occurring Cas9, Cas12a (Cpf1), Cas12b, Cas12c, Tnp-B-like, Cas13a (C2c2), Cas13b, or Cas14 nuclease, or a TALEN, meganuclease, or zinc finger nuclease. In embodiments, the nuclease is a Cas9 nuclease derived from S. pyogenes (SpCas9). In embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cas9 nuclease from S. pyogenes (SpCas9). In embodiments, the nuclease is Cas9 from Staphylococcus aureus (SaCas9). In embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cas9 from Staphylococcus aureus (SaCas9). In embodiments, the Cpfl is a Cpfl enzyme from Acidaminococcus (species BV3L6, UniProt accession number U2UMQ6). In embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cpf1 enzyme from Acidaminococcus (species BV3L6, UniProt accession number U2UMQ6).

[0182] In embodiments, the Cpfl is a Cpfl enzyme from Lachnospiraceae (species ND2006, UniProt accession number A0A182DWE3). In embodiments, the nuclease has at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to the Cpfl enzyme from Lachnospiraceae. In embodiments, the nuclease-encoding sequence is codon-optimized for expression in mammalian cells. In embodiments, the nuclease-encoding sequence is codon-optimized for expression in human or mouse cells.

[0183] In embodiments, the nuclease is a soluble protein.

[0184] In embodiments, the TM is a nucleotide sequence encoding a nuclease. In embodiments, the nucleic acid encoding the nuclease comprises a sequence encoding a promoter, which drives expression of the nuclease.

[0185] gRNA and nuclease combinations In embodiments, the compounds of the present disclosure comprise a gRNA and a nuclease or a nucleotide sequence encoding the nuclease as a TM. In embodiments, the nucleic acid encoding the nuclease and gRNA comprises a sequence encoding a promoter, which drives expression of the nuclease and gRNA. In embodiments, the nucleic acid encoding the nuclease and gRNA comprises two promoters, one controlling expression of the nuclease and the second controlling expression of the gRNA. In embodiments, the nucleic acid encoding the gRNA and nuclease encodes from about 1 to about 20 gRNAs, or from about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, or about 19 up to about 20 gRNAs. In embodiments, the gRNAs recognize different targets. In embodiments, the gRNAs recognize the same target.

[0186] In embodiments, the compounds of the present disclosure comprise a ribonucleoprotein (RNP) comprising a gRNA and a nuclease as a TM.

[0187] In embodiments, a composition comprising (a) a first compound comprising a gRNA TM, and (b) a second compound that is or comprises a nuclease is delivered to a cell. In embodiments, a composition comprising (a) a first compound that comprises a nuclease as a TM, CPP, and (b) a second molecule that is or comprises a gRNA is delivered to a cell. In embodiments, a composition comprising (a) a first compound that comprises a gRNA as a TM, and (b) a second compound that comprises a nuclease as a TM is delivered to a cell.

[0188] Genetic element of interest In embodiments, the compounds disclosed herein include a genetic element of interest as a TM. In embodiments, the genetic element of interest replaces a genomic DNA sequence cleaved by a nuclease. Non-limiting examples of genetic elements of interest include a gene, a single nucleotide polymorphism, a promoter, or a terminator.

[0189] Nuclease inhibitors In embodiments, the compounds disclosed herein include a nuclease inhibitor as a TM. A limitation of gene editing is potential off-target editing. Delivery of a nuclease inhibitor can limit off-target editing. In embodiments, the nuclease inhibitor is a polypeptide, polynucleotide, or small molecule. Nuclease inhibitors are described in U.S. Patent Application Publication No. 2020 / 087354, WO 2018 / 085288, U.S. Patent Application Publication No. 2018 / 0382741, WO 2019 / 089761, WO 2020 / 068304, WO 2020 / 041384, and WO 2019 / 076651, each of which is incorporated herein by reference in its entirety.

[0190] Endosomal escape vehicles (EEVs) Endosomal escape vehicles (EEVs) can be used to transport cargo across cell membranes, for example, to deliver cargo to the cytosol or nucleus of a cell. The cargo may include a TM. The EEV may include a cell-penetrating peptide (CPP), such as a cyclic cell-penetrating peptide (cCPP). In embodiments, the EEV includes a cCPP conjugated to an exocyclic peptide (EP). The EP may be interchangeably referred to as a regulatory peptide (MP). The EP may include a nuclear localization signal (NLS) sequence. The EP may be bound to the cargo. The EP may be bound to the cCPP. The EP may be bound to the cargo and the cCPP. The coupling between the EP, cargo, cCPP, or a combination thereof may be non-covalent or covalent. The EP may be bound to the N-terminus of the cCPP via a peptide bond. The EP may be bound to the C-terminus of the cCPP via a peptide bond. The EP may be bound to the cCPP via a side chain of an amino acid in the cCPP. The EP can be attached to the cCPP via a lysine side chain, which can be conjugated to the glutamine side chain in the cCPP. The EP can be conjugated to the 5' or 3' end of the oligonucleotide cargo. The EP can be attached to a linker. The exocyclic peptide can be conjugated to the amino group of the linker. The EP can be coupled to the linker via the C-terminus of the EP and cCPP through a side chain on the cCPP and / or the EP. For example, the EP can contain a terminal lysine, which can then be coupled to a glutamine-containing cCPP via an amide bond. If the EP contains a terminal lysine and the side chain of the lysine can be used to attach the cCPP, the C-terminus or N-terminus can be attached to the linker on the cargo.

[0191] exocyclic peptides The exocyclic peptide (EP) can contain 2 to 10 amino acid residues, e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10 amino acid residues (including all ranges and values ​​therebetween). The EP can contain 6 to 9 amino acid residues. The EP can contain 4 to 8 amino acid residues.

[0192] Each amino acid in the exocyclic peptide can be a natural or unnatural amino acid. The term "unnatural amino acid" refers to an organic compound that is a homolog of a natural amino acid in that it has a structure similar to a natural amino acid so as to mimic the structure and reactivity of the natural amino acid. An unnatural amino acid can be a modified amino acid and / or an amino acid analog that is not one of the 20 common natural amino acids or the rare natural amino acids selenocysteine ​​or pyrrolysine. An unnatural amino acid can also be a D-isomer of a natural amino acid. Examples of suitable amino acids include, but are not limited to, alanine, allosoleucine, arginine, citrulline, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, naphthylalanine, phenylalanine, proline, pyroglutamic acid, serine, threonine, tryptophan, tyrosine, valine, derivatives thereof, or combinations thereof. These and other amino acids, along with their abbreviations used herein, are listed in Table 1. For example, an amino acid can be A, G, P, K, R, V, F, H, NaI, or citrulline.

[0193] The EP may contain at least one positively charged amino acid residue, for example, at least one lysine residue and / or at least one amino acid residue containing a side chain containing a guanidine group or its protonated form. The EP may contain one or two amino acid residues containing a side chain containing a guanidine group or its protonated form. The amino acid residue containing a side chain containing a guanidine group may be an arginine residue. The protonated form may refer to a salt thereof throughout this disclosure.

[0194] The EP may contain at least two, at least three, or at least four or more lysine residues. The EP may contain two, three, or four lysine residues. The amino group on the side chain of each lysine residue may be substituted with a protecting group such as a trifluoroacetyl (-COCF3) group, an allyloxycarbonyl (Alloc) group, a 1-(4,4-dimethyl-2,6-dioxocyclohexylidene)ethyl (Dde) group, or a (4,4-dimethyl-2,6-dioxocyclohex-1-ylidene-3)-methylbutyl (ivDde) ​​group. The amino group on the side chain of each lysine residue may be substituted with a trifluoroacetyl (-COCF3) group. Protecting groups may be included to enable amide conjugation. The protecting groups may be removed after the EP is conjugated to the cCPP.

[0195] The EP may comprise at least two amino acid residues having hydrophobic side chains. The amino acid residues having hydrophobic side chains may be selected from valine, proline, alanine, leucine, isoleucine, and methionine. The amino acid residues having hydrophobic side chains may be valine or proline.

[0196] The EP may contain at least one positively charged amino acid residue, e.g., at least one lysine residue and / or at least one arginine residue. The EP may contain at least two, at least three, or at least four or more lysine and / or arginine residues.

[0197] EP is KK, KR, RR, HH, HK, HR, RH, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKH, KHK, HKK, HRR, HRH, HHR, HBH, HHH, HHHH (SEQ ID NO: 1), KHKK (SEQ ID NO: 2), KKHK (SEQ ID NO: 3), KKKH (SEQ ID NO: 4), KHKH (SEQ ID NO: 5), HKHK (SEQ ID NO: 6), KKKK (SEQ ID NO: 7), KKRK (SEQ ID NO: 8), KRKK (SEQ ID NO: 9), KRRK (SEQ ID NO: 10), RKKR ( SEQ ID NO: 6), KKKK (SEQ ID NO: 7), KKRK (SEQ ID NO: 8), KRKK (SEQ ID NO: 9), KRRK (SEQ ID NO: 10), RKKR (SEQ ID NO: 11), RRRR (SEQ ID NO: 12), KGKK (SEQ ID NO: 13), KKGK (SEQ ID NO: 14), HBHBH (SEQ ID NO: 15), HBKBH (SEQ ID NO: 16), RRRRR (SEQ ID NO: 17), RRRRR (SEQ ID NO: 17), KKKKK (SEQ ID NO: 18), KKKRK (SEQ ID NO: 19), RKKKK (SEQ ID NO: 20), KRKKK (SEQ ID NO: 21). No. 21), KKKRKK (SEQ ID NO: 22), KKKKR (SEQ ID NO: 23), KBKBK (SEQ ID NO: 24), RKKKG (SEQ ID NO: 25), KRKKKG (SEQ ID NO: 26), KRKRKKG (SEQ ID NO: 27), KKKKRG (SEQ ID NO: 28), RKKKKB (SEQ ID NO: 29), KRKKKB (SEQ ID NO: 30), KKRKKB (SEQ ID NO: 31), KKKKRB (SEQ ID NO: 32), KKKRKV (SEQ ID NO: 33), RRRRRR (SEQ ID NO: 34), HHHHHH (SEQ ID NO: 35), RHRH RH (SEQ ID NO:36), HRHRHR (SEQ ID NO:37), KRKRKR (SEQ ID NO:38), RKRKRK (SEQ ID NO:39), RBRBRB (40), KBKBKB (SEQ ID NO:41), PKKKRKV (SEQ ID NO:42), PGKKRKV (SEQ ID NO:43), PKGKRKV (SEQ ID NO:44), PKKGRKV (SEQ ID NO:45), PKKKGKV (SEQ ID NO:46), PKKKRGV (SEQ ID NO:47), or PKKKRKG (SEQ ID NO:48), wherein B is β-alanine. The amino acids in the EP can have D or L stereochemistry.

[0198] The EP can include KK, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK (SEQ ID NO:7), KKRK (SEQ ID NO:8), KRKK (SEQ ID NO:9), KRRK (SEQ ID NO:10), RKKR (SEQ ID NO:11), RRRR (SEQ ID NO:12), KGKK (SEQ ID NO:13), KKGK (SEQ ID NO:14), KKKKK (SEQ ID NO:18), KKKRK (SEQ ID NO:19), KBKBK (SEQ ID NO:24), KKKRKV (SEQ ID NO:33), PKKKRKV (SEQ ID NO:42), PGKKRKV (SEQ ID NO:43), PKGKRKV (SEQ ID NO:44), PKKGRKV (SEQ ID NO:45), PKKKGKV (SEQ ID NO:46), PKKKRGV (SEQ ID NO:47), or PKKKRKG (SEQ ID NO:48). The EP can include PKKKRKV (SEQ ID NO: 42), RR, RRR, RHR, RBR, RBRBR (SEQ ID NO: 49), RBHBR (SEQ ID NO: 50), or HBRBH (SEQ ID NO: 51), where B is β-alanine. The amino acids in the EP can have D or L stereochemistry.

[0199] EP is KK, KR, RR, KKK, KGK, KBK, KBR, KRK, KRR, RKK, RRR, KKKK (SEQ ID NO: 7), KKRK (SEQ ID NO: 8), KRKK (SEQ ID NO: 9), KRRK (SEQ ID NO: 10), RKKR (SEQ ID NO: 11), RRRR (SEQ ID NO: 12), KGKK (SEQ ID NO: 13), KKGK (SEQ ID NO: 14), KKKKK (SEQ ID NO: 18), KKKRK (SEQ ID NO: 19), KBKBK (SEQ ID NO: 24), KKKRKV (SEQ ID NO: 33), PKKKRKV (SEQ ID NO: 42), PGKKRKV (SEQ ID NO: 43), PGKKRKV (SEQ ID NO: 44), PGKKRKV (SEQ ID NO: 45), PGKKRKV (SEQ ID NO: 46), PGKKRKV (SEQ ID NO: 47), PGKKRKV (SEQ ID NO: 48), PGKKRKV (SEQ ID NO: 49), PGKKRKV (SEQ ID NO: 50), PGKKRKV (SEQ ID NO: 51), PGKKRKV (SEQ ID NO: 52), PGKKRKV (SEQ ID NO: 53), PGKKRKV (SEQ ID NO: 54), PGKKRKV (SEQ ID NO: 55), PGKKRKV (SEQ ID NO: 56), PGKKRKV (SEQ ID NO: 57), PGKKRKV (SEQ ID NO: 58), PGKKRKV (SEQ ID NO: 59), PGKKRKV (SEQ ID NO: 60), PGKKRKV (SEQ ID NO: 61), PGKKRKV (SEQ ID NO: 62), PG No. 4 3), PKGKRKV (SEQ ID NO: 44), PKKGRKV (SEQ ID NO: 45), PKKKGKV (SEQ ID NO: 46), PKKKRGV (SEQ ID NO: 47), or PKKKRKG (SEQ ID NO: 48). Great deal The EP can consist of PKKKRKV (SEQ ID NO: 42), RR, RRR, RHR, RBR, RBRBR (SEQ ID NO: 49), RBHBR (SEQ ID NO: 50), or HBRBH (SEQ ID NO: 51), where B is β-alanine. The amino acids in the EP can have D or L stereochemistry.

[0200] The EP may comprise an amino acid sequence identified in the art as a nuclear localization sequence (NLS). The EP may consist of an amino acid sequence identified in the art as a nuclear localization sequence (NLS). The EP may comprise an NLS comprising the amino acid sequence PKKKRKV (SEQ ID NO: 42). The EP may consist of an NLS comprising the amino acid sequence PKKKRKV (SEQ ID NO: 42). The EP can include an NLS comprising an amino acid sequence selected from NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 52), PAAKRVKLD (SEQ ID NO: 53), RQRRNELKRSF (SEQ ID NO: 54), RMRKFKNKGKDTAELRRRRVEVSVELR (SEQ ID NO: 55), KAKKDEQILKRRNV (SEQ ID NO: 56), VSRKRPRP (SEQ ID NO: 57), PPKKARED (SEQ ID NO: 58), PQPKKKPL (SEQ ID NO: 59), SALIKKKKKMAP (SEQ ID NO: 60), DRLRR (SEQ ID NO: 61), PKQKKRK (SEQ ID NO: 62), RKLKKKIKKL (SEQ ID NO: 63), REKKKFLKRR (SEQ ID NO: 64), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 65), and RKCLQAGMNLEARKTKK (SEQ ID NO: 66). The EP may consist of an NLS comprising an amino acid sequence selected from NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 52), PAAKRVKLD (SEQ ID NO: 53), RQRRNELKRSF (SEQ ID NO: 54), RMRKFKNKGKDTAELRRRRVEVSVELR (SEQ ID NO: 55), KAKKDEQILKRRNV (SEQ ID NO: 56), VSRKRPRP (SEQ ID NO: 57), PPKKARED (SEQ ID NO: 58), PQPKKKPL (SEQ ID NO: 59), SALIKKKKKMAP (SEQ ID NO: 60), DRLRR (SEQ ID NO: 61), PKQKKRK (SEQ ID NO: 62), RKLKKKIKKL (SEQ ID NO: 63), REKKKFLKRR (SEQ ID NO: 64), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 65), and RKCLQAGMNLEARKTKK (SEQ ID NO: 66).

[0201] All exocyclic sequences may also contain an N-terminal acetyl group. Thus, for example, the EP may have the structure: Ac-PKKKRKV (SEQ ID NO: 42).

[0202] Cell-penetrating peptides (CPPs) The cell-penetrating peptide (CPP) can contain 6 to 20 amino acid residues. The cell-penetrating peptide can be a cyclic cell-penetrating peptide (cCPP). The cCPP can penetrate the cell membrane. An exocyclic peptide (EP) can be conjugated to the cCPP, and the resulting construct can be called an endosomal escape vehicle (EEV). The cCPP can translocate a cargo (e.g., a therapeutic moiety (TM), such as an oligonucleotide, peptide, or small molecule) to penetrate the cell membrane. The cCPP can deliver the cargo to the cytosol of the cell. The cCPP can deliver the cargo to a cellular location where the target (e.g., pre-mRNA) is located. To conjugate the cCPP to a cargo (e.g., a peptide, oligonucleotide, or small molecule), at least one bond or lone pair on the cCPP can be replaced.

[0203] The total number of amino acid residues in a cCPP can range from 6 to 20 amino acid residues, e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acid residues, including all ranges and subranges therebetween. A cCPP can contain 6 to 13 amino acid residues. A cCPP disclosed herein can contain 6 to 10 amino acids. By way of example, a cCPP containing 6 to 10 amino acid residues can be represented by Formulas IA to IE:

[0204] [ka] wherein AA1, AA2, AA3, AA4, AA5, AA6, AA7, AA8, AA9, and AA 10 is an amino acid residue.

[0205] The cCPP may contain 6 to 8 amino acids. The cCPP may contain 8 amino acids.

[0206] Each amino acid in a cCPP may be a natural or unnatural amino acid. The term "unnatural amino acid" refers to an organic compound that is a homolog of a natural amino acid in that it has a structure similar to a natural amino acid so as to mimic the structure and reactivity of the natural amino acid. An unnatural amino acid may be a modified amino acid and / or an amino acid analog that is not one of the 20 common natural amino acids or the rare natural amino acids selenocysteine ​​or pyrrolysine. An unnatural amino acid may also be a D-isomer of a natural amino acid. Examples of suitable amino acids include, but are not limited to, alanine, allosoleucine, arginine, citrulline, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, naphthylalanine, phenylalanine, proline, pyroglutamic acid, serine, threonine, tryptophan, tyrosine, valine, derivatives thereof, or combinations thereof. These and other amino acids, along with their abbreviations used herein, are listed in Table 1.

[0207] [Table 2]

[0208] As used herein, "polyethylene glycol" and "PEG" are used interchangeably. m " has the formula HO(CO)-(CH2) n -(OCH2CH2) m In embodiments, n is 1 or 2. In embodiments, n is 1. In embodiments, n is 2. In embodiments, n is 1 and m is 2. In embodiments, n is 1 and m is 2. In embodiments, n is 1 and m is 4. In embodiments, n is 2 and m is 4. In embodiments, n is 1 and m is 12. In embodiments, n is 2 and m is 12.

[0209] As used herein, "miniPEGm" or "miniPEG m " has the formula HO(CO)-(CH2) n -(OCH2CH2) m -NH2 (where n is 1 and m is any integer from 1 to 23). For example, "miniPEG2" or "miniPEG2" is or is derived from (2-[2-[2-aminoethoxy]ethoxy]acetic acid), and "miniPEG4" or "miniPEG4" is HO(CO)-(CH2). n -(OCH2CH2) m is or is derived from -NH2, where n is 1 and m is 4.

[0210] A cCPP can contain 4 to 20 amino acids, where (i) at least one amino acid has a side chain containing a guanidine group or its protonated form, and (ii) at least one amino acid has no side chain, or

[0211] [ka] or its protonated form, and (iii) at least two amino acids have side chains that independently include an aromatic group or a heteroaromatic group.

[0212] At least two amino acids may have no side chains, or

[0213] [ka] or a side chain including its protonated form. As used herein, when no side chain is present, an amino acid has two hydrogen atoms on a carbon atom connecting the amine and carboxylic acid (e.g., -CH-).

[0214] The amino acid having no side chain may be glycine or β-alanine.

[0215] The cCPP can comprise 6 to 20 amino acid residues that form the cCPP, wherein (i) at least one amino acid can be a glycine, β-alanine, or 4-aminobutyric acid residue, (ii) at least one amino acid can have a side chain that includes an aryl or heteroaryl group, (iii) at least one amino acid can have a guanidine group,

[0216] [ka] or has a side chain containing its protonated form.

[0217] The cCPP can comprise 6 to 20 amino acid residues that form the cCPP, wherein (i) at least two amino acid residues can independently be glycine, β-alanine, or 4-aminobutyric acid residues, (ii) at least one amino acid can have a side chain that includes an aryl or heteroaryl group, (iii) at least one amino acid can have a guanidine group,

[0218] [ka] or has a side chain containing its protonated form.

[0219] The cCPP can comprise 6 to 20 amino acid residues that form the cCPP, wherein (i) at least three amino acids can be, independently, glycine, β-alanine, or 4-aminobutyric acid residues; (ii) at least one amino acid can have a side chain that includes an aromatic or heteroaromatic group; (iii) at least one amino acid can have a guanidine group;

[0220] [ka] or may have side chains containing the protonated form thereof.

[0221] Glycine and related amino acid residues The cCPP may contain (i) 1, 2, 3, 4, 5, or 6 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 2 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 3 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 4 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 5 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 6 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) 3, 4, or 5 glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) three or four glycine, β-alanine, 4-aminobutyric acid residues, or a combination thereof.

[0222] The cCPP may contain (i) 1, 2, 3, 4, 5, or 6 glycine residues. The cCPP may contain (i) 2 glycine residues. The cCPP may contain (i) 3 glycine residues. The cCPP may contain (i) 4 glycine residues. The cCPP may contain (i) 5 glycine residues. The cCPP may contain (i) 6 glycine residues. The cCPP may contain (i) 3, 4, or 5 glycine residues. The cCPP may contain (i) 3 or 4 glycine residues. The cCPP may contain (i) 2 or 3 glycine residues. The cCPP may contain (i) 1 or 2 glycine residues.

[0223] The cCPP may contain (i) three, four, five, or six glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) three glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) four glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) five glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) six glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) three, four, or five glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof. The cCPP may contain (i) three or four glycine, β-alanine, or 4-aminobutyric acid residues, or a combination thereof.

[0224] The cCPP may contain at least three glycine residues. The cCPP may contain (i) 3, 4, 5, or 6 glycine residues. The cCPP may contain (i) 3 glycine residues. The cCPP may contain (i) 4 glycine residues. The cCPP may contain (i) 5 glycine residues. The cCPP may contain (i) 6 glycine residues. The cCPP may contain (i) 3, 4, or 5 glycine residues. The cCPP may contain (i) 3 or 4 glycine residues.

[0225] In embodiments, none of the glycine, β-alanine, or 4-aminobutyric acid residues in the cCPP are adjacent. Two or three glycine, β-alanine, 4- or 4-aminobutyric acid residues may be adjacent. Two glycine, β-alanine, or 4-aminobutyric acid residues may be adjacent.

[0226] In embodiments, none of the glycine residues in the cCPP are adjacent. Each glycine residue in the cCPP may be separated by an amino acid residue that cannot be glycine. Two or three glycine residues may be adjacent. Two glycine residues may be adjacent.

[0227] Amino acid side chains containing aromatic or heteroaromatic groups A cCPP may comprise 2, 3, 4, 5, or 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 2 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 5 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP may comprise 2, 3, or 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group or a heteroaromatic group. A cCPP can include (ii) two or three amino acid residues that independently have side chains that include aromatic or heteroaromatic groups.

[0228] The cCPP may comprise 2, 3, 4, 5, or 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 2 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 5 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 6 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 2, 3, or 4 amino acid residues that (ii) independently have a side chain that includes an aromatic group. The cCPP may comprise 2 or 3 amino acid residues that (ii) independently have a side chain that includes an aromatic group.

[0229] The aromatic group can be a 6- to 14-membered aryl. The aryl can be phenyl, naphthyl, or anthracenyl, each of which is optionally substituted. The aryl can be phenyl or naphthyl, each of which is optionally substituted. The heteroaromatic group can be a 6- to 14-membered heteroaryl having 1, 2, or 3 heteroatoms selected from N, O, and S. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl.

[0230] The amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be bis(homonaphthylalanine), homonaphthylalanine, naphthylalanine, phenylglycine, bis(homophenylalanine), homophenylalanine, phenylalanine, tryptophan, 3-(3-benzothienyl)-alanine, 3-(2-quinolyl)-alanine, O-benzylserine, 3-(4-(benzyloxy)phenyl)-alanine, S-(4-methylbenzyl)cysteine, N-(naphthalen-2-yl)glutamine, 3-(1,1'-biphenyl-4-yl)-alanine, 3-(3-benzothienyl)-alanine, or tyrosine, each of which is optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be:

[0231] [ka] wherein H on the N-terminus and / or H on the C-terminus is replaced by a peptide bond.

[0232] Amino acid residues having a side chain containing an aromatic or heteroaromatic group can each independently be a residue of phenylalanine, naphthylalanine, phenylglycine, homophenylalanine, homonaphthylalanine, bis(homophenylalanine), bis-(homonaphthylalanine), tryptophan, or tyrosine, each of which is optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group may each independently be a residue of tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butylphenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, or 3-(9-anthryl)alanine. The amino acid residues having a side chain containing an aromatic group may each independently be a residue of phenylalanine, naphthylalanine, phenylglycine, homophenylalanine, or homophenylalanine, each of which is optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group may each independently be a residue of phenylalanine, naphthylalanine, homophenylalanine, homophenylalanine, bis(homonaphthylalanine), or bis(homonaphthylalanine), each of which is optionally substituted with one or more substituents. The amino acid residues having a side chain containing an aromatic group may each independently be a residue of phenylalanine or naphthylalanine, each of which is optionally substituted with one or more substituents. At least one amino acid residue having a side chain containing an aromatic group may be a residue of phenylalanine. At least two amino acid residues having side chains containing an aromatic group may be residues of phenylalanine. Each amino acid residue having a side chain containing an aromatic group may be a residue of phenylalanine.

[0233] In embodiments, none of the amino acids having a side chain containing an aromatic or heteroaromatic group are adjacent. Two amino acids having a side chain containing an aromatic or heteroaromatic group may be adjacent. Two adjacent amino acids may have opposite stereochemistry. Two adjacent amino acids may have the same stereochemistry. Three amino acids having a side chain containing an aromatic or heteroaromatic group may be adjacent. Three adjacent amino acids may have the same stereochemistry. Three adjacent amino acids may have alternate stereochemistry.

[0234] The amino acid residues containing aromatic or heteroaromatic groups may be L-amino acids. The amino acid residues containing aromatic or heteroaromatic groups may be D-amino acids. The amino acid residues containing aromatic or heteroaromatic groups may be a mixture of D- and L-amino acids.

[0235] An optional substituent can be, for example, any atom or group that does not significantly (e.g., by more than 50%) reduce the cytoplasmic delivery efficiency of the cCPP compared to an otherwise identical sequence lacking the substituent. An optional substituent can be a hydrophobic or hydrophilic substituent. An optional substituent can be a hydrophobic substituent. The substituent can increase the solvent accessible surface area (as defined herein) of the hydrophobic amino acid. A substituent can be halogen, alkyl, alkenyl, alkynylene, cycloalkyl, cycloalkenyl, cycloalkynyl, heterocyclyl, aryl, heteroaryl, alkoxy, aryloxy, acyl, alkylcarbamoyl, alkylcarboxamidyl, alkoxycarbonyl, alkylthio, or arylthio. A substituent can be halogen.

[0236] Without wishing to be bound by theory, it is believed that amino acids having aromatic or heteroaromatic groups with higher hydrophobicity values ​​(i.e., amino acids having side chains containing aromatic or heteroaromatic groups) can improve the cytoplasmic delivery efficiency of cCPPs compared to amino acids with lower hydrophobicity values. Each hydrophobic amino acid can independently have a hydrophobicity value greater than that of glycine. Each hydrophobic amino acid can independently have a hydrophobicity value greater than that of alanine. Each hydrophobic amino acid can independently have a hydrophobicity value equal to or greater than that of phenylalanine. Hydrophobicity can be measured using a hydrophobicity scale known in the art. Table 2 lists the hydrophobicity values ​​for various amino acids reported by Eisenberg and Weiss (Proc. Natl. Acad. Sci. USA 1984;81(1):140-144), Engleman et al. (Ann. Rev. of Biophys. Biophys. Chem. 1986;1986(15):321-53), Kyte and Doolittle (J. Mol. Biol. 1982;157(1):105-132), Hoop and Woods (Proc. Natl. Acad. Sci. USA 1981;78(6):3824-3828), and Janin (Nature. 1979;277(5696):491-492), each of which is incorporated herein by reference in its entirety. Hydrophobicity can be measured using the hydrophobicity scale reported in Engleman et al.

[0237] [Table 3]

[0238] The size of the aromatic or heteroaromatic group can be selected to improve the cytoplasmic delivery efficiency of cCPPs. Without wishing to be bound by theory, it is believed that a larger aromatic or heteroaromatic group on the side chain of an amino acid may improve the cytoplasmic delivery efficiency compared to an otherwise identical sequence having a smaller hydrophobic amino acid. The size of a hydrophobic amino acid can be measured in terms of the molecular weight of the hydrophobic amino acid, the steric effect of the hydrophobic amino acid, the solvent accessible surface area (SASA) of the side chain, or a combination thereof. The size of a hydrophobic amino acid can be measured in terms of the molecular weight of the hydrophobic amino acid, with larger hydrophobic amino acids having side chains with molecular weights of at least about 90 g / mol, or at least about 130 g / mol, or at least about 141 g / mol. The size of an amino acid can be measured in terms of the SASA of the hydrophobic side chain. A hydrophobic amino acid can have a side chain with a SASA equal to or greater than that of alanine, or equal to or greater than that of glycine. A larger hydrophobic amino acid can have a side chain with a SASA greater than that of alanine, or greater than that of glycine. The hydrophobic amino acid may have an aromatic or heteroaromatic group with a SASA of about piperidine-2-carboxylic acid or greater, about tryptophan or greater, about phenylalanine or greater, or about naphthylalanine or greater. H1 ) is at least about 200 Å 2 , at least about 210 Å 2 , at least about 220 Å 2 , at least about 240 Å 2 , at least about 250 Å 2 , at least about 260 Å 2 , at least about 270 Å 2 , at least about 280 Å 2 , at least about 290 Å 2 , at least about 300 Å 2 , at least about 310 Å 2 , at least about 320 Å 2 , or at least about 330 Å 2 The second hydrophobic amino acid (AA H2 ) is at least about 200 Å 2 , at least about 210 Å 2, at least about 220 Å 2 , at least about 240 Å 2 , at least about 250 Å 2 , at least about 260 Å 2 , at least about 270 Å 2 , at least about 280 Å 2 , at least about 290 Å 2 , at least about 300 Å 2 , at least about 310 Å 2 , at least about 320 Å 2 , or at least about 330 Å 2 The side chain may have a SASA of AA H1 and A.A. H2 The side chains of 2 , at least about 360 Å 2 and at least about 370 Å 2 , at least about 380 Å 2 , at least about 390 Å 2 and at least about 400 Å 2 , at least about 410 Å 2 and at least about 420 Å 2 , at least about 430 Å 2 and at least about 440 Å 2 and at least about 450 Å 2 , at least about 460 Å 2 and at least about 470 Å 2 , at least about 480 Å 2 and at least about 490 Å 2 , about 500Å 2 Greater than about 510 Å 2 , at least about 520 Å 2 , at least about 530 Å 2 , at least about 540 Å 2 , at least about 550 Å 2 , at least about 560 Å 2 , at least about 570 Å 2 , at least about 580 Å 2 , at least about 590 Å 2 , at least about 600 Å 2 , at least about 610 Å 2 , at least about 620 Å2 , at least about 630 Å 2 , at least about 640 Å 2 , approximately 650 Å 2 Greater than about 660 Å 2 , at least about 670 Å 2 , at least about 680 Å 2 , at least about 690 Å 2 , or at least about 700 Å 2 AA H2 AA H1 The hydrophobic amino acid residue may have a side chain with a SASA that is equal to or less than the SASA of the hydrophobic side chain of (I). By way of example, and not limitation, a cCPP having a NaI-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a Phe-Arg motif. A cCPP having a Phe-Nal-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a NaI-Phe-Arg motif, and a phe-Nal-Arg motif may exhibit improved cytoplasmic delivery efficiency compared to an otherwise identical cCPP having a NaI-Phe-Arg motif.

[0239] As used herein, "hydrophobic surface area" or "SASA" is the solvent-accessible surface area (reported as square angstroms) of an amino acid side chain; for example, SASA can be calculated using the "rolling ball" algorithm developed by Shrake & Rupley (J Mol Biol. 79(2):351-71), which is incorporated herein by reference in its entirety for all purposes. This algorithm uses a solvent "sphere" of a specific radius to probe the surface of the molecule. A typical value for the sphere is 1.4 Å, which approximates the radius of a water molecule.

[0240] The SASA values ​​for certain side chains are shown in Table 3 below. The SASA values ​​described herein are based on the theoretical values ​​listed in Table 3 below, as reported by Tien et al. (PLOS ONE 8(11):e80635, doi.org / 10.1371 / journal.pone.0080635), which is incorporated herein by reference in its entirety for all purposes.

[0241] [Table 4]

[0242] Amino acid residues having a side chain containing a guanidine group, a guanidine substituent, or their protonated forms As used herein, guanidine has the structure:

[0243] [ka] Refers to...

[0244] As used herein, the protonated form of guanidine has the structure:

[0245] [ka] Refers to...

[0246] A guanidine substituent refers to a functional group on the side chain of an amino acid that is positively charged at or above physiological pH, or a functional group that can mimic the hydrogen bond donating and accepting activity of a guanidinium group.

[0247] The guanidine substituents facilitate cell penetration and delivery of therapeutic agents while reducing toxicity associated with the guanidine group or its protonated form. The cCPP may comprise at least one amino acid having a side chain containing a guanidine or guanidinium substituent. The cCPP may comprise at least two amino acids having side chains containing guanidine or guanidinium substituents. The cCPP may comprise at least three amino acids having side chains containing guanidine or guanidinium substituents.

[0248] The guanidine or guanidinium group may be an isostere of guanidine or guanidinium. The guanidine or guanidinium substituent may be less basic than guanidine.

[0249] As used herein, a guanidine substituent is:

[0250] [ka] or its protonated form.

[0251] The present disclosure provides cCPPs comprising 4 to 20 amino acid residues, wherein (i) at least one amino acid has a side chain comprising a guanidine group or its protonated form, and (ii) at least one amino acid residue has no side chain, or

[0252] [ka] or a protonated form thereof, and (iii) at least two amino acid residues have side chains that independently include an aromatic group or a heteroaromatic group.

[0253] At least two amino acid residues may have no side chains, or

[0254] [ka] or a side chain including its protonated form. As used herein, when no side chain is present, an amino acid residue has two hydrogen atoms on the carbon atom connecting the amine and carboxylic acid (e.g., -CH2-).

[0255] cCPP consists of the following parts:

[0256] [ka] Or it may comprise at least one amino acid having a side chain comprising one of its protonated forms.

[0257] The cCPP can comprise at least two amino acids, each amino acid independently comprising the moiety:

[0258] [ka] or one of its protonated forms. At least two amino acids are

[0259] [ka] or its protonated form. At least one amino acid may have a side chain containing the same moiety selected from

[0260] [ka] or a side chain containing the protonated form thereof. At least two amino acids

[0261] [ka] or a side chain containing its protonated form.

[0262] [ka] or a side chain containing the protonated form thereof.

[0263] [ka] or the side chain may include its protonated form.

[0264] [ka] or may have side chains containing the protonated form thereof.

[0265] [ka] Or the protonated form can be attached to the terminus of an amino acid side chain.

[0266] [ka] can be attached to the terminus of an amino acid side chain.

[0267] The cCPP may comprise two, three, four, five, or six amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise two amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise three amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise four amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise five amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise six amino acid residues that independently have a side chain containing a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise two, three, four, or five amino acid residues that independently have a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise two, three, or four amino acid residues that independently have a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise two or three amino acid residues that independently have a side chain comprising a guanidine group, a guanidine substituent, or a protonated form thereof. The cCPP may comprise at least one amino acid residue that has a side chain comprising a guanidine group or a protonated form thereof. The cCPP may comprise two amino acid residues that have a side chain comprising a guanidine group or a protonated form thereof. The cCPP may comprise three amino acid residues that have a side chain comprising a guanidine group or a protonated form thereof.

[0268] The amino acid residues may independently have side chains containing non-adjacent guanidine groups, guanidine substituents, or their protonated forms. Two amino acid residues may independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms may be adjacent. Three amino acid residues may independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms may be adjacent. Four amino acid residues may independently have side chains containing guanidine groups, guanidine substituents, or their protonated forms may be adjacent. Adjacent amino acid residues may have the same stereochemistry. Adjacent amino acids may have alternate stereochemistry.

[0269] The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or a protonated form thereof can be L-amino acids. The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or a protonated form thereof can be D-amino acids. The amino acid residues independently having side chains containing a guanidine group, a guanidine substituent, or a protonated form thereof can be a mixture of L- and D-amino acids.

[0270] Each amino acid residue having a side chain containing a guanidine group or its protonated form can independently be a residue of arginine, homoarginine, 2-amino-3-propionic acid, 2-amino-4-guanidinobutyric acid, or its protonated form. Each amino acid residue having a side chain containing a guanidine group or its protonated form can independently be a residue of arginine or its protonated form.

[0271] Each amino acid having a side chain containing a guanidine substituent or its protonated form can independently

[0272] [ka] or its protonated form.

[0273] Without being bound by theory, it is hypothesized that the guanidine substituent has reduced basicity compared to arginine and, in some cases, is uncharged (e.g., -N(H)C(O)) at physiological pH, allowing it to maintain bidentate hydrogen-bonding interactions with phospholipids on the plasma membrane, which is believed to facilitate effective membrane binding and subsequent internalization. Removal of the positive charge is also believed to reduce the toxicity of cCPPs.

[0274] Those skilled in the art will understand that the N-terminus and / or C-terminus of the non-natural flavor hydrophobic amino acids form an amide bond upon incorporation into the peptides disclosed herein.

[0275] A cCPP can include a first amino acid having a side chain comprising an aromatic or heteroaromatic group and a second amino acid having a side chain comprising an aromatic or heteroaromatic group, where the N-terminus of the first glycine forms a peptide bond with the first amino acid having a side chain comprising an aromatic or heteroaromatic group, and the C-terminus of the first glycine forms a peptide bond with the second amino acid having a side chain comprising an aromatic or heteroaromatic group. By convention, the term "first amino acid" often refers to the N-terminal amino acid of a peptide sequence; however, as used herein, "first amino acid" is used to distinguish a reference amino acid from another amino acid (e.g., a "second amino acid") in a cCPP, and thus the term "first amino acid" can refer to the amino acid located at the N-terminus of a peptide sequence.

[0276] The cCPP can include an N-terminus of a second glycine that forms a peptide bond with an amino acid having a side chain that includes an aromatic or heteroaromatic group, and a C-terminus of the second glycine that forms a peptide bond with an amino acid having a side chain that includes a guanidine group or a protonated form thereof.

[0277] The cCPP can include a first amino acid having a side chain comprising a guanidine group or its protonated form, and a second amino acid having a side chain comprising a guanidine group or its protonated form, wherein the N-terminus of the third glycine forms a peptide bond with the first amino acid having a side chain comprising a guanidine group or its protonated form, and the C-terminus of the third glycine forms a peptide bond with the second amino acid having a side chain comprising a guanidine group or its protonated form.

[0278] The cCPP may comprise an asparagine, aspartic acid, glutamine, glutamic acid, or homoglutamine residue. The cCPP may comprise an asparagine residue. The cCPP may comprise a glutamine residue.

[0279] cCPPs can each independently comprise residues of tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butyl-phenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, 3-(9-anthryl)-alanine.

[0280] Without wishing to be bound by theory, it is believed that the chirality of amino acids in a cCPP can affect cytoplasmic uptake efficiency. A cCPP can contain at least one D amino acid. A cCPP can contain 1 to 15 D amino acids. A cCPP can contain 1 to 10 D amino acids. A cCPP can contain 1, 2, 3, or 4 D amino acids. A cCPP can contain 2, 3, 4, 5, 6, 7, or 8 adjacent amino acids with alternating D and L chirality. A cCPP can contain three adjacent amino acids with the same chirality. A cCPP can contain two adjacent amino acids with the same chirality. At least two amino acids can have opposite chirality. At least two amino acids with opposite chirality can be adjacent to each other. At least three amino acids can have alternate stereochemistry relative to each other. At least three amino acids with alternate chirality relative to each other can be adjacent to each other. At least four amino acids have alternate stereochemistry relative to one another. At least four amino acids of alternate chirality relative to one another can be adjacent to one another. At least two amino acids can have the same chirality. At least two amino acids of the same chirality can be adjacent to one another. At least two amino acids have the same chirality and at least two amino acids have opposite chirality. At least two amino acids of opposite chirality can be adjacent to at least two amino acids of the same chirality. Thus, adjacent amino acids in a cCPP can have any of the following sequences: DL, LD, DLLD, LDDL, LDLLD, DLDDL, DLLDL, or LDDLD. The amino acid residues forming the cCPP can all be L-amino acids. The amino acid residues forming the cCPP can all be D-amino acids.

[0281] At least two amino acids may have different chiralities. At least two amino acids with different chiralities may be adjacent to each other. At least three amino acids may have different chiralities relative to adjacent amino acids. At least four amino acids may have different chiralities relative to adjacent amino acids. At least two amino acids have the same chirality and at least two amino acids have different chiralities. One or more amino acid residues forming a cCPP may be achiral. A cCPP may contain a motif of 3, 4, or 5 amino acids, in which two amino acids with the same chirality may be separated by an achiral amino acid. A cCPP may contain the following sequences: DXD, DXDX, DXDXD, LXL, LXLX, or LXLXL (where X is an achiral amino acid). The achiral amino acid may be glycine.

[0282] an amino acid having a side chain comprising:

[0283] [ka] or its protonated form may be adjacent to an amino acid having a side chain comprising an aromatic or heteroaromatic group.

[0284] [ka] An amino acid having a side chain comprising guanidine or its protonated form may be adjacent to at least one amino acid having a side chain comprising guanidine or its protonated form. An amino acid having a side chain comprising guanidine or its protonated form may be adjacent to an amino acid having a side chain comprising an aromatic group or a heteroaromatic group. Two amino acids having side chains

[0285] [ka] Two amino acids having a side chain containing guanidine or its protonated form can be adjacent to each other. Two amino acids having a side chain containing guanidine or its protonated form can be adjacent to each other. A cCPP comprises at least two adjacent amino acids having side chains that can contain an aromatic or heteroaromatic group, and

[0286] [ka] or at least two non-adjacent amino acids having side chains containing an aromatic or heteroaromatic group, or a protonated form thereof. A cCPP may have at least two adjacent amino acids having side chains containing an aromatic or heteroaromatic group, and

[0287] [ka] or at least two non-adjacent amino acids having side chains containing the protonated form thereof. Adjacent amino acids may have the same chirality. Adjacent amino acids may have opposite chiralities. Other combinations of amino acids may have any arrangement of D and L amino acids, such as any of the sequences described in the previous paragraph.

[0288] At least two amino acids having the following side chains:

[0289] [ka] Or at least two amino acids having a side chain containing a guanidine group or its protonated form alternate with at least two amino acids having a side chain containing a guanidine group or its protonated form.

[0290] cCPP has the formula (A):

[0291] [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4, R5, R6, R7 are independently H or an amino acid side chain; at least one of R4, R5, R6, and R7 is a side chain of 3-guanidino-2-aminopropionic acid, 4-guanidino-2-aminobutanoic acid, arginine, homoarginine, N-methylarginine, N,N-dimethylarginine, 2,3-diaminopropionic acid, 2,4-diaminobutanoic acid, lysine, N-methyllysine, N,N-dimethyllysine, N-ethyllysine, N,N,N-trimethyllysine, 4-guanidinophenylalanine, citrulline, N,N-dimethyllysine, β-homoarginine, or 3-(1-piperidyl)alanine; AA SC is an amino acid side chain, q is 1, 2, 3 or 4.

[0292] In embodiments, the cyclic peptide of Formula (A) is not FfΦRrRrQ (SEQ ID NO: 67). In embodiments, the cyclic peptide of Formula (A) is FfΦRrRrQ (SEQ ID NO: 67).

[0293] cCPP has the formula (I):

[0294] [ka] Things, or a protonated form thereof, During the ceremony, R1, R2, and R3 may each independently be H or an amino acid residue having a side chain containing an aromatic group; at least one of R1, R2, and R3 is an aromatic or heteroaromatic side chain of an amino acid; R4 and R7 are independently H or an amino acid side chain; AA SC is an amino acid side chain, q is 1, 2, 3 or 4; Each m is independently an integer 0, 1, 2, or 3.

[0295] R1, R2, and R3 can each independently be H, -alkylene-aryl, or -alkylene-heteroaryl. R1, R2, and R3 can each independently be H, -C 1~3 Alkylene-aryl, or -C 1~3 R1, R2, and R3 can each independently be H or -alkylene-aryl. R1, R2, and R3 can each independently be H or -C 1~3 It can be alkylene-aryl. 1~3 The alkylene can be methylene. The aryl can be 6- to 14-membered aryl. The heteroaryl can be 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl. R1, R2, and R3 are each independently selected from H, -C, 1~3 Alkylene-Ph or -C 1~3 R1, R2, and R3 can each independently be H, -CH2Ph, or -CH2naphthyl. R1, R2, and R3 can each independently be H or -CH2Ph.

[0296] R1, R2, and R3 can each independently be the side chain of tyrosine, phenylalanine, 1-naphthylalanine, 2-naphthylalanine, tryptophan, 3-benzothienylalanine, 4-phenylphenylalanine, 3,4-difluorophenylalanine, 4-trifluoromethylphenylalanine, 2,3,4,5,6-pentafluorophenylalanine, homophenylalanine, β-homophenylalanine, 4-tert-butyl-phenylalanine, 4-pyridinylalanine, 3-pyridinylalanine, 4-methylphenylalanine, 4-fluorophenylalanine, 4-chlorophenylalanine, or 3-(9-anthryl)-alanine.

[0297] R1 can be the side chain of tyrosine. R1 can be the side chain of phenylalanine. R1 can be the side chain of 1-naphthylalanine. R1 can be the side chain of 2-naphthylalanine. R1 can be the side chain of tryptophan. R3 can be the side chain of 3-benzothienylalanine. R1 can be the side chain of 4-phenylphenylalanine. R1 can be the side chain of 3,4-difluorophenylalanine. R1 can be the side chain of 4-trifluoromethylphenylalanine. R1 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R1 can be the side chain of homophenylalanine. R1 can be the side chain of β-homophenylalanine. R1 can be the side chain of 4-tert-butyl-phenylalanine. R1 can be the side chain of 4-pyridinylalanine. R1 can be the side chain of 3-pyridinylalanine. R1 can be the side chain of 4-methylphenylalanine. R1 can be the side chain of 4-fluorophenylalanine. R1 can be the side chain of 4-chlorophenylalanine. R1 can be the side chain of 3-(9-anthryl)-alanine.

[0298] R2 can be the side chain of tyrosine. R2 can be the side chain of phenylalanine. R2 can be the side chain of 1-naphthylalanine. R1 can be the side chain of 2-naphthylalanine. R2 can be the side chain of tryptophan. R2 can be the side chain of 3-benzothienylalanine. R2 can be the side chain of 4-phenylphenylalanine. R2 can be the side chain of 3,4-difluorophenylalanine. R2 can be the side chain of 4-trifluoromethylphenylalanine. R2 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R2 can be the side chain of homophenylalanine. R2 can be the side chain of β-homophenylalanine. R2 can be the side chain of 4-tert-butyl-phenylalanine. R2 can be the side chain of 4-pyridinylalanine. R2 can be the side chain of 3-pyridinylalanine. R2 can be the side chain of 4-methylphenylalanine. R2 can be the side chain of 4-fluorophenylalanine. R2 can be the side chain of 4-chlorophenylalanine. R2 can be the side chain of 3-(9-anthryl)-alanine.

[0299] R3 can be the side chain of tyrosine. R3 can be the side chain of phenylalanine. R3 can be the side chain of 1-naphthylalanine. R3 can be the side chain of 2-naphthylalanine. R3 can be the side chain of tryptophan. R3 can be the side chain of 3-benzothienylalanine. R3 can be the side chain of 4-phenylphenylalanine. R3 can be the side chain of 3,4-difluorophenylalanine. R3 can be the side chain of 4-trifluoromethylphenylalanine. R3 can be the side chain of 2,3,4,5,6-pentafluorophenylalanine. R3 can be the side chain of homophenylalanine. R3 can be the side chain of β-homophenylalanine. R3 can be the side chain of 4-tert-butyl-phenylalanine. R3 can be the side chain of 4-pyridinylalanine. R3 can be the side chain of 3-pyridinylalanine. R3 can be the side chain of 4-methylphenylalanine. R3 can be the side chain of 4-fluorophenylalanine. R3 can be the side chain of 4-chlorophenylalanine. R3 can be the side chain of 3-(9-anthryl)-alanine.

[0300] R4 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1~3 Alkylene-aryl, or -C 1~3 R4 can be H or -alkylene-aryl. R4 can be H or -C 1~3 It can be alkylene-aryl. 1~3 The alkylene can be methylene. The aryl can be 6- to 14-membered aryl. The heteroaryl can be 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl. R4 can be H, -C 1~3 Alkylene-Ph or -C 1~3R4 can be alkylene-naphthyl. R4 can be H or the side chain of an amino acid in Table 1 or Table 3. R4 can be H or an amino acid residue having a side chain containing an aromatic group. R4 can be H, -CH2Ph, or -CH2naphthyl. R4 can be H or -CH2Ph.

[0301] R5 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1~3 Alkylene-aryl, or -C 1~3 R5 can be H or -alkylene-aryl. R5 can be H or -C 1~3 It can be alkylene-aryl. 1~3 The alkylene can be methylene. The aryl can be 6- to 14-membered aryl. The heteroaryl can be 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl. R5 can be H, -C 1~3 Alkylene-Ph or -C 1~3 R5 can be H or the side chain of an amino acid in Table 1 or Table 3. R4 can be H or an amino acid residue having a side chain containing an aromatic group. R5 can be H, -CH2Ph, or -CH2 naphthyl. R4 can be H or -CH2Ph.

[0302] R6 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1~3 Alkylene-aryl, or -C 1~3 R6 can be H or -alkylene-aryl. R6 can be H or -C 1~3 It can be alkylene-aryl. 1~3The alkylene can be methylene. The aryl can be 6- to 14-membered aryl. The heteroaryl can be 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl. R6 can be H, -C 1~3 Alkylene-Ph or -C 1~3 R6 can be alkylene-naphthyl. R6 can be H or the side chain of an amino acid in Table 1 or Table 3. R6 can be H or an amino acid residue having a side chain containing an aromatic group. R6 can be H, -CH2Ph, or -CH2 naphthyl. R6 can be H or -CH2Ph.

[0303] R7 can be H, -alkylene-aryl, -alkylene-heteroaryl. 1~3 Alkylene-aryl, or -C 1~3 R7 can be H or -alkylene-aryl. R7 can be H or -C 1~3 It can be alkylene-aryl. 1~3 The alkylene can be methylene. The aryl can be 6- to 14-membered aryl. The heteroaryl can be 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, and S. The aryl can be selected from phenyl, naphthyl, or anthracenyl. The aryl can be phenyl or naphthyl. The aryl can be phenyl. The heteroaryl can be pyridyl, quinolyl, or isoquinolyl. R7 can be H, -C 1~3 Alkylene-Ph or -C 1~3 R7 can be alkylene-naphthyl. R7 can be H or the side chain of an amino acid in Table 1 or Table 3. R7 can be H or an amino acid residue having a side chain containing an aromatic group. R7 can be H, -CH2Ph, or -CH2 naphthyl. R7 can be H or -CH2Ph.

[0304] One, two, or three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. One of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Two of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. At least one of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph. Up to four of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph.

[0305] One, two, or three of R1, R2, R3, and R4 are -CH2Ph. One of R1, R2, R3, and R4 is -CH2Ph. Two of R1, R2, R3, and R4 are -CH2Ph. Three of R1, R2, R3, and R4 are -CH2Ph. At least one of R1, R2, R3, and R4 is -CH2Ph.

[0306] One, two, or three of R1, R2, R3, R4, R5, R6, and R7 can be H. One of R1, R2, R3, R4, R5, R6, and R7 can be H. Two of R1, R2, R3, R4, R5, R6, and R7 can be H. Three of R1, R2, R3, R5, R6, and R7 can be H. At least one of R1, R2, R3, R4, R5, R6, and R7 can be H. Up to three of R1, R2, R3, R4, R5, R6, and R7 can be -CH2Ph.

[0307] One, two or three of R1, R2, R3, and R4 are H. One of R1, R2, R3, and R4 is H. Two of R1, R2, R3, and R4 are H. Three of R1, R2, R3, and R4 are H. At least one of R1, R2, R3, and R4 is H.

[0308] At least one of R4, R5, R6, and R7 can be the side chain of 3-guanidino-2-aminopropionic acid. At least one of R4, R5, R6, and R7 can be the side chain of 4-guanidino-2-aminobutanoic acid. At least one of R4, R5, R6, and R7 can be the side chain of arginine. At least one of R4, R5, R6, and R7 can be the side chain of homoarginine. At least one of R4, R5, R6, and R7 can be the side chain of N-methylarginine. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethylarginine. At least one of R4, R5, R6, and R7 can be the side chain of 2,3-diaminopropionic acid. At least one of R4, R5, R6, and R7 can be the side chain of 2,4-diaminobutanoic acid, lysine. At least one of R4, R5, R6, and R7 can be the side chain of N-methyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N-ethyllysine. At least one of R4, R5, R6, and R7 can be the side chain of N,N,N-trimethyllysine, 4-guanidinophenylalanine. At least one of R4, R5, R6, and R7 can be the side chain of citrulline. At least one of R4, R5, R6, and R7 can be the side chain of N,N-dimethyllysine, β-homoarginine. At least one of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.

[0309] At least two of R4, R5, R6, and R7 can be the side chains of 3-guanidino-2-aminopropionic acid. At least two of R4, R5, R6, and R7 can be the side chains of 4-guanidino-2-aminobutanoic acid. At least two of R4, R5, R6, and R7 can be the side chains of arginine. At least two of R4, R5, R6, and R7 can be the side chains of homoarginine. At least two of R4, R5, R6, and R7 can be the side chains of N-methylarginine. At least two of R4, R5, R6, and R7 can be the side chains of N,N-dimethylarginine. At least two of R4, R5, R6, and R7 can be the side chains of 2,3-diaminopropionic acid. At least two of R4, R5, R6, and R7 can be the side chains of 2,4-diaminobutanoic acid, lysine. At least two of R4, R5, R6, and R7 can be the side chains of N-methyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N-ethyllysine. At least two of R4, R5, R6, and R7 can be the side chains of N,N,N-trimethyllysine, 4-guanidinophenylalanine. At least two of R4, R5, R6, and R7 can be the side chains of citrulline. At least two of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine, β-homoarginine. At least two of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.

[0310] At least three of R4, R5, R6, and R7 can be the side chains of 3-guanidino-2-aminopropionic acid. At least three of R4, R5, R6, and R7 can be the side chains of 4-guanidino-2-aminobutanoic acid. At least three of R4, R5, R6, and R7 can be the side chains of arginine. At least three of R4, R5, R6, and R7 can be the side chains of homoarginine. At least three of R4, R5, R6, and R7 can be the side chains of N-methylarginine. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethylarginine. At least three of R4, R5, R6, and R7 can be the side chains of 2,3-diaminopropionic acid. At least three of R4, R5, R6, and R7 can be the side chains of 2,4-diaminobutanoic acid, lysine. At least three of R4, R5, R6, and R7 can be the side chains of N-methyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N-ethyllysine. At least three of R4, R5, R6, and R7 can be the side chains of N,N,N-trimethyllysine, 4-guanidinophenylalanine. At least three of R4, R5, R6, and R7 can be the side chains of citrulline. At least three of R4, R5, R6, and R7 can be the side chains of N,N-dimethyllysine, β-homoarginine. At least three of R4, R5, R6, and R7 can be the side chain of 3-(1-piperidinyl)alanine.

[0311] AA SC can be the side chain of an asparagine, glutamine, or homoglutamine residue. SC may be the side chain of a glutamine residue. SCFor example, a cCPP may further comprise a linker conjugated to an asparagine, glutamine, or homoglutamine residue. Thus, a cCPP may further comprise a linker conjugated to an asparagine, glutamine, or homoglutamine residue. A cCPP may further comprise a linker attached to a glutamine residue.

[0312] q can be 1, 2 or 3. q can be 1 or 2. q can be 1. q can be 2. q can be 3. q can be 4.

[0313] m can be 1 to 3. m can be 1 or 2. m can be 0. m can be 1. m can be 2. m can be 3.

[0314] The cCPP of formula (A) is a compound of formula (I)

[0315] [ka] or a protonated form thereof, wherein AA SC , R1, R2, R3, R 4- , R7, m and q are as defined herein.

[0316] The cCPP of formula (A) is represented by formula (Ia) or formula (Ib):

[0317] [ka] The structure of or its protonated form, AA SC , R1, R2, R3, R4, and m are as defined herein.

[0318] The cCPP of formula (A) is represented by formula (I-1), (I-2), (I-3) or (I-4):

[0319] [ka]

[0320] [ka] The structure of or a protonated form thereof, wherein AA SC and m is as defined herein.

[0321] The cCPP of formula (A) is represented by formula (I-5) or (I-6):

[0322] [ka] or a protonated form thereof, wherein AA SC is as defined herein.

[0323] The cCPP of formula (A) is represented by formula (I-1):

[0324] [ka] or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.

[0325] The cCPP of formula (A) is represented by formula (I-2):

[0326] [ka] or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.

[0327] The cCPP of formula (A) is represented by formula (I-3):

[0328] [ka] or a protonated form thereof, During the ceremony, A.A.SC and m is as defined herein.

[0329] The cCPP of formula (A) is represented by formula (I-4):

[0330] [ka] or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.

[0331] The cCPP of formula (A) is represented by formula (I-5):

[0332] [ka] or a protonated form thereof, During the ceremony, A.A. SC and m is as defined herein.

[0333] The cCPP of formula (A) is represented by formula (I-6):

[0334] [ka] or a protonated form thereof, wherein AA SC and m is as defined herein.

[0335] The cCPP may comprise one of the following sequences: FGFGRGR (SEQ ID NO: 68), GfFGrGr (SEQ ID NO: 69), FfΦGRGR (SEQ ID NO: 70), FfFGRGR (SEQ ID NO: 71), or FfΦGrGr (SEQ ID NO: 72). The cCPP may have one of the following sequences: FGFΦΦ (SEQ ID NO: 73), GfFGrGrQ (SEQ ID NO: 74), FfΦGRGRQ (SEQ ID NO: 75), FfFGRGRQ (SEQ ID NO: 76), or FfΦGrGrQ (SEQ ID NO: 77).

[0336] The present disclosure also provides a compound of formula (II):

[0337] [ka] In the context of a cCPP having the structure: AA SC is an amino acid side chain, R 1a , R 1b , and R 1c are each independently a 6- to 14-membered aryl or a 6- to 14-membered heteroaryl; R 2a , R 2b , R 2c and R 2d are independently amino acid side chains, R 2a , R 2b , R 2c and R 2d At least one of the

[0338] [ka] or its protonated form, R 2a , R 2b , R 2c and R 2d at least one of is guanidine or its protonated form; each n″ is independently an integer 0, 1, 2, 3, 4, or 5; each n' is independently an integer from 0, 1, 2, or 3; If n' is 0, then R 2a , R 2b , R 2b or R 2d does not exist.

[0339] R 2a , R 2b , R 2c and R 2d At least two of the

[0340] [ka] or its protonated form. 2a , R 2b , R 2c and R 2d Two or three of them are

[0341] [ka] or its protonated form. 2a , R 2b , R 2c and R 2d One of them is

[0342] [ka] or its protonated form. 2a , R 2b , R 2c and R 2d At least one of the

[0343] [ka] or its protonated form, R 2a , R 2b , R 2c and R 2d The remainder of R may be guanidine or its protonated form. 2a , R 2b , R 2c and R 2d At least two of the

[0344] [ka] or its protonated form. 2a , R 2b , R 2c and R 2d The remainder may be guanidine or its protonated form.

[0345] R 2a , R2b , R 2c and R 2d All of this is

[0346] [ka] or its protonated form. 2a , R 2b , R 2c and R 2d At least one of

[0347] [ka] or its protonated form, R 2a , R 2b , R 2c and R 2d The remainder of R may be a guaninide or its protonated form. 2a , R 2b , R 2c and R 2d The base is

[0348] [ka] or its protonated form, R 2a , R 2b , R 2c and R 2d The remainder is guanidine or its protonated form.

[0349] R 2a , R 2b , R 2c and R 2d can each independently be 2,3-diaminopropionic acid, 2,4-diaminobutyric acid side chain, ornithine, lysine, methyllysine, dimethyllysine, trimethyllysine, homo-lysine, serine, homo-serine, threonine, allo-threonine, histidine, 1-methylhistidine, 2-aminobutanedioic acid, aspartic acid, glutamic acid, or homo-glutamic acid.

[0350] AA SC teeth

[0351] [ka] where t can be an integer from 0 to 5. SC teeth

[0352] [ka] In the formula, t can be an integer from 0 to 5. t can be 1 to 5. t can be 2 or 3. t can be 2. t can be 3

[0353] R 1a , R 1b , and R 1c Each R can independently be a 6- to 14-membered aryl. 1a , R 1b , and R 1c R may each independently be a 6- to 14-membered heteroaryl having one or more heteroatoms selected from N, O, or S. 1a , R 1b , and R 1c Each R may be independently selected from phenyl, naphthyl, anthracenyl, pyridyl, quinolyl, or isoquinolyl. 1a , R 1b , and R 1c Each R may be independently selected from phenyl, naphthyl, or anthracenyl. 1a , R 1b , and R 1c Each R may independently be phenyl or naphthyl. 1a , R 1b , and R 1c may each independently be selected from pyridyl, quinolyl, or isoquinolyl.

[0354] Each n' can independently be 1 or 2. Each n' can be 1. Each n' can be 2. At least one n' can be 0. At least one n' can be 1. At least one n' can be 2. At least one n' can be 3. At least one n' can be 4. At least one n' can be 5.

[0355] Each n" can independently be an integer from 1 to 3. Each n" can independently be 2 or 3. Each n" can be 2. Each n" can be 3. At least one n" can be 0. At least one n" can be 1. At least one n" can be 2. At least one n" can be 3.

[0356] Each n" can independently be 1 or 2, and each n' can independently be 2 or 3. Each n" can be 1, and each n' can independently be 2 or 3. Each n" can be 1, and each n' can be 2. Each n" is 1, and each n' is 3.

[0357] The cCPP of formula (II) is represented by the formula (II-1):

[0358] [ka] may have the structure In the formula, R 1a , R 1b , R 1c , R 2a , R 2b , R 2c , R 2d , A.A. SC , n' and n'' are as defined herein.

[0359] The cCPP of formula (II) has the formula (IIa):

[0360] [ka] may have the structure In the formula, R 1a , R1b , R 1c , R 2a , R 2b , R 2c , R 2d , A.A. SC and n' are as defined herein.

[0361] The cCPP of formula (II) may be represented by the formula (IIb):

[0362] [ka] may have the structure In the formula, R 2a , R 2b , A.A. SC and n' is as defined herein.

[0363] cCPP has the formula (IIb):

[0364] [ka] or a protonated form thereof, During the ceremony, AA SC and n' are as defined herein.

[0365] The cCPP of formula (IIa) has the following structure:

[0366] [ka] wherein AA SC and n is as defined herein.

[0367] The cCPP of formula (IIa) has the following structure:

[0368] [ka] wherein AA SC and n is as defined herein.

[0369] The cCPP of formula (IIa) has the following structure:

[0370] [ka] wherein AA SC and n is as defined herein.

[0371] The cCPP of formula (II) has the structure:

[0372] [ka] may have:

[0373] The cCPP of formula (II) has the structure:

[0374] [ka] may have:

[0375] cCPP has the structure of formula (III):

[0376] [ka] and During the ceremony, AA SC is an amino acid side chain, R 1a , R 1b , and R 1c are each independently a 6- to 14-membered aryl or a 6- to 14-membered heteroaryl; R 2a and R 2c are each independently H,

[0377] [ka] or its protonated form, R 2band R 2d are each independently guanidine or its protonated form; each n'' is independently an integer from 1 to 3; each n' is independently an integer from 1 to 5; Each p' is independently an integer from 0 to 5.

[0378] The cCPP of formula (III) is represented by the formula (III-1):

[0379] [ka] may have the structure During the ceremony, AA SC , R 1a , R 1b , R 1c , R 2a , R 2c , R 2b , R 2d , n', n'' and p' are as defined herein.

[0380] The cCPP of formula (III) has the formula (IIIa):

[0381] [ka] may have the structure During the ceremony, AA SC , R 2a , R 2c , R 2b , R 2d , n', n'', and p' are as defined herein.

[0382] In formulas (III), (III-1), and (IIIa), R a and R c can be H. R a and R c can be H, and R b and R d R can each independently be guanidine or its protonated form. acan be H. R b can be H. p' can be 0. R a and R c may be H and each p' may be 0.

[0383] In formulas (III), (III-1) and (IIIa), R a and R c can be H, and R b and R d can each independently be guanidine or its protonated form, n'' can be 2 or 3, and each p' can be 0.

[0384] p' can be 0. p' can be 1. p' can be 2. p' can be 3. p' can be 4. p' can be 5.

[0385] cCPP has the structure:

[0386] [ka] may have:

[0387] The cCPP of formula (A) can be selected from:

[0388] [Table 5]

[0389] The cCPP of formula (A) can be selected from:

[0390] [Table 6]

[0391] In embodiments, the cCPP is selected from the following:

[0392] [Table 7] where Φ = L-naphthylalanine, φ = D-naphthylalanine, Ω = L-norleucine

[0393] In embodiments, the cCPP is not selected from the following:

[0394] [Table 8] where Φ = L-naphthylalanine, φ = D-naphthylalanine, Ω = L-norleucine

[0395] AA SC can be conjugated to a linker.

[0396] Linker The cCPP of the present disclosure can be conjugated to a linker. The linker can link a cargo to the cCPP. The linker can be attached to the side chain of an amino acid of the cCPP, and the cargo can be attached to an appropriate position on the linker.

[0397] The linker may be any suitable moiety capable of conjugating the cCPP to one or more additional moieties, such as an exocyclic peptide (EP) and / or cargo. Prior to attachment to the cCPP and one or more additional moieties, the linker has two or more functional groups, each capable of independently forming a covalent bond to the cCPP and one or more additional moieties. When the cargo is an oligonucleotide, the linker may be covalently attached to the 5'-end of the cargo or the 3'-end of the cargo. The linker may be covalently attached to the 5'-end of the cargo. The linker may be covalently attached to the 3'-end of the cargo. When the cargo is a peptide, the linker may be covalently attached to the N-terminus or C-terminus of the cargo. The linker may be covalently attached to the backbone of the oligonucleotide or peptide cargo. The linker may be any suitable moiety capable of conjugating the cCPP described herein to a cargo such as an oligonucleotide, peptide, or small molecule.

[0398] The linker may comprise a hydrocarbon linker.

[0399] The linker may comprise a cleavage site, which may be a disulfide or a caspase cleavage site (e.g., Val-Cit-PABC).

[0400] The linker may be (i) one or more D or L amino acids, each of which is optionally substituted; (ii) an optionally substituted alkylene; (iii) an optionally substituted alkenylene; (iv) an optionally substituted alkynylene; (v) an optionally substituted carbocyclyl; (vi) an optionally substituted heterocyclyl; (vii) one or more -(R 1- JR 2 )z″-subunits, where R 1 and R 2 each, at each occurrence, is independently selected from alkylene, alkenylene, alkynylene, carbocyclyl, and heterocyclyl; and each J is independently selected from C, NR 3 , -NR 3 C(O)—, S, and O, where R 3 is independently selected from H, alkyl, alkenyl, alkynyl, carbocyclyl, and heterocyclyl, each of which is optionally substituted, and z" is an integer from 1 to 50; 1 (J)z- or -(JR 1 )z-, wherein each R 1 is, at each occurrence, independently alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl; and each J is independently C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3 is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl, each of which is optionally substituted, and z″ is an integer from 1 to 50; or (ix) the linker may include one or more of (i) through (x).

[0401] The linker may comprise one or more D or L amino acids and / or -(R 1- JR 2 )z″-, wherein R 1 and R 2 each, at each occurrence, is independently alkylene; and each J is independently C, NR 3 , -NR 3 C(O)—, S, and O, where R 4 is independently selected from H and alkyl, and z″ is an integer from 1 to 50, or a combination thereof.

[0402] The linker is -(OCH2CH2) z’ - (e.g., as a spacer), where z' is an integer from 1 to 23, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. "-(OCH2CH2)z" can also be referred to as polyethylene glycol (PEG).

[0403] The linker may comprise one or more amino acids. The linker may comprise a peptide. The linker may comprise -(OCH2CH2) z’ - (wherein z' is an integer from 1 to 23), and a peptide. The peptide may contain 2 to 10 amino acids. The linker may further contain a functional group (FG) that can react via click chemistry. FG may be an azide or an alkyne, and a triazole is formed when the cargo is conjugated to the linker.

[0404] The linker consists of (i) β-alanine and lysine residues, (ii) -(JR 1 )z-, or (iii) combinations thereof. 1 can independently be alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl, and each J can independently be C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl, each of which is optionally substituted, and z″ can be an integer from 1 to 50. Each R 1 can be alkylene and each J can be O.

[0405] The linker may be (i) a residue of β-alanine, glycine, lysine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid, or a combination thereof, and (ii) -(R 1- J)z”-or-(JR 1 )z″-. Each R 1 can independently be alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl, and each J can independently be C, NR 3 , -NR 3 C(O)—, S, or O, wherein R 3 is H, alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl, each of which is optionally substituted, and z″ can be an integer from 1 to 50. Each R 1 can be alkylene and each J can be O. The linker can include glycine, beta-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid, or combinations thereof.

[0406] The linker can be a trivalent linker. The linker has the structure:

[0407] [ka] wherein A1, B1, and C1 can independently be a hydrocarbon linker (e.g., NRH-(CH2) n -COOH), PEG linker (e.g., NRH-(CHO) n -COOH, where R is H, methyl, or ethyl), or one or more amino acid residues, and Z is independently a protecting group. Linkers also include disulfides [NH-(CHO) n -SS-(CH2O) n-COOH], or a cleavage site such as a caspase cleavage site (Val-Cit-PABC) can be incorporated.

[0408] The carbohydrate may be a residue of glycine or beta-alanine.

[0409] The linker is bivalent and can link the cCPP to a cargo. The linker is bivalent and can link the cCPP to an exocyclic peptide (EP).

[0410] The linker may be trivalent and may link the cCPP to the cargo and the EP.

[0411] The linker is a divalent or trivalent C1-C 50 and alkylene, wherein 1 to 25 methylene groups are optionally and independently replaced by -N(H)-, -N(C-C alkyl)-, -N(cycloalkyl)-, -O-, -C(O)-, -C(O)O-, -S-, -S(O)-, -S(O)-, -S(O)N(C-C alkyl)-, -S(O)N(cycloalkyl)-, -N(H)C(O)-, -N(C-C alkyl)C(O)-, -N(cycloalkyl)C(O)-, -C(O)N(H)-, -C(O)N(C-C alkyl), -C(O)N(cycloalkyl), aryl, heterocyclyl, heteroaryl, cycloalkyl, or cycloalkenyl. The linker may be a divalent or trivalent C-C 50 It can be alkylene, wherein 1 to 25 methylene groups are optionally and independently replaced by -N(H)-, -O-, -C(O)N(H)-, or combinations thereof.

[0412] The linker has the structure:

[0413] [ka] wherein each AA is independently an amino acid residue; * AA SC AA SCis a side chain of an amino acid residue of cCPP. x is an integer from 1 to 10, y is an integer from 1 to 5, and z is an integer from 1 to 10. X can be an integer from 1 to 5. X can be an integer from 1 to 3. X can be 1. Y can be an integer from 2 to 4. Y can be 4. Z can be an integer from 1 to 5. Z can be an integer from 1 to 3. Z can be 1. Each AA can be independently selected from glycine, β-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, and 6-aminohexanoic acid.

[0414] The cCPP can be linked to the cargo via a linker ("L"), which can be conjugated to the cargo via a linking group ("M").

[0415] The linker has the structure:

[0416] [ka] wherein x is an integer from 1 to 10, y is an integer from 1 to 5, z is an integer from 1 to 10, and each AA is independently an amino acid residue; * AA SC AA SC is the side chain of an amino acid residue of the cCPP, and M is a linking group as defined herein.

[0417] The linker has the structure:

[0418] [ka] and In the formula, x' is an integer of 1 to 23, y is an integer of 1 to 5, and z' is an integer of 1 to 23. * AA SC AA SC is the side chain of an amino acid residue of the cCPP, and M is a linking group as defined herein.

[0419] The linker has the structure:

[0420] [ka] and In the formula, x' is an integer of 1 to 23, y is an integer of 1 to 5, and z' is an integer of 1 to 23. * AA SC AA SC is the side chain of an amino acid residue in cCPP.

[0421] x can be an integer from 1 to 10, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, including all ranges and subranges therebetween.

[0422] z' can be an integer from 1 to 23, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 (including all ranges and subranges therebetween). X' can be an integer from 5 to 15. X' can be an integer from 9 to 13. X' can be an integer from 1 to 5. X' can be 1.

[0423] y can be an integer from 1 to 5, for example, 1, 2, 3, 4, or 5 (including all ranges and subranges therebetween). Y can be an integer from 2 to 5. Y can be an integer from 3 to 5. Y can be 3 or 4. Y can be 4 or 5. Y can be 3. Y can be 4. Y can be 5.

[0424] z can be an integer from 1 to 10, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, including all ranges and subranges therebetween.

[0425] z' can be an integer from 1 to 23, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23 (including all ranges and subranges therebetween). Z' can be an integer from 5 to 15. Z' can be an integer from 9 to 13. Z' can be 11.

[0426] As described above, the linker or M (where M is part of the linker) can be covalently attached to the cargo at any suitable position on the cargo. The linker or M (where M is part of the linker) can be covalently attached to the 3'-end of the oligonucleotide cargo or the 5'-end of the oligonucleotide cargo. The linker or M (where M is part of the linker) can be covalently attached to the N-terminus or C-terminus of the peptide cargo. The linker or M (where M is part of the linker) can be covalently attached to the backbone of the oligonucleotide or peptide cargo.

[0427] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the cCPP, or a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the cCPP.

[0428] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the peptide cargo, or to a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the peptide cargo.

[0429] The linker has the structure:

[0430] [ka] and During the ceremony, M is a group that conjugates L to a cargo, e.g., an oligonucleotide; AA s is the side chain or terminus of an amino acid on the cCPP, Each AA x are independently amino acid residues, o is an integer from 0 to 10; p is an integer from 0 to 5.

[0431] The linker has the structure:

[0432] [ka] and During the ceremony, M is a group that conjugates L to a cargo, e.g., an oligonucleotide; AA s is the side chain or terminus of an amino acid on the cCPP, Each AA x are independently amino acid residues, o is an integer from 0 to 10; p is an integer from 0 to 5.

[0433] M can include alkylene, alkenylene, alkynylene, carbocyclyl, or heterocyclyl, each of which is optionally substituted.

[0434] [ka] wherein R is alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl.

[0435] M is

[0436] [ka] You can choose from

[0437] In the formula, R 10 is alkylene, cycloalkyl, or

[0438] [ka] where a is 0 to 10.

[0439] M is

[0440] [ka] R 10 teeth

[0441] [ka] where a is from 0 to 10. M can be

[0442] [ka] It could be.

[0443] M is a heterobifunctional crosslinker, e.g.,

[0444] [ka] which is disclosed in Williams et al. Curr. Protoc Nucleic Acid Chem. 2010, 42, 4.41.1-4.41.20, which is incorporated herein by reference in its entirety.

[0445] M can be —C(O)—.

[0446] AA s can be the side chain or terminus of an amino acid on the cCPP. s Non-limiting examples of AA include aspartic acid, glutamic acid, glutamine, asparagine, or lysine, or a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). s is AA as defined herein SC It could be.

[0447] Each AA x are independently natural or unnatural amino acids. x can be a natural amino acid. xmay be an unnatural amino acid. x may be a β-amino acid. The β-amino acid may be β-alanine.

[0448] o can be an integer from 0 to 10, for example, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. O can be 0, 1, 2, or 3. O can be 0. O can be 1. O can be 2. O can be 3.

[0449] p can be 0 to 5, for example, 0, 1, 2, 3, 4, or 5. P can be 0. P can be 1. P can be 2. P can be 3. P can be 4. P can be 5.

[0450] The linker has the structure:

[0451] [ka] and In the formula, M, AA s , each -(R 1- JR 2 )z″-, o, and z″ are defined herein. r can be 0 or 1.

[0452] r can be 0. R can be 1.

[0453] The linker has the structure:

[0454] [ka] and wherein M, AA s , o, p, q, r, and z may each be as defined herein.

[0455] z can be an integer from 1 to 50, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50, including all ranges and values ​​therebetween. Z" can be an integer from 5 to 20. Z" can be an integer from 10 to 15.

[0456] The linker has the structure:

[0457] [ka] and During the ceremony, M, A.A. s and o are as defined herein.

[0458] Other non-limiting examples of suitable linkers include:

[0459] [ka]

[0460] [ka] In the formula, M and AA s is as defined herein.

[0461] Provided herein is a compound comprising a cCPP and an AC complementary to a target in a pre-mRNA sequence, further comprising L, wherein a linker is conjugated to the AC via a linking group (M), wherein M is

[0462] [ka] A compound is provided, wherein:

[0463] Provided herein is a compound comprising a cCPP and a cargo comprising an antisense compound (AC), e.g., an antisense oligonucleotide complementary to a target in a pre-mRNA sequence, further comprising L, wherein the linker is conjugated to the AC via a linking group (M), wherein M is

[0464] [ka] Selected from R 1 is alkylene, cycloalkyl, or

[0465] [ka] wherein t' is 0 to 10, each R is independently alkyl, alkenyl, alkynyl, carbocyclyl, or heterocyclyl, and R 1 teeth

[0466] [ka] A compound is provided wherein t' is 2.

[0467] The linker has the structure:

[0468] [ka] and During the ceremony, A.A. s is as defined herein, and m' is 0-10.

[0469] The linker has the formula:

[0470] [ka] It can be of the following type.

[0471] The linker has the formula:

[0472] [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.

[0473] The linker has the formula:

[0474] [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.

[0475] The linker has the formula:

[0476] [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.

[0477] The linker has the formula:

[0478] [ka] where "base" is the nucleobase at the 3' end of the cargo phosphorodiamidate morpholino oligomer.

[0479] The linker has the formula:

[0480] [ka] It can be of the following type.

[0481] The linker can be covalently attached to the cargo at any suitable position on the cargo. The linker is covalently attached to the 3' end of the cargo or the 5' end of an oligonucleotide cargo. The linker can be covalently attached to the backbone of the cargo.

[0482] The linker can be attached to the side chain of aspartic acid, glutamic acid, glutamine, asparagine, or lysine on the cCPP, or a modified side chain of glutamine or asparagine (e.g., a reduced side chain bearing an amino group). The linker can be attached to the side chain of lysine on the cCPP.

[0483] cCPP-linker conjugates The cCPP may be conjugated to a linker as defined herein. The linker may be a linker between the AA SC can be conjugated to

[0484] The linker is -(OCH2CH2) z’ subunits (e.g., as spacers), where z' is an integer from 1 to 23, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, or 23. z’ is also referred to as PEG. The cCPP-linker conjugate may have a structure selected from Table 4.

[0485] [Table 9]

[0486] The linker is -(OCH2CH2) z’ -subunits (where z' is an integer from 1 to 23) and peptide subunits. The peptide subunits may comprise 2 to 10 amino acids. The cCPP-linker conjugate may have a structure selected from Table 5.

[0487] [Table 10]

[0488] An EEV is provided that includes a cyclic cell penetrating peptide (cCPP), a linker, and an exocyclic peptide (EP). The EEV has the formula (B):

[0489] [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; R4 and R7 are independently H or an amino acid side chain; EP is an exocyclic peptide as defined herein; each m is independently an integer from 0 to 3; n is an integer from 0 to 2, x' is an integer from 1 to 20, y is an integer from 1 to 5, q is 1 to 4; 68. The compound according to claim 66 or 67, wherein z' is an integer of 1 to 23.

[0490] R1, R2, R3, R4, R7, EP, m, q, y, x', z' are as described herein.

[0491] n can be 0. n can be 1. n can be 2.

[0492] EEV has the formula (Ba) or (Bb):

[0493] [ka] or a protonated form thereof, wherein EP (denoted as "PE"), R 1 , R 2 , R 3 , R 4 , m and z' are as defined above in formula (B).

[0494] EEV is calculated using the formula (Bc):

[0495] [ka] or a protonated form thereof, 1 , R 2 , R 3 , R 4 and m are as defined above in formula (B), AA is an amino acid as defined herein, M is as defined herein, n is an integer from 0 to 2, x is an integer from 1 to 10, y is an integer from 1 to 5, and z is an integer from 1 to 10.

[0496] EEV is represented by the formula (B-1), (B-2), (B-3), or (B-4):

[0497] [ka]

[0498] [ka] or a protonated form thereof, where EP is as defined above in formula (B).

[0499] The EEV can comprise the formula (B), having the structure: Ac-PKKKRKVAEEA-K(cyclo[FGFGRGRQ])-PEG 12 —OH(Ac-SEQ ID NO:132-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) or Ac-PK-KKR-KV-AEEA-K(cyclo[GfFGrGrQ])-PEG 12 —OH(Ac-SEQ ID NO:133-K(cyclo[SEQ ID NO:83])-PEG 12 -OH).

[0500] EEV is a function of the formula:

[0501] [ka] The cCPP may include:

[0502] The EEV may comprise the formula: Ac-PKKKRKV-miniPEG2-Lys(cyclo(FfFGRGRQ)-miniPEG2-K(N3)(Ac-SEQ ID NO:42-PEG2-Lys(cyclo(SEQ ID NO:81)-PEG2-K(N3)).

[0503] EEV is

[0504] [ka] It could be.

[0505] EEV is

[0506] [ka] It could be.

[0507] EEV is Ac-PK(Tfa)-K(Tfa)-K(Tfa)-RK(Tfa)-V-miniPEG2-K(cyclo-Ff-Nal-GrGrQ)-PEG 12 —OH(Ac-SEQ ID NO: 134)-miniPEG2-K(cyclo-SEQ ID NO: 135)-PEG 12 -OH).

[0508] EEV is

[0509] [ka] It could be.

[0510] EEV was generated using Ac-PKKKRKV-miniPEG2-K(cyclo(Ff-Nal-GrGrQ)-PEG 12 —OH(Ac-SEQ ID NO:42-PEG2-K(cyclo(SEQ ID NO:135)-PEG 12 -OH).

[0511] EEV is

[0512] [ka] It could be.

[0513] EEV is

[0514] [ka] It could be.

[0515] EEV is

[0516] [ka] It could be.

[0517] EEV is

[0518] [ka] It could be.

[0519] EEV is

[0520] [ka] It could be.

[0521] EEV is

[0522] [ka]

[0523] EEV is

[0524] [ka] It could be.

[0525] EEV is

[0526] [ka] It could be.

[0527] EEV is

[0528] [ka] It could be.

[0529] EEV is

[0530] [ka] It could be.

[0531] EEV is

[0532] [ka] It could be.

[0533] The EEV may be selected from the following:

[0534] [Table 11-1]

[0535] [Table 11-2]

[0536] The EEV may be selected from the following: Ac-PKKKRKV-Lys(cyclo[FfΦGrGrQ])-PEG 12 -K(N3)-NH2 (Ac-SEQ ID NO:42-Lys(cyclo[SEQ ID NO:80])-PEG12 -K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FfΦGrGrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 80])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRGRQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 82])-miniPEG2-K(N3)-NH2) Ac-KR-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-KR-PEG2-K(cyclo[SEQ ID NO: 82])-PEG2-K(N3)-NH2) Ac-PKKKGKV-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:46-PEG2-K(cyclo[SEQ ID NO:82])-PEG2-K(N3)-NH2) Ac-PKKKRKG-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:48-PEG2-K(cyclo[SEQ ID NO:82])-PEG2-K(N3)-NH2) Ac-KKKRK-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:19-PEG-K(cyclo[SEQ ID NO:82])-PEG-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FFΦGRGRQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42 mini-PEG2-Lys(cyclo[SEQ ID NO: 80])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[βhFfΦGrGrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 142])-miniPEG2-K(N3)-NH2) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FfΦSrSrQ])-miniPEG2-K(N3)-NH2 (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 143])-miniPEG2-K(N3)-NH2).

[0537] The EEV may be selected from the following: Ac-PKKKRKV-miniPEG2-Lys(cyclo(GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo(SEQ ID NO: 133))-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFKRKRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 144])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFRGRGQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 145])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 146])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRrRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 147])-PEG 12 -OH) Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 84])-PEG 12 -OH) and Ac-PKKKRKV-miniPEG2-Lys(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-miniPEG2-Lys(cyclo[SEQ ID NO: 85])-PEG 12 -OH).

[0538] The EEV may be selected from the following: Ac-KKKRKG-miniPEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 148-miniPEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-KKKRK-miniPEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 19-miniPEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-KKRKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:22-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) Ac-KRKKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:21-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) Ac-KKKKR-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:23-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH) Ac-RKKKK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO:20-PEG4-K(cyclo[SEQ ID NO:82])-PEG 12 -OH)and Ac-KKKRK-PEG4-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 19-PEG4-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH).

[0539] The EEV may be selected from the following: Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:42-PEG-K(cyclo[SEQ ID NO:82])-PEG-K(N3)-NH2) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG2-K(N3)-NH2 (Ac-SEQ ID NO:42-PEG2-K(cyclo[SEQ ID NO:133])-PEG2-K(N3)-NH2) and Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH).

[0540] The cargo may be a protein and the EEV may be Ac-PKKKRKV-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO:42-PEG2-K(cyclo[SEQ ID NO:80])-PEG12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-PKKKRKV-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 42-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FfF-GRGRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rrr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rrr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rhr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rhr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-rbr-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFGRGRQ])-PEG12-OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbrbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 138-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-rbhbr-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 149-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfΦGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 80])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfΦCit-r-Cit-rQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 79])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FfFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 81])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFGRGRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 82])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[GfFGrGrQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 133])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFGRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 84])-PEG 12 -OH) Ac-hbrbh-PEG2-K(cyclo[FGFRRRRQ])-PEG 12 -OH (Ac-SEQ ID NO: 141-PEG2-K(cyclo[SEQ ID NO: 85])-PEG 12 -OH) wherein b is beta-alanine and the exocyclic sequence can be D or L stereochemistry.

[0541] cargo A cell penetrating peptide (CPP), for example, a cyclic cell penetrating peptide (e.g., cCPP), can be conjugated to a cargo. As used herein, "cargo" refers to a compound or moiety that is desired to be delivered to a cell. The cargo can be conjugated to the terminal carbonyl group of the linker. At least one atom of the cyclic peptide can be substituted by the cargo, or at least one lone pair of electrons can form a bond to the cargo. The cargo can be conjugated to the cCPP via a linker. The cargo can be linked to the AA SC At least one atom of the cCPP can be replaced with a therapeutic moiety, or at least one lone pair of electrons of the cCPP can form a bond to a therapeutic moiety. A hydroxyl group on an amino acid side chain of the cCPP can be replaced with a bond to a cargo. A hydroxyl group on a glutamine side chain of the cCPP can be replaced with a bond to a cargo. The cargo can be conjugated to the cCPP by a linker. The cargo can be linked to an AA SC can be conjugated to

[0542] In embodiments, the amino acid side chain comprises a chemically reactive group to which a linker or cargo is conjugated. The chemically reactive group may comprise an amine group, a carboxylic acid group, an amide group, a hydroxyl group, a sulfhydryl group, a guanidinyl group, a phenol group, a thioether group, an imidazolyl group, or an indolyl group. In embodiments, the amino acid of the cCPP to which a cargo is conjugated comprises lysine, arginine, aspartic acid, glutamic acid, asparagine, glutamine, homoglutamine, serine, threonine, tyrosine, cysteine, arginine, tyrosine, methionine, histidine, or tryptophan.

[0543] The cargo may comprise one or more detectable moieties, one or more therapeutic moieties (TM), one or more targeting moieties, or any combination thereof. In embodiments, the cargo comprises a TM. In embodiments, the TM comprises an antisense compound (AC). In embodiments, the AC binds to at least a portion of a splice element (SE) of a target gene transcript, or is sufficiently proximal to the SE of the target gene transcript to modulate splicing of the target gene transcript. In embodiments, the AC binds to at least a portion of the SE of a target IRF-5, DPMK, or DUX4 gene transcript. In embodiments, the AC binds to at least a portion of the SE of a target IRF-5, DPMK, or DUX4 gene transcript in sufficient proximity to the SE of the target IRF-5, DPMK, or DUX4 gene transcript to modulate splicing of the target IRF-5, DPMK, or DUX4 gene transcript.

[0544] Cyclic cell-penetrating peptides (cCPPs) conjugated to cargo moieties Cyclic cell-penetrating peptides (cCPPs) can be conjugated to cargo moieties.

[0545] The cargo moiety is conjugated to the linker at the terminal carbonyl group to form the following structure:

[0546] [ka] wherein

[0547] EP is an exocyclic peptide, M, AA SC , Cargo, x', y and z' are as defined above; * AA SC x' can be 1. y can be 4. z' can be 11. -(OCH2CH-2) x’ -and / or-(OCH2CH-2) z’ The - can be independently replaced with one or more amino acids, such as, for example, glycine, beta-alanine, 4-aminobutyric acid, 5-aminopentanoic acid, 6-aminohexanoic acid, or combinations thereof.

[0548] Endosomal escape vehicles (EEVs) can comprise a cyclic cell-penetrating peptide (cCPP), an exocyclic peptide (EP), and a linker, conjugated to a cargo, and have the structure of formula (C):

[0549] [ka] or a protonated form thereof, During the ceremony, R1, R2, and R3 may each independently be H or an amino acid residue having a side chain containing an aromatic group; R4 is H or an amino acid side chain; EP is an exocyclic peptide as defined herein; Cargo is a moiety as defined herein; each m is independently an integer from 0 to 3; n is an integer from 0 to 2, x' is an integer from 2 to 20, y is an integer from 1 to 5, q is an integer from 1 to 4, z' is an integer of 2 to 20.

[0550] R1, R2, R3, R4, EP, cargo, m, n, x', y, q, and z' are as defined herein.

[0551] The EEV can be conjugated to a cargo, the EEV-conjugate having the structure of formula (Ca) or (Cb):

[0552] [ka] or a protonated form thereof, where EP, m and z are as defined above in formula (C).

[0553] The EEV can be conjugated to a cargo, the EEV-conjugate having the formula (Cc):

[0554] [ka] or a protonated form thereof, wherein EP, R 1 , R 2 , R 3 , R 4 and m are as defined above in formula (III), AA can be an amino acid as defined herein, n can be an integer from 0 to 2, x can be an integer from 1 to 10, y can be an integer from 1 to 5, and z can be an integer from 1 to 10.

[0555] The EEV can be conjugated to an oligonucleotide cargo, the EEV-oligonucleotide conjugate having the structure of formula (C-1), (C-2), (C-3), or (C-4):

[0556] [ka]

[0557] [ka] may include:

[0558] The EEV can be conjugated to an oligonucleotide cargo, the EEV conjugate having the structure:

[0559] [ka] may include:

[0560] Cytoplasmic delivery efficiency Modifications to cyclic cell-penetrating peptides (cCPPs) can improve cytoplasmic delivery efficiency. Improved cytoplasmic uptake efficiency can be measured by comparing the cytoplasmic delivery efficiency of a cCPP having a modified sequence with a control sequence. The control sequence does not contain certain substituted amino acid residues (such as, but not limited to, arginine, phenylalanine, and / or glycine) in the modified sequence, but is otherwise identical.

[0561] As used herein, cytoplasmic delivery efficiency refers to the ability of a cCPP to cross the cell membrane and enter the cytoplasm of a cell. The cytoplasmic delivery efficiency of a cCPP does not necessarily depend on the receptor or cell type. Cytoplasmic delivery efficiency may refer to absolute cytoplasmic delivery efficiency or relative cytoplasmic delivery efficiency.

[0562] Absolute cytoplasmic delivery efficiency is the ratio of the cytosolic concentration of cCPP (or cCPP-cargo conjugate) to the concentration of cCPP (or cCPP-cargo conjugate) in the growth medium. Relative cytoplasmic delivery efficiency refers to the concentration of cCPP in the cytosol compared to the concentration of a control cCPP in the cytosol. Quantification can be achieved by fluorescently labeling the cCPP (e.g., with FITC dye) and measuring fluorescence intensity using techniques well known in the art.

[0563] The relative cytoplasmic delivery efficiency is determined by comparing (i) the amount of the cCPP of the present invention internalized by a cell type (e.g., HeLa cells) with (ii) the amount of a control cCPP internalized by the same cell type. To measure the relative cytoplasmic delivery efficiency, the cell type may be incubated in the presence of the cCPP for a specific period (e.g., 30 minutes, 1 hour, 2 hours, etc.), and then the amount of cCPP internalized by the cells is quantified using methods known in the art, such as fluorescence microscopy. Separately, the same concentration of the control cCPP is incubated in the presence of the cell type for the same period, and the amount of the control cCPP internalized by the cells is quantified.

[0564] Relative cytoplasmic delivery efficiency is calculated by the IC of cCPPs with modified sequences for intracellular targets 50 Measure the IC of cCPP with modified sequences 50 can be determined by comparing with a control sequence (as described herein).

[0565] The relative cytoplasmic delivery efficiency of cCPPs compared to cyclo(FfΦRrRrQ, SEQ ID NO: 150) ranges from about 50% to about 450%, e.g., about 60%, about 70%, about 80%, about 90%, about 100%, 200%, about 110%, 210%, about 120%, 220%, about 130%, about 140%, about 150%, about 160%, about 170%, about 180%, about 190%, about 300%, about 310%, about 320%, about 230%, 330%, about 240%, about 250%, about 260%, about 270%, about 380%, about 390%, about 400%, about 410%, about 420%, about 430%, about 440%, about 450%, about 460%, about 470%, about 480%, about 490%, about 500%, about 510%, about 520%, about 530%, about 540%, about 550%, about 560%, about 570%, about 580%, about 590%, about 600%, about 610%, about 620%, about 630%, about 640%, about 650%, about 660%, about 670%, about 680%, about 690%, about 700%, about 710%, about 720%, about 730%, about 740%, about 750%, about 760%, about 770%, about 780%, about 790%, about 800%, about 810%, about 820%, about 830%, about 840%, about 850%, about 860%, about 870%, about 880 The relative cytoplasmic delivery efficiency of the cCPP may be improved by more than about 600% compared to a cyclic peptide comprising the cyclic (FfφRrRrQ, SEQ ID NO: 150).

[0566] The absolute cytoplasmic delivery efficiency is from about 40% to about 100%, e.g., about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% (including all values ​​and subranges therebetween).

[0567] The cCPPs of the present disclosure may increase cytoplasmic delivery efficiency by about 1.1 fold to about 30 fold, e.g., about 1.2 fold, about 1.3 fold, about 1.4 fold, about 1.5 fold, about 1.6 fold, about 1.7 fold, about 1.8 fold, about 1.9 fold, about 2.0 fold, about 2.5 fold, about 3.0 fold, about 3.5 fold, about 4.0 fold, about 4.5 fold, about 5.0 fold, about 5.5 fold, about 6.0 fold, about 6.5 fold, about 7.0 fold, about 7.5 fold, about 8.0 fold, about 8.5 fold, about 9.0 fold, about 10 fold, about 10.5 fold, about 11.0 fold, about 11.5 fold, about 12.0 fold, about 12.5 fold, about 13.0 fold, about 13.5 fold, about 14 fold, about 15 fold, about 16 fold, about 17 fold, about 18 fold, about 19 fold, about 20 fold, about 21 fold, about 22 fold, about 23 fold, about 24 fold, about 25 fold, about 26 fold, about 27 fold, about 28 fold, about 29 fold, about 30 fold, about 31 fold, about 32 fold, about 33 fold, about 34 fold, about 35 fold, about 36 fold, about 37 fold, about 38 fold, about 39 fold, about 40 fold, about 41 fold, about 42 fold, about 43 fold, about 44 fold, about The improvement may be about 0.0 fold, about 14.5 fold, about 15.0 fold, about 15.5 fold, about 16.0 fold, about 16.5 fold, about 17.0 fold, about 17.5 fold, about 18.0 fold, about 18.5 fold, about 19.0 fold, about 19.5 fold, about 20 fold, about 20.5 fold, about 21.0 fold, about 21.5 fold, about 22.0 fold, about 22.5 fold, about 23.0 fold, about 23.5 fold, about 24.0 fold, about 24.5 fold, about 25.0 fold, about 25.5 fold, about 26.0 fold, about 26.5 fold, about 27.0 fold, about 27.5 fold, about 28.0 fold, about 28.5 fold, about 29.0 fold, or about 29.5 fold (including all values ​​and subranges therebetween).

[0568] Detectable Part In embodiments, the compounds disclosed herein comprise a detectable moiety. In embodiments, the detectable moiety can be attached to the cell-penetrating peptide at an amino group, carboxylate group, or side chain of any amino acid in the cell-penetrating peptide portion (e.g., at an amino group, carboxylate group, or side chain of any amino acid in the CPP). In embodiments, the therapeutic moiety comprises a detectable moiety. The detectable moiety can comprise any detectable label. Examples of suitable detectable labels include, but are not limited to, a UV-Vis label, a near-infrared label, a luminescent group, a phosphorescent group, a magnetic spin resonance label, a photosensitizer, a photocleavable moiety, a chelating center, a heavy atom, a radioisotope, an isotopically detectable spin resonance label, a paramagnetic moiety, a chromophore, or any combination thereof. In embodiments, the label is detectable without the addition of additional reagents.

[0569] In embodiments, the detectable moiety may be a biocompatible detectable moiety, such that the compounds may be suitable for use in a variety of biological applications. "Biocompatible" and "biologically compatible," as used herein, generally refer to compounds (together with any of their metabolic or degradation products) that are generally non-toxic to cells and tissues and do not cause any significant adverse effects to cells and tissues when the cells and tissues are incubated (e.g., cultured) in their presence.

[0570] The detectable moiety may comprise a luminophore, such as a fluorescent or near-infrared label. Examples of suitable luminophores include, but are not limited to, metalloporphyrins, benzoporphyrins, azabenzoporphyrins, napthoporphyrins, phthalocyanines, polycyclic aromatic hydrocarbons (such as perylenedimine and pyrene), azo dyes, xanthene dyes, boron dipyrromethenes, aza-boron dipyrromethenes, cyanine dyes, metal-ligand complexes (such as bipyridines, bipyridyls, phenanthrolines, and coumarins), and ruthenium and iridium acetylacetonates, acridines, oxazine derivatives (such as benzophenoxazines), aza-annulenes, squaraines, 8-hydroxyquinolines, polymethines, luminescence-generating nanoparticles (such as quantum dots and nanocrystals), carbostyrils, terbium complexes, inorganic fluorophores, ionophores (such as crown ether series or derivatized dyes), or combinations thereof. Specific examples of suitable luminophores include Pd(II) octaethylporphyrin, Pt(II)-octaethylporphyrin, Pd(II) tetraphenylporphyrin, Pt(II) tetraphenylporphyrin, Pd(II) meso-tetraphenylporphyrin tetrabenzoporphine, Pt(II) meso-tetraphenylmetrylbenzoporphyrin, Pd(II) octaethylporphyrin ketone, Pt(II) octaethylporphyrin ketone, Pd(II) meso-tetra(pentafluorophenyl)porphyrin, Pt(II) meso-tetra(pentafluorophenyl)porphyrin, Ru(II) tris(4,7-diphenyl-1,10-phenanthroline) (Ru(dpp)3), Ru(II) tris(1,10-phenanthroline) (Ru(phen)3), tris(2,2'-Bipyridine)ruthenium(II) chloride hexahydrate (Ru(bpy)3), erythrosin B, fluorescein, fluorescein isothiocyanate (FITC), eosin, iridium(III) ((N-methyl-benzimidazol-2-yl)-7-(diethylamino)-coumarin)), (enzothiazole) ((benzothiazol-2-yl)-7-(diethylamino)-coumarin))-2-(acetylacetonate), lumogen dye , Macroflex Fluorescent Red, Macrolex Fluorescent Yellow, Texas Red, Rhodamine B, Rhodamine 6G, Sulfur Rhodamine, m-Cresol, Thymol Blue, Xylenol Blue, Cresol Red, Chlorophenol Blue, Bromocresol Green, Bromocresol Red, Bromothymol Blue, Cy2, Cy3, Cy5, Cy5.5, Cy7, 4-Nitrophenol, Alizarin, Phenolphthalein, o-Cresolphthalein, Chlorophenol Red , Calmagite, Bromo-xylenol, Phenol Red, Neutral Red, Nitrazine, 3,4,5,6-Tetrabromophenolphthalein, Congo Red, Fluor'sc'in, Eosin, 2',7'-Dichlorofluorescein, 5(6)-Carboxy-fluorescein, Carboxynaphthofluorescein, 8-Hydroxypyrene-1,3,6-trisulfonic acid, Seminaphthofluorescein, Seminaphthofluorescein, Tris(4,7-diphenyl-1,10- Examples of suitable fluorescein-modifying agents include, but are not limited to, (4,7-diphenyl-1,10-phenanthroline)ruthenium(II) dichloride, (4,7-diphenyl-1,10-phenanthroline)ruthenium(II) tetraphenylborate, platinum(II) octaethylporphyrin, dialkylcarbocyanine, dioctadecylcyclooxacarbocyanine, fluorenylmethyloxycarbonyl chloride, 7-amino-4-methylcoumarin (Amc), green fluorescent protein (GFP), and derivatives or combinations thereof.

[0571] In some examples, the detection moiety may include rhodamine B (Rho), fluorescein isothiocyanate (FITC), 7-amino-4-methylcoumarin (Amc), green fluorescent protein (GFP), or derivatives or combinations thereof.

[0572] Manufacturing method The compounds described herein can be prepared by various methods known to those skilled in the art of organic synthesis, or variations thereof as will be appreciated by those skilled in the art. The compounds described herein can be prepared from readily available starting materials. Optimum reaction conditions may vary with the particular reactants or solvents used, but such conditions can be determined by one skilled in the art.

[0573] Modification of the compounds described herein includes the addition, removal, or movement of various components as described for each compound. Similarly, if one or more chiral centers are present in the molecule, the chirality of the molecule can be changed. Furthermore, the synthesis of the compounds can include the protection and deprotection of various chemical groups. Those skilled in the art can determine the use of protection and deprotection, and the selection of appropriate protecting groups. The chemistry of protecting groups can be found, for example, in Wuts and Greene, Protective Groups in Organic Synthesis, 4th Ed., Wiley & Sons, 2006, which is incorporated herein by reference in its entirety.

[0574] Starting materials and reagents used in preparing the compounds and compositions of the present disclosure are available from Aldrich Chemical Corporation (Milwaukee, Wisconsin), Acros Organics (Morris Plains, New Jersey), Fisher Scientific (Pittsburgh, Pennsylvania), Sigma (St. Louis, Missouri), Pfizer (New York, New York), GlaxoSmithKline (Raleigh, North Carolina), Merck (Whitehouse Station, New Jersey), Johnson & Johnson (New Brunswick, New Jersey), Abe These compounds are available from commercial suppliers such as Pharma (Bridgewater, NJ), AstraZeneca (Wilmington, DE), Novartis (Basel, Switzerland), Wyeth (Madison, NJ), Bristol-Myers Squibb (New York, NY), Roche (Basel, Switzerland), Lilly (Indianapolis, IN), Abbott (Abbott Park, IL), Schering-Plough (Kenilworth, NJ), or Boehringer Ingelheim (Ingelheim, Germany), or are commercially available from Fieser and Fieser's Reagents for Organic Synthesis, Volumes 1-17 (John Wiley and Sons, 1991), Rodd's Chemistry of Carbon Compounds, Volumes 1-5 and Supplementals (Elsevier Science Publishers, 1989), Organic Reactions, Volumes 1-40 (John Wiley and Sons, 1991), March's Advanced Organic Chemistry, (John Wiley and Sons, 4th Edition), and Larock's Comprehensive Organic Transformations (VCH Publishers Inc., 1989), or are prepared by methods known to those skilled in the art according to procedures described in references such asOther materials, such as the pharmaceutical carriers disclosed herein, can be obtained from commercial sources.

[0575] The reactions to produce the compounds described herein can be carried out in a solvent that can be selected by one skilled in the art of organic synthesis. The solvent can be substantially non-reactive with the starting materials (reactants), intermediates, or products under the conditions, e.g., temperature and pressure, at which the reaction is carried out. The reaction can be carried out in one solvent or a mixture of two or more solvents. The formation of the product or intermediate can be monitored according to any suitable method known in the art. For example, the formation of the product can be monitored by spectroscopic means, e.g., nuclear magnetic resonance spectroscopy (e.g., 1 H or 13 C) It can be monitored by infrared spectroscopy, spectrophotometry (eg, UV-visible), or mass spectrometry, or by chromatography, for example, high performance liquid chromatography (HPLC) or thin layer chromatography.

[0576] The compounds of the present disclosure can be prepared by solid-phase peptide synthesis in which the α-N-terminal amino acid is protected with an acid or base protecting group. Such protecting groups should have the properties of being stable to the conditions of peptide bond formation while being easily removable without disrupting the growing peptide chain or racemizing any of the chiral centers contained therein. Suitable protecting groups include 9-fluorenylmethyloxycarbonyl (Fmoc), t-butyloxycarbonyl (Boc), benzyloxycarbonyl (Cbz), biphenylisopropyloxycarbonyl, t-amyloxycarbonyl, isobornyloxycarbonyl, α,α-dimethyl-3,5-dimethoxybenzyloxycarbonyl, o-nitrophenylsulfenyl, 2-cyano-t-butyloxycarbonyl, and the like. The 9-fluorenylmethyloxycarbonyl (Fmoc) protecting group is particularly preferred for synthesizing the compounds of the present disclosure. Other preferred side chain protecting groups are 2,2,5,7,8-pentamethylchroman-6-sulfonyl (pmc), nitro, p-toluenesulfonyl, 4-methoxybenzenesulfonyl, Cbz, Boc, and adamantyloxycarbonyl for side chain amino groups such as lysine and arginine; benzyl, o-bromobenzyloxycarbonyl, 2,6-dichlorobenzyl, isopropyl, t-butyl (t-Bu), cyclohexyl, cyclopenyl, and acetyl (Ac) for tyrosine; t-butyl, benzyl, and tetrahydropyranyl for serine; trityl, benzyl, Cbz, p-toluenesulfonyl, and 2,4-dinitrophenyl for histidine; formyl for tryptophan; benzyl and t-butyl for aspartic acid and glutamic acid; and triphenylmethyl (trityl) for cysteine.

[0577] In solid-phase peptide synthesis, the α-C-terminal amino acid is attached to a suitable solid support or resin. Suitable solid supports useful for the above synthesis are those that are inert to the reagents and reaction conditions of the stepwise condensation-deprotection reactions and insoluble in the media used. Solid supports for the synthesis of α-C-terminal carboxypeptides are 4-hydroxymethylphenoxymethyl-copoly(styrene-1% divinylbenzene) or 4-(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxyacetamidoethyl resin, available from Applied Biosystems (Foster City, CA). The α-C-terminal amino acid is coupled to the resin by mediated coupling using N,N'-dicyclohexylcarbodiimide (DCC), N,N'-diisopropylcarbodiimide (DIC), or O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU), with or without 4-dimethylaminopyridine (DMAP), 1-hydroxybenzotriazole (HOBT), benzotriazol-1-yloxy-tris(dimethylamino)phosphonium hexafluorophosphate (BOP), or bis(2-oxo-3-oxazolidinyl)phosphine chloride (BOPCl), in a solvent such as dichloromethane or DMF, at a temperature between 10°C and 50°C for about 1 to about 24 hours. When the solid support is a 4-(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxy-acetamidoethyl resin, the Fmoc group is cleaved with a secondary amine, preferably piperidine, before coupling with the α-C-terminal amino acid as described above. One method for coupling to the deprotected 4(2',4'-dimethoxyphenyl-Fmoc-aminomethyl)phenoxy-acetamidoethyl resin is O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU, 1 equivalent) and 1-hydroxybenzotriazole (HOBT, 1 equivalent) in DMF. The coupling of successive protected amino acids can be carried out in an automated polypeptide synthesizer. In one example, the α-N-terminus of the amino acid of the growing peptide chain is protected with Fmoc.Removal of the Fmoc protecting group from the α-N-terminus of the growing peptide is accomplished by treatment with a secondary amine, preferably piperidine. Each protected amino acid is then introduced in approximately 3-fold molar excess, and coupling is preferably carried out in DMF. The coupling agents can be O-benzotriazol-1-yl-N,N,N',N'-tetramethyluronium hexafluorophosphate (HBTU, 1 equivalent) and 1-hydroxybenzotriazole (HOBT, 1 equivalent). At the end of solid-phase synthesis, the polypeptide is removed from the resin and deprotected, either sequentially or in a single operation. Removal and deprotection of the polypeptide can be accomplished in a single operation by treating the resin-bound polypeptide with a cleavage reagent containing thioanisole, water, ethanedithiol, and trifluoroacetic acid. If the α-C-terminus of the polypeptide is an alkylamide, the resin is cleaved by aminolysis with an alkylamine. Alternatively, the peptide can be removed by, for example, transesterification with methanol followed by aminolysis or direct transamidation. The protected peptide can be purified at this point or carried directly to the next step. Removal of side chain protecting groups can be achieved using the cleavage cocktail described above. The fully deprotected peptide can be purified by a series of chromatographic steps using any or all of the following types: ion exchange on weakly basic resins (acetate form), hydrophobic adsorption chromatography on underivatized polystyrene-divinylbenzene (e.g., Amberlite XAD), silica gel adsorption chromatography, ion exchange chromatography on carboxymethylcellulose, partition chromatography on e.g., Sephadex G-25, LH-20, or countercurrent distribution, high-performance liquid chromatography (HPLC), particularly reverse-phase HPLC on octyl- or octadecylsilyl-silica bonded phase column packings.

[0578] The polymer (e.g., PEG group) can be attached to an oligonucleotide (e.g., AC) under any suitable conditions. Any means known in the art can be used, including acylation, reductive alkylation, Michael addition, thiol alkylation, or other chemoselective conjugation / ligation methods via reactive groups on the PEG moiety (e.g., aldehyde, amino, ester, thiol, α-haloacetyl, maleimide, or hydrazino groups) on the AC. Activated groups that can be used to link water-soluble polymers to one or more proteins include, but are not limited to, sulfone, maleimide, sulfhydryl, thiol, triflate, tresylate, aziridine, oxirane, 5-pyridyl, and α-halogenated acyl groups (e.g., α-iodoacetic acid, α-bromoacetic acid, α-chloroacetic acid). When conjugated to AC by reductive alkylation, the selected polymer should have a single reactive aldehyde so that the degree of polymerization can be controlled. See, for example, Kinstler et al., Adv. Drug. Delivery Rev. (2002), 54: 477-485; Roberts et al., Adv. Drug Delivery Rev. (2002), 54: 459-476; and Zalipsky et al., Adv. Drug Delivery Rev. (1995), 16: 157-182.

[0579] To directly and covalently attach an AC or linker to a CPP, appropriate amino acid residues of the CPP may be reacted with an organic derivatizing agent capable of reacting with selected side chains or the N- or C-terminus of an amino acid. Reactive groups on the peptide or conjugate moiety include, for example, aldehyde, amino, ester, thiol, α-haloacetyl, maleimide, or hydrazino groups. Derivatizing agents include, for example, maleimidobenzoyl sulfosuccinimide ester (conjugation via cysteine ​​residues), N-hydroxysuccinimide (via lysine residues), glutaraldehyde, succinic anhydride, or other agents known in the art.

[0580] Methods for making ACs and conjugating ACs to linear CPPs are generally described in U.S. Patent Application Publication No. 2018 / 0298383, which is incorporated herein by reference for all purposes. This method can be applied to the cyclic CPPs disclosed herein.

[0581] The synthesis scheme is shown in FIGS. 5A to 5D and 6.

[0582] Non-limiting examples of compounds containing reactive groups useful for conjugation to CPPs and ACs are shown in Table 6. Examples of linker groups are also provided. Exemplary reactive groups include tetrafluorophenyl esters (TFP), free carboxylic acids (COOH), and azides (N3). In Table 6, n is an integer from 0 to 20. Pipa6 is AcRXRRBRRXRYQFLIRXRBRXRB, where B is β-alanine, X is aminohexanoic acid, Dap is 2,3-diaminopropionic acid, NLS is a nuclear localization sequence, βA is beta-alanine, -ss- is a disulfide, PABC is poly(A) binding protein C-terminal domain, and C x where x is an alkyl chain of length x and BCN is bicyclo[6.1.0]nonyne.

[0583] [Table 12]

[0584] In embodiments, the CPP has a free carboxylic acid group available for conjugation to an AC. In embodiments, the EEV has a free carboxylic acid group available for conjugation to an AC.

[0585] The following structure is an azide

[0586] [ka] is a 3'-cyclooctyne-modified PMO used in click reactions with compounds containing

[0587] An exemplary scheme for conjugation of a CPP and a linker to the 3' end of an AC via an amide bond is shown below.

[0588] [ka]

[0589] An exemplary scheme for the conjugation of a CPP and a linker to a 3'-cyclooctyne-modified PMO via strain-promoted azide-alkyne cycloaddition is shown below.

[0590] [ka]

[0591] Examples of conjugation chemistries used to connect the AC and CPP with additional linkers containing polyethylene glycol moieties are shown below.

[0592] [ka]

[0593] An example of the conjugation of a CPP-linker to a 5'-cyclooctyne-modified PMO via strain-promoted azide-alkyne cycloaddition (click chemistry) is shown below.

[0594] [ka]

[0595] Methods for synthesizing oligomeric antisense compounds are known in the art. The present disclosure is not limited by the method for synthesizing the AC. In embodiments, provided herein are compounds having reactive phosphorus groups useful for forming internucleoside linkages, including, for example, phosphodiester and phosphorothioate internucleoside linkages. The methods for preparing and / or purifying precursors or antisense compounds are not limitations of the compositions or methods provided herein. Methods for synthesizing and purifying DNA, RNA, and antisense compounds are well known to those skilled in the art.

[0596] Oligomerization of modified and unmodified nucleosides is usually carried out according to the following literature procedures for DNA (Protocols for Oligonucleotides and Analogs, Ed. Agrawal (1993), Humana Press) and / or RNA (Scaringe, Methods (2001), 23, 206-217; Gait et al., Applications of Chemically Synthesized RNA in RNA: Protein Interactions, Ed. Smith (1998), 1-36; Gallo et al., Tetrahedron (2001), 57, 5707-5713).

[0597] The antisense compounds provided herein can be easily and routinely produced through the well-known technique of solid phase synthesis.Devices for such synthesis are sold by several vendors, including, for example, Applied Biosystems (Foster City, CA).Any other means for such synthesis known in the art can additionally or alternatively be used.It is well known to use similar techniques to prepare oligonucleotides such as phosphorothioates and alkylated derivatives.The present disclosure is not limited by the method of antisense compound synthesis.

[0598] Methods for purifying and analyzing oligonucleotides are known to those skilled in the art. Analytical methods include capillary electrophoresis (CE) and electrospray mass spectrometry. Such synthesis and analysis methods can be performed in multi-well plates. The method of the present invention is not limited by the method of oligomer purification.

[0599] disease In some embodiments, various diseases or conditions can be treated, prevented, or ameliorated by administration of a composition comprising one or more of the compounds described herein. In embodiments, the diseases treated, prevented, or ameliorated with the compositions of the present disclosure are associated with dysregulation of splicing, protein expression, and / or protein activity.

[0600] In embodiments, the compounds disclosed herein are used to treat, prevent, or ameliorate a disease or condition. Exemplary diseases or conditions that may be treated, prevented, or modulated using the compounds of the present disclosure include, but are not limited to, cancers including, for example, acute myeloid leukemia, B-cell leukemia / lymphoma, bladder cancer, breast cancer, chronic lymphocytic leukemia, colon cancer, colorectal cancer, Duchenne muscular dystrophy, esophageal squamous cell carcinoma, Fanconi anemia, gastric cancer, glioblastoma, hepatocellular carcinoma, lung cancer, Lynch syndrome, mantle cell lymphoma, melanoma, nasopharyngeal carcinoma, neuroblastoma, ovarian cancer, pancreatic ductal adenocarcinoma, proliferative conditions, prostate cancer, and small intestinal neuroendocrine carcinoma; cardiovascular conditions including, for example, atherosclerosis, cardiac hypertrophy, dilated cardiomyopathy, hypertension, ischemia / reperfusion injury, thrombosis (deep vein), and thrombosis (venous); congenital abnormalities including microphthalmia, Müllerian duct dysplasia, bone fragility (osteogenesis imperfecta), and rickets; neonatal diabetes mellitus and type 2 diabetes mellitus. immunological disorders including IPEX syndrome, nasal polyps, severe combined immunodeficiency, systemic lupus erythematosus, and Wiskott-Aldrich syndrome; pulmonary diseases including pulmonary fibrosis; musculoskeletal conditions including muscle fibrosis, facioscapulohumeral muscular dystrophy, oculopharyngeal muscular dystrophy, myotonic dystrophy, and oculopharyngeal muscular dystrophy; neurological conditions including Alzheimer's disease, amyotrophic lateral sclerosis, anxiety disorders, Fabry disease, Fragile X syndrome, Friedrich ataxia, Huntington's disease, metachromatic leukodystrophy, pseudodeficiency syndrome, neuropsychiatric disorders, Parkinson's disease, and suicidal behavior; stress; glycogen storage diseases such as Zellweger syndrome and Pompe disease, or a combination thereof.

[0601] In embodiments, the compounds disclosed herein are used to treat, prevent, or ameliorate diseases associated with aberrant gene transcription, splicing, and / or translation. In embodiments, the compounds disclosed herein are used to treat, prevent, or ameliorate diseases associated with aberrant IRF-5, GYS1, and / or DUX4 transcription, splicing, and / or translation. In embodiments, the compounds disclosed herein are used to treat, prevent, or ameliorate diseases associated with upregulation of IRF-5, GYS1, and / or DUX4, IRF-5, GYS1, and / or DUX4 polymorphisms, accumulation of mutant IRF-5, GYS1, and / or DUX4, or a combination thereof.

[0602] glycogen storage disease Glycogen synthesis and degradation are multistep processes involving many different enzymatic reactions. For example, alpha-glucosidase (GAA) catalyzes the hydrolysis of glycogen by cleaving α-1,4 and α-1,6 glycosidic bonds, liberating glucose into the cytoplasm. In the absence of GAA, glycogen accumulates in the lysosomes of various tissues, primarily cardiac and skeletal muscles. Conditions caused by deficiencies of this protein are called glycogen storage diseases (GSDs) (Douillard-Guilloux et al., Hum. Mol. Genet. (2010), 19(4):684-96).

[0603] GSDs are inherited metabolic disorders of glycogen metabolism. There are over 12 types of glycogen storage diseases, classified based on the enzyme deficiency and the tissues affected (primarily liver or muscle). Type 0 GSD is caused by a deficiency of glycogen synthase. Type I is caused by a deficiency of glucose-6-phosphatase alpha. Type II is caused by a deficiency of alpha-glucosidase (GAA). Type III is caused by a deficiency of glycogen debranching enzyme (GDE). Type IV is caused by a deficiency of glycogen branching activity. Type V is caused by a deficiency of the muscle isoform of glycogen phosphorylase (encoded by PYGM). Type VI is caused by a deficiency of the hepatic isoform of glycogen phosphorylase (encoded by PYGL). Little information is available about the remaining GSDs, and some former GSDs have been classified as other disorders. A list of glycogen storage diseases is provided in Table 7 (Ellingwood S. et al., (2018), J. Endocrinol. 238(3):R131-R141. doi:10.1530 / JOE-18-0120).

[0604] [Table 13] * OMIM (Online Mendelian Inheritance in Man);#In early sources, GSD type XI was associated with a deficiency of lactate dehydrogenase A (OMIM 612933).

[0605] Glycogen storage disease type II (GSDII), or Pompe disease, is an autosomal recessive lysosomal storage disorder caused by mutations in the gene encoding the glucosidase protein (GAA), which results in a lack or deficiency of the GAA protein, which is essential for the breakdown of the complex sugar glycogen. Normally, the body uses GAA to break down the complex carbohydrate glycogen and convert it into glucose. Failure to achieve proper breakdown and abnormalities in glycogen metabolism lead to excessive accumulation of glycogen in the body's cells, particularly cardiac, smooth, and skeletal muscle cells, which can result in the impairment and degradation of normal tissue and organ function. Patients with Pompe disease experience severe muscle-related problems, including progressive muscle weakness throughout the body, particularly in the legs, trunk, and diaphragm. As the disorder progresses, respiratory problems can lead to respiratory failure. To date, more than 300 pathogenic mutations in GAA have been identified. Pompe disease is generally estimated to affect a total of 5,000 to 10,000 patients in the United States and Europe. However, the advent of newborn screening suggests that the disease is underdiagnosed.

[0606] Based on the age of onset and severity of symptoms, Pompe disease is typically classified as either infantile-onset Pompe disease (IOPD) or late-onset Pompe disease (LOPD). IOPD is characterized by severe muscle weakness and abnormally reduced muscle tone and usually appears within the first few months of life. If left untreated, IOPD is often fatal due to malnutrition caused by progressive heart failure, respiratory distress, or feeding difficulties. LOPD presents in childhood, adolescence, or adulthood. Patients with LOPD typically have milder symptoms, such as decreased mobility and respiratory problems. Patients with LOPD experience progressive difficulty walking and respiratory depression. Early symptoms of LOPD are subtle and may remain unrecognized for years.

[0607] At the time of filing, the only currently approved therapies for Pompe disease are alglucosidase alfa (Lumizyme in the United States, Myozyme elsewhere) and avalaglucosidase alfa-ngpt (Nexviazyme in the United States), both of which are forms of enzyme replacement therapy (ERT) delivered via IV infusion. While infant patients treated with ERT for Pompe disease have shown improved survival, ERT is not curative, and many patients in long-term observational studies continue to be at increased risk for both cardiomyopathy and heart failure. These patients also experience residual muscle weakness, including difficulty swallowing and the associated increased risk of aspiration. ERT is particularly limited in its ability to improve skeletal muscle myopathy and respiratory disorders, primarily due to its inability to penetrate critical tissues affected by the disease, its lack of activity in the cytosol, and its potential immunogenicity. Despite the availability of ERT, there remains a significant unmet medical need in patients with either IOPD or LOPD.

[0608] GAA catalyzes the hydrolysis of glycogen by cleaving α-1,4 and α-1,6 glycosidic bonds, liberating glucose into the cytoplasm. In the absence of GAA, glycogen accumulates in the lysosomes of various tissues, primarily cardiac and skeletal muscles. Conditions caused by deficiencies of this protein are called glycogen storage diseases (GSDs) (Douillard-Guilloux (2010) Hum. Mol. Genet. 19(4):684-96).

[0609] One way glycogen storage diseases can be treated is by downregulating glycogen synthesis, for example, by downregulating the expression and / or activity of glycogen synthase. Glycogen synthase has two major isozymes, GYS1 and GYS2. GYS1 is ubiquitously expressed in skeletal and cardiac muscle (NCBI Reference No. 2997). GYS2 is expressed primarily in the liver and adipose tissue (NCBI Gene Reference No. 2998). GYS1 functions to break down ingested glucose to provide glycogen energy stores for muscle. In contrast, GYS2 functions to maintain blood glucose levels. Alignment of GYS1 and GYS2 mRNAs shows that the two isozymes share 54% and 71% homology.

[0610] Downregulation of glycogen synthase (GYS1) expression has been shown to result in the reversal of glycogen accumulation (Douillard-Guilloux et al., Hum. Mol. Genet. (2010), 19(4):684-96). The structure and mechanism of action of GYS1 have been reviewed (Palm, DC et al., FEBS (2013), 280(1), 2-27; and Baskaran S. et al., Proc. Natl. Acad. Sci. USA (2010) 107, 17563-17568). Due to the differences in function between GYS1 and GYS2, it is important to selectively target GYS1 for downregulation.

[0611] In embodiments, a method for treating glycogen storage disease is provided. In embodiments, the method comprises administering a compound that downregulates glycogen synthesis. In embodiments, the method comprises administering a compound that downregulates expression of glycogen synthase. In embodiments, the method comprises administering a compound that downregulates expression of muscle-type glycogen synthase (GYS1). In embodiments, the compound comprises an AC. The AC may be any AC and may have any AC properties as described elsewhere herein. In embodiments, the AC may bind to at least a portion of an SE or SRE of a target transcript, as described elsewhere herein. In embodiments, the AC may bind proximal to an SE or SRE of a target transcript, as described elsewhere herein. In embodiments, the AC is an ASO. In embodiments, the ASO is a PMO. The AC may bind to any splice element of a GYS1 target transcript, as described elsewhere herein.

[0612] In embodiments, methods are provided for treating glycogen storage diseases. In embodiments, methods are provided for treating glycogen storage diseases associated with glycogen accumulation in muscle tissue. In embodiments, methods are provided for treating glycogen storage diseases associated with glycogen accumulation in cardiac muscle tissue. In embodiments, methods are provided for treating glycogen storage diseases associated with glycogen accumulation in skeletal muscle tissue. In embodiments, methods are provided for treating type II glycogen storage disease. In embodiments, methods are provided for treating Pompe disease. In embodiments, methods are provided for treating Andersen disease. In embodiments, methods are provided for treating McArdle disease. In embodiments, methods are provided for treating Lafora disease. In embodiments, methods are provided for treating Tarui disease.

[0613] In embodiments, GYS1 is encoded by a nucleotide sequence encoding isoform 1 or isoform 2. The nucleotide sequences are available from the online NCBI database (Isoform 1=NM_002103.5; Isoform 2=NM_001161587.2). In embodiments, the nucleotide sequence encoding GYS1 differs from the nucleotide sequence encoding isoform 1 or isoform 2 by one or more nucleic acids. In embodiments, the nucleotide sequence encoding GYS1 differs by one or more polymorphisms (e.g., single nucleotide polymorphisms (SNPs)). In embodiments, the nucleotide sequence encoding GYS1 shares less than 100% sequence identity with the nucleotide sequence encoding isoform 1 or isoform 2. In embodiments, GYS1 is encoded by a nucleotide sequence that is at least about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% identical to a nucleic acid sequence encoding isoform 1 or isoform 2. In embodiments, GYS1 is encoded by a nucleotide sequence that is 80% to 100%, 90% to 100%, 95% to 100%, or 99% to 100% identical to a nucleic acid sequence encoding isoform 1 or isoform 2.

[0614] In embodiments, the method includes administering a compound that induces exon skipping of one or more exons in a GYS1 target transcript. In embodiments, the method includes administering a compound comprising an antisense compound (AC) that induces skipping of one or more exons in a GYS1 target transcript. In embodiments, hybridization of the AC to a target nucleotide sequence in a GYS1 transcript results in inclusion or skipping of one or more exons in the target transcript. In embodiments, the skipping or inclusion of one or more exons induces a frameshift in the GYS1 target transcript. In embodiments, the frameshift results in a GYS1 transcript encoding a glycogen synthase with reduced activity. In embodiments, the frameshift results in a truncated or non-functional glycogen synthase. In embodiments, the frameshift results in the introduction of a premature stop codon in the GYS1 transcript. In embodiments, the introduction of a premature stop codon results in degradation of the GYS1 mRNA transcript by nonsense-mediated decay.

[0615] In embodiments, the compound comprises an antisense compound (AC) that induces skipping of one or more of exons 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 of human and / or mouse GYS1. In embodiments, the compound comprises an AC that induces skipping of one or more exons to produce an out-of-frame frameshift, leading to the GYS1 target transcript being degraded (e.g., nonsense-mediated decay) or translated into a GYS1 protein with reduced or no activity. In embodiments, the compound comprises an AC that induces skipping of one or more of exons 2, 5, 6, 7, 8, 10, 12, and / or 14 to produce an out-of-frame frameshift. In embodiments, the compound comprises an AC that induces skipping of one or more exons to generate an in-frame deletion in the GYS1 target transcript. In embodiments, the compound comprises an AC that induces skipping of one or more of exons 3, 4, 9, 11, 13, and / or 15. In an embodiment, the compound comprises an AC that induces skipping of one or more of exons 3, 4, 9, 11, 13 and / or 15, resulting in an in-frame deletion in the GYS1 target transcript.

[0616] In embodiments, the compound comprises an AC that binds to one or more exon / intron and / or intron / exon junctions to induce exon skipping. In embodiments, the AC compound comprises any of the following sequences in Table 8, where uppercase letters indicate exon nucleotides and lowercase letters indicate intron nucleotides. In Table 8, SEQ ID NOs: 151-247 are designed to induce exon skipping to result in a frameshift modification. In embodiments, the frameshift modification results in a premature stop codon. In embodiments, the frameshift modification results in nonsense-mediated decay of the GYS1 target transcript. In Table 8, SEQ ID NOs: 249-318 are designed to induce exon skipping to result in an in-frame deletion. The ACs listed in Table 8 are designed to bind to target nucleotide sequences that include exons, exon / intron junctions, and / or intron / exon junctions.

[0617] [Table 14-1]

[0618] [Table 14-2]

[0619] [Table 14-3]

[0620] [Table 14-4]

[0621] In some embodiments, the AC comprises a PMO sequence from U.S. Patent Application No. 16 / 867,261 and / or Clayton et al., Molecular Therapy-Nucleic Acids (2014) 3, e206, such as those listed in Table 9, or a portion thereof.

[0622] The PMO sequences are designed to induce exon skipping, resulting in a frameshift modification. In embodiments, the frameshift modification results in a premature stop codon, leading to nonsense-mediated decay of the GYS1 target transcript. SEQ ID NOs: 321-327 are designed to bind to target nucleotide sequences including intron / exon and / or exon / intron junctions of the GYS1 target transcript. SEQ ID NOs: 319 and 327 are designed to bind to target nucleotide sequences including intron sequences of the target GYS1 transcript.

[0623] [Table 15]

[0624] fruitIn embodiments, AC comprises 15 to 25 or 15 to 20 contiguous bases of any sequence in Table 8 and / or Table 9. In embodiments, AC comprises 20 to 25 contiguous bases of any sequence in Table 8 and / or Table 9.

[0625] In embodiments, a mouse model using mouse GYS1 is used to assess the effects of compounds that induce exon skipping in GYS1 target transcripts. test Mouse and human GYS1 share 97% homology on chromosome 19. Furthermore, both mouse and human GYS1 contain 16 exons and share the same splicing pattern, resulting in a full-length protein that is 737 amino acids long.

[0626] In embodiments, the compounds include antisense compounds (ACs) that induce downregulation of human and / or mouse GYS1 by targeting its start codon. Examples of such sequences include those in Table 10.

[0627] [Table 16]

[0628] Interferon regulatory factor-5 (IRF-5) In embodiments, compounds are provided for modulating the activity of interferon regulatory factor-5 (IRF-5). , a member of the IRF family of transcription factors that is highly expressed in monocytes, macrophages, B cells, and dendritic cells. Its expression can be induced in other cell types by type I interferons (Almuttaqi and Udalova, FEBS J. (2018), 286:1624-1637). IRF-5 is involved in innate and adaptive immunity, antiviral defense, production of inflammatory cytokines, macrophage polarization, and cell proliferation. proliferation adjustment, Then It is also involved in differentiation and apoptosis.

[0629] Abnormal IRF-5 expression is associated with various diseases I'mFurthermore, increased IRF5 mRNA levels strongly correlate with disease pathology. For example, IRF-5 of Upregulation To do It is effective in treating autoimmune diseases, infectious diseases, cancer, and obesity. Satisfied This may result in increased production of IFN, which is associated with the development of numerous inflammatory diseases, including do.

[0630] This can be delivered intracellularly by EEV. The EEV-conjugate can be an EEV-conjugate of formula (C):

[0631] IRF-5 exists in multiple isoforms generated by three alternative non-coding 5' exons and at least nine alternatively spliced ​​mRNAs. The sequences of IRF-5 isoforms are publicly available, for example, through the online NCBI database. Isoforms exhibit cell-type-specific expression, subcellular localization, and function. Some isoforms are associated with the risk of autoimmune disease. For example, isoform 2 is associated with overexpression of IRF-5 and susceptibility to autoimmune diseases such as systemic lupus erythematosus. Furthermore, polymorphisms, including single nucleotide polymorphisms, in the gene encoding IRF-5 that result in higher mRNA expression have been associated with many autoimmune diseases (Krausgruber et al., Nat. Immunol. (2010), 12(3):231-238; Kozyrev et al., Arthritis and Rheumatology (2007), 56(4):1234-1241).

[0632] IRF-5 activation, mechanisms of action, signaling pathways, and regulatory elements have been reviewed (Song et al., J. Clin. Invest. (2020), 130(12):6700-6717; Almutaqqi and Udalova FEBS J. (2018), 286:1624-1637; Banga et al., Sci. Adv. (2020), 6:eaay1057; Thompson et al., Front. Immunol. (2018), 9:2622).

[0633] The gene encoding IRF-5 contains nine exons (exon 1, exon 2, exon 3, exon 4, exon 5, exon 6, exon 7, exon 8, and exon 9). Exon 1 is located in the 5'-untranslated region (5'-UTR) and has four variants: exon 1A, exon 1B, exon 1C, and exon 1D. The predominant isoform contains exon 1A. Exon 1B is associated with IRF-5 hyperactivation and disease progression. Single-nucleotide polymorphisms (SNPs) introducing donor splice sites (e.g., rs2004640) can result in increased expression of exon 1B transcripts and decreased expression of exon 1C-derived transcripts. Other SNPs (e.g., rs2280714) have also been associated with elevated IRF-5 expression (Kozyrev et al., Arthritis and Rheumatology (2007), 56(4):1234-1241).

[0634] The six isoforms of IRF-5 are provided below.

[0635] Human interferon regulatory factor-5 (IRF-5) (isoform 1) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYDGPRDMPPQPYKIYEVCSNGPA PTDSQPPEDYSFGAGEEEEEEEELQRMLPSLSLTEDVKWPPTLQPPTLRPTLQPPTLQPPVVLGPPAPPDPSPLAPPPGNPAGFRELLSEVLEPGPLPASLPPAGEQLLPDLLISPHMLPLTDLEIK FQYRGRPPRALTISNPHGCRLFYSQLEATQEQVELFGPISLEQVRFPSPEDIPSDKQRFYTNQLLDVLDRGLILQLQGQDLYAIRLCQCKVFWSGPCASAHDSCPNPIQREVKTKLFSLEHFLNELILFQKGQTNTPPPFEIFFCFGEEWPDRKPREKKLITVQVVPVAARLLLEMFSGELSWSADSIRLQISNPDLKDRMVEQFKELHHIWQSQQRLQPVAQAPPGAGLGVGQGPWPMHPAGM (SEQ ID NO: 334)

[0636] Human interferon regulatory factor-5 (IRF-5) (isoform 2) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYDGPRDMPPQPYKIYEVCSNGPAPTDS QPPEDYSFGAGEEEEEEEELQRMLPSLSLTDAVQSGPHMTPYSLLKEDVKWPPTLQPPTLRPPTLQPPTLQPPVVLGPPAPDPPSPLAPPPGNPAGFRELLSEVLEPGPLPASLPPAGEQLLPDLLISPHML PLTDLEIKFQYRGRPPRALTISNPHGCRLFYSQLEATQEQVELFGPISLEQVRFPSPEDIPSDKQRFYTNQLLDVLDRGLILQLQGQDLYAIRLCQCKVFWSGPCASAHDSCPNPIQREVKTKLFSLEHFLNELILFQKGQTNTPPPFEIFFCFGEEWPDRKPREKKLITVQVVPVAARLLLEMFSGELSWSADSIRLQISNPDLKDRMVEQFKELHHIWQSQQRLQPVAQAPPGAGLGVGQGPWPMHPAGMQ (SEQ ID NO: 335)

[0637] Human interferon regulatory factor-5 (IRF-5) (isoform 3) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYDGPRDMPPQPYKIYEVCSNGPAPT DSQPPEDYSFGAGEEEEEEEELQRMLPSLSLTDAVQSGPHMTPYSLLKEDVKWPPTLQPPTLQPPVVLGPPAPPDPSPLAPPPGNPAGFRELLSEVLEPGPLPASLPPAGEQLLPDLLISPHMLPLTDL EIKFQYRGRPPRALTISNPHGCRLFYSQLEATQEQVELFGPISLEQVRFPSPEDIPSDKQRFYTNQLLDVLDRGLILQLQGQDLYAIRLCQCKVFWSGPCASAHDSCPNPIQREVKTKLFSLEHFLNELILFQKGQTNTPPPFEIFFCFGEEWPDRKPREKKLITVQVVPVAARLLLEMFSGELSWSADSIRLQISNPDLKDRMVEQFKELHHIWQSQQRLQPVAQAPPGAGLGVGQGPWPMHPAGMQ (SEQ ID NO: 336)

[0638] Human interferon regulatory factor-5 (IRF-5) (isoform 4) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYDGPRDMPPQPYKIYEVCSNG PAPTDSQPPEDYSFGAGEEEEEEEELQRMLPSLSLTEDVKWPPTLQPPTLQPPVVLGPPAPDPPSPLAPPPGNPAGFRELLSEVLEPGPLPASLPPAGEQLLPDLLISHMLPLTDLEIKFQYRG RPPRALTISNPHGCRLFYSQLEATQEQVELFGPISLEQVRFPSPEDIPSDKQRFYTNQLLDVLDRGLILQLQGQDLYAIRLCQCKVFWSGPCASAHDSCPNPIQREVKTKLFSLEHFLNELILFQKGQTNTPPPFEIFFCFGEEWPDRKPREKKLITVQVVPVAARLLLEMFSGELSWSADSIRLQISNPDLKDRMVEQFKELHHIWQSQQRLQPVAQAPPGAGLGVGQGPWPMHPAGMQ (SEQ ID NO: 337)

[0639] Human interferon regulatory factor-5 (IRF-5) (isoform 5) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYD GPRDMPPQPYKIYEVCSNGPAPTDSQPPEDYSFGAGEEEEEEEELQRMLPSLSLTVTDLEIKFQYRGRPPRALTISNPHGCRLFYSQLEATQEQVELFGPISLEQ VRFPSPEDIPSDKQRFYTNQLLDVLDRGLILQLQGQDLYAIRLCQCKVFWSGPCASAHDSCPNPIQREVKTKLFSLEHFLNELILFQKGQTNTPPPFEIFFCFGEEWPDRKPREKKLITVQVVPVAARLLLEMFSGELSWSADSIRLQISNPDLKDRMVEQFKELHHIWQSQQRLQPVAQAPPGAGLGVGQGPWPMHPAGMQ (SEQ ID NO: 338)

[0640] Human interferon regulatory factor-5 (IRF-5) (isoform 6) MNQSIPVAPTPPRRVRLKPWLVAQVNSCQYPGLQWVNGEKKLFCIPWRHATRHGPSQDGDNTIFKAWAKETGKYTEGVDEADPAKWKANLRCALNKSRDFRLIYDGPRDMPPQPYKIYETPSPLRITLLVQERRRKKRKSCRGCCQA (SEQ ID NO: 339)

[0641] In embodiments, IRF-5 is encoded by a nucleotide sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6. In embodiments, the nucleotide sequence encoding IRF-5 differs from the nucleotide sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6 by one or more nucleic acids. In embodiments, the nucleotide sequence encoding IRF-5 differs by one or more polymorphisms (e.g., single nucleotide polymorphisms (SNPs). In embodiments, the nucleotide sequence encoding IRF-5 shares less than 100% sequence identity with a nucleotide sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6. In embodiments, IRF-5 differs by one or more polymorphisms (e.g., single nucleotide polymorphisms (SNPs)). In embodiments, the nucleotide sequence encoding IRF-5 shares less than 100% sequence identity with a nucleotide sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6. In some embodiments, IRF-5 is encoded by a nucleotide sequence that is at least about 80%, about 85%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% identical to a nucleic acid sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6. In some embodiments, IRF-5 is encoded by a nucleotide sequence that is 80% to 100%, 90% to 100%, 95% to 100%, or 99% to 100% identical to a nucleic acid sequence encoding IRF-5 isoform 1, IRF-5 isoform 2, IRF-5 isoform 3, IRF-5 isoform 4, IRF-5 isoform 5, or IRF-5 isoform 6.

[0642] IRF-5 has been shown to influence inflammatory macrophage phenotypes (Almuttaqi and Udalova, FEBSJ. (2018), 286:1624-1637). Macrophages can be classified as M1 (classically activated) or M2 (alternatively activated) macrophages, and can convert between these two types depending on the tissue microenvironment. There are three classes of alternatively activated macrophages: M2a, M2b, and M2c. In normal tissues, the ratio of M1 to M2 macrophages is highly regulated. An imbalance between M1 and M2 macrophages can lead to pathologies such as asthma, chronic lung disease, atherosclerosis, or osteoclastogenesis in rheumatoid arthritis. IRF-5 is a key regulator of pro-inflammatory M1 macrophage polarization (Weiss et al. Mediators of Inflammation (2013) Dx.doi.org / 10.1155 / 2013 / 245804).

[0643] Exposure of naive monocytes or recruited macrophages to the Th1 cytokines IFN-γ, TNF, or LPS promotes the development o...

Claims

1. (a) Formula: 【Chemistry 1】 or a protonated form or salt thereof, wherein: R 1 , R 2 , and R 3 are each independently H or an aromatic or heteroaromatic side chain of an amino acid; R 1 , R 2 , and R 3 at least two of which are side chains of phenylalanine; R 4 and R 7 are each independently H or an amino acid side chain; each m is independently an integer from 0 to 3; A.A. sc is an amino acid side chain; and a cyclic cell-penetrating peptide, or a protonated form or salt thereof, wherein q is 1, 2, 3, or 4; (b) an exocyclic peptide containing 2 to 10 amino acid residues, where 2, 3, or 4 of the residues are lysine residues; (c) Formula: 【Chemistry 2】 wherein: x' is an integer from 1 to 23; y is an integer from 1 to 5; z is an integer from 1 to 23; * denotes AA of cyclic peptide SC is a point of attachment to M is a linking group; (d) an antisense compound (AC) comprising a nucleotide sequence complementary to a target nucleotide sequence of a target transcript of a target gene, wherein the AC specifically hybridizes to the target nucleotide sequence and modulates splicing of the target transcript to down-regulate the expression or activity of a protein expressed from the target transcript, wherein the AC comprises about 10-50 nucleotides in length; A compound comprising:

2. The compound of claim 1 , wherein the exocyclic peptide comprises at least two amino acid residues having hydrophobic side chains.

3. The compound of claim 1 or 2, wherein the exocyclic peptide comprises TPKKKRKV.

4. 3. The compound of claim 1 or 2, wherein z' is 11, x' is 1, or both.

5. The M is 【Chemistry 3】 3. The compound according to claim 1 or 2,

6. The M is 【Chemistry 4】 3. The compound of claim 1 or 2, comprising:

7. The cyclic cell-penetrating peptide 【Chemistry 5】 【Chemistry 6】 【Chemistry 7】 【Chemistry 8】 【Chemistry 9】 【Chemistry 10】 【Chemistry 11】 【Chemistry 12】 or a protonated form or salt thereof, where if the structure does not include AASC, then at least one atom of the amino acid side chain is replaced by a linker or at least one unshared electron pair forms a bond to a linker.

8. Ac-PKKKRKV-PEG 2 -K(シシ[Ff-Nal-Cit-r-Cit-rQ])-PEG 12 -Lys(N 3 )-NH 2 ; Ac-PKKKRKV-K(cyclo[Ff-Nal-GrGrQ])-PEG 12 -Lys(N 3 )-NH 2 ; Ac-PKKKRKV-miniPEG 2 -Lys(シクロ(FfFGRGRQ)-PEG 2 -K(N 3 )-NH 2 ;; Ac-PKKKRKV-PEG 2 -K(cyclo[FGFGRGRQ])-PEG 2 -Lys(N 3 )-NH 2 ; Ac-PKKKRKV-PEG 2 -K(cyclo[GfFGrGrQ])-PEG 2 -Lys(N 3 )-NH 2 ; Ac-PKKKRKV-PEG 2 -Lys(シクGRQ)-miniPEG2-K(N 3 )-NH 2 ; Ac-PKKKRKV-PEG 2 -K(cyclo(Ff-Nal-GrGrQ)-PEG 12 -OH; Ac-PKKKRKV-PEG 2 -K(cyclo[FGFGRGRQ])-PEG 12 -OH; Ac-PKKKRKV-PEG 2 -K(cyclo[GfFGrGrQ])-PEG 12 -OH; Ac-PKKKRKV-PEG 2 -K(cyclo[FGFGRRRQ])-PEG 12 -OH; and Ac-PKKKRKV-PEG 2 -K(cyclo[FGFRRRRQ])-PEG 12 -OH The compound of claim 1, comprising an oligonucleotide bound to an endosomal escape vehicle having a sequence selected from:

9. The compound of claim 1 or 2, wherein the antisense compound comprises a phosphorodiamidate morpholino nucleotide.

10. 3. The compound of claim 1 or 2, wherein the antisense compound induces exon skipping, which induces a frameshift in the target transcript or introduces a premature stop codon into the target transcript.

11. 3. The compound of claim 1 or 2, wherein the target transcript is a glycogen synthase transcript and the frameshift results in a GYS1 transcript encoding a truncated GYS1 protein.

12. The compound of claim 1 or 2, wherein the target transcript is GYS1 and the antisense compound comprises any one of SEQ ID NO:151 to SEQ ID NO:

318.

13. 3. The compound of claim 1 or 2, wherein the target transcript is a double homeobox 4 (DUX4) transcript and the antisense compound downregulates expression of DUX4-fl.

14. 3. The compound of claim 1 or 2, wherein the target transcript is a DUX4 transcript and the antisense compound comprises any one of SEQ ID NOs: 344 to 364.

15. 3. A pharmaceutical composition comprising a compound according to claim 1 or 2 and a pharma- ceutically acceptable carrier.

16. 13. Use of a compound according to claim 1 or 2 in the manufacture of a medicament for treating a disease or disorder associated with a target gene in a patient, wherein the target gene comprises glycogen synthase (GYS1) and the disease comprises glycogen storage disease type II, Pompe disease, Anderson disease, McArdle disease, Lafra disease, or Tarui disease.

17. 16. The pharmaceutical composition of claim 15 for treating a disease or disorder associated with a target gene in a patient, wherein the target gene comprises glycogen synthase (GYS1) and the disease comprises glycogen storage disease type II, Pompe disease, Anderson disease, McArdle disease, Lafra disease, or Tariu disease.