Compositions and methods for treating CAG repeat diseases
Patent Information
- Application Number
- JP2024521180
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-06
- Filing Date
- 2022-10-05
- Publication Date
- 2025-10-14
AI Technical Summary
Current strategies for treating repeat expansion disorders, such as Huntington's disease, face challenges in selectively inhibiting mutant alleles without affecting wild-type genes, as they often require detailed population-specific genetic studies and may exclude certain individuals or populations due to variations in single nucleotide polymorphisms (SNPs).
Development of double-stranded RNA molecules that target CAG repeat regions in RNA with specific mismatches at positions 8 to 16, allowing for allele-selective inhibition of mutant proteins by modifying the design to enhance proper processing and predictability, using recombinant expression vectors and viral delivery vehicles.
The double-stranded RNA effectively reduces translation of disease-associated CAG repeat-containing RNAs, providing selective inhibition of mutant proteins while minimizing impact on wild-type genes, thus offering a targeted therapeutic approach for repeat expansion disorders.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 253,070, filed October 6, 2021, and U.S. Provisional Patent Application No. 63 / 339,363, filed May 6, 2022, which applications are incorporated herein by reference in their entireties.
[0002] Incorporation by Reference of Electronically Submitted Materials The Sequence Listing is provided herein as Sequence Listing XML "IRIS-001WO_SEQ_LIST", created on October 4, 2022, and having a size of 1,298 KB. The contents of the Sequence Listing XML are incorporated herein by reference in their entirety. [Background technology]
[0003] background
[0004] Repeat expansion disorders are autosomal dominant genetic disorders caused by the expansion of DNA repeats. DNA repeats can consist of a single nucleotide to a dodecamer or more. The threshold at which repeat expansions become symptomatic varies depending on the specific disorder. There are over 50 different disorders caused by repeat expansion. Repeat expansions occur in coding or non-coding regions of genes. Repeat expansions can cause defects in the proteins encoded by genes; alter the regulation of gene expression; produce toxic RNA or cause chromosomal instability.
[0005] Inhibiting both mutant and wild-type expression of repeat-containing genes can cause significant side effects. Therefore, suppressing mutant repeat expansion alleles is a desirable therapeutic strategy for repeat expansion disorders. Current strategies for mutant allele-specific inhibition include targeting disease-associated single nucleotide polymorphisms (SNPs) or deletions with antisense oligonucleotides or RNA interference agents. However, identifying SNPs associated with repeat expansion mutations requires detailed population-specific genetic studies in large clinical cohorts. Furthermore, specific affected individuals or populations can be excluded depending on the frequency of the target SNP or its location on the mutant repeat expansion allele.
[0006] overview
[0007] The present disclosure provides a double-stranded RNA comprising a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand. The first strand comprises i) a first mismatch to the target CAG repeat region; and ii) at least a second mismatch to the target CAG repeat region. The present disclosure provides a DNA molecule comprising a nucleotide sequence encoding the first strand of the double-stranded RNA, wherein the nucleotide sequence is operably linked to a promoter functional in eukaryotic cells. The present disclosure provides a recombinant nucleic acid comprising a) the double-stranded RNA of the present disclosure; and b) a microRNA scaffold. The present disclosure also provides a recombinant expression vector comprising a nucleotide sequence encoding such a recombinant nucleic acid. The present disclosure provides a DNA molecule comprising a nucleotide sequence encoding a recombinant nucleic acid comprising a) the double-stranded RNA of the present disclosure and b) a microRNA scaffold. The present disclosure also provides a recombinant expression vector comprising such a DNA molecule. The present disclosure also provides a recombinant expression vector comprising such a DNA molecule. The present disclosure provides viral and non-viral delivery vehicles comprising the recombinant expression vectors of the present disclosure; and pharmaceutical compositions comprising such delivery vehicles.The present disclosure provides a method for selectively reducing the translation of disease-related CAG repeat-containing RNA.
[0008] In some embodiments, the disclosure provides a double-stranded RNA comprising a 5'-3' sequence: (a) a 5' leader sequence; (b) a 5' stem comprising a passenger or guide sequence; (c) a 5' linker of 1-6 bases; (d) a terminal loop; a 3' linker of 1-6 bases; (f) a 3' stem comprising: (i) a guide sequence if the 5' stem comprises a passenger sequence; or (ii) a passenger sequence if the 5' stem comprises a guide sequence; and (g) a 3' trailer sequence; wherein the guide sequence targets a CAG repeat region of a CAG repeat-containing RNA (e.g., an mRNA or pre-mRNA) and comprises a 1-5 base mismatch relative to the CAG repeat region, the base mismatch being located at positions 8-16 of the guide sequence.
[0009] The patent or application file contains at least one drawing in color. Copies of this patent or patent application publication will be provided by the Office upon request and payment of the necessary fee. [Brief explanation of the drawings]
[0010] [Figure 1] Figures 1A-1D: Cloning and design of shRNAs expressed from custom U6 promoter-driven constructs. (Figure 1A) Plasmid map for customized pxTRC-EGFP-puro. (Figure 1B) Initial shRNA expression cassette (SEQ ID NO: 760). (Figure 1C) Modification of shRNA construct to improve accuracy and quantity of shRNA processing (SEQ ID NO: 761). (Figure 1D) shHD1L-1 design (SEQ ID NO: 762). For Figures 1B-1D, green markings indicate Drosha cleavage on the 5' leader and 3' trailer sequences, and Dicer cleavage in the 5' upper stem and 3' upper stem; mismatch positions in the guide strand are boxed in purple. [Figure 2]Figures 2A-2C: miRNA miR33 scaffold design for shRNA expression. (Figure 2A) miR33 miRNA structure and sequence elements (SEQ ID NO: 763). (Figure 2B) Generalized miR33 scaffold for shRNA cloning and expression (SEQ ID NO: 764). (Figure 2C) shHD-33 full mimic design (SEQ ID NO: 765). For Figures 2A-2C, green markings indicate Drosha cleavage on the 5' leader and 3' trailer sequences, and Dicer cleavage in the 5' upper stem and 3' upper stem; and mismatch positions in the guide strand are circled in purple. [Figure 3] Figures 3A-3B: Cell-based evaluation of select shRNAs targeting the CAG repeat expansion of HTT. (Figure 3A) HEK293T luciferase assay results for several shRNAs targeting the CAG repeat. Wild-type (wt) and mutant (mut) constructs are shown, and percent luciferase activity was normalized to the scrambled shRNA control. Means of two independent replicates are shown. Error bars are standard error of the mean (SEM). Statistical significance is indicated by a p-value of less than 0.05 (*) or 0.01 (**). (Figure 3B) Quantification of knockdown of HTT wt and mut proteins in patient-derived cells, assayed by Western blot and fitted to the Hill plot equation. Means of single or two replicates are shown. Error bars are SEM. [Figure 4] 1 is a graph comparing shRNA dosage for a virus-encoded shRNA with a guide sequence that perfectly matches the CAG repeat region of HTT mRNA versus a virus-encoded shRNA with a guide sequence that contains a mismatch to the CAG repeat region of HTT mRNA, where the mismatch is located at positions 8-16 of the guide sequence. [Figure 5]Figure 5 is a graph showing that lentiviral constructs encoding allele-selective shRNAs significantly reduced the expression of the pathogenic HTT allele (mut-HTT) in fibroblasts transduced with shHD-33full and shHD33-fullmimic compared with the normal HTT allele. [Figure 6] FIG. 6 shows colocalization of DARP32 and GFP staining in the striatum of zQ175 mice. [Figure 7] 7A-7B show Western blot analysis of wild-type (WT) and mut HTT proteins in the striatum of zQ175 mice in the AAV9-shScr-treated group versus the AAV9-shHD33-Full-Mimic-treated group. [Figure 8] FIG. 8 shows dose-dependent AAV delivery and GFP expression in the striatum. [Figure 9] FIG. 9 shows allele-selective knockdown in vivo by small binding RNAs (sbRNAs) delivered via recombinant AAV vectors. [Figure 10] FIG. 10 shows transduction efficiency in the cerebellum following ICV administration of recombinant AAV virions containing recombinant AAV encoding sbRNA. [Figure 11] Figure 11 shows in vivo allele-selective knockdown by small binding RNAs (sbRNAs) delivered via recombinant AAV vectors in an SCA2 mouse model (left panel) and partial restoration of expression of key cerebellar genes that are molecular markers of pathology in ATXN2-Q127 mice (right panel). [Figure 12] Figure 12 shows the effect of sbRNA on preserving wild-type (WT) gene expression in ATXN2-Q mice. Protein levels were unchanged for non-target genes containing CAG repeats. [Figure 13]Figures 13A-13E show the effect of registry on knockdown (Figure 13(A) from top to bottom: guide strand column of SEQ ID NOs: 406, 873, 406, 406, 406, 874, 875; loop column of SEQ ID NOs: 876, 876, 877, 878, 876, 879, 876; passenger strand column of SEQ ID NOs: 881, 880, 880, 880, 880, 882, 883, 884) (Figure 13(E) from top to bottom: SEQ ID NOs: 868, 869, 870, 871, and 872). [Figure 14] FIG. 14 is a schematic diagram of a guide sequence screening system. [Figure 15] Figure 15 shows knockdown and allele selectivity using various guide sequences. [Figure 16] FIG. 16 shows guide sequence screening and allele selectivity. [Figure 16] FIG. 17 shows the effect of the number of mismatches to the target CAG repeat region on knockdown and allele selectivity. [Figure 17] FIG. 17 shows the effect of the number of mismatches to the target CAG repeat region on knockdown and allele selectivity. [Figure 18] FIG. 18 shows the effect of a single mismatch on knockdown and allele selectivity when the mismatch is at position 8, 9, 10, or 11. [Figure 19] The effect of the distance between the first and second mismatch on knockdown and allele selectivity when the first mismatch is at position 8, 9, 10, or 11 is shown. [Figure 20] Figure 20 shows the effect of distance between mismatches in guide sequences with three (left panel) or four (right panel) mismatches to the target CAG on knockdown and allele selectivity. [Figure 21]Figure 21 shows the effect of distance between mismatches in a guide sequence with three mismatches to the target CAG on knockdown and allele selectivity when the first mismatch is at position 9, 10, or 11. [Figure 22] FIG. 22 provides Table 6. [Figure 23] FIG. 23 provides Table 7. [Figure 24] Figure 24 provides Table 8. [Figure 25] Figure 25 provides the nucleotide sequence of the sbRNA comprising the 5' and 3' flanking polynucleotides of miR451. DETAILED DESCRIPTION OF THE INVENTION
[0011] Repeat expansion disorders pose a major obstacle to selectively inhibiting disease alleles over normal alleles. The present disclosure provides double-stranded RNAs that utilize differences in repeat number to achieve allele-selective inhibition of repeat-containing proteins. The double-stranded RNA targets the repeat region of a repeat-containing target RNA molecule (e.g., mRNA or pre-mRNA) and contains 1 to 5 (e.g., 1, 2, 3, 4, or 5) nucleobase mismatches with the repeat region of the target mRNA or pre-mRNA at positions 8 to 16 (e.g., 8, 9, 10, 11, 12, 13, 14, 15, or 16) of the guide sequence, which enhances the ability of the double-stranded RNA to selectively inhibit expression of mutant proteins over wild-type proteins. Standard designs of double-stranded RNAs encoded by vectors targeting repeat regions of repeat-containing target mRNAs or pre-mRNAs and containing one to five mismatches to the repeat region at positions 8 to 16 of the guide sequence revealed a positional shift in the processing of the 5' cleavage site. Due to the shift in processing, the positions of the mismatches in the guide sequence were also shifted, placing them in offset or undesired positions. Double-stranded RNAs with mismatches in offset or undesired positions may not have the desired function. The double-stranded RNA designs disclosed herein have been modified for vector expression to place mismatches in desired positions, enhancing proper processing to provide a more predictable 5' cleavage site.
[0012] Before describing this disclosure in more detail, it may be helpful to an understanding thereof to provide definitions of certain terms used herein. Additional definitions are set forth throughout this disclosure.
[0013] As used herein, any concentration range, percentage range, ratio range, or integer range should be understood to include any integer value within the recited range, and, where appropriate, fractions thereof (such as tenths and hundredths of integers), unless otherwise indicated. Also, any numerical range recited herein for any physical characteristic, such as polymer subunits, size, or thickness, should be understood to include any integer within the recited range, unless otherwise indicated. As used herein, the term "about" means ±20% of the indicated range, value, or structure, unless otherwise indicated. As used herein, the terms "a" and "an" should be understood to refer to "one or more" of the recited components. Use of the alternative (e.g., "or") should be understood to mean either, both, or any combination thereof. As used herein, the terms "comprise," "have," and "comprise" are used interchangeably, and these terms and variations thereof are intended to be non-limiting.
[0014] As used herein, the term "nucleic acid" or "polynucleotide" refers to any nucleic acid polymer composed of covalently linked nucleotide subunits, such as polydeoxyribonucleotides or polyribonucleotides. Examples of nucleic acids include RNA and DNA.
[0015] As used herein, the term "RNA" refers to a molecule containing one or more ribonucleotides, including double-stranded RNA, single-stranded RNA, isolated RNA, synthetic RNA, recombinant RNA, and modified RNA that differs from naturally occurring RNA by the addition, deletion, substitution, and / or modification of one or more nucleotides. The nucleotides of an RNA molecule can include standard nucleotides or non-standard nucleotides, such as non-naturally occurring nucleotides or chemically synthesized nucleotides.
[0016] As used herein, "DNA" refers to a molecule containing one or more deoxyribonucleotides, and includes double-stranded DNA, single-stranded DNA, isolated DNA, synthetic DNA, recombinant DNA, and modified DNA that differs from naturally occurring DNA by the addition, deletion, substitution, and / or alteration of one or more nucleotides. The nucleotides of a DNA molecule can include standard nucleotides or non-standard nucleotides (e.g., non-naturally occurring nucleotides or chemically synthesized nucleotides).
[0017] As used herein, "nucleoside" refers to a compound comprising a nucleobase moiety and a sugar moiety. Nucleosides include, but are not limited to, naturally occurring nucleosides (as found in DNA and RNA) and modified nucleosides. Nucleosides can be linked to a phosphate moiety.
[0018] As used herein, "nucleotide" refers to a nucleoside that further comprises a phosphate linking group. As used herein, "linked nucleosides" may or may not be linked by a phosphate bond, and thus include, but are not limited to, "linked nucleotides." As used herein, "linked nucleosides" are nucleosides that are linked in a contiguous sequence (i.e., there are no additional nucleosides between the linked nucleosides).
[0019] As used herein, "nucleobase" or "base" means a group of atoms that can be attached to a sugar moiety to form a nucleoside that can be incorporated into an oligonucleotide, and that can be attached to a complementary naturally occurring nucleobase of another oligonucleotide or nucleic acid. The nucleobase can be naturally occurring or modified.
[0020] As used herein, "oligonucleotide" refers to a compound containing multiple linked nucleosides. In some embodiments, an oligonucleotide contains one or more unmodified ribonucleosides (RNA) and / or unmodified deoxyribonucleosides (DNA) and / or one or more modified nucleosides.
[0021] As used herein, "oligomeric compound" refers to a polymeric structure comprising two or more substructures. In some embodiments, the oligomeric compound comprises an oligonucleotide. In some embodiments, the oligomeric compound comprises one or more conjugated groups and / or terminal groups. In some embodiments, the oligomeric compound comprises an oligonucleotide. Oligomeric compounds also include naturally occurring nucleic acids.
[0022] As used herein, "single-stranded" means an oligomeric compound that will not hybridize to its complement and lacks sufficient self-complementarity to form a stable self-duplex.
[0023] As used herein, "duplex" refers to an oligomeric compound that partially or completely hybridizes with its complement to form a stable double-stranded molecule. A double-stranded oligomeric compound can consist of two separate strands of complementary oligomeric compounds hybridized to each other, or a single oligomeric compound with sufficient self-complementarity to form a stable self-duplex. A stable self-duplex can contain a stem-loop structure and / or a bulge.
[0024] "Isolated" refers to a material that is isolated from its natural environment or that is artificially produced. As used herein with respect to a cell, "isolated" refers to a cell that has been isolated from its natural environment (e.g., from a subject, organ, tissue, or bodily fluid). As used herein with respect to a nucleic acid, "isolated" refers to a nucleic acid that has been isolated from its natural environment or purified (e.g., from a cell, organelle, or cytoplasm), recombinantly produced, amplified, or synthesized. In embodiments, an isolated nucleic acid includes a nucleic acid contained within a vector.
[0025] As used herein, the term "wild-type" or "non-mutant" form of a gene refers to a nucleic acid that encodes a protein associated with normal or non-pathogenic activity (e.g., a protein lacking a mutation, such as a repeat region expansion, that confers a higher risk of onset, development, or progression of a neurodegenerative disease).
[0026] As used herein, the term "mutation" refers to any change in the structure of a gene, e.g., gene sequence, that results in an altered form of the gene that may or may not be passed on to subsequent generations (heritable mutations). Genetic mutations include substitutions, insertions, deletions, or rearrangements of multiple bases or larger sections of a gene or chromosome, including single base substitutions, insertions, or deletions in DNA, or repeat expansions.
[0027] As used herein, the term "inhibitory nucleic acid" refers to a nucleic acid comprising a guide strand sequence that hybridizes to at least a portion of a target nucleic acid, such as a target RNA, mRNA, or pre-mRNA, and inhibits its expression or activity. The inhibitory nucleic acid may target a protein-coding region (e.g., an exon) or a non-coding region (e.g., a 5'UTR, a 3'UTR, an intron, etc.) of the target nucleic acid. In some embodiments, the inhibitory nucleic acid is a single-stranded or double-stranded molecule. The inhibitory nucleic acid may further comprise a passenger strand sequence on a separate strand (e.g., a duplex) or in the same strand (e.g., a single-stranded, self-annealing duplex structure). In some embodiments, the inhibitory nucleic acid is an RNA molecule such as an siRNA, shRNA, pri-miRNA, pre-miRNA, or miRNA. In some embodiments, the inhibitory nucleic acid is a double-stranded RNA (dsRNA) such as a pri-miRNA, pre-miRNA, miRNA, or shRNA.
[0028] As used herein, "microRNA" or "miRNA" refers to a small non-coding RNA molecule that can mediate target gene silencing by target mRNA cleavage, target mRNA translational repression, target mRNA degradation, or a combination thereof. Typically, miRNAs are transcribed as hairpin or stem-loop (e.g., with a self-complementary single-stranded backbone) duplex structures called primary miRNAs (pri-miRNAs), which are enzymatically processed into pre-miRNAs (e.g., by Drosha, DGCR8, Pasha, etc.). The pre-miRNAs are transported to the cytoplasm, where they are enzymatically processed by Dicer to generate miRNA duplexes with passenger strands, and then single-stranded mature miRNA molecules, which are then loaded into the RNA-induced silencing complex (RISC). Reference to miRNAs may include synthetic or artificial miRNAs.
[0029] As used herein, "synthetic miRNA" or "artificial miRNA" or "amiRNA" or "small binding RNA" (sbRNA) refers to an endogenous, modified, or synthetic pri-miRNA or pre-miRNA (e.g., miRNA scaffold or scaffold) in which the endogenous miRNA guide sequence and passenger sequence within the stem sequence have been replaced with heterologous guide sequence and passenger sequence that direct highly efficient RNA silencing of target genes (see, e.g., Eamens et al. (2014), Methods Mol. Biol. 1062:211-224). In certain embodiments, the nature of complementarity of the guide sequence and passenger sequence (e.g., number of bases, position of mismatch, type of bulge, etc.) can be similar to or different from the nature of complementarity of the guide sequence and passenger sequence in the endogenous miRNA scaffold from which the synthetic miRNA is constructed.
[0030] As used herein, the terms "microRNA backbone," "miR backbone," "microRNA scaffold," or "miR scaffold" refer to a pri-miRNA or pre-miRNA scaffold in which the stem sequence is replaced with a heterologous RNA of interest, producing a functional mature miRNA that directs RNA silencing at the gene targeted by the miRNA of interest. In some cases, the miR backbone includes a 5' flanking region (also referred to herein as a "5' flanking polynucleotide" or "5' leader"), a loop motif region (also referred to herein as a "loop polynucleotide"), and a 3' flanking region (also referred to herein as a "3' flanking polynucleotide" or "3' trailer"). In some cases, the miR backbone includes the 5' flanking region and the 3' flanking region (but not the loop motif region). The miR backbone can be derived entirely or partially from a wild-type miRNA scaffold, or can be a completely artificial sequence.
[0031] As used herein, the term "short hairpin RNA" or "shRNA" includes conventional stem-loop shRNAs that form precursor miRNAs (pre-miRNAs). "shRNA" also includes microRNA-embedded shRNAs (miRNA-based shRNAs), in which the guide strand and passenger strand of the miRNA duplex are incorporated into existing (or natural) miRNAs or modified or synthetic (designed) miRNAs. When transcribed, conventional shRNAs form a structure very similar to primary miRNAs (pri-miRNAs) or natural pri-miRNAs. The pri-miRNAs are then processed into pre-miRNAs by Drosha and its cofactors. Therefore, the term "shRNA" includes pri-miRNA molecules and pre-miRNA molecules.
[0032] "Stem-loop structure" refers to a nucleic acid having a secondary structure that includes a region of nucleotides known or predicted to form a double-stranded or self-duplex (the stem portion) connected on one side by a region of primarily single-stranded nucleotides (the terminal loop portion). The terms "hairpin," "self-duplex," and "fold-back" structure are also used herein to refer to stem-loop structures. Such structures are well known in the art, and the terms are used consistently with their well-known meanings. As is well known in the art, the secondary structure does not require exact base pairing. Thus, the stem can contain one or more base mismatches or bulges. Alternatively, the base pairing can be exact, i.e., it does not contain any mismatches.
[0033] As used herein, the term inhibitory nucleic acid (e.g., at least 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary) refers to a sequence that is substantially complementary to a region of about 10-50 nucleotides (e.g., about 15-30, 16-25, 18-23, or 19-22 nucleotides) of an mRNA or pre-mRNA targeted for silencing. The guide sequence is sufficiently complementary to the target mRNA sequence to direct target-specific silencing, e.g., to induce destruction of the target mRNA by the RNAi machinery or process or to reduce translation of the target mRNA. In some embodiments, the guide strand sequence refers to the mature guide sequence remaining after cleavage by Dicer.
[0034] As used herein, the term "passenger strand sequence" of inhibitory nucleic acid refers to the sequence that is homologous to target mRNA or pre-mRNA and is partially or completely complementary to the guide strand sequence of inhibitory nucleic acid.The guide strand sequence and passenger strand sequence of inhibitory nucleic acid hybridize to form a double-stranded structure (for example, form a double-stranded duplex or a single-stranded self-annealing double-stranded structure).In some embodiments, the guide strand sequence and passenger strand sequence refer to the mature sequence that remains after Dicer cleavage.
[0035] As used herein, the term "5' arm" or "5' stem" refers to a portion of a double-stranded RNA (eg, shRNA, pre-miRNA, pri-mRNA) that includes the guide strand or the passenger strand.
[0036] As used herein, the term "3' arm" or "3' stem" refers to the portion of a double-stranded RNA that includes the passenger strand relative to the guide strand of the 5' stem, or the guide strand relative to the passenger strand of the 5' stem.
[0037] As used herein, "duplex," when used in reference to an inhibitory nucleic acid, refers to two nucleic acid strands (e.g., a guide strand and a passenger strand) that hybridize to each other to form a double-stranded structure. A duplex can be formed by two separate nucleic acid strands or by a single nucleic acid strand with a self-complementary region (e.g., a hairpin or stem loop).
[0038] As used herein, "target nucleic acid" refers to a nucleic acid molecule to which an antisense compound hybridizes. The target nucleic acid may be an mRNA (target mRNA) or pre-mRNA (target pre-mRNA) encoded by a target gene.
[0039] As used herein, "targeting" or "targeted" refers to the association of an antisense compound with a specific target nucleic acid molecule or a specific region of a target nucleic acid molecule. A double-stranded RNA targets a target nucleic acid if it is sufficiently complementary to the target nucleic acid and allows hybridization under physiological conditions.
[0040] As used herein, the term "complementary" refers to the ability of polynucleotides to base pair with each other. Base pairs are typically formed by hydrogen bonds between nucleotide subunits in antiparallel polynucleotide strands or a single self-annealing polynucleotide strand. Complementary polynucleotide strands can base pair in a Watson-Crick manner (e.g., A to T, A to U, C to G) or any other manner that allows for the formation of a double strand. In some embodiments, complementary nucleotides include G and U (wobble base pair). As will be apparent to those skilled in the art, when using RNA as opposed to DNA, uracil, rather than thymine, is the base considered complementary to adenosine. Furthermore, when "U" is indicated in the context of the present invention, its ability to substitute for "T" is understood unless otherwise specified. Complementarity also encompasses Watson-Crick base pairing between unmodified and modified nucleobases (e.g., 5-methylcytosine instead of cytosine). Full complementarity, perfect complementarity, or 100% complementarity between two polynucleotide strands is when each nucleotide in one polynucleotide strand can form a hydrogen bond with a nucleotide unit in a second polynucleotide strand. % complementarity refers to the number of nucleotides in a contiguous nucleotide sequence in a nucleic acid molecule that is complementary to an aligned reference sequence (e.g., target mRNA, passenger strand) divided by the total number of nucleotides, multiplied by 100. In such an alignment, nucleobases / nucleotides that do not form base pairs are called mismatches. When calculating the % complementarity of a contiguous nucleotide sequence, insertions and deletions are not considered. Those skilled in the art will understand that chemical modifications to nucleobases are not taken into account when calculating complementarity, as long as the nucleobase's Watson-Crick base pairing ability is maintained (e.g., 5-methylcytosine is considered the same as cytosine for the purpose of calculating % complementarity).
[0041] As used herein, "non-complementary" with respect to nucleobases means nucleobase pairs that do not form hydrogen bonds with each other.
[0042] As used herein, "mismatch" refers to the nucleobase of a first oligomeric compound that cannot pair with the nucleobase at the corresponding position of a second oligomeric compound when the first and second oligomeric compounds are aligned. Either or both of the first and second oligomeric compounds can be oligonucleotides. Nucleotides that do not form base pairs include self-pairing nucleotides (AA, TT, UU, CC, and GG), A and C, C and U, C and T, and A and G. In some embodiments, mismatch does not include GU wobble base pairs.
[0043] "Percent identity" between two or more nucleic acid sequences refers to the percentage of nucleotides in the contiguous nucleotide sequence in a nucleic acid molecule that are shared by a reference sequence (i.e., % identity = number of identical nucleotides / total number of nucleotides in the aligned region (e.g., contiguous nucleotide sequence) × 100). Insertions and deletions are not allowed in calculating the percent identity of contiguous nucleotide sequences. It is understood by those skilled in the art that, in calculating identity, chemical modifications to nucleobases are not taken into account, as long as the nucleobases retain their Watson-Crick base pairing ability (e.g., 5-methylcytosine is considered the same as cytosine for purposes of calculating % identity).
[0044] As used herein, the term "hybridizing" or "hybridizing" refers to two nucleic acid strands that form hydrogen bonds between base pairs on antiparallel strands, thereby forming a duplex. Although not limited to a specific mechanism, the most common mechanism of pairing involves hydrogen bonding, which can be Watson-Crick, Hoogsteen, or reversed Hoogsteen hydrogen bonds between complementary nucleic acid bases. The strength of hybridization between two nucleic acid strands can be described by the melting temperature (Tm), which is defined as the temperature at which 50% of the target sequence hybridizes to a complementary polynucleotide at a given ionic strength and pH.
[0045] As used herein, "heterologous" refers to a nucleic acid that is not found in natural (naturally occurring) nucleic acids. For example, with respect to components of microRNA (e.g., 5' flanking polynucleotide, loop polynucleotide, 3' flanking polynucleotide), heterologous guide sequences and heterologous passenger sequences comprise nucleotide sequences that are not associated with natural microRNA. As used herein, "guide sequence" is interchangeable with the double-stranded RNA (or "targeting strand," where "targeting strand" hybridizes to target RNA), regardless of orientation.
[0046] As used herein, "expression cassette" refers to any type of genetic construct containing a nucleic acid (e.g., a transgene) from which part or all of a nucleic acid coding sequence can be transcribed. In some embodiments, expression includes transcription of the nucleic acid, for example, to produce a biologically active polypeptide product or inhibitory RNA (e.g., siRNA, shRNA, miRNA) from the transcribed gene. In some embodiments, the transgene is operably linked to an expression control sequence.
[0047] As used herein, the term "transgene" refers to an exogenous nucleic acid that is introduced into another cell, either naturally or by genetic engineering means, and that can be transcribed and optionally translated.
[0048] As used herein, the term "gene expression" refers to the process by which nucleic acids are transcribed from nucleic acid molecules and often translated into peptides or proteins. The process can include transcription, post-transcriptional control, post-transcriptional modification, translation, post-translational control, post-translational modification, or any combination thereof. Reference to measuring "gene expression" can refer to measuring transcription products (e.g., RNA or mRNA) or translation products (e.g., peptides or proteins).
[0049] As used herein, the term "inhibiting gene expression" means decreasing, down-regulating, suppressing, blocking, reducing, or stopping the expression of a gene. The expression product of a gene can be an RNA molecule (e.g., mRNA) transcribed from the gene, or a polypeptide translated from the mRNA transcribed from the gene. A decrease in the level of mRNA results in a decrease in the level of the polypeptide translated therefrom. In some embodiments, inhibiting expression reduces the level of a polypeptide without substantially affecting the production of the mRNA encoding the polypeptide. The level of expression can be determined using standard techniques for measuring mRNA or protein.
[0050] As used herein, a "vector" refers to a genetic construct capable of transporting nucleic acid molecules (e.g., a transgene encoding an inhibitory nucleic acid) between cells and expressing the nucleic acid molecule when operably linked to appropriate expression control sequences. Expression control sequences can include transcription initiation, termination, promoter, and enhancer sequences; efficient RNA processing signals such as splicing and polyadenylation (polyA) signals; sequences that stabilize cytoplasmic mRNA; sequences that increase translation efficiency (i.e., Kozak consensus sequences); sequences that increase protein stability; and, optionally, sequences that promote secretion of the encoded product. Vectors can be plasmids, phage particles, transposons, cosmids, phagemids, chromosomes, artificial chromosomes, viruses, virions, lipid nanoparticles, and the like. Once transformed into a suitable host cell, the vector can replicate and function independently of the host genome or, in some cases, can be integrated into the genome itself.
[0051] As used herein, "host cell" refers to any cell that contains or is capable of containing a composition of interest, e.g., an inhibitory nucleic acid. In embodiments, the host cell is a mammalian cell, such as a rodent cell (e.g., a mouse or rat) or a primate cell (e.g., a monkey, chimpanzee, human). In embodiments, the host cell may be in vitro or in vivo. In embodiments, the host cell may be derived from an established cell line or a primary cell. In embodiments, the host cell may be obtained from a patient having or suspected of having a repeat expansion disease or disorder. In embodiments, the host cell is a non-CNS cell, such as a fibroblast. In embodiments, the host cell is a cell of the CNS, such as a neuron, glial cell, astrocyte, or microglial cell.
[0052] As used herein, "expanded repeat-containing gene" or "expanded repeat-containing RNA" refers to a mutant gene or RNA molecule (e.g., a pre-mRNA or mRNA) encoded by a mutant gene having a base sequence containing a repetitive region (e.g., CAG repeat), where the repetitive region is extended beyond a predetermined number or range of base repeats typically present in a "normal" expanded repeat-containing gene or RNA encoded by the gene. The presence or length of the repeat region can affect the normal processing, function, or activity of the RNA or the encoded protein, potentially causing a "repeat expansion" or "expanded repeat" disease or disorder. The expanded repeat may be an unstable (dynamic) mutation that changes size with successive generations. The expanded repeat may be a dinucleotide repeat, trinucleotide repeat, tetranucleotide repeat, pentanucleotide repeat, hexanucleotide repeat, etc. In some embodiments, the repeat is a CAG repeat or polyglutamine. The expanded repeat-containing gene or RNA encoded by the expanded repeat-containing gene may also be referred to as a "pathological allele" or "pathogenic allele." In some embodiments, the pathological or pathogenic allele of the CAG repeat-containing gene or RNA encoded by the gene has 30 or more consecutive CAG repeats.
[0053] "Repeat expansion disease or disorder" or "expanded repeat disease or disorder" refers to a disease or disorder caused by the expansion of a base repeat sequence beyond a predetermined number or range of base repeats typically present in the "normal" expanded repeat-containing gene or RNA encoded by the gene. Repeat expansion diseases or disorders can manifest with significantly varying phenotypes depending on the size of the repeat expansion. Repeat expansion diseases or disorders are primarily neurodegenerative diseases. Some repeat expansion diseases are ophthalmological diseases. In some embodiments, the repeat expansion or disorder is a polyglutamine disease.
[0054] As used herein, "neurodegenerative disease" or "neurodegenerative disorder" refers to a disease or disorder that exhibits neuronal cell death as a pathological condition. Neurodegenerative diseases can exhibit chronic neurodegeneration, e.g., slow, progressive neuronal cell death over several years, or acute neurodegeneration, e.g., sudden onset or death of neuronal cells. Examples of chronic neurodegenerative diseases include Alzheimer's disease, Parkinson's disease, Huntington's disease, spinocerebellar ataxia types 1-8 (SCA1-8), frontotemporal dementia (FTD), and amyotrophic lateral sclerosis (ALS). Neurodegenerative diseases can exhibit primarily the death of one type of neuron or multiple types of neurons. As used herein, the terms "subject," "patient," and "individual" are used interchangeably herein and refer to a living organism (e.g., a mammal) that is treated or selected for treatment. Examples of subjects include humans and non-human mammals such as primates (monkeys, chimpanzees), cows, horses, sheep, dogs, cats, rats, mice, guinea pigs, pigs, and transgenic species thereof.
[0055] A. Double-stranded RNA
[0056] The present disclosure provides artificial double-stranded RNA that functions as artificial microRNA or shRNA.The double-stranded RNA of the present disclosure regulates the expression of target RNA (for example, mRNA or mRNA precursor) transcripts.Double-stranded RNA includes precursor molecules that are processed in cells before regulation.Double-stranded RNA can be encoded in plasmids, vectors, genomes, or other nucleic acid expression vectors for delivery to cells.
[0057] In some embodiments, the artificial double-stranded RNA comprises, 5' to 3': (a) a 5' leader sequence; (b) a 5' stem comprising or substantially comprising a passenger or guide sequence; (c) a 5' linker of 1 to 6 bases; (d) a terminal loop; (e) a 3' linker of 1 to 6 bases; (f) a 3' stem comprising or substantially comprising: (i) a guide sequence whose 5' stem comprises or substantially comprises a passenger sequence; or (ii) a passenger sequence whose 5' stem comprises or substantially comprises a guide sequence; and (g) a 3' trailer sequence; wherein the guide sequence targets a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA and comprises 1 to 5 base mismatches relative to the CAG repeat region, the base mismatches being located at positions 8 to 16 of the guide sequence.
[0058] In some embodiments, double-stranded RNA refers to a single RNA oligonucleotide compound that has at least partial self-complementarity and forms stable self-duplex.Unless otherwise specified, the numbering of the nucleotide position in double-stranded RNA is counted from 5' to 3' on single-stranded RNA.Similarly, unless otherwise specified, nucleotide sequence is read from 5' to 3' on single-stranded RNA.
[0059] The artificial double-stranded RNA of the present disclosure comprises a 5' leader, also referred to as a 5' flanking sequence. The 5' leader sequence may be derived from a wild-type microRNA sequence, or may be wholly or partially derived from a wild-type microRNA sequence, or may be wholly or partially artificial. In some embodiments, the 5' leader sequence is derived from a flanking sequence of a wild-type pre-miRNA scaffold or pri-miRNA scaffold, or may be wholly or partially derived from a flanking sequence of a wild-type pre-miRNA scaffold.
[0060] The 5' leader sequence is contiguous with a 5' stem that includes or substantially includes a passenger or guide sequence. The 5' leader sequence may be any length. In some embodiments, the 5' leader sequence is about 1 nucleotide to about 1,000 nucleotides in length, about 1 nucleotide to about 900 nucleotides, about 1 nucleotide to about 800 nucleotides, about 1 nucleotide to about 700 nucleotides, about 1 nucleotide to about 600 nucleotides, about 1 nucleotide to about 500 nucleotides, about 1 nucleotide to about 400 nucleotides, about 1 nucleotide to about 300 nucleotides, about 1 nucleotide to about 200 nucleotides, about 1 nucleotide to about 100 nucleotides, about 1 nucleotide to about 75 nucleotides, about 1 nucleotide to about 50 nucleotides, about 1 nucleotide to about 25 nucleotides, about 1 nucleotide to about 20 nucleotides, about 1 nucleotide to about 15 nucleotides, or about 1 nucleotide to about 10 nucleotides in length.
[0061] In some embodiments, the 5' leader comprises a 5' bulge sequence. As used herein, the term "bulge sequence" refers to a region of nucleic acid that is non-complementary to the nucleic acid opposite it in a duplex. For example, a duplex can comprise a region of complementary nucleic acid, followed by a region of non-complementary nucleic acid, followed by a second region of complementary nucleic acid. The regions of complementary nucleic acid bind to each other, but the central non-complementary region does not, thereby forming a "bulge." In some embodiments, the two strands of nucleic acid located between the two complementary regions are of different lengths, thereby forming a "bulge."
[0062] The artificial double-stranded RNA of the present disclosure comprises a 3' trailer, also referred to as a 3' flanking sequence. The 3' trailer sequence may be derived from a wild-type microRNA sequence, or may be wholly or partially derived from a wild-type microRNA sequence, or may be wholly or partially artificial. In some embodiments, the 3' trailer sequence is derived from a flanking sequence of a wild-type pre-miRNA scaffold or pri-miRNA scaffold, or may be wholly or partially derived from a flanking sequence of a wild-type pre-miRNA scaffold or pri-miRNA scaffold.
[0063] The 3' trailer sequence is linked adjacent to a 3' stem that includes or substantially includes a guide sequence or passenger sequence. The 3' trailer sequence can be any length. In some embodiments, the 3' trailer sequence is about 1 nucleotide to about 1,000 nucleotides, about 1 nucleotide to about 900 nucleotides, about 1 nucleotide to about 800 nucleotides, about 1 nucleotide to about 700 nucleotides, about 1 nucleotide to about 600 nucleotides, about 1 nucleotide to about 500 nucleotides, about 1 nucleotide to about 400 nucleotides, about 1 nucleotide to about 300 nucleotides, about 1 nucleotide to about 200 nucleotides, about 1 nucleotide to about 100 nucleotides, about 1 nucleotide to about 75 nucleotides, about 1 nucleotide to about 50 nucleotides, about 1 nucleotide to about 25 nucleotides, about 1 nucleotide to about 20 nucleotides, about 1 nucleotide to about 15 nucleotides, or about 1 nucleotide to about 10 nucleotides in length. In some embodiments, the 3' trailer includes a 3' bulge sequence.
[0064] In some embodiments, the 3' trailer comprises a poly-U (poly-uridine) tail. In some embodiments, the 3' trailer comprises 3 to 6 uridines, e.g., 3 uridines, 4 uridines, 5 uridines, or 6 uridines. In some embodiments, the poly-U tail is immediately adjacent to the guide sequence or passenger sequence in the 3' stem. In some embodiments, an artificial double-stranded RNA having a 3' trailer comprising a poly-U tail is expressed using a Pol III promoter.
[0065] In some embodiments, the 3' trailer comprises a polyadenylation (pA) signal sequence. Suitable polyadenylation signals include, but are not limited to, the SV40 late pA signal, the BGH pA signal, and the like. In some embodiments, an artificial double-stranded RNA having a 3' trailer comprising a pA signal sequence is expressed using a Pol II promoter.
[0066] In some embodiments, the 5' leader sequence and the 3' trailer sequence have the same number of nucleotides. In some embodiments, the 5' leader sequence and the 3' trailer sequence have different lengths.
[0067] In certain embodiments, the 5' leader sequence and the 3' trailer sequence are obtained or derived, in whole or in part, from the same miRNA scaffold, for example, the same wild-type pre-miRNA scaffold or the same pri-miRNA scaffold. In certain embodiments, the 5' leader sequence and the 3' trailer sequence are both obtained or derived, in whole or in part, from a miR-33 scaffold. In some embodiments, the 5' leader sequence and the 3' trailer sequence are both obtained or derived, in whole or in part, from a pri-miR-33 scaffold. In some embodiments, the 5' leader sequence and the 3' trailer sequence are both obtained or derived, in whole or in part, from a pri-miR-33 scaffold. In some embodiments, the 5' leader sequence and the 3' trailer sequence are selected from Table E.
[0068] In some embodiments, the 5' leader sequence is not complementary to the 3' trailer sequence. In some embodiments, the 5' leader sequence is partially complementary to the 3' trailer sequence. In some embodiments, the 5' leader sequence contains one, two, or more C mismatches to the uridines in the poly-U tail in the 3' trailer sequence (or CT mismatches for a DNA sequence encoding a double-stranded RNA).
[0069] In some embodiments, the 5' leader and 3' trailer sequences contain sequences that enable recognition and cleavage by Drosha. The canonical pathway of miRNA biogenesis in mammals is initiated by the Drosha-DGCR8 (DiGeorge syndrome critical region gene 8) complex (microprocessor), which processes a long primary miRNA (pri-miRNA) into a ~60 nt pre-miRNA, which is further processed by Dicer into a ~22 nt long duplex. In some embodiments, the primary miRNA sequence used as, or as part of, the 5' leader and / or 3' trailer sequence can direct Drosha cleavage of the double-stranded RNA. Guide: Methods for using precursor miRNAs as scaffolds for the selected expression of passenger duplexes are provided in U.S. Patent Publication No. 2008 / 0226553 and Liu et al. (2008) Nucleic Acids Res. 36:2811-24, each of which is incorporated by reference in its entirety.
[0070] In some embodiments, the artificial double-stranded RNA is processed by a Drosha-independent / Dicer-dependent pathway. In some embodiments, splicing, 3'-5' exoribonuclease, or pol III termination can be an alternative to Drosha cleavage.
[0071] In some embodiments, the artificial double-stranded RNA is processed by a Drosha-dependent / Dicer-independent pathway. An example of Drosha-dependent / Dicer-independent pathway processing is provided by pri-miR-451, which is processed by Drosha to produce pre-miR-451, which is then cleaved by Ago2 (argonaute 2), ac-pre-miR-451, which is further excised by an unknown mechanism to generate mature miR-451.
[0072] In some embodiments, the 5' leader sequence and / or the 3' trailer sequence comprise or consist of a nucleotide sequence set forth in Table A. In some embodiments, the artificial double-stranded RNA comprises a 5' leader sequence comprising or consisting of CCGG and a 3' trailer sequence comprising or consisting of UUUUG. In some embodiments, the artificial double-stranded RNA comprises a 5' leader sequence comprising or consisting of CC and a 3' trailer sequence comprising or consisting of UUUUG. In some embodiments, the artificial double-stranded RNA comprises a 5' leader sequence comprising or consisting of GCUG and a 3' trailer sequence comprising or consisting of ga uuuuug. In some embodiments, the artificial double-stranded RNA comprises a 5' leader sequence comprising or consisting of SEQ ID NO:7 and a 3' trailer sequence comprising or consisting of SEQ ID NO:13.
[0073] TIFF2024543782000002.tif100164
[0074] In some embodiments, the 5' stem (or 5' arm) of the double-stranded RNA comprises a passenger sequence, also referred to as a sense sequence. The passenger sequence has identity to the target mRNA transcript. The passenger sequence can be about 15 to 30 nucleotides in length, for example, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the passenger sequence can be about 19 to 24 nucleotides in length.
[0075] In some embodiments, the 3' stem (or 3' arm) of the double-stranded RNA comprises a guide sequence, also known as an antisense sequence. The guide sequence is complementary to the target mRNA transcript. The guide strand can be approximately 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the guide sequence can be approximately 19 to 24 nucleotides in length.
[0076] In some embodiments, the 5' stem comprises a guide sequence and the 3' stem comprises a passenger sequence of double-stranded RNA.
[0077] The guide sequence and passenger sequence have sufficient complementarity to form a double-stranded siRNA molecule during processing in the host cell, which serves as a suitable substrate for the RNA interference machinery so that the guide sequence from the 3' stem (or 5' stem) is recognized by the RISC complex and targets its specific mRNA transcript. In some embodiments, the guide sequence and passenger sequence have 100% complementarity. In some embodiments, the guide sequence and passenger sequence are substantially complementary to each other, for example, about 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary. In some embodiments, the passenger sequence can contain 1-10 or 1-5 base mismatches or bulges.
[0078] The guide sequence has perfect or near-perfect Watson-Crick complementarity to the target mRNA sequence and can include a seed sequence located at positions 1-7, 2-7, 1-8, or 2-8 of the guide sequence relative to the first 5' nucleotide of the guide strand. The seed region is important for efficient gene silencing by double-stranded RNA.
[0079] The guide sequence targets the CAG repeat region of the CAG repeat-containing mRNA and contains one to five mismatches to the CAG repeat region, where the base mismatches are located at positions 8-16 of the guide sequence. Mismatches include self-pairing nucleotides (AA, UU, TT, CC, and GG), A to C pairings, C to U pairings, C to T pairings, and A to G pairings. In some embodiments, mismatches include purine mismatches, such as introducing an adenosine base into the guide strand.
[0080] In some embodiments, a guide sequence targeting a CAG repeat region contains about 1-5 base mismatches relative to the CAG repeat region of a CAG repeat-containing mRNA. In some embodiments, a guide sequence targeting a CAG repeat region contains about 1-4 base mismatches, about 1-3 base mismatches, about 1-2 base mismatches, about 2-5 base mismatches, about 3-4 base mismatches, about 3-5 base mismatches, or about 4-5 base mismatches relative to the CAG repeat region of a CAG repeat-containing mRNA. In some embodiments, a guide sequence targeting a CAG repeat region contains about 1 mismatch, about 2 mismatches, about 3 mismatches, about 4 mismatches, or about 5 mismatches relative to the CAG repeat region of a CAG repeat-containing mRNA.
[0081] In some embodiments, at least one mismatch (1, 2, 3, 4, or 5) to the CAG repeat region may be located at positions 8-12 of the guide sequence. In some embodiments, at least one mismatch (1, 2, or 3) to the CAG repeat region may be located at positions 9-11 of the guide sequence. In some embodiments, two or more mismatches to the CAG repeat region located at positions 8-16 of the guide sequence are contiguous or adjacent to each other. In some embodiments, at least one mismatch to the CAG repeat region located at positions 8-16 is not adjacent to another mismatch located at positions 8-16.
[0082] In some embodiments, one mismatch to the CAG repeat region is located at position 8 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 9 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 10 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 11 of the guide sequence.
[0083] In some embodiments, one mismatch to the CAG repeat region is located at position 8 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 9-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 10-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 10 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 11-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 11 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 12-16 of the guide sequence.
[0084] In some embodiments, one mismatch to the CAG repeat region is located at position 8 of the guide sequence, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 10-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, one mismatch to the CAG repeat region is located at position 10 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 11-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 10 of the guide sequence, one mismatch to the CAG repeat region is located at position 11 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 12-16 of the guide sequence. In some embodiments, one mismatch to the CAG repeat region is located at position 11 of the guide sequence, one mismatch to the CAG repeat region is located at position 12 of the guide sequence, and one mismatch to the CAG repeat region is located at any of positions 13-16 of the guide sequence.
[0085] In some embodiments, one mismatch to the CAG repeat region is located at position 8 of the guide sequence, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, one mismatch to the CAG repeat region is located at any of positions 12-16 of the guide sequence, and one mismatch to the CAG repeat region is located at position 14, 15, or 16. In some embodiments, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, one mismatch to the CAG repeat region is located at position 10 of the guide sequence, one mismatch to the CAG repeat region is located at any of positions 11-16 of the guide sequence, and one mismatch to the CAG repeat region is located at position 15 or 16. In some embodiments, one mismatch to the CAG repeat region is located at position 9 of the guide sequence, one mismatch to the CAG repeat region is located at position 11 of the guide sequence, one mismatch to the CAG repeat region is located at any of positions 12-16 of the guide sequence, and one mismatch to the CAG repeat region is located at position 13, 14, 15 or 16.
[0086] In some embodiments, in a guide sequence targeting a CAG repeat region within a CAG register, one mismatch to the CAG repeat region is at position 9 of the guide sequence, one mismatch to the CAG repeat region is at position 10 of the guide sequence, one mismatch to the CAG repeat region is at position 11 of the guide sequence, one mismatch to the CAG repeat region is at position 15 of the guide sequence, and one mismatch to the CAG repeat region is at position 16 of the guide sequence. In some embodiments, the mismatches at positions 9, 10, and 11 are A, A, and A, respectively, and the mismatches at positions 15 and 16 are AA, AU, UA, or UU.
[0087] In some embodiments, guide sequences targeting a CAG repeat region in the UGC register have one mismatch to the CAG repeat region at position 9 of the guide sequence, one mismatch to the CAG repeat region at position 10 of the guide sequence, one mismatch to the CAG repeat region at position 12 of the guide sequence, one mismatch to the CAG repeat region at position 15 of the guide sequence, and one mismatch to the CAG repeat region at position 16 of the guide sequence. In some embodiments, the mismatches at positions 9, 10, and 12 are A, C, and A, respectively, and the mismatches at positions 15 and 16 are AA, UA, or UU.
[0088] In some embodiments, a guide sequence targeting a CAG repeat region in a GCU register comprises one mismatch to the CAG repeat region at position 9 of the guide sequence, one mismatch to the CAG repeat region at position 10 of the guide sequence, one mismatch to the CAG repeat region at position 11 of the guide sequence, one mismatch to the CAG repeat region at position 15 of the guide sequence, and one mismatch to the CAG repeat region at position 16 of the guide sequence. In some embodiments, the mismatches at positions 9, 10, and 11 are A, U, and A, respectively, and the mismatches at positions 15 and 16 are AA or AU.
[0089] In some embodiments, the guide sequence targeting the CAG repeat region comprises or consists of any sequence selected from Tables B1-B2. In some embodiments, the double-stranded RNA comprises a guide sequence selected from Tables B1-B2 and a corresponding passenger sequence that is either fully complementary to the guide sequence or has 1 to 10 or 1 to 5 mismatches or bulges compared to the selected guide sequence.
[0090] TIFF2024543782000003.tif109164 Nucleotide mismatches in the guide sequence relative to the CAG repeat region of the CAG repeat-containing mRNA are shown in bold and underlined.
[0091] TIFF2024543782000004.tif212163TIFF2024543782000005.tif227163TIFF202 4543782000006.tif226163TIFF2024543782000007.tif226159TIFF2024543782 000008.tif225164TIFF2024543782000009.tif226162TIFF2024543782000010. tif226162TIFF2024543782000011.tif225162TIFF2024543782000012.tif26163
[0092] In some embodiments, the specificity or selectivity of a guide sequence targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA increases with the number of base mismatches to the CAG repeat region, where the base mismatches are located at positions 8-16 of the guide sequence. The specificity (or off-target activity) of a guide sequence can be determined by detecting potential off-target matches in the human unspliced transcriptome (Ensembl database, release 100). For example, for a 21-mer guide sequence that has perfect complementarity to the CAG repeat (or AGC or GCA if using a guide sequence that targets a repeat in a different register) in the seed sequence (nucleotides 1-7 from the 5' end) and targets a CAG repeat with a mismatch at positions 8, 9, 10, or 11 of the guide sequence, the following steps can be used to measure off-target activity: measure the frequency of off-target genes that are perfect matches to the guide sequence; measure the frequency of off-target genes that are perfect 17-mer matches to the guide sequence within positions 1-21; and measure the frequency of off-target genes that are matched to the guide sequence with 0, 1, 2, 3, or 4 mismatches between positions 8 and 21. The off-target frequencies of perfect matches, perfect 17-mer matches across positions 1-21, matches with guides with 0 mismatches, matches with guides with 1 mismatch, off-target matches with guides with 2 mismatches, off-target matches with guides with 3 mismatches, and off-target matches with guides with 4 mismatches can be tallied.
[0093] In some embodiments, the specificity of a guide sequence targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA is increased by 2 to 5 base mismatches relative to the CAG repeat region, with the base mismatches located at positions 8 to 16 of the guide sequence. In some embodiments, the specificity of a guide sequence targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA is increased by 3 to 5 base mismatches relative to the CAG repeat region, with the base mismatches located at positions 8 to 16 of the guide sequence.
[0094] In some embodiments, guide sequences targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA with 1, 2, 3, 4, or 5 base mismatches to the CAG repeat region have zero predicted perfect match off-target transcripts.
[0095] In some embodiments, a guide sequence targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA has 1, 2, 3, 4, or 5 base mismatches to the CAG repeat region, where the base mismatches are located at positions 8-16 of the guide sequence, and has 0-1 predicted off-target transcripts with a perfect 17-mer match within positions 1-21.
[0096] In some embodiments, a guide sequence targeting a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA with 1, 2, 3, 4, or 5 base mismatches to the CAG repeat region, where the base mismatches are located at positions 8-16 of the guide sequence, and 0-2 predicted off-target transcripts with 1 mismatch.
[0097] In some embodiments, a guide sequence that targets a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA with a 1, 2, 3, 4, or 5 base mismatch to the CAG repeat region, where the base mismatch is located at positions 8-16 of the guide sequence, and 0-2 predicted off-target transcripts have perfect 17-mer matches within positions 1-21.
[0098] In some embodiments, a guide sequence that targets a CAG repeat region of a CAG repeat-containing mRNA or pre-mRNA with 1, 2, 3, 4, or 5 base mismatches to the CAG repeat region, where the base mismatches are located at positions 8-16 of the guide sequence, and 0-66 predicted off-target transcripts with 1 mismatch.
[0099] TIFF2024543782000013.tif232155TIFF2024543782000014.tif235157TIFF2024543782000015.tif234157 TIFF2024543782000016.tif236158TIFF2024543782000017.tif235156TIFF2024543782000018.tif232159 TIFF2024543782000019.tif233157TIFF2024543782000020.tif232158TIFF2024543782000021.tif231161 TIFF2024543782000022.tif231158TIFF2024543782000023.tif232158TIFF2024543782000024.tif232161
[0100] Table B4 shows exemplary filtering criteria for off-target activity of guide sequences targeting CAG repeat regions of CAG repeat-containing mRNAs or pre-mRNAs from Table B3. In some embodiments, the specificity of guide sequences targeting CAG repeat regions of CAG repeat-containing mRNAs or pre-mRNAs follows any one or combination of the thresholds set out in Table B4.
[0101] TIFF2024543782000025.tif217165TIFF2024543782000026.tif228164TIFF2024543782000027.tif84165
[0102] In some embodiments, a linker is present in the artificial double-stranded RNA and connects the stem and loop of the artificial double-stranded RNA. In some embodiments, a 5' linker connects the 5' stem and loop of the artificial double-stranded RNA. In some embodiments, a 3' linker connects the 3' stem and loop of the artificial double-stranded RNA. In some embodiments, a 5' linker connects the 5' stem and loop of the artificial double-stranded RNA, and a 3' linker connects the 3' stem and loop of the artificial double-stranded RNA. In some embodiments, the 5' linker and / or 3' linker has about 1 to 6 nucleotides, 1 to 5 nucleotides, 1 to 4 nucleotides, 1 to 3 nucleotides, 1 to 2 nucleotides, 2 to 6 nucleotides, 3 to 6 nucleotides, 4 to 6 nucleotides, 5 to 6 nucleotides, 2 to 5 nucleotides, or 2 to 4 nucleotides. In some embodiments, the 5' linker and / or 3' linker has about 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, or 6 nucleotides. In some embodiments, the 5' linker has the same number of nucleotides as the 3' linker. In some embodiments, the 5' linker is 100% complementary to the 3' linker. In some embodiments, the 5' linker comprises or consists of the nucleotide sequence CAGC, and / or the 3' linker comprises or consists of the nucleotide sequence GCUG. In some cases, the sbRNAs of the present disclosure do not comprise a linker.
[0103] In some embodiments, the 5' linker and the 3' linker each comprise at least 4 nucleotides, and optionally, at least 75% of the 5' linker nucleotides are complementary to the 3' linker nucleotides.
[0104] Examples of 5' and 3' linkers are provided in Table C. In some embodiments, the double-stranded RNA has a 5' linker of SEQ ID NO:15 and a 3' linker of SEQ ID NO:23.
[0105] TIFF2024543782000028.tif106157
[0106] In certain embodiments, a terminal loop separates the 5' linker and the 3' linker of the artificial double-stranded RNA. The terminal loop sequence may be of any length, and may be from about 4 nucleotides to about 1,000 nucleotides, from about 4 nucleotides to about 900 nucleotides, from about 4 nucleotides to about 800 nucleotides, from about 4 nucleotides to about 700 nucleotides, from about 4 nucleotides to about 600 nucleotides, from about 4 nucleotides to about 500 nucleotides, from about 4 nucleotides to about 400 nucleotides, from about 4 nucleotides to about 300 nucleotides, from about 4 nucleotides to about 200 nucleotides, from about 4 nucleotides to about 100 nucleotides, from about 4 nucleotides to about 90 nucleotides, from about 4 nucleotides to about 80 nucleotides, from about 4 nucleotides to about 70 nucleotides, from about 4 nucleotides to about 50 nucleotides, from about 4 nucleotides to about 40 nucleotides, from about 4 nucleotides to about 30 nucleotides, from about 4 nucleotides to about 20 nucleotides, from about 4 nucleotides to about 15 nucleotides, or from about 4 nucleotides to about 10 nucleotides. In some embodiments, the terminal loop sequence has about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides. In some embodiments, the terminal loop comprises a palindromic sequence. In some embodiments, the terminal loop comprises an asymmetric sequence. In some embodiments, the terminal loop comprises or consists of the nucleotide sequence ACCUGC. Examples of loop sequences are provided in Table D.
[0107] TIFF2024543782000029.tif61162
[0108] Further embodiments of the 5' leader sequence, 3' trailer sequence, and terminal loop sequence are provided in Table E. In some embodiments, the double-stranded RNA is an artificial miRNA comprising 5'-3':5' leader sequence, passenger or guide sequence, terminal loop, guide or passenger sequence, and 3' trailer sequence, where the guide sequence targets the CAG repeat region of the CAG repeat-containing RNA and comprises a 1-5 base mismatch relative to the CAG repeat region, where the base mismatch is located at positions 8-16 of the guide sequence. In some embodiments, the 5' leader sequence, 3' trailer sequence, and terminal loop of the artificial miRNA are selected from Table E. In some embodiments, the guide sequence is selected from Tables B1-B2.
[0109] TIFF2024543782000030.tif143167TIFF2024543782000031.tif225164TIFF2024543782000032.tif178164
[0110] In some embodiments, a 5' leader of Table E is extended at the 5' end with four nucleotides that are non-complementary to the 5'-terminal four nucleotides of the corresponding 5' leader sequence, and / or a corresponding 3' trailer of Table E is extended at the 3' end, for improved processing. In some embodiments, a 5' leader of Table E is extended at the 5' end with four Us, and / or a corresponding 3' trailer of Table E is extended at the 3' end with four Us, for improved processing.
[0111] It should be understood that the double-stranded RNA of the present disclosure that comprises a guide sequence that targets the CAG repeat region of CAG repeat-containing mRNA also comprises a double-stranded RNA that comprises a guide sequence that targets a different frame, also called register, of the CAG repeat region of CAG repeat-containing mRNA.Therefore, targeting CAG repeat includes a guide sequence that has a frame shift of +1 for targeting AGC repeat, or a frame shift of +2 for targeting GCA repeat and targeting the same CAG repeat-containing mRNA transcript.
[0112] In some embodiments, cleavage by Drosha and / or Dicer determines the sequence and function of the siRNA produced from the double-stranded RNA. In some embodiments, the double-stranded RNA is cleaved by Drosha within the 5' leader sequence and / or 3' trailer sequence to produce an shRNA comprising a guide sequence and a passenger sequence. In some embodiments, the shRNA is cleaved by Dicer to produce an siRNA. The siRNA is loaded onto an RNA-induced silencing complex (RISC). In some embodiments, the double-stranded RNA is a pri-miRNA-like molecule that is cleaved by Drosha to produce a pre-miRNA. The pre-miRNA molecule is an shRNA-like molecule that can then be processed by Dicer to produce an siRNA-like duplex. In some embodiments, cleavage of the double-stranded RNA is Dicer-independent (e.g., cleaved by Ago2). In some embodiments, the shRNA produced from the double-stranded RNA has a 5' overhang and / or a 3' overhang, e.g., 1 to 6 nucleotides. In some embodiments, the shRNA produced from the double-stranded RNA has a 2-3 nucleotide overhang at the 5' and / or 3' end. In some embodiments, the shRNA produced from the double-stranded RNA has a dinucleotide overhang at the 5' and / or 3' end. In some embodiments, the shRNA produced from the double-stranded RNA has a length of about 38-300 nucleotides, 38-250 nucleotides, 38-200 nucleotides, 38-150 nucleotides, 38-100 nucleotides, 38-75 nucleotides, or 38-50 nucleotides.
[0113] In some embodiments, siRNAs produced or processed from double-stranded RNA in mammalian cells contain a 1-5 base mismatch to the CAG repeat region at their predicted positions within positions 8-16 of the guide sequence. The presence of a 1-5 base mismatch to the CAG repeat region at their predicted positions within positions 8-16 of the guide sequence may reflect correct 5' processing of the guide strand. In some embodiments, siRNAs produced or processed from double-stranded RNA in mammalian cells that have a 1-5 base mismatch to the CAG repeat region at their predicted position within positions 8-16 of a guide sequence are at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 99% or more abundant in mammalian cells (in vitro or in vivo) compared to other siRNAs produced or processed from the same double-stranded RNA.
[0114] The guide strand guides RISC to the cognate target mRNA in a sequence-specific manner. In some embodiments, the guide strand induces cleavage of the target mRNA transcript. In some embodiments, the guide strand induces translational repression and / or post-repression through mRNA decay.
[0115] In some cases, a double-stranded RNA of the disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand, wherein the first strand i) comprises a first mismatch to the target CAG repeat region; and ii) comprises at least a second mismatch to the target CAG repeat region, wherein 1) the first mismatch is at position 8 based on the numbering of SEQ ID NO:743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743)), SEQ ID NO:744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:744)), SEQ ID NO:745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:745)), SEQ ID NO:866 (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO:866)), or SEQ ID NO:867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO:867)). In some cases, the second mismatch is 1 to 8 bases 3' to the first mismatch; 2) if the first mismatch is SEQ ID NO:743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743)), SEQ ID NO:744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:744)), or SEQ ID NO:745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:745), the second mismatch is at position 9 based on the first numbering, the second mismatch is 1 to 7 bases 3' to the mismatch; 3) if the first mismatch is at position 10 based on the numbering of SEQ ID NO:743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743)), SEQ ID NO:744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:744)), or (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743)), the second mismatch is 1 to 6 bases 3' to the first mismatch. and 4) when the first mismatch is at position 11 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), the second mismatch is 1 to 5 bases 3' to the first mismatch.In some cases, the first strand contains two or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches with the first strand (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches). In some cases, the second strand contains one to four mismatches with the first strand. In some cases, the second strand contains one, two, three, or four or fewer mismatches with the first strand. In some cases, the second strand contains one or fewer mismatches with the first strand. In some cases, the second strand contains two or fewer mismatches with the first strand. In some cases, the second strand contains three or fewer mismatches with the first strand. In some cases, the second strand contains four or fewer mismatches relative to the first strand. In some cases, the second strand contains five or fewer mismatches relative to the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, the guide sequence comprises any of SEQ ID NOs: 298-375 (see Table 6; Figure 22) and has a length of 21 nucleotides. In some cases, each mismatch is generated by substituting a nucleotide (e.g., a nucleotide present in CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), a nucleotide present in GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), or a nucleotide present in UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867)) with a different nucleotide.In some cases, each mismatch is generated by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G.
[0116] First mismatch at position 8; two mismatches
[0117] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA; and b) a second strand that hybridizes to the first strand, the first strand comprising: i) a first mismatch to the target CAG repeat region, the first mismatch being at position 8 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), or SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745); and ii) a second mismatch to the target CAG repeat region, the second mismatch being 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence: CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), where the first substitution creates a first mismatch and is at position 8 based on the numbering of CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), and the second substitution creates a second mismatch and is 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence: GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 8 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), and the second substitution creates a second mismatch and is 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) comprising at least a first and a second substitution, wherein the first substitution creates a first mismatch and is at position 8 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), and the second substitution creates a second mismatch and is 1 to 8 bases 3' to the first mismatch. In some cases, the first strand comprises no more than two mismatches with the target CAG repeat region. In some cases, the first strand comprises no more than three mismatches with the target CAG repeat region.In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains one to ten mismatches (e.g., one to four, three to five, five to seven, or five to ten mismatches) with the first strand. In some cases, the second strand contains one to four mismatches with the first strand. In some cases, the second strand contains one, two, three, or four or fewer mismatches with the first strand. In some cases, the second strand contains one or fewer mismatches with the first strand. In some cases, the second strand contains two or fewer mismatches with the first strand. In some cases, the second strand contains three or fewer mismatches with the first strand. In some cases, the second strand contains four or fewer mismatches with the first strand. In some cases, the second strand contains five or fewer mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from the following: a) G to A, U, or C; b) U to A, G, or C; and c) C to A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 317-324. In some cases, the first strand comprises the nucleotide sequence CUGCUGCAACUGCUGCUGCUG (SEQ ID NO: 317; the RNA guide strand sequence for "CUG_NA_B" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGAUGCUGCUGCUG (SEQ ID NO: 318; RNA guide strand sequence for "CUG_NA_C" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCAGCUGCUGCUG (SEQ ID NO: 319; RNA guide strand sequence for "CUG_NA_D" in Table 6).In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCUACUGCUGCUG (SEQ ID NO: 320; the RNA guide strand sequence of "CUG_NA_E" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCUGAUGCUGCUG (SEQ ID NO: 321; the RNA guide strand sequence of "CUG_NA_F" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCUGCAGCUGCUG (SEQ ID NO: 322; the RNA guide strand sequence of "CUG_NA_G" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCUGCUACUGCUG (SEQ ID NO: 323; the RNA guide strand sequence of "CUG_NA_H" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCAGCUGCUGAUGCUG (SEQ ID NO: 324; the RNA guide strand sequence of "CUG_NA_I" in Table 6). In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 804-819 (as shown in Table 8; Figure 24).
[0118] First mismatch at position 8; three mismatches
[0119] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, wherein the first strand comprises a first mismatch to the target CAG repeat region, wherein the first mismatch is at position 8 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 866 (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or SEQ ID NO: 867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867)); and b) a second strand that hybridizes to the first strand, wherein the first strand comprises a second mismatch and a third mismatch to the target CAG repeat region, wherein the second and third mismatches are 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, and third substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the first substitution creates a first mismatch and is at position 8 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the second substitution creates a second mismatch, and the third substitution creates a third mismatch, wherein the second and third substitutions are 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866) that includes at least a first substitution, a second substitution, and a third substitution, wherein the first substitution creates a first mismatch and is at position 8 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, and the second and third substitutions are 1 to 8 bases 3' to the first mismatch.In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) containing at least a first substitution, a second substitution, and a third substitution, wherein the first substitution creates a first mismatch and is located at position 8 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, with the second and third substitutions being 1 to 8 bases 3' to the first mismatch. In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches with the first strand (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches). In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G.
[0120] First mismatch at position 8; four mismatches
[0121] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, the first strand comprising a first mismatch to the target CAG repeat region, the first mismatch being at position 8 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 867 (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 867)), or SEQ ID NO: 866 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 866)); and b) a second strand that hybridizes to the first strand, the first strand comprising a second mismatch, a third mismatch, and a fourth mismatch to the target CAG repeat region, the second, third, and fourth mismatches being 1 to 8 bases 3' of the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743), wherein the first substitution creates a first mismatch and is at position 8 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 8 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the first substitution creates a first mismatch and is at position 8 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 8 bases 3' to the first mismatch.In some cases, the first strand is a variant containing at least a first, second, third, and fourth substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the first substitution creates a first mismatch and is at position 8 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, where the second, third, and fourth substitutions are 1 to 8 bases 3' to the first mismatch. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G.
[0122] First mismatch at position 9; two mismatches
[0123] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA; and b) a second strand that hybridizes to the first strand, wherein the first strand has i) a first mismatch to the target CAG repeat region, and the first strand has a sequence selected from the group consisting of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745), SEQ ID NO: 866 (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or SEQ ID NO: 867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867); and ii) is a second mismatch to the target CAG repeat region, wherein the second mismatch is 1 to 7 bases 3' to the first mismatch. In some cases, the first strand is a variant of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743) comprising at least a first and a second substitution, wherein the first substitution creates the first mismatch. In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 9 based on the numbering of GCUGCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), and the second substitution creates a second mismatch and is 1 to 7 bases 3' to the first mismatch. The first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the first substitution creates a first mismatch and is at position 9 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), and the second substitution creates a second mismatch and is 1 to 7 bases 3' to the first mismatch. Optionally, the first strand contains no more than two mismatches with the target CAG repeat region.In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) to the first strand. In some cases, the second strand contains 1 to 4 mismatches to the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches to the first strand. In some cases, the second strand contains no more than 1 mismatch to the first strand. In some cases, the second strand contains no more than 2 mismatches to the first strand. In some cases, the second strand contains no more than 3 mismatches to the first strand. In some cases, the second strand contains no more than 4 mismatches to the first strand. In some cases, the second strand contains no more than 5 mismatches to the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA is 21 nucleotides in length. In some cases, the double-stranded RNA is 22 nucleotides in length. In some cases, the double-stranded RNA is 23 nucleotides in length. In some cases, the double-stranded RNA is 24 nucleotides in length. In some cases, the double-stranded RNA is 25 nucleotides in length. In some cases, each mismatch is generated by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 298-304.
[0124] First mismatch at position 9; three mismatches
[0125] In some cases, a double-stranded RNA of the present disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand. wherein the first mismatch is a variant of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745), (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or SEQ ID NO: 867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867)), wherein the first strand comprises a second mismatch and a third mismatch relative to the target CAG repeat region, and the second and third mismatches are 1 to 7 bases 3' to the first mismatch. Optionally, the first strand is a variant comprising at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743). wherein the first substitution creates a first mismatch and is at position 9 based on numbering of CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO:743), wherein the second substitution creates a second mismatch and the third substitution creates a third mismatch, and wherein the second and third substitutions are 1 to 7 bases 3' to the first mismatch. In some cases, the first strand is a variant including at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO:866), wherein the first substitution creates a first mismatch and is at position 9 based on numbering of GCUGCUGCUGCUGCUGCUGCUGCU (SEQ ID NO:866), wherein the second substitution creates a second mismatch and the third substitution creates a third mismatch, and wherein the second and third substitutions are 1 to 7 bases 3' to the first mismatch.In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) containing at least a first substitution, a second substitution, and a third substitution, where the first substitution creates a first mismatch and is located at position 9 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, with the second and third substitutions being 1 to 7 bases 3' to the first mismatch. In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains 1 to 4 mismatches relative to the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches relative to the first strand. In some cases, the second strand contains no more than 1 mismatch relative to the first strand. In some cases, the second strand contains no more than 2 mismatches relative to the first strand. In some cases, the second strand contains no more than 3 mismatches relative to the first strand. In some cases, the second strand contains no more than 4 mismatches relative to the first strand. In some cases, the second strand contains no more than 5 mismatches relative to the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G.In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 325-332 and 336-338. In some cases, the first strand comprises the nucleotide sequence CUGCUGCUAAAGCUGCUGCUG (SEQ ID NO: 325; the RNA guide strand sequence of "CUG_307" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCUAAUACUGCUGCUG (SEQ ID NO: 326; the RNA guide strand sequence of "CUG_334" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUAAUGAUGCUGCUG (SEQ ID NO: 327; the RNA guide strand sequence of "CUG_361" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUAAUGCAGCUGCUG (SEQ ID NO: 328; the RNA guide strand sequence of "CUG_388" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUAAUGCUACUGCUG (SEQ ID NO: 329; the RNA guide strand sequence of "CUG_415" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUACAGAUGCUGCUG (SEQ ID NO: 330; the RNA guide strand sequence of "CUG_631" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUACAGCAGCUGCUG (SEQ ID NO: 331; the RNA guide strand sequence of "CUG_658" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUACAGCUGAUGCUG (SEQ ID NO: 332; the RNA guide strand sequence of "CUG_712" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUAAUGCUGAUGCUG (SEQ ID NO: 336; the RNA guide strand sequence of "CUG_442" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUACAACUGCUGCUG (SEQ ID NO: 337; the RNA guide strand sequence of "CUG_604" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUACAGCUACUGCUG (SEQ ID NO: 338; the RNA guide strand sequence of "CUG_685" in Table 6).
[0126] First mismatch at position 9; four mismatches
[0127] In some cases, the double-stranded RNA of the present disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand, wherein the first mismatch is SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or SEQ ID NO: 867 UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), wherein the first strand comprises a second mismatch, a third mismatch, and a fourth mismatch relative to the target CAG repeat region, and the second, third, and fourth mismatches are 1 to 7 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the first substitution creates a first mismatch and is at position 9 based on the numbering of CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 7 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the first substitution creates a first mismatch and is at position 9 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 7 bases 3' to the first mismatch.In some cases, the first strand is a variant containing at least a first, second, third, and fourth substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the first substitution creates a first mismatch and is located at position 9 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, where the second, third, and fourth substitutions are 1 to 7 bases 3' to the first mismatch. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA is 18 to 25 nucleotides in length. In some cases, the double-stranded RNA is 20 nucleotides in length. In some cases, the double-stranded RNA is 21 nucleotides in length. In some cases, the double-stranded RNA is 22 nucleotides in length. In some cases, the double-stranded RNA is 23 nucleotides in length. In some cases, the double-stranded RNA is 24 nucleotides in length. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from: a) G with A, U, or C; b) U with A, G, or C; and c) C with A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 341-344 and 347-367.
[0128] First mismatch at position 10; two mismatches
[0129] In some cases, the double-stranded RNA of the present disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand, wherein the first strand has i) a first mismatch to the target CAG repeat region and is selected from the group consisting of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745), (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or the sequence Number 867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867); and ii) is a second mismatch to the target CAG repeat region, the second mismatch being 1 to 6 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), where the first substitution generates the first mismatch and is CUGCUGCUGCUGCUGCUGCU In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 10 based on the numbering of G (SEQ ID NO: 743), and the second substitution creates a second mismatch and is 1 to 6 bases 3' to the first mismatch. The substitution creates a second mismatch and is 1 to 6 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the first substitution creates a first mismatch and is at position 10 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), and the second substitution creates a second mismatch and is 1 to 6 bases 3' to the first mismatch.In some cases, the first strand contains two or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches with the first strand (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches). In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or 4 or fewer mismatches with the first strand. In some cases, the second strand contains one or fewer mismatches with the first strand. In some cases, the second strand contains two or fewer mismatches with the first strand. In some cases, the second strand contains three or fewer mismatches with the first strand. In some cases, the second strand contains four or fewer mismatches with the first strand. In some cases, the second strand contains five or fewer mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from the following: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 305-310.
[0130] First mismatch at position 10; three mismatches
[0131] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA; and b) a second strand that hybridizes to the first strand, wherein a first mismatch is at position 10 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), or (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), and the first strand comprises a second mismatch and a third mismatch relative to the target CAG repeat region, and the second and third mismatches are 1 to 6 bases 3' to the first mismatch, wherein the first strand comprises: In some cases, the first strand is a variant that includes at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), where the first substitution creates a first mismatch and is at position 10 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, and the second and third substitutions are 1 to 6 bases 3' to the first mismatch. In some cases, the first strand is a variant that includes at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 10 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the second substitution creates a second mismatch, and the third substitution creates a third mismatch, where the second substitution and the third substitution are 1 to 6 bases 3' to the first mismatch.In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) containing at least a first substitution, a second substitution, and a third substitution, wherein the first substitution creates a first mismatch and is located at position 10 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, with the second and third substitutions being 1 to 6 bases 3' to the first mismatch. In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) with respect to the first strand. In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA is 20 nucleotides long. In some cases, the double-stranded RNA is 21 nucleotides long. In some cases, the double-stranded RNA is 22 nucleotides long. In some cases, the double-stranded RNA is 23 nucleotides long. In some cases, the double-stranded RNA is 24 nucleotides long. In some cases, the double-stranded RNA is 25 nucleotides long. In some cases, each mismatch is generated by a substitution independently selected from the following: a) G with A, U, or C; b) U with A, G, or C; and c) C with A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 333-335, 339, and 340.In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUGAAGAUGCUGCUG (SEQ ID NO: 333; the RNA guide strand sequence of "CUG_2116" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUGAAGCAGCUGCUG (SEQ ID NO: 334; the RNA guide strand sequence of "CUG_2143" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUGAAGCUACUGCUG (SEQ ID NO: 335; the RNA guide strand sequence of "CUG_2170" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUGAAACUGCUGCUG (SEQ ID NO: 339; the RNA guide strand sequence of "CUG_2089" in Table 6). In some cases, the first strand comprises the following nucleotide sequence: CUGCUGCUGAAGCUGAUGCUG (SEQ ID NO: 340; the RNA guide strand sequence of "CUG_2197" in Table 6).
[0132] First mismatch at position 10; four mismatches
[0133] In some cases, the double-stranded RNA of the present disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand, wherein the first mismatch is SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), SEQ ID NO: 745 (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866)), or SEQ ID NO: 867 (UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867)), the first strand comprises a second mismatch, a third mismatch, and a fourth mismatch relative to the target CAG repeat region, the second, third, and fourth mismatches being 1 to 6 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the first substitution creates a first mismatch and is at position 10 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 6 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the first substitution creates a first mismatch and is at position 10 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 6 bases 3' to the first mismatch.In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) containing at least a first, second, third, and fourth substitution, wherein the first substitution creates a first mismatch and is located at position 10 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 6 bases 3' of the first mismatch. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA is 18 to 25 nucleotides in length. In some cases, the double-stranded RNA is 20 nucleotides in length. In some cases, the double-stranded RNA is 21 nucleotides in length. In some cases, the double-stranded RNA is 22 nucleotides in length. In some cases, the double-stranded RNA is 23 nucleotides in length. In some cases, the double-stranded RNA is 24 nucleotides in length. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from: a) G with A, U, or C; b) U with A, G, or C; and c) C with A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 345, 346, and 368-375.
[0134] First mismatch at position 11; two mismatches
[0135] In some cases, the double-stranded RNA of the present disclosure comprises a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA, and b) a second strand that hybridizes to the first strand, wherein the first strand i) has a first mismatch to the target CAG repeat region, wherein the first mismatch is SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745), (GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866) or (UGCUGCUGCUGCUGCUGCU GC (SEQ ID NO: 867); and ii) a second mismatch to the target CAG repeat region, the second mismatch being 1 to 5 bases 3' to the first mismatch. Optionally, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), where the first substitution creates a first mismatch and is at position 11 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743). and the second substitution creates a second mismatch and is 1 to 5 bases 3' to the first mismatch. Optionally, the first strand is a variant comprising at least a first and a second substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 11 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), and the second substitution creates a second mismatch and is 1 to 5 bases 3' to the first mismatch. The first strand is a variant containing at least a first and a second substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 743), where the first substitution creates a first mismatch and is at position 11 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 743), and the second substitution creates a second mismatch and is 1 to 5 bases 3' to the first mismatch. In some cases, the first strand contains no more than two mismatches with the target CAG repeat region.In some cases, the first strand contains three or fewer mismatches with the target CAG repeat region. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains one to ten mismatches (e.g., one to four, three to five, five to seven, or five to ten mismatches) with the first strand. In some cases, the second strand contains one to four mismatches with the first strand. In some cases, the second strand contains one, two, three, or four or fewer mismatches with the first strand. In some cases, the second strand contains one or fewer mismatches with the first strand. In some cases, the second strand contains two or fewer mismatches with the first strand. In some cases, the second strand contains three or fewer mismatches with the first strand. In some cases, the second strand contains four or fewer mismatches with the first strand. In some cases, the second strand contains five or fewer mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from the following: a) substitution of G with A, U, or C; b) substitution of U with A, G, or C; and c) substitution of C with A, U, or G. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 311-315. In some cases, the first strand comprises a nucleotide sequence selected from SEQ ID NOs: 793 to 803 (as shown in Table 8; Figure 24). In some cases, the first strand comprises the nucleotide sequence CUGCUGCUGCAACUGCUGCUG (SEQ ID NO: 311; RNA guide strand sequence for "CUG_217" in Table 6). In some cases, the first strand comprises the nucleotide sequence CUGCUGCUGCAGAUGCUGCUG (SEQ ID NO: 312; RNA guide strand sequence for "CUG_226" in Table 6).In some cases, the first strand includes the nucleotide sequence CUGCUGCUGCAGCAGCUGCUG (SEQ ID NO: 313; the RNA guide strand sequence of "CUG_235" in Table 6). In some cases, the first strand includes the following nucleotide sequence: CUGCUGCUGCAGCUACUGCUG (SEQ ID NO: 314; the RNA guide strand sequence of "CUG_244" in Table 6). In some cases, the first strand includes the following nucleotide sequence: CUGCUGCUGCAGCUGAUGCUG (SEQ ID NO: 315; the RNA guide strand sequence of "CUG_253" in Table 6).
[0136] First mismatch at position 11; three mismatches
[0137] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA; and b) a second strand that hybridizes to the first strand, wherein a first mismatch is at position 11 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), or (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), and the first strand comprises a second mismatch and a third mismatch relative to the target CAG repeat region, and the second and third mismatches are 1 to 5 bases 3' to the first mismatch, wherein the first strand comprises: In some cases, the first strand is a variant that includes at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), where the first substitution creates a first mismatch and is at position 11 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, and the second and third substitutions are 1 to 5 bases 3' to the first mismatch. In some cases, the first strand is a variant that includes at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), where the first substitution creates a first mismatch and is at position 11 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, and the second and third substitutions are 1 to 5 bases 3' to the first mismatch.In some cases, the first strand is a variant comprising at least a first substitution, a second substitution, and a third substitution of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the first substitution creates a first mismatch and is located at position 11 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), the second substitution creates a second mismatch, and the third substitution creates a third mismatch, and the second and third substitutions are 1 to 5 bases 3' to the first mismatch. In some cases, the first strand has no more than three mismatches with the target CAG repeat region. In some cases, the first strand has no more than four mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains 1 to 4 mismatches with the first strand. In some cases, the second strand contains 1, 2, 3, or no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 1 mismatch with the first strand. In some cases, the second strand contains no more than 2 mismatches with the first strand. In some cases, the second strand contains no more than 3 mismatches with the first strand. In some cases, the second strand contains no more than 4 mismatches with the first strand. In some cases, the second strand contains no more than 5 mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA is 21 nucleotides long. In some cases, the double-stranded RNA is 22 nucleotides long. In some cases, the double-stranded RNA is 23 nucleotides long. In some cases, the double-stranded RNA is 24 nucleotides long. In some cases, the double-stranded RNA is 25 nucleotides in length. In some cases, each mismatch is caused by a substitution independently selected from: a) G for A, U, or C; b) U for A, G, or C; and c) C for A, U, or G.
[0138] First mismatch at position 11; four mismatches
[0139] In some cases, the double-stranded RNA of the disclosure comprises: a) a first strand that hybridizes to a target CAG repeat region of a CAG repeat-containing RNA; and b) a second strand that hybridizes to the first strand, wherein the first mismatch is at position 11 based on the numbering of SEQ ID NO: 743 (CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743)), SEQ ID NO: 744 (GCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 744)), or (UGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 745)), wherein the first strand comprises a second mismatch, a third mismatch, and a fourth mismatch relative to the target CAG repeat region, wherein the second, third, and fourth mismatches are 1 to 5 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence CUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the first substitution creates a first mismatch and is at position 11 based on the numbering of CUGCUGCUGCUGCUGCUGCUGCUG (SEQ ID NO: 743), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 5 bases 3' to the first mismatch. In some cases, the first strand is a variant comprising at least a first, second, third, and fourth substitution of the nucleotide sequence GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the first substitution creates a first mismatch and is at position 11 based on the numbering of GCUGCUGCUGCUGCUGCUGCU (SEQ ID NO: 866), wherein the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, wherein the second, third, and fourth substitutions are 1 to 5 bases 3' to the first mismatch.In some cases, the first strand is a variant of the nucleotide sequence UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867) containing at least a first, second, third, and fourth substitution, where the first substitution creates a first mismatch and is located at position 11 based on the numbering of UGCUGCUGCUGCUGCUGCUGC (SEQ ID NO: 867), where the second substitution creates a second mismatch, the third substitution creates a third mismatch, and the fourth substitution creates a fourth mismatch, where the second, third, and fourth substitutions are 1 to 5 bases 3' to the first mismatch. In some cases, the first strand contains four or fewer mismatches with the target CAG repeat region. In some cases, the second strand is 100% complementary to the first strand. In some cases, the second strand contains 1 to 10 mismatches (e.g., 1 to 4, 3 to 5, 5 to 7, or 5 to 10 mismatches) relative to the first strand. In some cases, the second strand contains one to four mismatches with the first strand. In some cases, the second strand contains one, two, three, or no more than four mismatches with the first strand. In some cases, the second strand contains no more than one mismatch with the first strand. In some cases, the second strand contains no more than two mismatches with the first strand. In some cases, the second strand contains no more than three mismatches with the first strand. In some cases, the second strand contains no more than four mismatches with the first strand. In some cases, the second strand contains no more than five mismatches with the first strand. In some cases, the double-stranded RNA has a length of 18 to 25 nucleotides. In some cases, the double-stranded RNA has a length of 20 nucleotides. In some cases, the double-stranded RNA has a length of 21 nucleotides. In some cases, the double-stranded RNA has a length of 22 nucleotides. In some cases, the double-stranded RNA has a length of 23 nucleotides. In some cases, the double-stranded RNA has a length of 24 nucleotides. In some cases, the double-stranded RNA has a length of 25 nucleotides. In some cases, each mismatch is generated by a substitution independently selected from the following: a) G to A, U, or C; b) U to A, G, or C; and c) C to A, U, or G.
[0140] A. Target nucleic acid
[0141] The double-stranded RNA of the present disclosure can target any gene or nucleic acid construct that contains target repeat region.In some embodiments, the gene (DNA or mRNA) that codes for human or primate protein is targeted.In some embodiments, the non-coding gene is targeted.In some embodiments, the coding region of gene is targeted.In some embodiments, the non-coding region of gene is targeted.
[0142] In some embodiments, the double-stranded RNA of the present disclosure targets CAG repeat-containing (polyglutamine) genes.In some embodiments, the CAG repeat-containing genes are selected from HTT, ataxin 1, ataxin 2, ataxin 3, CACNA1A, ataxin 7, PPP2R2B, TBP, androgen receptor, Atrophin, MLLT3, BMP2K, THAP11, ZFHX3, POU3F2, MAML2, SMARCA2, MAML3, ORC4, RUNX2, MED12, EP400, MAGI1, UMAD1, DM1-AS, AC007161.3, IRF2BPL and MAB21L1.CAG repeat expansion is associated with many dominant genetic disorders called polyglutamine (polyQ) diseases. Although CAG repeat-containing protein is ubiquitously expressed throughout the body, the pathology of polyglutamine disease mainly appears in, but is not limited to, nervous tissue.Therefore, as used herein, the term polyglutamine disease refers to any disease or disorder associated with the expansion of CAG repeat, including but not limited to neurodegenerative disease.
[0143] Huntingtin (HTT), also known as interesting transcript 15 (IT15), is the gene encoding the huntingtin protein. While the exact function of huntingtin is unknown, it is involved in axonal transport. An example of a huntingtin transcript sequence is shown in NCBI Reference Sequence NM_002111.8 (SEQ ID NO: 820). Typically, the polyglutamine tract of huntingtin contains 10 to 35 CAG repeats. Expansion of the polyglutamine tract to 36 to 120 or more CAG repeats results in Huntington's disease. Early signs and symptoms include irritability, depression, small involuntary movements, loss of coordination, and difficulty learning new information and making decisions. Many patients with Huntington's disease develop involuntary jerking movements known as chorea. As the disease progresses, these movements become more pronounced. Patients may experience problems walking, speaking, and swallowing. People with Huntington's disease also experience personality changes and a decline in thinking and reasoning skills.
[0144] Ataxin 1 (ATXN1), also known as spinocerebellar ataxia type 1 (SCA1), refers to a gene encoding a polyglutamine-containing protein that is primarily expressed in the nucleus, binds to chromatin, and functions as a transcriptional repressor. An example of the sequence of the ATXN1 transcript is provided in the NCBI reference sequence NM_001128164.2 (SEQ ID NO: 821). Mutant forms of ataxin 1, containing expansions of polyglutamine tracts, typically to approximately 40-83 repeats, cause the movement disorder spinocerebellar ataxia type 1 (SCA1) through a toxic gain-of-function mechanism in the cerebellum. Cerebellar dysfunction is progressive and persistent. Patients with this disease initially experience problems with coordination and balance (ataxia). Other signs and symptoms of SCA1 include difficulty speaking and swallowing, muscle stiffness (spasticity), and weakness of the muscles controlling eye movement (ophthalmoplegia). Weakness of the eye muscles causes rapid, involuntary eye movements (nystagmus). People with SCA1 may have difficulty processing information, learning, and remembering (cognitive impairment). Over time, people with SCA1 may develop numbness, tingling, or pain in the arms and legs (sensory neuropathy); uncontrollable muscle tension (dystonia); muscle wasting (atrophy); and muscle spasms (fasciculations). Rarely, stiffness, tremors, and involuntary jerking movements (chorea) have been reported in people with long-standing disease.
[0145] Ataxin 2 (ATXN2), also known as spinocerebellar ataxia type 2 (SCA2), refers to a gene encoding a polyglutamine-containing RNA-binding protein that targets cis-regulatory elements in the 3'UTR to stabilize a subset of mRNAs and increase protein expression. An example of the sequence of the ATXN2 transcript is provided by NCBI Reference Sequence: NM_001372574.1 (SEQ ID NO: 822). Polyglutamine repeat expansions (e.g., typically about 33 or more repeats) in ATXN2 can cause the signs and symptoms of spinocerebellar ataxia type 2 (SCA2). Patients with SCA2 initially experience problems with coordination and balance (ataxia). Other early signs and symptoms of SCA2 include further motor impairments, difficulty speaking and swallowing, and weakness of the muscles controlling eye movement (ophthalmoplegia). Weakness of the eye muscles causes involuntary back-and-forth eye movements (nystagmus) and a decreased ability to perform rapid eye movements (saccadic hypoostosis). Over time, patients with SCA2 may develop loss of sensation and muscle weakness in the limbs (peripheral neuropathy), muscle atrophy (atrophy), uncontrollable muscle tension (dystonia), and involuntary jerking movements (chorea). Some patients with SCA2 develop a group of movement abnormalities known as parkinsonism, which includes abnormally slow movements (bradykinesia), involuntary shaking (tremor), and muscle stiffness (rigidity). Patients with SCA2 may have problems with short-term memory, planning, and problem-solving, or experience a general decline in intellectual function (dementia). The intermediate polyglutamine expansion (27–33 CAG repeats) in ATXN2 also increases the risk of amyotrophic lateral sclerosis (ALS). ALS is a neurodegenerative neuromuscular disease that causes the gradual loss of motor neurons that control voluntary muscles. Early symptoms of ALS include muscle stiffness, muscle spasms, and gradually progressive muscle weakness and atrophy. Limb-onset ALS begins with muscle weakness in the arms and legs, while bulbar-onset ALS begins with difficulty speaking and swallowing. Half of ALS patients experience at least mild impairment in thinking and behavior, and approximately 15% develop frontotemporal dementia. Most people experience pain. Motor neuron loss continues until patients lose the ability to eat, speak, move, and eventually breathe.ALS eventually leads to paralysis and premature death, usually from respiratory failure.
[0146] Ataxin 3 (ATXN3), also known as spinocerebellar ataxia type 3 (SCA3), refers to a gene encoding a polyglutamine-containing deubiquitinase. An example of the ATXN3 transcript sequence is provided by NCBI Reference Sequence NM_004993.6 (SEQ ID NO: 823). Expansion of polyglutamine repeats from the normal 13–36 repeats to 50 or more causes Machado-Joseph disease (MJD), which is also known as Machado-Joseph Azorean disease, Machado disease, Joseph disease, or spinocerebellar ataxia type 3 (SCA3). Individuals with this condition initially experience problems with coordination and balance (ataxia). Other early signs and symptoms of SCA3 include speech impairment, uncontrollable muscle tension (dystonia), muscle stiffness (spasticity), rigidity, tremors, bulging eyes, and diplopia. People with this condition may experience sleep disorders such as restless legs syndrome and REM sleep behavior disorder. Over time, people with SCA3 may develop loss of sensation and muscle weakness in the hands and feet (peripheral neuropathy), muscle spasms, muscle fasciculations, and difficulty swallowing. People with SCA3 may have problems with memory, planning, and problem-solving.
[0147] Calcium voltage-gated channel subunit α1A (CACNA1A), also known as spinocerebellar ataxia type 6 (SCA6), encodes the α1A pore-forming subunit of the neuronal calcium channel P / Q. An example of the CACNA1A transcript sequence is provided by NCBI reference sequence NM_000068.4 (SEQ ID NO: 824). Expansion of the CACNA1A gene's polyglutamine tract to a typical 19-33 repeat causes spinocerebellar ataxia type 6 (SCA6). Individuals with this condition initially experience problems with coordination and balance (ataxia). Other early signs and symptoms of SCA6 include difficulty speaking, involuntary eye movements (nystagmus), and diplopia. Over time, individuals with SCA6 may develop loss of coordination in their arms, tremors, and uncontrollable muscle tension (dystonia).
[0148] Ataxin 7 (ATXN7), also known as spinocerebellar ataxia type 7 (SCA7), encodes a polyglutamine-containing protein that is an essential subunit of the GCN5 (general control of amino acid synthesis-5; KAT2A)-containing SAGA family of histone acetyltransferase (HAT) complexes. An example of the sequence of the ATXN7 transcript is provided by NCBI reference sequence NM_001377405.1 (SEQ ID NO: 825). Polyglutamine expansion in ATXN7 causes spinocerebellar ataxia type 7 (SCA7), which is characterized by progressive cerebellar ataxia, retinal degeneration or blindness due to cone-rod dystrophy, and mild sensory or reflex changes. Subsequent symptoms include loss of motor control, slurred speech (dysarthria), and difficulty swallowing (dysphagia).
[0149] Protein phosphatase 2 regulatory subunit Bβ (PPP2R2B), also known as spinocerebellar ataxia type 12 (SCA12), encodes the B regulatory subunit of protein phosphatase 2, a serine / threonine phosphatase. An example of the PPP2R2B transcript sequence is provided by NCBI reference sequence NM_181674.3 (SEQ ID NO: 826). Expansion of polyglutamine repeats from the typical normal range of approximately 7-28 to approximately 55-78 causes spinocerebellar ataxia type 12. Age of onset ranges from 8 to 55 years, but onset is most common in the fourth decade. Symptoms typically begin with tremor and progress to cerebellar ataxia. Signs of dementia have also been reported in association with SCA.
[0150] TATA box-binding protein (TBP), also known as spinocerebellar ataxia type 17 (SCA17), encodes the TATA-binding protein, a component of transcription factor IID (TFIID). An example of the sequence of the TBP transcript is provided by NCBI reference sequence NM_003194.5 (SEQ ID NO: 827). TBP typically contains 25-42 polyglutamine repeats, and expansions to 45-66 repeats are associated with spinocerebellar ataxia type 17 (SCA). Patients with this condition typically experience symptoms such as ataxia, dementia, involuntary movements such as chorea and dystonia, and pyramidal signs such as rigidity, spasticity, weakness, slowed rapid alternating movements, and hyperreflexia.
[0151] The androgen receptor (AR) encodes a steroid hormone-activated transcription factor. An example of the sequence of the AR transcript is provided by the NCBI reference sequence NM_000044.6 (SEQ ID NO: 828). Expansion of polyglutamine repeats from the typical 9-34 repeats to 38-62 repeats causes spinal-bulbar muscular atrophy (SBMA), also known as Kennedy disease. SBMA is characterized by muscle weakness and atrophy that worsens over time, leading to spasticity and difficulty walking, swallowing, and speaking. SBMA can also cause gynecomastia and infertility.
[0152] Atrophin 1 (ATN1) encodes a protein hypothesized to be a transcriptional corepressor that recruits nuclear receptor subfamily 2 group E member 1 (NR2E1) to repress transcription. An example of the ATN1 transcript sequence is provided by NCBI reference sequence NM_001007026.2 (SEQ ID NO: 829). Dentatorubral-pallidoluysian atrophy (DRPLA) is a rare neurodegenerative disorder associated with an expansion of the ATN1 polyglutamine repeat sequence from the typical 7-35 copies to 49-93 copies. When DRPLA manifests before approximately 20 years of age, it is typically accompanied by myoclonus, ataxia, seizures, behavioral changes, and intellectual disability. When it manifests after approximately 20 years of age, it is accompanied by ataxia, choreoathetosis, delusions, and dementia.
[0153] Myeloid / Lymphoid Or Mixed-Lineage Leukemia Translocated To Chromosome 3 (MLLT3), also known as AF-9, encodes a component of the super-elongation complex (SEC) required to increase the catalytic rate of RNA polymerase II transcription. An example of the sequence of the MLLT3 transcript is provided in NCBI Reference Sequence NM_004529.4 (SEQ ID NO: 830). MLLT3 contains unstable polyglutamine repeats, and genetic abnormalities involving MLLT3 are associated with leukemia, neuromotor developmental delay, cerebellar ataxia, and epilepsy.
[0154] Bone Morphogenic Protein 2 Inducible Kinase (BMP2K) encodes a protein involved in skeletal development and pattern formation. An example of the sequence of BMP2K transcript is listed in NCBI reference sequence NM_198892.2 (SEQ ID NO: 831). BMP2K contains polyglutamine repeats and is associated with myopia and cancer, particularly cancer-related gene dysregulation.
[0155] THAP domain-containing 11 (THAP) encodes a transcriptional repressor involved in embryonic development. An example of the sequence of a THAP11 transcript is provided by NCBI reference sequence NM_020457.3 (SEQ ID NO: 832). THAP11 typically contains approximately 29 copies of polyglutamine repeats, but can range from 20 to over 40 copies. For example, an increase in the number of polyglutamine repeats to 38 copies is associated with neurodegenerative diseases. Polyglutamine expansion in THAP11 is also associated with intracellular aggregation of THAP11, cytotoxicity, growth inhibition, G0 / G1 arrest, and inhibition of transcriptional activity.
[0156] Zinc Finger Homeobox 3 (ZFHX3) encodes a transcription factor that regulates myogenesis and neuronal differentiation. It also functions as a tumor suppressor in some cancers and is associated with atrial fibrillation. An example of the sequence of the ZFHX3 transcript is provided by NCBI reference sequence NM_006885.4 (SEQ ID NO: 833). ZFHX3 contains a polyglutamine repeat sequence. Individuals with expanded polyglutamine repeat sequences, for example, 19 copies, are associated with coronary heart disease, hypertension, diabetes, or dyslipidemia compared to individuals with fewer repeat sequences, for example, 17 copies.
[0157] POU class 3 homeobox 2 (POU3F2) encodes a transcription factor involved in neuronal differentiation. An example of the sequence of the POU3F2 transcript is provided by NCBI reference sequence NM_005604.4 (SEQ ID NO: 834). POU3F2 contains polyglutamine tracts and is associated with bipolar disorder, obesity, developmental delay and intellectual disability.
[0158] Mastermind-like transcriptional coactivator 2 (MAML2) encodes a transcriptional coactivator for NOTCH proteins and promotes the turnover of β-catenin. An example of the sequence of the MAML2 transcript is provided by NCBI reference sequence NM_032427.4 (SEQ ID NO: 835). MAML2 contains a polyglutamine tract with observed variability and is associated with cancers such as mucoepidermoid carcinoma, hidradenoma, B-cell lymphoma, and chronic lymphocytic leukemia.
[0159] Mastermind-like transcriptional coactivator 3 (MAML3) encodes a transcriptional coactivator for NOTCH proteins. An example of the sequence of the MAML3 transcript is provided by NCBI reference sequence NM_018717.5 (SEQ ID NO: 836). MAML3 contains polyglutamine tracts and is associated with cancers such as Schneider's carcinoma and ossifying fibromyxoid tumor.
[0160] SWI / SNF Related, Matrix Associated, Actin Dependent Regulator of Chromatin, Subfamily A, Member 2 (SMARCA2) encodes a component of the SWI / SNF complex, which is involved in transcriptional regulation by chromatin remodeling. SMARCA2 is also involved in neurogenesis. An example of the sequence of the SMARCA2 transcript is provided by NCBI Reference Sequence NM_003070.5 (SEQ ID NO: 837). SMARCA2 contains a polymorphic polyglutamine tract and is associated with disorders such as Nicolaides-Baraitser syndrome and blepharoptosis-mental developmental disability syndrome. The SMARCA2 gene is also located in a chromosomal region associated with schizophrenia and bipolar disorder.
[0161] Origin Recognition Complex Subunit 4 (ORC4) encodes a component of the six-subunit origin recognition complex (ORC) required for the initiation of DNA replication. An example of the ORC4 transcript sequence is provided by NCBI reference sequence NM_001190879.3 (SEQ ID NO: 838). ORC4 contains a region of polymorphic trinucleotide CAG repeats located upstream of the coding sequence and is associated with Meier-Gorlin syndrome 1 and Meier-Gorlin syndrome 2.
[0162] RUNX family transcription factor 2 (RUNX2) encodes a nuclear protein involved in osteoblast differentiation and skeletal morphogenesis. An example of the sequence of the RUNX2 transcript is provided by the NCBI reference sequence NM_001024630.4 (SEQ ID NO: 839). RUNX2 contains polyglutamine and polyalanine tracts. Expansion of the polyglutamine tract from the typical 23 residues to, for example, 27-30 residues, causes cleidocranial dysplasia, decreased bone mineral density, and decreased RUNX2 transactivation activity. Cleidocranial dysplasia (CCD) is a disease affecting the skull, bones, and teeth. Signs and symptoms include absent or underdeveloped clavicles, delayed closure of the skull fontanelle, dental abnormalities, short stature, decreased bone mineral density, hearing loss, and other bone abnormalities.
[0163] Mediator complex subunit 12 (MED12) encodes a component of the preinitiation complex involved in the control of transcription initiation.An example of the sequence of MED12 transcript is provided by NCBI reference sequence NM_005120.3 (SEQ ID NO: 840).MED12 has a polyglutamine tract and is associated with tumorigenesis such as Opitz-Kaveggia syndrome, Lujan-Fryns syndrome, Ohdo syndrome, X-linked syndrome, and uterine leiomyoma.
[0164] E1A-binding protein P400 (EP400) encodes a component of the NuA4 histone acetyltransferase complex, which is involved in transcriptional activation. An example of the sequence of the EP400 transcript is provided by NCBI reference sequence NM_015409.5 (SEQ ID NO: 841). EP400 typically contains approximately 32 CAG repeats. EP400 is involved in ossifying fibromyxoid tumors and epilepsy, familial temporal lobe 1.
[0165] Membrane-associated guanylate kinase, WW and PDZ domain-containing 1 (MAGI1) encodes a protein involved in the assembly of multiprotein complexes at cell-cell contact regions. An example of the sequence of the MAGI1 transcript is provided by NCBI reference sequence NM_015520.2 (SEQ ID NO: 842). MAGI1 contains a polymorphic polyglutamine tract and is associated with conditions such as cervical large cell neuroendocrine carcinoma and microscopic colitis.
[0166] UBAP1-MVB12-associated (UMA) domain-containing 1 (UMAD1) is a gene that encodes a protein.An example of the sequence of the transcript of UMAD1 is provided by NCBI reference sequence NM_001302348.2 (SEQ ID NO: 843).UMAD1 contains a polymorphic trinucleotide CAG repeat region upstream of the start codon and is associated with retinitis pigmentosa.
[0167] DM1 Locus Antisense RNA (DM1-AS) is an RNA gene. An example of the RNA sequence of DM1-AS is provided by NCBI reference sequence NR_147193.1 (SEQ ID NO: 844). DM1-AS contains a polymorphic trinucleotide CAG repeat region within an intron and is associated with myotonic dystrophy 1 and branchio-oto-renal syndrome 2.
[0168] AC007161.3, also known as ENSG 00000283549, is an RNA gene and contains CAG repeats.
[0169] Interferon regulatory factor 2-binding protein-like (IRF2BPL) encodes a transcription factor involved in the development of the central nervous system, the maintenance of neurons, and the regulation of female reproductive function. An example of the sequence of the IRF2BPL transcript is provided by NCBI reference sequence NM_024496.4 (SEQ ID NO: 845). IRF2BPL contains polyglutamine tracts and is associated with neurodevelopmental disorders with regression, abnormal movements, speech loss, and seizures, and neurological problems such as Irf2bpl-associated degenerative neurodevelopmental disorder-dystonia-seizure syndrome.
[0170] Mab-21Like 1 (MAB21L1) encodes a protein related to eye and cerebellar development. An example of the sequence of the MAB21L1 transcript is provided by NCBI reference sequence NM_005584.5 (SEQ ID NO: 846). MAB21L1 is associated with cerebellar syndrome, ocular syndrome, craniofacial syndrome, and genital syndrome and hydrophthalmia. MAB21L1 contains a polymorphic trinucleotide CAG repeat in the 5' untranslated portion of the gene, which is associated with psychiatric conditions such as bipolar disorder.
[0171] In some embodiments, the pathogenic or pathological allele of the CAG repeat-containing gene or RNA encoded by the CAG repeat-containing gene comprises at least about 30 consecutive CAG repeats.
[0172] C. Expression Cassettes and Vectors
[0173] In some embodiments, the double-stranded RNA of the present disclosure is encoded by a nucleic acid molecule, for example, a DNA sequence. The double-stranded RNA sequence provided herein can be converted into a DNA format by replacing each uracil base "U" with a thymine "T" base.
[0174] In certain embodiments, the nucleic acid molecule (e.g., DNA) encoding the double-stranded RNA is contained within an expression cassette.
[0175] In some embodiments, the expression cassette further comprises one or more expression control sequences (regulatory sequences) operably linked to the transgene. "Operably linked" sequences include expression control sequences that are adjacent to the transgene or that act in trans or remote from the transgene to regulate its expression. Examples of expression control sequences include transcription initiation sequences, termination sequences, promoter sequences, enhancer sequences, repressor sequences, splice site sequences, polyadenylation (polyA) signal sequences, or any combination thereof.
[0176] In some embodiments, the promoter is an endogenous promoter, a synthetic promoter, a hybrid promoter, a constitutive promoter, an inducible promoter, a tissue-specific promoter (e.g., CNS-specific), or a cell-specific promoter (neuron, glial cell, astrocyte). Examples of constitutive promoters include the Rous sarcoma virus (RSV) LTR promoter (optionally with an RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with a CMV enhancer), the SV40 promoter, and the dihydrofolate reductase promoter. Examples of inducible promoters include the zinc-inducible sheep metallothionine (MT) promoter, the dexamethasone (Dex)-inducible mouse mammary tumor virus (MMTV) promoter, the T7 polymerase promoter system, the ecdysone insect promoter, the tetracycline-suppressible system, the tetracycline-inducible system, the RU486-inducible system, and the rapamycin-inducible system. Further examples of promoters that can be used include, for example, chicken beta actin promoter (CBA promoter), CAG promoter, H1 promoter, CD68 promoter, JeT promoter, synapsin promoter, RNA polymerase II promoter, or RNA polymerase III promoter (e.g., U6, H1, etc.).
[0177] In some embodiments, the promoter is an RNA pol II promoter. Examples of pol II promoters include PGK, CBA, U1, CMV, EIF1α, EF1α, CAG, or synaptophysin promoters. In some embodiments, the promoter is a tissue-specific RNA pol II promoter. In some embodiments, the tissue-specific RNA pol II promoter is derived from a gene that exhibits neuron-specific expression. In some embodiments, the expression cassette comprises a pol II promoter and a poly(A) tail, e.g., a DNA sequence encoding a double-stranded RNA flanked at the 5' end by a pol II promoter and at the 3' end by a poly(A) tail.
[0178] In some embodiments, the promoter is a neuron-specific promoter. Examples of neuron-specific promoters include those from neuron-specific enolase (NSE), human synapsin 1, human synapsin 2 promoter, caMK kinase, and tubulin.
[0179] In some embodiments, the promoter is an RNA pol III promoter. Examples of pol III promoters include U6, H1, 7SK, Y, RPR, MRP, and selenocysteine tRNA. In some embodiments, the expression cassette comprises a pol III promoter and a poly(T) tail, for example, with a DNA sequence encoding a double-stranded RNA flanked at the 5' end by a pol III promoter and at the 3' end by a poly(T) tail.
[0180] In certain embodiments, the promoter is an RNA pol I promoter. In some embodiments, the expression cassette comprises a pol I promoter and a 3' box, e.g., a DNA sequence encoding a double-stranded RNA flanked at its 5' end by a pol I promoter and at its 3' end by a 3' box.
[0181] Expression cassettes for double-stranded RNA are known in the art, see, for example, ter Brake et al. Mol. Ther. (2008) 16:557; Maczuga et al., BMC Biotechnol. (2012) 12:42; and Bofill-De Ros and Gu (2016) 103:157.
[0182] In some embodiments, the DNA sequence encoding the double-stranded RNA of the present disclosure is located in the untranslated region of the expression cassette. In some embodiments, the sequence encoding the inhibitory nucleic acid of the present disclosure is located in an intron, 5' untranslated region (5'UTR), or 3' untranslated region (3'UTR) of the expression cassette. In some embodiments, the sequence encoding the inhibitory nucleic acid of the present disclosure is located in an intron downstream of the promoter and upstream of the expressed gene.
[0183] In some embodiments, the DNA sequence encoding the double-stranded RNA of the present disclosure is flanked by two AAV inverted terminal repeats (ITRs) (e.g., a 5' ITR and a 3' ITR) within the expression cassette. In some embodiments, each AAV ITR is a full-length ITR (e.g., approximately 145 bp in length and includes a functional Rep binding site (RBS) and terminal resolution site (trs)). In some embodiments, one of the ITRs is truncated (e.g., shortened or not full-length). In some embodiments, the truncated ITR lacks a functional terminal resolution site (trs) and is used to produce a self-complementary AAV vector (scAAV vector).
[0184] In some embodiments, the double-stranded RNA described herein can be encoded by a vector such as a plasmid, a non-viral vector, or a viral vector.The use of a vector to express the double-stranded RNA of the present disclosure can allow the continuous or controlled expression of double-stranded RNA in a subject, rather than administering double-stranded RNA to the subject multiple times.The present disclosure provides a vector comprising an isolated nucleic acid comprising an expression cassette that encodes the double-stranded RNA described herein.
[0185] Viral vectors include, but are not limited to, herpesvirus (HSV) vectors, retrovirus vectors, adenovirus vectors, adeno-associated virus (AAV) vectors, lentivirus vectors, baculovirus vectors, and the like.
[0186] In some embodiments, the vector encoding the double-stranded RNA of the present disclosure is a retroviral vector. In some embodiments, the retroviral vector is a murine stem cell virus, murine leukemia virus (e.g., Moloney murine leukemia virus vector), feline leukemia virus, feline sarcoma virus, or avian reticuloendotheliosis virus vector. In some embodiments, the vector encoding the double-stranded RNA of the present disclosure is a lentivirus or lentivirus-based vector. In some embodiments, the lentiviral vector is a human immunodeficiency virus, including HIV type 1 and HIV type 2, equine infectious anemia virus, feline immunodeficiency virus (FIV), bovine immunodeficiency virus (BIV), and simian immunodeficiency virus (SIV), equine infectious anemia virus, or Maedi-Visna virus vector. Methods for expressing shRNA using lentiviral-engineered cells are known in the art. For example, Stegmeier et al. Proc. Natl. Acad. Sci. USA (2005) 102:13212-13217; Klinghoffer et al. RNA (2010) 16:879-884. Production of replication-incompetent recombinant lentivirus can be achieved by co-transfection of an expression vector and a packaging plasmid using, for example, a commercially available packaging cell line such as TLA-HEK293™ and a packaging plasmid (Thermo Scientific / Open Biosystems (Huntsville, AL)).
[0187] In some embodiments, the vector encoding the double-stranded RNA of the present disclosure is an adeno-associated virus (AAV) vector, such as a recombinant rAAV vector produced by recombinant methods. AAV is a single-stranded, non-enveloped DNA virus with a genome encoding a replication protein (rep) and a capsid (Cap), flanked by two ITRs that function as the origin of replication of the viral genome. AAV also contains a packaging sequence that enables packaging of the viral genome into an AAV capsid. In some embodiments, the AAV vector contains an expression cassette encoding the double-stranded RNA of the present disclosure flanked by two cis-acting AAV ITRs (5'ITR and 3'ITR). The functional ITR sequence is used for the rescue, replication, and packaging of AAV viral particles. Therefore, an AAV vector is defined herein as at least including the sequences required in cis for viral replication and packaging (e.g., one or two functional ITRs and a packaging sequence). In some embodiments, each AAV ITR is a full-length ITR (e.g., approximately 145 bp in length and including a functional Rep binding site (RBS) and terminal resolution site (trs)). In some embodiments, where the ITR provides functional rescue, replication, and packaging, one or both of the ITRs are modified, for example, by insertion, deletion, or substitution. In some embodiments, the modified ITR lacks a functional terminal resolution site (trs) and is used to produce a self-complementary AAV vector (scAAV vector). In some embodiments, the ITR is selected from any one of serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV.Rh10, AAV11, and variants thereof. In some cases, the ITR is derived from AAV2.
[0188] Other expression control sequences may be present in the rAAV vector operably linked to the DNA sequence encoding the double-stranded RNA, including one or more of a transcription initiation sequence, termination sequence, promoter sequence, enhancer sequence, repressor sequence, splice site sequence, polyadenylation (polyA) signal sequence, or any combination thereof.
[0189] The rAAV vector may have one or more AAV wild-type genes deleted in whole or in part. In some embodiments, the rAAV vector is replication-deficient. In some embodiments, the rAAV vector lacks functional Rep protein and / or capsid protein.
[0190] Methods for packaging recombinant AAV vectors into AAV capsids using host cell culture are known in the art. In some embodiments, one or more components required for packaging rAAV vectors (e.g., Rep sequence, cap sequence, and / or accessory functions) can be provided by stable host cells engineered to contain one or more required components (e.g., by a vector). Expression of the components required for AAV packaging can be under the control of an inducible or constitutive promoter in the host packaging cell. AAV helper vectors are commonly used to provide transient expression of AAV rep and / or cap genes that function in trans to complement missing AAV functions required for AAV replication. In some embodiments, AAV helper vectors lack AAV ITRs and cannot replicate or package themselves. AAV helper vectors can be in the form of a plasmid, phage, transposon, cosmid, virus, or virion.
[0191] The recombinant AAV vector of the present disclosure can be encapsulated by an AAV capsid to form rAAV particles. "rAAV particle" or "rAAV virion" refers to an infectious, replication-defective virus that contains an AAV protein shell and encapsulates a transgene of interest flanked by AAV ITRs. rAAV particles are produced in suitable host cells that incorporate rAAV vector-specific sequences, AAV helper functions, and accessory functions, allowing the host cell to encode the AAV polypeptides required for packaging the rAAV vector into infectious rAAV particles for subsequent gene delivery to target cells.
[0192] In some embodiments, rAAV particles can be produced using a triple transfection method (see, e.g., U.S. Patent No. 6,001,650, the entire contents of which are incorporated herein by reference). In this approach, rAAV particles are produced by transfecting a host cell with an rAAV vector (containing a transgene) to be packaged into the rAAV particle, an AAV helper vector, and an accessory function vector. In some embodiments, the AAV helper function vector supports efficient AAV vector production without producing detectable wild-type AAV virions (e.g., AAV virions containing functional rep and cap genes). The accessory function vector encodes nucleotide sequences for non-AAV-derived viral and / or cellular functions on which AAV depends for replication (e.g., "accessory functions"). Accessory functions include functions required for AAV replication, including, but not limited to, moieties involved in AAV gene transcription activation, stage-specific AAV mRNA splicing, AAV DNA replication, cap expression product synthesis, and AAV capsid assembly. Viral-based accessory functions can be derived from any of the known helper viruses, such as adenovirus, herpesvirus (other than herpes simplex virus type 1), and vaccinia virus. In some embodiments, rAAV particles are produced using a double transfection method in which AAV helper and accessory functions are cloned onto a single vector.
[0193] AAV capsid is an important factor in determining the tissue specificity of rAAV particles.Therefore, it is possible to select rAAV particles with specific capsid tissue specificity.In some embodiments, rAAV particles contain a capsid selected from an AAV serotype selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV.Rh10, AAV11, and their variants.In some embodiments, the AAV capsid is selected from a serotype that can cross the blood-brain barrier, such as AAV9, AAVrh.10, or its variant.In some embodiments, the AAV capsid is a chimeric AAV capsid.
[0194] In some embodiments, the rAAV vector is a mammalian serotype AAV vector (for example, AAV genome and ITR from mammalian serotype AAV), including primate serotype AAV vector or human serotype AAV vector.In some embodiments, the AAV vector is derived from any one of serotypes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV.Rh10, AAV11, and their variants.In some embodiments, the AAV vector is a chimeric AAV vector.In some embodiments, the rAAV vector can be a vector that comprises the AAV genome and AAV capsid from the same AAV serotype.In some embodiments, the rAAV vector is pseudotyped, which means that the rAAV vector comprises the AAV genome from one AAV serotype and the AAV capsid from at least a part of a different AAV serotype.
[0195] In some embodiments, the rAAV vector is an AAV9 serotype. In some embodiments, the rAAV comprises an AAV9 capsid protein (e.g., SEQ ID NO: 2 in U.S. Patent No. 7,198,951), an AAV9 rep protein (e.g., SEQ ID NO: 3 in U.S. Patent No. 7,198,951), or both. In some embodiments, the rAAV comprises (i) an AAV9 capsid protein (e.g., SEQ ID NO: 2 in U.S. Patent No. 7,198,951), and (ii) an AAV2 ITR.
[0196] In some embodiments, rAAV particles can transduce cells of the central nervous system (CNS). In some embodiments, rAAV particles can transduce non-neuronal or neuronal cells of the CNS. In some embodiments, the CNS cells are neurons, glial cells, astrocytes, or microglial cells.
[0197] In some embodiments, the rAAV vector is a self-complementary AAV (scAAV) vector. The scAAV vector comprises two complementary DNA strands in the form of a dimeric inverted repeat genome. The two complementary strands in the dimeric inverted repeat genome anneal together to form a single double-stranded DNA that is ready for immediate replication and transcription, thus bypassing the need for host cell DNA synthesis. Self-complementary AAV vectors are described in U.S. Patent Nos. 7,465,583; 7,790,154; 8,361,457; and 8,784,799.
[0198] The present disclosure also provides a host cell transfected with an rAAV comprising a DNA sequence encoding the double-stranded RNA described herein. In some embodiments, the host cell is a prokaryotic cell or a eukaryotic cell. In some embodiments, the host cell is a mammalian cell (e.g., HEK293T, COS cell, HeLa cell, KB cell), a bacterial cell (E. coli), a yeast cell, an insect cell (Sf9, Sf21, Drosophila, mosquito), or the like. In some embodiments, the host cell is obtained or derived from a human subject. In some embodiments, the host cell is a fibroblast.
[0199] A DNA molecule that encodes one or both strands of double-stranded RNA
[0200] The present disclosure provides a DNA molecule comprising a nucleotide sequence encoding a first strand of a double-stranded RNA of the present disclosure. Optionally, the nucleotide sequence encoding the first strand is operably linked to a promoter. Optionally, the nucleotide sequence encoding the first strand is operably linked to a promoter functional in eukaryotic cells. The present disclosure provides a DNA molecule comprising a nucleotide sequence encoding: i) the first strand of the double-stranded RNA of the present disclosure; and ii) the second strand of the double-stranded RNA of the present disclosure. Optionally, the nucleotide sequences encoding the first strand and the second strand are operably linked to a promoter. Optionally, the promoter is a Pol II promoter. Optionally, the promoter is a U6 promoter. Optionally, the promoter is a CAG promoter. Optionally, the promoter is a CBA promoter. Optionally, the promoter is a CMV promoter. Optionally, the promoter is an EF1α promoter. Optionally, the promoter is an H1 promoter. Optionally, the DNA molecule of the present disclosure comprises a nucleotide sequence encoding any one of SEQ ID NOs: 295-375.
[0201] Recombinant RNA molecules
[0202] The present disclosure provides a recombinant nucleic acid (e.g., recombinant RNA; which may be referred to as an "artificial microRNA" or "small binding RNA" (sbRNA)) comprising: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide (also referred to herein as a "5' leader"), a loop polynucleotide, and a 3' flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5' flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) the loop polynucleotide; (iv) a second strand of the double-stranded RNA; iii) the 3' trailer polynucleotide; and at least one of the 5' flanking polynucleotide, the loop polynucleotide, and the 3' flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. The present disclosure provides a recombinant nucleic acid (e.g., recombinant RNA) comprising: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5'-flanking polynucleotide, a loop polynucleotide, and a 3'-flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5'-flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a loop polynucleotide; (iv) a first strand of the double-stranded RNA; and iii) a 3'-flanking polynucleotide; at least one of the 5'-flanking polynucleotide, the loop polynucleotide, and the 3'-flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5'-flanking polynucleotide, the loop polynucleotide, and the 3'-flanking polynucleotide are derived from miR33.
[0203] The present disclosure provides a recombinant nucleic acid (e.g., recombinant RNA; which may be referred to as an "artificial microRNA" or "small binding RNA" (sbRNA)) comprising: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide (also referred to herein as a "5' leader") and a 3' flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5' flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) a second strand of the double-stranded RNA; and iv) the 3' trailer polynucleotide; one or both of the 5' flanking polynucleotide and the 3' flanking polynucleotide are heterologous to the first strand and / or second strand of the double-stranded RNA. The present disclosure provides a recombinant nucleic acid (e.g., recombinant RNA) comprising: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5'-flanking polynucleotide and a 3'-flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5'-flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a first strand of the double-stranded RNA; and iv) a 3'-flanking polynucleotide; one or both of the 5'-flanking polynucleotide and the 3'-flanking polynucleotide are heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5'-flanking polynucleotide and the 3'-flanking polynucleotide are derived from miR451.
[0204] Cassettes encoding recombinant RNA molecules
[0205] The present disclosure provides a DNA molecule (e.g., a "cassette" that can be inserted into an expression vector to generate a recombinant expression vector) comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5'-flanking polynucleotide (also referred to herein as a "5' leader"), a loop polynucleotide, and a 3'-flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5'-flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) the loop polynucleotide; (iv) a second strand of the double-stranded RNA; iii) the 3'-trailer polynucleotide; and at least one of the 5'-flanking polynucleotide, the loop polynucleotide, and the 3'-flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. The present disclosure provides a DNA molecule (e.g., a "cassette" that can be inserted into an expression vector to generate a recombinant expression vector) comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, the recombinant RNA molecule comprising: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide, a loop polynucleotide, and a 3' flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5' flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a loop polynucleotide; iv) a first strand of the double-stranded RNA; and iii) a 3' flanking polynucleotide; at least one of the 5' flanking polynucleotide, the loop polynucleotide, and the 3' flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5' flanking polynucleotide, the loop polynucleotide, and the 3' flanking polynucleotide are derived from miR33. In some cases, the cassette comprises a Pol3 transcription sequence. For example, in some cases, the cassette comprises the nucleotide sequence TTTTTG 3' of a nucleotide sequence encoding a 3' trailer polynucleotide.In some cases, the cassette includes 10 nucleotide sequences Tn (e.g., n is 5, 6, 7, 8, 9, or 10). In some cases, the cassette has a length of about 110 nucleotides to about 150 nucleotides. In some cases, the cassette includes a Pol II transcription sequence. For example, in some cases, the cassette includes a polyadenylation sequence 3' of the nucleotide sequence encoding the 3' flanking polynucleotide.
[0206] The present disclosure provides DNA molecules (e.g., "cassettes" that can be inserted into an expression vector to generate a recombinant expression vector) comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide (also referred to herein as a "5' leader") and a 3' flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5' flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) a second strand of the double-stranded RNA; and iv) the 3' trailer polynucleotide; wherein one or both of the 5' flanking polynucleotide and the 3' flanking polynucleotide are heterologous to the first strand and / or second strand of the double-stranded RNA. The present disclosure provides a DNA molecule (e.g., a "cassette" that can be inserted into an expression vector to generate a recombinant expression vector) comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide and a 3' flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5' flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a first strand of the double-stranded RNA; and iv) a 3' flanking polynucleotide; one or both of the 5' flanking polynucleotide and the 3' flanking polynucleotide are heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5' flanking polynucleotide and the 3' flanking polynucleotide are derived from miR451. In some cases, the cassette comprises a Pol3 transcription sequence. For example, in some cases, the cassette comprises the nucleotide sequence TTTTTG 3' of a nucleotide sequence encoding a 3' trailer polynucleotide. In some cases, the cassette comprises 10 nucleotide sequences Tn (eg, n is 5, 6, 7, 8, 9, or 10).In some cases, the cassette has a length of about 110 nucleotides to about 650 nucleotides (e.g., 110 nucleotides (nt) to 115 nt, 115 nt to 120 nt, 500 nt to 600 nt, or 600 nt to 610 nt). In some cases, the cassette comprises a Pol II transcription sequence. For example, in some cases, the cassette comprises a 3' polyadenylation sequence of a nucleotide sequence encoding a 3' flanking polynucleotide.
[0207] In some instances, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence tgcacacctcctggggcagctg (SEQ ID NO: 738). In some instances, the portion of the cassette encoding the loop polynucleotide comprises the nucleotide sequence tgtctggcaatacctg (SEQ ID NO: 739). In some instances, the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence ggggaggcctgccctgactgcccac (SEQ ID NO: 740). In some instances, the cassette comprises a Pol 3 transcription sequence. For example, in some instances, the cassette comprises the nucleotide sequence TTTTTG 3' of the nucleotide sequence encoding the 3' trailer polynucleotide. In some instances, the cassette has a length of from about 110 nucleotides to about 150 nucleotides. In some instances, the cassette comprises a Pol II transcription sequence. For example, in some instances, the cassette comprises a polyadenylation sequence 3' of the nucleotide sequence encoding the 3' flanking polynucleotide.
[0208] In some instances, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence acctactgactgccagggcacttgggaatggcaagg (SEQ ID NO: 854). In some instances, the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaac (SEQ ID NO: 855). In some instances, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence acctactgactgccagggcacttgggaatggcaagg (SEQ ID NO: 854). Also, the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaac (SEQ ID NO: 855).
[0209] In some cases, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence gctcctgggcaacgtgctggttattgtgctgtctcatcattttggcaaagaattaagggcgaattcgagctcggtacctcgcgaatgcatctagatatcggcgctatgcttcctgtgcccccagtggggccctggctgggatTtcatcatatactgtaagtttgcgatgagacactacagtatagatgatgtactagtccgggcacccccagctctggagcctgacaaggaggacaggagagatgctgcaagcccaagaagctctctgctcagcctgtcacaacctactgactgccagggcacttgggaatggcaagg (SEQ ID NO: 856). In some cases, the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaactcaggaccctgaagcagactactggaagggagactccagctcaaacaaggcaggggtgggggcgtgggattgggggtaggggagggaatagatacattttctctttcctgttgtaaagaaataaagataagccaggcacagtggctcacgcctgtaatcccaccactttcagaggccaaggcgctggatccagatctcgagcggccgcccg (SEQ ID NO: 857).Optionally, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence gctcctgggcaacgtgctggttattgtgctgtctcatcattttggcaaagaattaagggcgaattcgagctcggtacctcgcgaatgcatctagatatcggcgctatgcttcctgtgcccccagtggggccctggctgggatTtcatcatatactgtaagtttgcgatgagacactacagtatagatgatgtactagtccgggcacccccagctctggagcctgacaaggaggacaggagagatgctgcaagcccaagaagctctctgctcagcctgtcacaacctactgact gccagggcacttgggaatggcaagg (SEQ ID NO:856); and the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaactcaggaccctgaagcagactactggaagggagactccagctcaaacaaggcaggggtgggggcgtgggattgggggtaggggagggaatagatacattttctctttcctgttgtaaagaaataaagataagccaggcacagtggctcacgcctgtaatcccaccactttcagaggccaaggcgctggatccagatctcgagcggccgcccg (SEQ ID NO:857).
[0210] In some cases, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence gctcctgggcaacgtgctggttattgtgctgtctcatcattttggcaaagaattaagggcgaattcgagctcggtacctcgcgaatgcatctagatatcggcgctatgcttcctgtgcccccagtggggccctggctgggatAtcatcatatactgtaagtttgcgatgagacactacagtatagatgatgtactagtccgggcacccccagctctggagcctgacaaggaggacaggagagatgctgcaagcccaagaagctctctgctcagcctgtcacaacctactgactgccagggcacttgggaatggcaagg (SEQ ID NO: 858). In some cases, the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaactcaggaccctgaagcagactactggaagggagactccagctcaaacaaggcaggggtgggggcgtgggattgggggtaggggagggaatagatacattttctctttcctgttgtaaagaaataaagataagccaggcacagtggctcacgcctgtaatcccaccactttcagaggccaaggcgctggatccagatctcgagcggccgccc (SEQ ID NO: 859).In some cases, the portion of the cassette encoding the 5' flanking polynucleotide comprises the nucleotide sequence gctcctgggcaacgtgctggttattgtgctgtctcatcattttggcaaagaattaagggcgaattcgagctcggtacctcgcgaatgcatctagatatcggcgctatgcttcctgtgcccccagtggggccctggctgggatAtcatcatatactgtaagtttgcgatgagacactacagtatagatgatgtactagtccgggcacccccagctctggagcctgacaaggaggacaggagagatgctgcaagcccaagaagctctctgctcagcctgtcacaacctactga the portion of the cassette encoding the 3' flanking polynucleotide comprises the nucleotide sequence tcttgctatacccagaaaacgtgccaggaagagaactcaggaccctgaagcagactactggaagggagactccagctcaaacaaggcaggggtgggggcgtgggattgggggtaggggagggaatagatacattttctctttcctgttgtaaagaaataaagataagccaggcacagtggctcacgcctgtaatcccaccactttcagaggccaaggcgctggatccagatctcgagcggccgccc (SEQ ID NO: 859).
[0211] The following are non-limiting examples of cassettes. In the following cassette: (i) tgcacacctcctggcgggcagctctg (SEQ ID NO:738) encodes the 5' leader polynucleotide; (ii) the first uppercase sequence encodes the first strand of the double-stranded RNA; (iii) tgttctggcaatacctg (SEQ ID NO:739) encodes the loop polynucleotide; (iv) the second uppercase sequence encodes the second strand of the double-stranded RNA; (v) gggaggcctgccctgactgcccac (SEQ ID NO:740) encodes the 3' trailer polynucleotide; and (vi) TTTTTG is a Pol3 transcription termination sequence. In some cases, the cassette does not include the 3' TTTTTG sequence. In some cases, the cassette includes the nucleotide sequence Tn (e.g., n is 5, 6, 7, 8, 9, or 10) instead of the 3' TTTTTG sequence. In some cases, the cassette includes the nucleotide sequence TTTTT instead of the 3' TTTTTG sequence.
[0212] 1) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 579; "CUG-10" in Table 7; Figure 23). In some cases, the cassette does not comprise the 3'TTTTTG sequence. 2) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 580; "CUG_19" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3'TTTTTG sequence; 3) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTACTACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 581; "CUG_28" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 4) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 582; "CUG_37" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 5) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 583; "CUG_46" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 6) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTACTGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 584; "CUG_55" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 7) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTACTGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 585; "CUG_64" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 8) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 586; "CUG_118" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3'TTTTTG sequence; 9) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGATACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 587; "CUG_127" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 10) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGATGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 588; "CUG_136" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 11) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 589; "CUG_145" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 12) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGATGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 590; "CUG_154" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 13) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 591; "CUG_163" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 14) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGCAACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 592; "CUG_217" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 15) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGCAGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 593; "CUG_226" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 16) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 594; "CUG_235" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 17) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 595; "CUG_244" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 18) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 596; "CUG_253" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 19) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCAGCTGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 597; "CUG_NA-A" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 20) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCAACTGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 598; "CUG_NA_B" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 21) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCAGATGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 599; "CUG_NA_C" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette does not include a 3' TTTTTG sequence; 22) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCAGCAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 600; "CUG_NA_D" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 23) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCAGCTACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 601; "CUG_NA_E" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 24) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCAGCTGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 602; "CUG_NA_F" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 25) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCAGCTGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 603; "CUG_NA_G" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 26) In some cases, the cassette has the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCAGCTGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 604; "CUG_NA_H" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 27) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCAGCTGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 605; "CUG_NA_I" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 28) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAAAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 606; "CUG_307" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 29) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 607; "CUG_334" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 30) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 608; "CUG_361" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 31) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 609; "CUG_388" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 32) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 610; "CUG_415" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 33) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 611; "CUG_631" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 34) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 612; "CUG_658" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 35) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 613; "CUG_712" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 36) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 614; "CUG_2116" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 37) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 615; "CUG_2143" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 38) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 616; "CUG_2170" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 39) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTAATGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 617; "CUG_442" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 40) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 618; "CUG_604" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 41) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 619; "CUG_685" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 42) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 620; "CUG_2089" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 43) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 621; "CUG_2197" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 44) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAAAGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 622; "CUG_4870" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 45) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGAAGCTGCTGtgttctggcaatacctgCAGCAGCAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 623; "CUG_9973" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 46) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGATGATGCTGtgttctggcaatacctgCAGCAACAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 624; "CUG_1013" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 47) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCAGATGCTGtgttctggcaatacctgCAGCAACAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 625; "CUG_1070" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 48) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGATACTGCTGtgttctggcaatacctgCAGCAGAAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 626; "CUG_2341" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 49) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCAACTGCTGtgttctggcaatacctgCAGCAGAAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 627; "CUG_2398" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 50) In some cases, the cassette comprises the nucleotide sequence: tgcacacctcctggcgggcagctctgCTGCTGCTAAAACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 628; "CUG_4789" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 51) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAAAGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 629; "CUG_4951" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 52) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAAAGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 630; "CUG_5032" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 53) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAAAGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 631; "CUG_5113" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 54) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATAATGCTGCTGtgttctggcaatacctgCAGCAGCAAAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 632; "CUG_5599" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 55) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATACAGCTGCTGtgttctggcaatacctgCAGCAGCAGAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 633; "CUG_5680" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 56) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATACTACTGCTGtgttctggcaatacctgCAGCAGAAGAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 634; "CUG_5761" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 57) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATACTGATGCTGtgttctggcaatacctgCAGCAACAGAAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 635; "CUG_5842" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 58) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGAAGCTGCTGtgttctggcaatacctgCAGCAGCAACAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 636; "CUG_6328" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 59) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGATACTGCTGtgttctggcaatacctgCAGCAGAAACAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 637; "CUG_6409" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 60) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGATGATGCTGtgttctggcaatacctgCAGCAACAACAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 638; "CUG_6490" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 61) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCAACTGCTGtgttctggcaatacctgCAGCAGAAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 639; "CUG_6976" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 62) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCAGATGCTGtgttctggcaatacctgCAGCAACAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 640; "CUG_7057" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 63) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCTAATGCTGtgttctggcaatacctgCAGCAAAAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 641; "CUG_7543" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 64) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAAATGCTGCTGtgttctggcaatacctgCAGCAGCAAAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 642; "CUG_9244" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 65) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAACAGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 643; "CUG_9325" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 66) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAACTACTGCTGtgttctggcaatacctgCAGCAGAAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 644; "CUG_9406" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 67) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAACTGATGCTGtgttctggcaatacctgCAGCAACAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 645; "CUG_9487" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 68) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGATACTGCTGtgttctggcaatacctgCAGCAGAAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 646; "CUG_10054" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 69) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCAACTGCTGtgttctggcaatacctgCAGCAGAAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 647; "CUG_10621" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 70) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCTAATGCTGtgttctggcaatacctgCAGCAAAAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 648; "CUG_11188" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 71) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAAATGCTGCTGtgttctggcaatacctgCAGCAGCAAAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 649; "CUG_222609" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 72) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAACAGCTGCTGtgttctggcaatacctgCAGCAGCAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 650; "CUG_22690" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 73) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAACTACTGCTGtgttctggcaatacctgCAGCAGAAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 651; "CUG_22771" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 74) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAACTGATGCTGtgttctggcaatacctgCAGCAACAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 652; "CUG_22852" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 75) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGAAGCTGCTGtgttctggcaatacctgCAGCAGCAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 653; "CUG_233338" in Table 7; Figure 23). In some cases, the cassette does not include the 3'TTTTTG sequence; 76) In some cases, the cassette comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGATGATGCTGtgttctggcaatacctgCAGCAACAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 654; "CUG_23500" in Table 7; Figure 23). In some cases, the cassette does not include a 3'TTTTTG sequence; 77) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCAGATGCTGtgttctggcaatacctgCAGCAACAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 655; "CUG_24067" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence; 78) In some cases, the cassette has the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCTAATGCTGtgttctggcaatacctgCAGCAAAAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 656; "CUG_24553" in Table 7; Figure 23). In some cases, the cassette does not include the 3' TTTTTG sequence.
[0213] The following are non-limiting examples of cassettes. In the following cassettes, the pairs of 5'-flanking polynucleotides and 3'-flanking polynucleotides include i) SEQ ID NO:854 and SEQ ID NO:855; ii) SEQ ID NO:856 and SEQ ID NO:857; and iii) SEQ ID NO:858 and SEQ ID NO:859. Optionally, the cassette includes a 3'TTTTTG transcription termination sequence. Optionally, the cassette does not include a 3'TTTTTG sequence. Optionally, the cassette includes a nucleotide sequence Tn at the 3' end, where n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette includes a nucleotide sequence TTTTT at the 3' end. Non-limiting examples include the cassette shown in FIG. 25.
[0214] 1) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:379; iii) a sequence complementary to SEQ ID NO:379, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:379; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:379; iii) a sequence complementary to SEQ ID NO:379, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:379; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:379; iii) a sequence complementary to SEQ ID NO:379, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:379; and iv) SEQ ID NO:859.
[0215] 2) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:380; iii) a sequence complementary to SEQ ID NO:380, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:380; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:380; iii) a sequence complementary to SEQ ID NO:380, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:380; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:380; iii) a sequence complementary to SEQ ID NO:380, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:380; and iv) SEQ ID NO:859.
[0216] 3) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:381; iii) a sequence complementary to SEQ ID NO:381, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:381; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:381; iii) a sequence complementary to SEQ ID NO:381, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:381; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:381; iii) a sequence complementary to SEQ ID NO:381, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:381; and iv) SEQ ID NO:859.
[0217] 4) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:382; iii) a sequence complementary to SEQ ID NO:382, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:382; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:382; iii) a sequence complementary to SEQ ID NO:382, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:382; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:382; iii) a sequence complementary to SEQ ID NO:382, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:382; and iv) SEQ ID NO:859.
[0218] 5) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:383; iii) a sequence complementary to SEQ ID NO:383, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:383; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:383; iii) a sequence complementary to SEQ ID NO:383, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:383; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:383; iii) a sequence complementary to SEQ ID NO:383, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:383; and iv) SEQ ID NO:859.
[0219] 6) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:384; iii) a sequence complementary to SEQ ID NO:384, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:384; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:384; iii) a sequence complementary to SEQ ID NO:384, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:384; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:384; iii) a sequence complementary to SEQ ID NO:384, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:384; and iv) SEQ ID NO:859.
[0220] 7) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:385; iii) a sequence complementary to SEQ ID NO:385, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:385; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:385; iii) a sequence complementary to SEQ ID NO:385, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:385; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:385; iii) a sequence complementary to SEQ ID NO:385, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:385; and iv) SEQ ID NO:859.
[0221] 8) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:386; iii) a sequence complementary to SEQ ID NO:386, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:386; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:386; iii) a sequence complementary to SEQ ID NO:386, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:386; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:386; iii) a sequence complementary to SEQ ID NO:386, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:386; and iv) SEQ ID NO:859.
[0222] 9) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:387; iii) a sequence complementary to SEQ ID NO:387, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:387; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:387; iii) a sequence complementary to SEQ ID NO:387, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:387; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:387; iii) a sequence complementary to SEQ ID NO:387, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:387; and iv) SEQ ID NO:859.
[0223] 10) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:388; iii) a sequence complementary to SEQ ID NO:388, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:388; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:388; iii) a sequence complementary to SEQ ID NO:388, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:388; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:388; iii) a sequence complementary to SEQ ID NO:388, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:388; and iv) SEQ ID NO:859.
[0224] 11) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:389; iii) a sequence complementary to SEQ ID NO:389, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:389; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:389; iii) a sequence complementary to SEQ ID NO:389, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:389; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:389; iii) a sequence complementary to SEQ ID NO:389, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:389; and iv) SEQ ID NO:859.
[0225] 12) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:390, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:390, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:309, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:859.
[0226] 13) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:859.
[0227] 14) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:390, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:390, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:390; iii) a sequence complementary to SEQ ID NO:390, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:390; and iv) SEQ ID NO:859.
[0228] 15) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:391; iii) a sequence complementary to SEQ ID NO:391, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:391; and iv) SEQ ID NO:859.
[0229] 16) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:392; iii) a sequence complementary to SEQ ID NO:392, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:392; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:392; iii) a sequence complementary to SEQ ID NO:392, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:392; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:392; iii) a sequence complementary to SEQ ID NO:392, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:392; and iv) SEQ ID NO:859.
[0230] 17) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:393; iii) a sequence complementary to SEQ ID NO:393, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:393; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:393; iii) a sequence complementary to SEQ ID NO:393, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:393; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:393; iii) a sequence complementary to SEQ ID NO:393, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:393; and iv) SEQ ID NO:859.
[0231] 18) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:394; iii) a sequence complementary to SEQ ID NO:394, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:394; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:394; iii) a sequence complementary to SEQ ID NO:394, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:394; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:394; iii) a sequence complementary to SEQ ID NO:394, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:394; and iv) SEQ ID NO:859.
[0232] 19) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:395; iii) a sequence complementary to SEQ ID NO:395, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:395; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:395; iii) a sequence complementary to SEQ ID NO:395, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:395; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:395; iii) a sequence complementary to SEQ ID NO:395, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:395; and iv) SEQ ID NO:859.
[0233] 20) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:396; iii) a sequence complementary to SEQ ID NO:396, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:396; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:396; iii) a sequence complementary to SEQ ID NO:396, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:396; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:396; iii) a sequence complementary to SEQ ID NO:396, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:396; and iv) SEQ ID NO:859.
[0234] 21) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:397; iii) a sequence complementary to SEQ ID NO:397, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:397; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:397; iii) a sequence complementary to SEQ ID NO:397, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:397; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:397; iii) a sequence complementary to SEQ ID NO:397, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:397; and iv) SEQ ID NO:859.
[0235] 22) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:398; iii) a sequence complementary to SEQ ID NO:398, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:398; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:398; iii) a sequence complementary to SEQ ID NO:398, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:398; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:398; iii) a sequence complementary to SEQ ID NO:398, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:398; and iv) SEQ ID NO:859.
[0236] 23) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:399; iii) a sequence complementary to SEQ ID NO:399, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:399; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:399; iii) a sequence complementary to SEQ ID NO:399, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:399; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:399; iii) a sequence complementary to SEQ ID NO:399, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:399; and iv) SEQ ID NO:859.
[0237] 24) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:400; iii) a sequence complementary to SEQ ID NO:400, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:400; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:400; iii) a sequence complementary to SEQ ID NO:400, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:400; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:400; iii) a sequence complementary to SEQ ID NO:400, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:400; and iv) SEQ ID NO:859.
[0238] 25) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:401; iii) a sequence complementary to SEQ ID NO:401, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:401; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:401; iii) a sequence complementary to SEQ ID NO:401, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:401; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:401; iii) a sequence complementary to SEQ ID NO:401, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:401; and iv) SEQ ID NO:859.
[0239] 26) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:402; iii) a sequence complementary to SEQ ID NO:402, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:402; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:402; iii) a sequence complementary to SEQ ID NO:402, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:402; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:402; iii) a sequence complementary to SEQ ID NO:402, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:402; and iv) SEQ ID NO:859.
[0240] 27) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:403; iii) a sequence complementary to SEQ ID NO:403, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:403; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:403; iii) a sequence complementary to SEQ ID NO:403, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:403; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:403; iii) a sequence complementary to SEQ ID NO:403, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:403; and iv) SEQ ID NO:859.
[0241] 28) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:404; iii) a sequence complementary to SEQ ID NO:404, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:404; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:404; iii) a sequence complementary to SEQ ID NO:404, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:404; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:404; iii) a sequence complementary to SEQ ID NO:404, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:404; and iv) SEQ ID NO:859.
[0242] 29) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:405; iii) a sequence complementary to SEQ ID NO:405, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:405; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:405; iii) a sequence complementary to SEQ ID NO:405, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:405; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:405; iii) a sequence complementary to SEQ ID NO:405, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:405; and iv) SEQ ID NO:859.
[0243] 30) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:406; iii) a sequence complementary to SEQ ID NO:406, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:406; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:406; iii) a sequence complementary to SEQ ID NO:406, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:406; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:406; iii) a sequence complementary to SEQ ID NO:406, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:406; and iv) SEQ ID NO:859.
[0244] 31) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:407; iii) a sequence complementary to SEQ ID NO:407, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:407; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:407; iii) a sequence complementary to SEQ ID NO:407, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:407; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:407; iii) a sequence complementary to SEQ ID NO:407, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:407; and iv) SEQ ID NO:859.
[0245] 32) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:408; iii) a sequence complementary to SEQ ID NO:408, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:408; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:408; iii) a sequence complementary to SEQ ID NO:408, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:408; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:408; iii) a sequence complementary to SEQ ID NO:408, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:408; and iv) SEQ ID NO:859.
[0246] 33) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:409; iii) a sequence complementary to SEQ ID NO:409, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:409; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:409; iii) a sequence complementary to SEQ ID NO:409, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:409; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:409; iii) a sequence complementary to SEQ ID NO:409, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:409; and iv) SEQ ID NO:859.
[0247] 34) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:410; iii) a sequence complementary to SEQ ID NO:410, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:410; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:410; iii) a sequence complementary to SEQ ID NO:410, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:410; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:410; iii) a sequence complementary to SEQ ID NO:410, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:410; and iv) SEQ ID NO:859.
[0248] 35) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:411; iii) a sequence complementary to SEQ ID NO:411, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:411; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:411; iii) a sequence complementary to SEQ ID NO:411, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:411; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:411; iii) a sequence complementary to SEQ ID NO:411, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:411; and iv) SEQ ID NO:859.
[0249] 36) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:412; iii) a sequence complementary to SEQ ID NO:412, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:412; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:412; iii) a sequence complementary to SEQ ID NO:412, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:412; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:412; iii) a sequence complementary to SEQ ID NO:412, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:412; and iv) SEQ ID NO:859.
[0250] 37) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:413; iii) a sequence complementary to SEQ ID NO:413, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:413; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:413; iii) a sequence complementary to SEQ ID NO:413, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:413; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:413; iii) a sequence complementary to SEQ ID NO:413, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:413; and iv) SEQ ID NO:859.
[0251] 38) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:414; iii) a sequence complementary to SEQ ID NO:414, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:414; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:414; iii) a sequence complementary to SEQ ID NO:414, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:414; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:414; iii) a sequence complementary to SEQ ID NO:414, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:414; and iv) SEQ ID NO:859.
[0252] 39) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:415; iii) a sequence complementary to SEQ ID NO:415, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:415; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:415; iii) a sequence complementary to SEQ ID NO:415, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:415; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:415; iii) a sequence complementary to SEQ ID NO:415, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:415; and iv) SEQ ID NO:859.
[0253] 40) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:416; iii) a sequence complementary to SEQ ID NO:416, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:416; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:416; iii) a sequence complementary to SEQ ID NO:416, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:416; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:416; iii) a sequence complementary to SEQ ID NO:416, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:416; and iv) SEQ ID NO:859.
[0254] 41) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:417; iii) a sequence complementary to SEQ ID NO:417, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:417; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:417; iii) a sequence complementary to SEQ ID NO:417, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:417; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:417; iii) a sequence complementary to SEQ ID NO:417, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:417; and iv) SEQ ID NO:859.
[0255] 42) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:418; iii) a sequence complementary to SEQ ID NO:418, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:418; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:418; iii) a sequence complementary to SEQ ID NO:418, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:418; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:418; iii) a sequence complementary to SEQ ID NO:418, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:418; and iv) SEQ ID NO:859.
[0256] 43) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:419; iii) a sequence complementary to SEQ ID NO:419, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:419; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:419; iii) a sequence complementary to SEQ ID NO:419, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:419; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:419; iii) a sequence complementary to SEQ ID NO:419, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:419; and iv) SEQ ID NO:859.
[0257] 44) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:420; iii) a sequence complementary to SEQ ID NO:420, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:420; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:420; iii) a sequence complementary to SEQ ID NO:420, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:420; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:420; iii) a sequence complementary to SEQ ID NO:420, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:420; and iv) SEQ ID NO:859.
[0258] 45) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:421; iii) a sequence complementary to SEQ ID NO:421, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:421; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:421; iii) a sequence complementary to SEQ ID NO:421, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:421; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:421; iii) a sequence complementary to SEQ ID NO:421, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:421; and iv) SEQ ID NO:859.
[0259] 46) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:422; iii) a sequence complementary to SEQ ID NO:422, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:422; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:422; iii) a sequence complementary to SEQ ID NO:422, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:422; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:422; iii) a sequence complementary to SEQ ID NO:422, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:422; and iv) SEQ ID NO:859.
[0260] 47) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:423; iii) a sequence complementary to SEQ ID NO:423, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:423; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:423; iii) a sequence complementary to SEQ ID NO:423, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:423; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:423; iii) a sequence complementary to SEQ ID NO:423, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:423; and iv) SEQ ID NO:859.
[0261] 48) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:424; iii) a sequence complementary to SEQ ID NO:424, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:424; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:424; iii) a sequence complementary to SEQ ID NO:424, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:424; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:424; iii) a sequence complementary to SEQ ID NO:424, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:424; and iv) SEQ ID NO:859.
[0262] 49) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:425; iii) a sequence complementary to SEQ ID NO:425, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:425; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:425; iii) a sequence complementary to SEQ ID NO:425, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:425; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:425; iii) a sequence complementary to SEQ ID NO:425, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:425; and iv) SEQ ID NO:859.
[0263] 50) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:426; iii) a sequence complementary to SEQ ID NO:426, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:426; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:426; iii) a sequence complementary to SEQ ID NO:426, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:426; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:426; iii) a sequence complementary to SEQ ID NO:426, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:426; and iv) SEQ ID NO:859.
[0264] 51) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:427; iii) a sequence complementary to SEQ ID NO:427, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:427; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:427; iii) a sequence complementary to SEQ ID NO:427, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:427; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:427; iii) a sequence complementary to SEQ ID NO:427, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:427; and iv) SEQ ID NO:859.
[0265] 52) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:428; iii) a sequence complementary to SEQ ID NO:428, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:428; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:428; iii) a sequence complementary to SEQ ID NO:428, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:428; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:428; iii) a sequence complementary to SEQ ID NO:428, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:428; and iv) SEQ ID NO:859.
[0266] 53) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:429; iii) a sequence complementary to SEQ ID NO:429, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:429; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:429; iii) a sequence complementary to SEQ ID NO:429, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:429; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:429; iii) a sequence complementary to SEQ ID NO:429, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:429; and iv) SEQ ID NO:859.
[0267] 54) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:430; iii) a sequence complementary to SEQ ID NO:430, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:430; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:430; iii) a sequence complementary to SEQ ID NO:430, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:430; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:430; iii) a sequence complementary to SEQ ID NO:430, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:430; and iv) SEQ ID NO:859.
[0268] 55) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:431; iii) a sequence complementary to SEQ ID NO:431, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:431; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:431; iii) a sequence complementary to SEQ ID NO:431, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:431; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:431; iii) a sequence complementary to SEQ ID NO:431, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:431; and iv) SEQ ID NO:859.
[0269] 56) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:432; iii) a sequence complementary to SEQ ID NO:432, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:432; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:432; iii) a sequence complementary to SEQ ID NO:432, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:432; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:432; iii) a sequence complementary to SEQ ID NO:432, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:432; and iv) SEQ ID NO:859.
[0270] 57) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:433; iii) a sequence complementary to SEQ ID NO:433, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:433; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:433; iii) a sequence complementary to SEQ ID NO:433, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:433; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:433; iii) a sequence complementary to SEQ ID NO:433, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:433; and iv) SEQ ID NO:859.
[0271] 58) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:434; iii) a sequence complementary to SEQ ID NO:434, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:434; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:434; iii) a sequence complementary to SEQ ID NO:434, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:434; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:434; iii) a sequence complementary to SEQ ID NO:434, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:434; and iv) SEQ ID NO:859.
[0272] 59) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:435; iii) a sequence complementary to SEQ ID NO:435, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:435; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:435; iii) a sequence complementary to SEQ ID NO:435, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:435; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:435; iii) a sequence complementary to SEQ ID NO:435, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:435; and iv) SEQ ID NO:859.
[0273] 60) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:436; iii) a sequence complementary to SEQ ID NO:436, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:436; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:436; iii) a sequence complementary to SEQ ID NO:436, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:436; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:436; iii) a sequence complementary to SEQ ID NO:436, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:436; and iv) SEQ ID NO:859.
[0274] 61) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:437; iii) a sequence complementary to SEQ ID NO:437, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:437; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:437; iii) a sequence complementary to SEQ ID NO:437, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:437; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:437; iii) a sequence complementary to SEQ ID NO:437, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:437; and iv) SEQ ID NO:859.
[0275] 62) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:438; iii) a sequence complementary to SEQ ID NO:438, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:438; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:438; iii) a sequence complementary to SEQ ID NO:438, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:438; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:438; iii) a sequence complementary to SEQ ID NO:438, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:438; and iv) SEQ ID NO:859.
[0276] 63) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:439; iii) a sequence complementary to SEQ ID NO:439, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:439; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:439; iii) a sequence complementary to SEQ ID NO:439, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:439; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:439; iii) a sequence complementary to SEQ ID NO:439, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:439; and iv) SEQ ID NO:859.
[0277] 64) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:440; iii) a sequence complementary to SEQ ID NO:440, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:440; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:440; iii) a sequence complementary to SEQ ID NO:440, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:440; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:440; iii) a sequence complementary to SEQ ID NO:440, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:440; and iv) SEQ ID NO:859.
[0278] 65) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:441; iii) a sequence complementary to SEQ ID NO:441, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:441; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:441; iii) a sequence complementary to SEQ ID NO:441, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:441; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:441; iii) a sequence complementary to SEQ ID NO:441, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:441; and iv) SEQ ID NO:859.
[0279] 66) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:442; iii) a sequence complementary to SEQ ID NO:442, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:442; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:442; iii) a sequence complementary to SEQ ID NO:442, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:442; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:442; iii) a sequence complementary to SEQ ID NO:442, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:442; and iv) SEQ ID NO:859.
[0280] 67) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:443; iii) a sequence complementary to SEQ ID NO:443, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:443; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:443; iii) a sequence complementary to SEQ ID NO:443, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:443; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:443; iii) a sequence complementary to SEQ ID NO:443, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:443; and iv) SEQ ID NO:859.
[0281] 68) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:444; iii) a sequence complementary to SEQ ID NO:444, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:444; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:444; iii) a sequence complementary to SEQ ID NO:444, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:444; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:444; iii) a sequence complementary to SEQ ID NO:444, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:444; and iv) SEQ ID NO:859.
[0282] 69) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:445; iii) a sequence complementary to SEQ ID NO:445, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:445; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:445; iii) a sequence complementary to SEQ ID NO:445, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:445; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:445; iii) a sequence complementary to SEQ ID NO:445, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:445; and iv) SEQ ID NO:859.
[0283] 70) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:446; iii) a sequence complementary to SEQ ID NO:446, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:446; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:446; iii) a sequence complementary to SEQ ID NO:446, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:446; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:446; iii) a sequence complementary to SEQ ID NO:446, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:446; and iv) SEQ ID NO:859.
[0284] 71) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:447; iii) a sequence complementary to SEQ ID NO:447, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:447; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:447; iii) a sequence complementary to SEQ ID NO:447, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:447; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:447; iii) a sequence complementary to SEQ ID NO:447, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:447; and iv) SEQ ID NO:859.
[0285] 72) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:448; iii) a sequence complementary to SEQ ID NO:448, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:448; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:448; iii) a sequence complementary to SEQ ID NO:448, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:448; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:448; iii) a sequence complementary to SEQ ID NO:448, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:448; and iv) SEQ ID NO:859.
[0286] 73) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:449; iii) a sequence complementary to SEQ ID NO:449, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:449; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:449; iii) a sequence complementary to SEQ ID NO:449, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:449; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:449; iii) a sequence complementary to SEQ ID NO:449, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:449; and iv) SEQ ID NO:859.
[0287] 74) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:450; iii) a sequence complementary to SEQ ID NO:450, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:450; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:450; iii) a sequence complementary to SEQ ID NO:450, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:450; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:450; iii) a sequence complementary to SEQ ID NO:450, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:450; and iv) SEQ ID NO:859.
[0288] 75) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:451; iii) a sequence complementary to SEQ ID NO:451, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:451; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:451; iii) a sequence complementary to SEQ ID NO:451, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:451; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:451; iii) a sequence complementary to SEQ ID NO:451, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:451; and iv) SEQ ID NO:859.
[0289] 76) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:452; iii) a sequence complementary to SEQ ID NO:452, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:452; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:452; iii) a sequence complementary to SEQ ID NO:452, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:452; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:452; iii) a sequence complementary to SEQ ID NO:452, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:452; and iv) SEQ ID NO:859.
[0290] 77) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:453; iii) a sequence complementary to SEQ ID NO:453, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:453; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:453; iii) a sequence complementary to SEQ ID NO:453, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:453; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:453; iii) a sequence complementary to SEQ ID NO:453, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:453; and iv) SEQ ID NO:859.
[0291] 78) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:454; iii) a sequence complementary to SEQ ID NO:454, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:454; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:454; iii) a sequence complementary to SEQ ID NO:454, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:454; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:454; iii) a sequence complementary to SEQ ID NO:454, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:454; and iv) SEQ ID NO:859.
[0292] 79) Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:455; iii) a sequence complementary to SEQ ID NO:455, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:455; and iv) SEQ ID NO:855. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:455; iii) a sequence complementary to SEQ ID NO:455, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:455; and iv) SEQ ID NO:857. Optionally, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:455; iii) a sequence complementary to SEQ ID NO:455, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:455; and iv) SEQ ID NO:859.
[0293] 80) In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:854; ii) SEQ ID NO:456; iii) a sequence complementary to SEQ ID NO:456, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:456; and iv) SEQ ID NO:855. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:856; ii) SEQ ID NO:456; iii) a sequence complementary to SEQ ID NO:456, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:456; and iv) SEQ ID NO:857. In some cases, the cassette comprises, in 5' to 3' order, i) SEQ ID NO:858; ii) SEQ ID NO:456; iii) a sequence complementary to SEQ ID NO:456, which may include 0 to 10 mismatches (non-complementary nucleotides) relative to SEQ ID NO:456; and iv) SEQ ID NO:859.
[0294] [Recombinant expression vector encoding sbRNA] The present disclosure provides recombinant expression vectors comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure (wherein the recombinant RNA molecule may be referred to as an "artificial microRNA" or "sbRNA"). In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a promoter functional in a eukaryotic cell. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an RNA polymerase II promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an RNA polymerase III promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a CMV promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a CAG promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a CBA promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a U6 promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an EF1α promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an H1 promoter. In some cases, the recombinant expression vector comprises 5' adeno-associated virus (AAV) inverted terminal repeat (ITR) sequences and 3' AAV ITR sequences.
[0295] The present disclosure provides a recombinant expression vector comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide (also referred to herein as a "5' leader"), a loop polynucleotide, and a 3' flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5' flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) the loop polynucleotide; (iv) a second strand of the double-stranded RNA; and iii) the 3' trailer polynucleotide, wherein at least one of the 5' flanking polynucleotide, the loop polynucleotide, and the 3' flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. The present disclosure provides a recombinant expression vector comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5'-flanking polynucleotide, a loop polynucleotide, and a 3'-flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5'-flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a loop polynucleotide; iv) a first strand of the double-stranded RNA; and iii) a 3'-flanking polynucleotide, wherein at least one of the 5'-flanking polynucleotide, the loop polynucleotide, and the 3'-flanking polynucleotide is heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5'-flanking polynucleotide, the loop polynucleotide, and the 3'-flanking polynucleotide are derived from miR33.
[0296] The present disclosure provides recombinant expression vectors comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5' flanking polynucleotide (also referred to herein as a "5' leader") and a 3' flanking polynucleotide (also referred to herein as a "3' trailer"), wherein the recombinant nucleic acid comprises: i) the 5' flanking polynucleotide; ii) a first strand of the double-stranded RNA; iii) a second strand of the double-stranded RNA; and iv) a 3' trailer polynucleotide, wherein one or both of the 5' flanking polynucleotide and the 3' flanking polynucleotide are heterologous to the first strand and / or the second strand of the double-stranded RNA. The present disclosure provides a recombinant expression vector comprising a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure, wherein the recombinant RNA molecule comprises: a) a double-stranded RNA of the present disclosure; and b) a microRNA scaffold comprising a 5'-flanking polynucleotide and a 3'-flanking polynucleotide, wherein the recombinant nucleic acid comprises: i) a 5'-flanking polynucleotide; ii) a second strand of the double-stranded RNA; iii) a first strand of the double-stranded RNA; and iv) a 3'-flanking polynucleotide, wherein one or both of the 5'-flanking polynucleotide and the 3'-flanking polynucleotide are heterologous to the first strand and / or the second strand of the double-stranded RNA. In some cases, the 5'-flanking polynucleotide and the 3'-flanking polynucleotide are derived from miR451.
[0297] [Recombinant expression vector containing a cassette] The present disclosure provides a recombinant expression vector comprising a cassette ("DNA molecule") of the present disclosure, wherein the cassette comprises a nucleotide sequence encoding a recombinant RNA molecule of the present disclosure (wherein the recombinant RNA molecule may be referred to as an "artificial microRNA" or "sbRNA"). In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a promoter functional in a eukaryotic cell. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an RNA polymerase II promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an RNA polymerase III promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a CMV promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to a U6 promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an EF1α promoter. In some cases, the nucleotide sequence encoding the recombinant RNA molecule is operably linked to an H1 promoter. In some cases, the recombinant expression vector comprises 5' adeno-associated virus (AAV) inverted terminal repeat (ITR) sequences and 3' AAV ITR sequences.
[0298] The following are non-limiting examples of recombinant expression vectors that contain the cassettes of the present disclosure.
[0299] 1) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTAATGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAAAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 579; "CUG-10" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0300] 2) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 580; "CUG_19" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0301] 3) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 581; "CUG_28" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0302] 4) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 582; "CUG_37" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0303] 5) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 583; "CUG_46" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0304] 6) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 584; "CUG_55" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0305] 7) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTACTGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGAAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 585; "CUG_64" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0306] 8) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGAAGCTGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 586; "CUG_118" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0307] 9) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 587; "CUG_127" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0308] 10) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 588; "CUG_136" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0309] 11) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 589; "CUG_145" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0310] 12) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 590; "CUG_154" in Table 7; Figure 23). In some cases, the cassette does not comprise a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0311] 13) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGATGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAACAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 591; "CUG_163" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0312] 14) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAACTGCTGCTGtgttctggcaatacctgCAGCAGCAGAAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 592; "CUG_217" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0313] 15) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGATGCTGCTGtgttctggcaatacctgCAGCAGCAACAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 593; "CUG_226" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0314] 16) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCAGCTGCTGtgttctggcaatacctgCAGCAGCAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 594; "CUG_235" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0315] 17) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCTACTGCTGtgttctggcaatacctgCAGCAGAAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 595; "CUG_244" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the recombinant expression vector comprises a 5' AAV ITR sequence and a 3' AAV ITR sequence. Optionally, the AAV ITR is an AAV9 ITR. Optionally, the AAV ITR is an AAV2 ITR.
[0316] 18) In some cases, the recombinant expression vector comprises the nucleotide sequence tgcacacctcctggcgggcagctctgCTGCTGCTGCAGCTGATGCTGtgttctggcaatacctgCAGCAACAGCAGCAGCAGCAGCAgggaggcctgccctgactgcccacTTTTTG (SEQ ID NO: 596; "CUG_253" in Table 7; Figure 23). In some cases, the cassette does not include a 3' TTTTTG sequence. In some cases, the cassette comprises the nucleotide sequence T nwherein n is an integer between 5 and 10 (e.g., n is 5, 6, 7, 8, 9, or 10). Optionally, the cassette is operably linked to a promoter that functions in a eukaryotic cell. Optionally, the cassette is operably linked to an RNA polymerase II promoter. Optionally, the cassette is operably linked to an RNA polymerase III promoter. Optionally, the cassette is operably linked to a CAG promoter. Optionally, the cassette is operably linked to a CBA promoter. Optionally, the cassette is operably linked to a CMV promoter. Optionally, the cassette is operably linked to a U6 promoter. Optionally, the cassette is operably linked to an EF1α promoter. Optionally, the cassette is operably linked to an H1 promoter. Optionally, the reco...
Claims
[Claim 1] An invention as described in the specification or drawings.