Disrupted RAN protein in disease
By identifying and treating RAN protein diseases through the recognition of interrupted RAN proteins with discontinuous motifs, the challenges of detecting and managing RAN protein translation in neurodegenerative disorders are addressed, providing effective treatment strategies for conditions like Alzheimer's and ALS.
Patent Information
- Application Number
- JP2025525086
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-31
- Filing Date
- 2023-11-01
- Publication Date
- 2026-01-28
AI Technical Summary
Current technologies struggle to detect and address the challenges posed by repeat-associated non-ATG (RAN) protein translation, which is associated with various neurodegenerative disorders, due to the complexity of identifying repeat expansions and the molecular features of RNA foci and protein accumulation.
The identification and treatment of RAN protein diseases are facilitated by recognizing a novel class of interrupted RAN proteins, characterized by discontinuous poly-amino acid repeat motifs, through methods such as antibody-based assays and administration of anti-RAN protein agents to reduce transcription, translation, expression, aggregation, or accumulation of these proteins.
This approach enables effective identification and treatment of RAN proteinopathies, including Alzheimer's disease and ALS, by targeting disrupted RAN proteins, thereby mitigating the progression of these disorders.
Smart Images

Figure 2026503198000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Application No. 63 / 421,533, filed November 1, 2022, and U.S. Provisional Application No. 63 / 456,405, filed March 31, 2023, each of which is incorporated by reference in its entirety.
[0002] Electronic Sequence Listing Reference The contents of the electronic sequence listing (U120270128WO00-SEQ-KZM.txt, size: 183,398 bytes, and creation date: November 1, 2023) are incorporated herein by reference in their entirety.
[0003] Federal Funding Statement This invention was made with U.S. government support under Grant Nos. K99 AG065511, NS126536, and R01 NS098819 awarded by the National Institutes of Health, and Grant No. W81XWH-22-1-0592 awarded by the USA Army Medical Acquisition Activity. The U.S. government has certain rights in this invention. [Background technology]
[0004] Microsatellite repeat expansions are known to cause over 40 neurodegenerative disorders. Common molecular features of many of these disorders include the accumulation of RNA foci containing sense and antisense expanded transcripts, as well as protein accumulation from repeat-associated non-AUG (RAN) translation. RAN translation can occur across a wide range of repeat lengths, from premutation lengths (approximately 30–40 repeats) to full expansions (up to 10,000 repeats). Although repetitive elements comprise a large portion of the human genome, detecting repeat and repeat expansion mutations is challenging. Summary of the Invention
[0005] Aspects of the present disclosure relate to methods and compositions for identifying and / or treating diseases associated with repeat-associated non-ATG (RAN) protein translation (e.g., RAN protein diseases). A "RAN protein (e.g., a repeat-associated non-ATG translation protein)" is a polypeptide translated from an mRNA sequence carrying a nucleotide extension without an obvious AUG start codon. Generally, RAN proteins contain extended repeats of single, di-, tri-, or quad-amino acids (e.g., tetra-amino acids), referred to as poly-amino acid repeats. However, the present inventors have surprisingly discovered that a novel class of RAN proteins containing discontinuous or "interrupted" poly-amino acid repeat motifs is expressed in certain RAN protein diseases. As used herein, "interrupted RAN protein" or "discontiguous RNA protein" refers to a RAN protein translated from an RNA transcript containing multiple nucleotide extended repeat units (e.g., GGGGCT repeat units, GAAGGA repeat units, GGGAGA repeat units, etc.) with one or more amino acid changes that cause a frameshift and result in the production of a polypeptide containing a discontinuous, repeating pattern of poly-amino acid repeat units spanning the entire length of the nucleotide extension. In some embodiments, the RNA transcript encoding the interrupted RAN protein contains one or more nucleotides inserted between one or more extended repeat units and / or one or more nucleotide substitutions within one or more extended repeat units.
[0006] As further described in the Examples, the present inventors have recognized that truncated RAN proteins expressed from certain RNA transcripts are associated with RAN protein disorders, such as Alzheimer's disease and amyotrophic lateral sclerosis (ALS). In some embodiments, one or more genes that can be transcribed to produce a truncated RAN protein are listed in Table 1 or Table 6.
[0007] Generally, an interrupted RAN protein can contain between about 2 and about 10,000 discontinuous or "interrupted" amino acid repeats (total) ("RAN repeat units"). In some embodiments, the interrupted RAN protein contains between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 discontinuous or "interrupted" amino acid repeats, etc. (e.g., RAN repeat units). In some embodiments, the interrupted RAN protein contains one or more poly-amino acid repeats that are between 2 and 500, between 20 and 300, between 30 and 200, between 40 and 100, between 50 and 90, or between 60 and 80 amino acid residues in length. In some embodiments, the interrupted RAN protein comprises one or more poly-amino acid repeats that are at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or at least 200 amino acid residues in length. In some embodiments, the interrupted RAN protein has one or more poly-amino acid repeats that are more than 200 amino acid residues in length (e.g., 500, 1000, 5000, 10,000, etc.).
[0008] As will be appreciated, the poly-amino acid repeats (e.g., RAN repeat units) contained in the interrupted RAN proteins are separated by one or more non-repeated amino acids (see, e.g., FIG. 1B). In some embodiments, at least one of the interrupted RAN proteins contains at least one amino acid residue between each RAN repeat unit. In some embodiments, the poly-amino acid repeats (e.g., RAN repeat units) contained in the interrupted RAN proteins are separated by between about 2 and about 100 non-repeated amino acids. In some embodiments, at least one of the interrupted RAN proteins contains between 2 and 20 amino acid residues between each RAN repeat unit. In some embodiments, the poly-amino acid repeats contained in the interrupted RAN protein are about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, are separated by 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 non-repeated amino acids.
[0009] In some aspects, the present disclosure provides a method for identifying a subject as having a RAN proteinopathy based on the presence or absence of a disrupted RAN protein, as described herein. In some embodiments, the method includes detecting one or more disrupted RAN proteins in a biological sample obtained from the subject, each of the one or more disrupted RAN proteins comprising a plurality of RAN repeat units.
[0010] In some embodiments, the biological sample is tissue, blood, serum, or cerebrospinal fluid (CSF). In some embodiments, the tissue is brain tissue or spinal cord tissue.
[0011] Aspects of the present disclosure relate to methods of treating a subject having or suspected of having a RAN protein-related disease, the methods comprising administering one or more anti-RAN protein agents to the subject. In some embodiments, the subject is identified as having a RAN protein disease according to the methods for identifying a subject as having a RAN protein disease described herein. In some embodiments, the subject is identified as having a RAN protein disease according to the methods for identifying a subject as having a RAN protein disease described herein, and a therapeutic agent (e.g., an anti-RAN protein agent) is administered to the identified subject.
[0012] In some embodiments, the one or more anti-RAN protein agents target one or more interrupted RAN proteins. In some embodiments, the one or more anti-RAN protein agents reduce the transcription, translation, expression, aggregation, or accumulation of interrupted (e.g., discontinuous) RAN proteins comprising an interrupted poly-amino acid repeat motif (e.g., RAN repeat unit) relative to the level of transcription, translation, expression, aggregation, or accumulation of the RAN protein in the subject prior to administration of the one or more anti-RAN protein agents. In some embodiments, the one or more anti-RAN protein agents reduce the transcription of interrupted (e.g., discontinuous) RAN proteins comprising an interrupted poly-amino acid repeat motif (e.g., RAN repeat unit) relative to the level of transcription of the RAN protein in the subject prior to administration of the one or more anti-RAN protein agents. In some embodiments, the one or more anti-RAN protein agents reduce the translation of interrupted (e.g., discontinuous) RAN proteins comprising an interrupted poly-amino acid repeat motif (e.g., RAN repeat unit) relative to the level of translation of the RAN protein in the subject prior to administration of the one or more anti-RAN protein agents. In some embodiments, the one or more anti-RAN protein agents reduce expression of interrupted (e.g., discontinuous) RAN proteins comprising interrupted poly-amino acid repeat motifs (e.g., RAN repeat units) relative to expression in the subject prior to administration of the one or more anti-RAN protein agents. In some embodiments, the one or more anti-RAN protein agents reduce aggregation of interrupted (e.g., discontinuous) RAN proteins comprising interrupted poly-amino acid repeat motifs (e.g., RAN repeat units) relative to the level of aggregation of the RAN protein in the subject prior to administration of the one or more anti-RAN protein agents.In some embodiments, the one or more anti-RAN protein agents reduce the accumulation of interrupted (e.g., discontinuous) RAN protein containing an interrupted poly-amino acid repeat motif (e.g., RAN repeat unit) relative to the level of RAN protein accumulation in the subject prior to administration of the one or more anti-RAN protein agents.
[0013] In some embodiments, the RAN protein is an interrupted poly-glycine-alanine (poly-GA) or poly-glycine-arginine (poly-GR) repeat-containing RAN protein. The coding sequence for interrupted poly-GA and poly-GR RAN proteins can be found in a genome (e.g., the human genome) at one or more loci, including but not limited to, the loci and sequences set forth in Table 1 or Table 6. In some embodiments, a subject is characterized as having a mutation at one or more chromosomal loci or genes set forth in Table 1 or Table 6, wherein the mutation or mutations result in the translation of one or more interrupted poly-GA and / or poly(GR) RAN proteins. In some embodiments, the expanded repeat encoding the interrupted RAN protein is located in a protein-coding region (e.g., an exon region) of a gene. In some embodiments, the expanded repeat encoding the interrupted RAN protein is located in a non-coding region (e.g., an intron region, an untranslated region such as the 5'UTR or 3'UTR, etc.) of a gene. In some embodiments, the expanded repeat encoding the interrupted RAN protein is located within an intergenic region of a chromosome (eg, a nucleic acid sequence located between genes on a chromosome).
[0014] In some embodiments, one or more of the disrupted RAN proteins comprises a poly-GA-disrupted RAN protein. In some embodiments, at least one of the disrupted RAN proteins comprises (GGGGCT) x Expanded repeat, (GGGAGA) x Expanded repeat, or (GAAGGA) xIn some embodiments, one or more interrupted RAN proteins comprise a poly-GR interrupted RAN protein. In some embodiments, at least one of the interrupted RAN proteins comprises (GGGAGA) xTranslated from the expanded repeat. In some embodiments, x comprises an integer between 2 and 200. In some embodiments, x comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, x is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 8, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168 13, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, the poly-GA-interrupted RAN protein comprises between 10 and 50 (e.g., between 10% and 50%) GA repeat units over a 100 amino acid stretch.In some embodiments, the poly-GR-interrupted RAN protein comprises between 10 and 50 (eg, between 10% and 50%) GR repeat units over a 100 amino acid stretch.
[0015] In some embodiments, at least one of the interrupted RAN proteins is transcribed from a gene or chromosomal locus shown in Table 1. In some embodiments, at least one of the interrupted RAN proteins is transcribed from ARMCX4, ALK, and / or CASP8. In some embodiments, the ARMCX4, ALK, and / or CASP8 genes are transcribed to produce one or more interrupted RAN proteins. In some embodiments, the ARMCX4 gene is transcribed to produce one or more interrupted poly(GA)RAN proteins. In some embodiments, the ALK gene is transcribed to produce one or more interrupted poly(GA)RAN proteins. In some embodiments, the CASP8 gene is transcribed to produce one or more interrupted poly(GR)RAN proteins.
[0016] In some embodiments, the interrupted poly(GA) or interrupted poly(GR)RAN protein is translated from an mRNA transcript encoded by the ARMCX4, ALK, and / or CASP8 gene or their gene loci. In some embodiments, the interrupted poly(GA)RAN protein is translated from an mRNA transcript encoded by ARMCX4 or its gene loci. In some embodiments, the interrupted poly(GA)RAN protein is translated from an mRNA transcript encoded by ALK or its gene loci. In some embodiments, the interrupted poly(GR)RAN protein is translated from an mRNA transcript encoded by CASP8 or its gene loci. In some embodiments, the ARMCX4 gene is (GGGGCT) x Expanded repeat, (GGGAGA) x Expanded repeat, or (GAAGGA) xIn some embodiments, the ALK gene comprises an expanded repeat, where x represents the number of repeat units present. x Expanded repeat, (GGGAGA) x Expanded repeat, or (GAAGGA) x In some embodiments, the CASP8 gene comprises an expanded repeat, where x represents the number of repeat units present. x Expanded repeat, (GGGAGA) x Expanded repeat, or (GAAGGA) x It includes extended repeats, where x represents the number of repeat units present.
[0017] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human. In some embodiments, the subject is a non-human mammal, such as a mouse, rat, dog, cat, or pig.
[0018] In some embodiments, detecting one or more disrupted RAN proteins comprises performing an assay on the biological sample. In some embodiments, the assay comprises an antibody-based capture assay, a binding assay, a hybridization assay (e.g., Fluorescence In Situ Hybridization (FISH)), immunoblot analysis, Western blot analysis, immunohistochemistry, dCas9-based enrichment, label-free immunoassay, immunoquantitative PCR, mass spectrometry, bead-based immunoassay, immunoprecipitation, immunostaining, immunoelectrophoresis, and / or ELISA.
[0019] In some embodiments of the methods described by the present disclosure, a sample (e.g., a biological sample) is processed by an antibody-based capture process to isolate and / or detect one or more disrupted RAN proteins in the sample. Typically, the antibody-based capture method involves contacting the sample with one or more (e.g., two, three, four, five, or more) anti-RAN protein antibodies (e.g., anti-disrupted RAN protein antibodies). In some embodiments, the one or more anti-RAN antibodies are conjugated to a solid support (e.g., a scaffold, resin beads, etc.). In some embodiments, the antibody-based capture method involves physically separating and / or isolating the RAN protein bound by the anti-RAN antibody, for example, eluting the RAN protein by a chromatographic method such as affinity chromatography or ion exchange chromatography.
[0020] Detection of disrupted RAN proteins in biological samples may be performed by Western blotting. Western blotting generally involves the use of a detection agent or probe to identify the presence of a protein or peptide. In some embodiments, detection of one or more disrupted RAN proteins is performed by immunoblotting (e.g., dot blot, 2-D gel electrophoresis, Western blot, etc.), immunohistochemistry (IHC), ELISA (e.g., RCA-based ELISA or RT-PCR-based ELISA), label-free immunoassays such as surface plasmon resonance biolayer interferometry, immunoquantitative PCR, mass spectrometry (e.g., GC-MS, LC-MS, MALDI-TOF-MS), bead-based immunoassays, immunoprecipitation, immunostaining, or immunoelectrophoresis. In some embodiments, the detection agent is an antibody. In some embodiments, the antibody is an anti-RAN protein antibody (e.g., an anti-disrupted RAN protein antibody), such as an anti-poly(GA) antibody or an anti-poly(GR) antibody. In some embodiments, an anti-RAN protein antibody targets (e.g., specifically binds to) an amino acid repeat region of a RAN protein (e.g., GAGAGAGAGAGAGAGA (SEQ ID NO: 1)). In some embodiments, an anti-RAN protein antibody targets (e.g., specifically binds to) an epitope comprising amino acids at the characteristic reading frame-specific C-terminus translated 3' to the repeated amino acids. In some embodiments, an anti-RAN protein antibody targets (e.g., specifically binds to) an epitope comprising amino acids bridging the C-terminus of the amino acid repeat region and the N-terminus of the characteristic reading frame-specific C-terminus translated 3' to the repeated amino acids.
[0021] In some embodiments, the anti-RAN antibody targets (e.g., specifically binds to) a portion of a truncated RAN protein that does not contain a poly-amino acid repeat, such as the C-terminus of a truncated RAN protein (e.g., the C-terminus of a truncated poly(GA) or poly(GR) repeat protein). Examples of anti-RAN antibodies that target RAN protein poly-amino acid repeats are disclosed, for example, in International Application Publication No. WO2014 / 159247, the entire contents of which are incorporated herein by reference. Examples of anti-RAN antibodies that target the C-terminus of RAN protein are disclosed, for example, in U.S. Publication No. 2013 / 0115603, the entire contents of which are incorporated herein by reference.
[0022] In some embodiments, the RAN proteinopathy is amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), or spinocerebellar ataxia type 7 (SCA8). Spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), or Fuchs' corneal dystrophy.
[0023] In some embodiments, the method of identifying a subject as having a RAN protein disorder further comprises administering to the subject one or more anti-RAN protein agents. In some embodiments, the one or more anti-RAN protein agents comprise a protein, peptide, nucleic acid, or small molecule.
[0024] In some embodiments, the protein comprises an antibody. In some embodiments, the antibody is an anti-poly-GA antibody or an anti-poly-GR antibody. In some embodiments, the anti-poly-GA antibody specifically binds to the poly-GA repeat region of the RAN protein in the subject. In some embodiments, the anti-poly-GR antibody specifically binds to the poly-GR repeat region of the RAN protein in the subject.
[0025] Anti-RAN antibodies can be polyclonal or monoclonal. In some embodiments, anti-poly-GA or anti-poly-GR antibodies are polyclonal. Typically, polyclonal antibodies are produced by inoculation of a suitable mammal, such as a mouse, rabbit, or goat. Large mammals are often preferred because of the large amount of serum that can be recovered. By injecting an antigen into a mammal, B lymphocytes are induced to produce IgG immunoglobulins specific to the antigen. This polyclonal IgG is purified from the mammal's serum. In some embodiments, anti-poly-GA or anti-poly-GR antibodies are monoclonal. Monoclonal antibodies are generally produced from a single cell line (e.g., a hybridoma cell line). In some embodiments, anti-RAN antibodies are purified (e.g., isolated from serum). In some embodiments, the antigen is 12 to 20 amino acids. In the case of an antibody against a repeat motif, the antigen is a repeat sequence. In the case of an antibody against the C-terminal sequence of the RAN protein, the antigen is a C-terminal specific sequence. In some embodiments, the antigen is a portion of the C-terminal sequence, e.g., a fragment of the C-terminal sequence that is 3-5 or 5-10 or more amino acids in length, e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, or 50 amino acids in length.
[0026] In some embodiments, the present disclosure provides a method for producing an antibody, comprising administering to a subject a peptide antigen comprising a RAN protein repeat sequence, e.g., a poly(GA) or poly(GR) repeat sequence. In some embodiments, the subject is a mammal, e.g., a non-human primate, a rodent (e.g., a rat, a hamster, a guinea pig, etc.). In some embodiments, the subject is a human (e.g., a human subject is injected with a peptide antigen to elicit a host antibody response against the peptide antigen, e.g., a RAN protein). In some embodiments, the antibody is produced by expressing one or more interrupted RAN proteins or interrupted RAN protein repeat sequences in cells (e.g., B cells, hybridoma cells, etc.).
[0027] Numerous methods can be used to obtain anti-RAN antibodies. For example, antibodies can be produced using recombinant DNA methods. Monoclonal antibodies can also be produced by generating hybridomas according to known methods [see, e.g., Kohler and Milstein (1975) Nature, 256:495-499]. The hybridomas thus formed are then screened using standard methods, such as enzyme-linked immunosorbent assays (ELISAs; e.g., RCA-based ELISAs or rtPCR-based ELISAs) and surface plasmon resonance (e.g., OCTET or BIACORE) analysis, to identify one or more hybridomas that produce antibodies that specifically bind to the antigen. Any form of the antigen (e.g., truncated RAN protein), such as recombinant antigens, naturally occurring forms, or any mutant or fragment thereof, can be used as an immunogen. One exemplary method for generating antibodies involves screening a protein expression library, such as a phage display library or ribosome display library, that expresses antibodies or fragments thereof (e.g., scFvs). Phage display is described, for example, in U.S. Pat. No. 5,223,409 to Ladner et al., Smith (1985) Science 228:1315-1317, Clackson et al. (1991) Nature, 352:624-628, Marks et al. (1991) J. Mol. Biol., 222:581-597, WO92 / 18619, WO91 / 17271, WO92 / 20791, WO92 / 15679, WO93 / 01288, WO92 / 01047, WO92 / 09690, and WO90 / 02809.
[0028] In addition to using a display library, the antigens described above (e.g., one or more disrupted RAN proteins) can be used to immunize a non-human animal, such as a rodent, e.g., a mouse, hamster, or rat. In one embodiment, the non-human animal is a mouse.
[0029] In another embodiment, monoclonal antibodies are obtained from non-human animals and then modified, e.g., chimerized, using recombinant DNA techniques known in the art. Various techniques for producing chimeric antibodies have been described. See, for example, Morrison et al., Proc. Natl. Acad. Sci. USA 81:6851, 1985; Takeda et al., Nature 314:452, 1985; U.S. Patent No. 4,816,567 to Cabilly et al.; U.S. Patent No. 4,816,397 to Boss et al.; European Patent Publication Nos. EP 171496 and 0173494 to Tanaguchi et al.; and British Patent No. GB 2177096B.
[0030] Antibodies can be humanized by methods known in the art. For example, monoclonal antibodies with desired binding specificity can be commercially humanized (Scotgene, Scotland and Oxford Molecular, Palo Alto, Calif.). Fully humanized antibodies, such as those expressed in transgenic animals, are within the scope of the present invention (see, for example, Green et al. (1994) Nature Genetics 7, 13, and U.S. Patent Nos. 5,545,806 and 5,569,825).
[0031] For additional antibody production techniques, see Antibodies: A Laboratory Manual, Second Edition. Edited by Edward A. Greenfield, Dana-Farber Cancer Institute, ©2014. This disclosure is not necessarily limited to a particular source, method of production, or other particular characteristics of the antibodies.
[0032] In some embodiments, the one or more anti-RAN protein agents comprise a nucleic acid. In some embodiments, the nucleic acid is an inhibitory nucleic acid. In some embodiments, the inhibitory nucleic acid is an interfering RNA. In some embodiments, the nucleic acid is a double-stranded RNA (dsRNA), a small interfering RNA (siRNA), a short hairpin RNA (shRNA), a microRNA (miRNA), an artificial microRNA (amiRNA), an aptamer, or an antisense oligonucleotide (ASO). In some embodiments, the inhibitory nucleic acid is a nucleic acid aptamer (e.g., an RNA aptamer or a DNA aptamer). Generally, inhibitory RNA molecules can be unmodified or modified. Because phosphorothioate modifications, 2'-O-methyl modifications, and the like are recognized in the art to improve the stability of oligonucleotides in vivo, in some embodiments, the inhibitory RNA molecule comprises one or more modified oligonucleotides, such as oligonucleotides with the aforementioned modifications.
[0033] In some embodiments, the nucleic acid comprises a region of complementarity to a nucleic acid sequence encoding a poly-GA or poly-GR repeat expansion in a subject. The region of complementarity can comprise between 2 and 50 nucleotides. In some embodiments, the region of complementarity comprises between 2 and 10, between 2 and 15, between 5 and 20, or between 10 and 30 nucleotides in length. In some embodiments, the nucleic acid comprises a region of complementarity to a nucleic acid sequence present at a chromosomal locus or gene set forth in Table 1 or Table 6. In some embodiments, the inhibitory nucleic acid is selected from the group consisting of PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8, and the like. ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, AC TA1, DNAH14, PFN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A3, GP R137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDIN2, AIFM2, MGMT, LINC01164, KNDC1, ANK3, MLL T10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX4L2, AGAP12P, C10orf53, SMPD1, IFITM10, BUD13, TSPAN18 , CD82, OTUB1, NADSYN1, MIR4492, CHID1, SMUG1, LINC00938, LINC01257, SLC15A4, ASIC1, DCP1B, TMT C2, TNS2, LOC100240735, SOX1, LATS2, RAB20, ANKRD20A9P, FLT1, RCBTB1, ELK2AP, STON2, FOXN3, TTLL 5, BCL11B, BRMS1L, SMAD3, RPLP1, BAHD1, MYO5A, DNM1P46, DNM1P35, TUBGCP4, C16orf95, OSGIN1, LIN C00311, MIR4718, RBFOX1, SBK1, MIR4722, BANP, C16orf78, MIR5189, ADGRG5, NPRL3, ZDHHC1, MIR662,LINC00482, MRPL12, TBC1D3H, WSCD1, TBC1D3B, TBC1D3, TBC1D3C, METRNL, D NAH9, ASGR1, FOXK2, NPEPPS, SARM1, CLUH, TIAF1, LOC440434, PHOSPHO1, TB CD, CYP4F35P, CXADRP3, LINC00668, MEX3C, COX7A1, SCAF1, RFPL4AL1, SIX5 DIRAS1, POLRMT, ZNF554, MED16, SIPA1L3, DOT1L, KMT5C, PDE4C, ZNF480, CE BPA, PTPN18, HAAO, LOC654342, RGPD2, RAB3GAP1, TNS1, FAM95A, NTSR1, FRG 1BP, FAM182B, CDH4, PRNP, MIR1257, MIR4758, OGFR, SRC, COL9A3, ZNF512B, P.S ICSAR, SIK1, CYP4F29P, MX1, LARGE1, CRELD2, UPK3A, RRP7A, MIR4762, SHAN K3, SHISA8, CCDC188, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CAC NA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10, TNIP 3. ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2 SLC34A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, ME A1, RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, A BCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051.L RRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf5 0, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA( These include MRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2.The inhibitory nucleic acid comprises a region complementary to an RNA transcript encoded by a gene or gene locus selected from the group consisting of FAM226A, ALK, and CASP8. In some embodiments, the inhibitory nucleic acid comprises a region complementary to an RNA transcript encoded by a chromosomal locus beginning at 100748986 and ending at 100749205 in chrX. In some embodiments, the inhibitory nucleic acid comprises a region complementary to an RNA transcript encoded by ARMCX4. In some embodiments, the inhibitory nucleic acid comprises a region complementary to an RNA transcript encoded by ALK. In some embodiments, the inhibitory nucleic acid comprises a region complementary to an RNA transcript encoded by CASP8.
[0034] In some embodiments, the one or more anti-RAN protein agents comprise a small molecule. In some embodiments, the small molecule inhibits the expression or activity of one or more RAN proteins. In some embodiments, the small molecule is an inhibitor of eIF3 (or an eIF3 subunit). Examples of small molecule inhibitors of eIF3 include, but are not limited to, mTOR inhibitors (e.g., rapamycin, PP242), S6 kinase (S6K) inhibitors, and the like. In some embodiments, the small molecule inhibits the expression or activity of eukaryotic initiation factor 2A (eIF2A) or eIF2α. Examples of small molecule inhibitors of eIF2A include, but are not limited to, salubrinal, Sal003, ISRIB, and the like. In some embodiments, the small molecule is an inhibitor of TARBP2. Examples of TARBP2 inhibitors include anti-TARBP2 antibodies, interfering RNAs (e.g., dsRNA, siRNA, shRNA, miRNA, etc.) targeting TARBP2, peptide inhibitors of TARBP2, and small molecule inhibitors of TARBP2. In some embodiments, the small molecule is metformin, also known as N,N-dimethylbiguanide (IUPAC name N,N-dimethylimidodicarbonimidic diamide and CAS 657-24-9), or chloroguanide [1-[amino-(4-chloroanilino)methylidene]-2-propan-2-yl-guanidine, CAS 500-92-5], chlorproguanil [1-[amino-(3,4-dichloroanilino)methylidene]-2-propan-2-yl guanidine, CAS 537-21-3], buformin [N-butylimidodicarbonimidic diamide, CAS 692-13-7], or phenformin [2-(N-phenethylcarbamimidoyl)guanidine, CAS 114-86-3], or a pharmaceutically acceptable salt, co-crystal, tautomer, stereoisomer, solvate, hydrate, polymorph, isotopically enriched derivative, or prodrug of any of the biguanides. In some embodiments, the small molecule is metformin.
[0035] The accompanying drawings are not intended to be drawn to scale. In the drawings, identical or nearly identical components shown in various figures are each represented by a like numeral. For purposes of clarity, not every component is shown in every figure. [Brief explanation of the drawings]
[0036] [Figure 1A] Figures 1A and 1B show predicted ARMCX4 poly-GA repeat proteins. Figure 1A shows alternative splicing variants, indicating that the ARMCX4 repeat expansion can be intronic (left) or exonic (right). Figure 1B shows predicted GA-rich proteins produced by ARMCX4 repeat expansion. The example amino acid sequence shown in Figure 1B (SEQ ID NO: 4) shows an interrupted GA repeat motif. [Figure 1B] Same as above. [Figure 2A] Figures 2A and 2B show predicted ALK poly-GA repeat proteins. Figure 2A shows the predicted coding region. The expanded allele has an expanded GGA repeat, containing approximately 143-156 repeats. Figure 2B shows the predicted GA-rich protein produced by ALK repeat expansion. The example amino acid sequence shown in Figure 2B (SEQ ID NO: 7) shows the interrupted GA repeat motif. [Figure 2B] Same as above. [Figure 3A] Figures 3A and 3B show representative data demonstrating that anti-GA antibodies recognize GA-rich proteins expressed from ARMCX4 and ALK repeat expansions. Figure 3A shows examples of designed and cloned ARMCX4-RE and ALK-RE plasmids. The FLAG tag is expressed in frame with the ARMCX4 and ALK GA-rich proteins. Figure 3B shows data demonstrating that anti-GA antibody staining colocalizes with the FLAG tag in ARMCX4-RE-overexpressing HEK293T cells and ALK-RE-overexpressing HEK293T cells. [Figure 3B] Same as above. [Figure 4A] Figures 4A-4C show predicted CASP8 poly-GR repeat proteins. Figure 4A shows the predicted coding region. Figure 4B shows a predicted GR-rich protein produced by CASP8 repeat expansion. The example amino acid sequence shown in Figure 4B (SEQ ID NO: 8) shows an interrupted GR repeat motif. Figure 4C shows a predicted GR-rich protein produced by CASP8 repeat expansion. The example amino acid sequence shown in Figure 4C (SEQ ID NO: 9) shows an interrupted GR repeat motif. [Figure 4B] Same as above. [Figure 4C] Same as above. [Figure 5A]Figures 5A-5I show representative data demonstrating distinct poly-GR accumulation correlated with p-Tau levels in AD- and tauopathy-related autopsy brains. Figure 5A shows examples of immunohistochemical (IHC) staining (red) of poly-GR detected by rabbit polyclonal α-poly-GR antibody in hippocampal sections (HC) from AD and control cases. Figure 5B shows a diagram illustrating the location of poly-GR aggregates (red dots) found in hippocampal sections from AD cases. HC = hippocampus, erc = entorhinal cortex. Figure 5C shows quantification of poly-GR aggregates in the hippocampus (shown by the red box in Figure 5B) from AD (n = 80) and control (n = 18) cases. Figure 5D shows an example of dot blot analysis of poly-GR detected by rat monoclonal α-poly-GR antibody in protein extracts of frozen frontal cortex from AD (n = 65) and control (n = 20) cases. Figure 5E shows quantification of polyGR levels in frontal cortex protein lysates from AD and control cases, as determined by dot blot analysis. Figure 5F shows an example of IHC staining analysis for polyGR, p-tau (S202 and T205) (detected with AT8 antibody), Aβ plaques, and p-TDP43 in the cornu ammonis (CA) and dentate gyrus (DG) of the hippocampus. Positive staining is shown in red; scale bar = 20 μm. Figure 5G shows double IHC staining analysis showing polyGR (pink) detected in brain regions containing both high and low p-tau (S202 and T205) (brown) subregions in the same AD brain section. Black arrows indicate cells showing both polyGR and p-tau signals; white arrows indicate cells showing polyGR staining. Figure 5H shows plots of polyGR and p-tau (S202 and T205) staining detected in consecutive slides from 21 randomly selected AD cases. Data represent the mean ± SEM. Two-tailed unpaired t-test. ****p<0.0001. Figure 5I shows α-polyGR staining in autopsy brain tissue from patients with tauopathy-related disorders. IHC staining was performed using tissue samples from patients with the dominant pathological signatures of progressive supranuclear palsy (PSP, n=10), Pick's disease (Pick's, n=10), and dementia with Lewy bodies (LBD, n=9), as well as AD (n=3) as a positive control, and control cases without dementia pathology (n=8).Several cases had mixed pathologies, including case 9 (C9-LBD / PD / AD), which had a mixed pathology of LBD, Parkinson's disease (PD), and AD. This figure includes representative staining images from all cases with polyGR staining and several cases with negative polyGR staining. All remaining cases (not shown) were negative for polyGR staining. Positive polyGR staining was detected in some PSP, Pick's disease, and LBD cases, which accumulated in different patterns compared with C6-AD, C7-AD, and C8-AD. The two non-AD cases that showed strong polyGR staining were C14-LBD and C19-Pick's disease. In C9-LBD / PD / AD, C12-LBD / AD, C16-Pick's disease, C18-Pick's disease, and C24-PSP, polyGR staining was detected in only a few cells and / or was very faint. [Figure 5B] Same as above. [Figure 5C] Same as above. [Figure 5D] Same as above. [Figure 5E] Same as above. [Figure 5F] Same as above. [Figure 5G] Same as above. [Figure 5H] Same as above. [Figure 5I] Same as above. [Figure 6A]Figures 6A-6F show an example of a CRISPR-inactivated Cas9-based repeat enrichment and detection (dCas9READ) strategy for pulling down repeat expansion mutations. Figure 6A is a schematic diagram illustrating how the dCas9READ method works. Figure 6B shows an example of primers used in a qPCR assay to measure the levels of C9orf72 flanking sequences in dCas9READ-enriched DNA samples. Figure 6C shows quantification of C9orf72 flanking sequence levels in dCas9READ-enriched C9 GGGGCCexp(+)[C9(+)] (n=4) compared to C9 GGGGCCexp(-)[C9(-)] (n=3). Figure 6D shows an enrichment plot of the C9orf72 GGGGCC flanking sequence in control assays with and without the G4C2 sgRNA (n=3). Figure 6E shows Illumina short-read sequencing read mapping and total read counts at the C9orf72 G4C2 locus for dCas9-enriched C9 (n=4) and control (n=4) samples, along with quantification of the enriched locus and flanking regions. Figure 6F shows Illumina short-read sequencing read mapping for C9 and control samples enriched with a mixture of 8 or 24 sgRNAs containing a GR-encoding repeat motif. (E, F) Scale bars indicate number of reads. Data represent mean ± SEM. Unpaired two-tailed t-test. **p<0.01, ***p<0.001. [Figure 6B] Same as above. [Figure 6C] Same as above. [Figure 6D] Same as above. [Figure 6E] Same as above. [Figure 6F] Same as above. [Figure 7A]Figures 7A-7F show the detection of the GGGAGA·TCTCCC (SEQ ID NO: 10) repeat expansion mutation using dCas9READ. Figure 7A shows an example of a method for identifying novel repeat expansions using the polyGR protein signature and dCas9READ and determining their association with AD. Figure 7B shows a diagram of the genomic location of the CASP8 GGGAGA repeat expansion within the SVA-E retrotransposon element, the reference repeat sequence, and the repeat-primed PCR (RP-PCR) primers (SEQ ID NOs: 11 and 12) used to characterize the repeat expansion. Figure 7C shows an example of dCas9READ enrichment folds at 10 loci using genomic DNA from a polyGR(-) control (Cntl), a polyGR(-) AD (AD#1), and five polyGR(+) AD cases. Figure 7D shows mapping of Illumina short-read sequencing reads at the CASP8 locus, demonstrating enhanced enrichment in AD#2 and AD#3 cases. The red arrow indicates the location of the repeat expansion in CASP8. Figure 7E shows RP-PCR data showing two positive repeat expansion patterns [approximately 64 repeats (SEQ ID NO: 13) and 44 repeats (SEQ ID NO: 14)] for the (GAGAGG)2GAGACG (SEQ ID NO: 15) repeat primer at the CASP8 locus. Figure 7F shows the percentage of long [approximately 64 repeats (SEQ ID NO: 13)], medium [approximately 44 repeats (SEQ ID NO: 14)] CASP8 GGGAGA repeats, and non-expanded CASP8 GGGAGA repeats in AD and control populations. [Figure 7B] Same as above. [Figure 7C] Same as above. [Figure 7D] Same as above. [Figure 7E] Same as above. [Figure 7F] Same as above. [Figure 8A]Figures 8A-8E show representative data on elevated levels of cleaved caspase-8 and the accumulation of extended proteins expressed from CASP8 GGGAGAexp in AD autopsy brain tissue. Figure 8A shows cleaved caspase-8 detected in frontal cortex tissue from CASP8-GGGAGAexp(+) AD, CASP8-GGGAGAexp(-) AD, and non-AD control cases (no AD pathology). Figure 8B shows quantification of cleaved caspase-8 in Figure 8A. Figure 8C shows a diagram of extended proteins translated from sense and antisense transcripts from the CASP8 GGGAGAexp locus. The amino acid sequences shown in red were used to generate frame-specific C-terminal (CT) antibodies. S is sense, AS is antisense, f1-3 are reading frames 1-3, and * indicates a stop codon. Figure 8D shows IHC detection of CASP8 RAN Sf3 protein aggregate staining (red) in the hippocampus of a CASP8-GGGAGAexp(+) AD patient detected by α-CT-f3S antibody. Figure 8E shows dual IF analysis of colocalization of polyGR (red) and α-CT-f3S (green) staining in the frontal cortex of a CASP8-GGGAGAexp(+) AD patient. Data represent mean ± SEM. One-way ANOVA Holm-Sidak multiple comparison test. **p<0.01. [Figure 8B] Same as above. [Figure 8C] Same as above. [Figure 8D] Same as above. [Figure 8E] Same as above. [Figure 9A]Figures 9A-9J show the effects of stress on CASP8 RAN translation, as well as the effects of CASP8 GGGAGA and polyGR on cellular and tau phosphorylation. Figure 9A shows a representative protein blot of FLAG-tagged RAN protein expressed from the 6XStop-CASP8-RE-3T minigene in HEK293T cells. The 6XStop-CASP8-RE-3T minigenes were: CASP8-hi64-3T [highly interrupted with 64 GGGAGA repeats (SEQ ID NO: 13)], CASP8-i44-3T [interrupted with 44 GGGAGA repeats (SEQ ID NO: 14)], and CASP8-i64-3T [interrupted with 64 GGGAGA repeats (SEQ ID NO: 13)]. Figure 9B shows the effects of thapsigargin (Tg) and metformin (M) on CASP8 RAN protein expression (n = 3-4). Figure 9C shows the effect of the CASP8 GGGAGAexp minigene on T98 cell survival and viability, as measured by LDH and MTT assays (n = 4), respectively. Figure 9D shows increased tau phosphorylation (S202 and T205) in polyGR-overexpressing SH-SY5Y cells overexpressing FLAG-GR60 using alternative codon minigenes. Figure 9E shows quantification of p-tau in polyGR(+) (n = 46) and polyGR(-) SH-SY5Y cells (n = 287). Figure 9F shows a diagram of the RAN repeat unit and corresponding flanking sequences containing an insertion in the VNTR sequence identified in CASP8, which is associated with a high risk of AD. This variant is associated with an increased risk of AD with an odds ratio of 2.3 (p = 0.0001). This particular repeat configuration can be used for diagnostic screening and therapeutic targeting using ASOs, RNAi, CRISPR / Cas, and other methods. Figure 9G shows a schematic of the assay used to analyze poly-GR aggregates in HEK293T cells transfected with CASP8 SVA cloned from an AD or control case. Figure 9H shows representative confocal images showing poly-GR staining (red) in cells transfected with AD SVA or control SVA (Cntl SVA) plasmid.Figure 9I shows quantification of polyGR detected in HEK293T cells transfected with AD SVA (n=7) and control SVA (n=3) plasmids. Figure 9J shows a schematic representation illustrating the contribution of CASP8 and other repeat expansion mutations in the pathogenesis of AD and pathogenic tau phosphorylation. Data represent mean ± SEM. Figures 9B-9C: One-way ANOVA Holm-Sidak multiple comparison test and unpaired two-tailed t-test (E). *p<0.05, **p<0.01, ***p<0.001, ****p<0.0001. Figure 9I: Data represent mean ± SEM. Unpaired two-tailed t-test **p<0.01. [Figure 9B] Same as above. [Figure 9C] Same as above. [Figure 9D] Same as above. [Figure 9E] Same as above. [Figure 9F] Same as above. [Figure 9G] Same as above. [Figure 9H] Same as above. [Figure 9I] Same as above. [Figure 9J] Same as above. [Figure 10] FIG. 10 shows repeat-primer PCR analysis of the Casp8 repeat gene sequence encoding a poly(GR)-containing interrupted dipeptide, performed using primers containing (GGGAGA) 3 (SEQ ID NO: 23). [Figure 11A]Figures 11A-11C show the pathology of poly-GR protein in autopsy brain tissue from AD patients. Figure 11A shows an example of wide-field microscopy analysis of IHC staining of poly-GR detected by rabbit polyclonal α-poly-GR antibody in AD sections of the hippocampus (HC) but not in control sections. Figure 11B shows an example of a 3xFLAG-(GR)60 construct and immunofluorescence analysis using rabbit polyclonal α-poly-GR antibody to stain T98 cells overexpressing the 3xFLAG-(GR)60 plasmid. Figure 11C shows a dot blot analysis of poly-GR detected by rat monoclonal poly-GR antibody. Total protein control was detected using the LICOR Revert™ 700 Total Protein Staining Kit. [Figure 11B] Same as above. [Figure 11C] Same as above. [Figure 12A] Figures 12A-12D show an example of a CRISPR-inactive Cas9-based repeat enrichment and detection (dCas9READ) strategy for pull-down of repeat expansion mutations. Figure 12A shows a qPCR assay for measuring the levels of the CNBP CCTG flanking sequence in dCas9READ-enriched DNA samples and a diagram of the quantification of CNBP CCTG flanking sequence levels in dCas9READ-enriched DM2 samples (n = 4) compared to DM2(-) controls (n = 3). Figure 12B shows Illumina short-read sequencing read mapping and total repeat counts at the CNBP CCTG locus in dCas9-enriched DM2 (n = 4) and control samples (n = 4). Figure 12C shows a representative fragment analysis plot showing the molecular weight range of DNA fragments enriched in a dCas9READ pull-down assay using 24 GR-coding repeat sgRNAs. LM: low abundance marker. Figure 12D shows a summary graph showing the proportion of enriched GR repeat loci in intergenic, intronic, and exonic regions. Data represent mean ± SEM. Unpaired two-tailed t-test. **p<0.01, ****p<0.0001. [Figure 12B] Same as above. [Figure 12C]Same as above. [Figure 12D] Same as above. [Figure 13A-1] Figures 13A-13E show GGGAGA·TCTCCC (SEQ ID NO: 10) repeat expansions within SVA retrotransposon elements detected by dCas9READ. Figure 13A shows repeat-primed PCR (RP-PCR) analysis of a homozygous case, showing an example of a negative repeat expansion pattern for the (GAGAGG)4 (SEQ ID NO: 24) repeat primer (left) and an example of a biallelic CASP8 GGGAGAexp repeat pattern for the (GAGAGG)2GAGACG (SEQ ID NO: 15) repeat primer (right). Figure 13B shows the genomic locations and RP-PCR patterns of five repeat expansion loci for samples with (+) or without (-) repeat expansions at specific loci. Figure 13C shows RP-PCR analysis of repeat patterns for genomic DNA (gDNA) samples extracted from blood monocytes or frontal cortex tissue from the same individual. Figure 13D shows the long-range PCR (LR-PCR) primers used to amplify CASP8 GGGAGA. Figure 13E shows EtBR gel analysis of LR-PCR products using gDNA extracted from blood monocytes or frontal cortex tissue from three different cases. Cases A and B have approximately 70 CASP8-GGGAGA repeats. There is no amplification product for case C, indicating that this case does not have an SVA insertion at the CASP8 locus. Normal allele size = 586 bp [approximately 10 GGGAGA repeats (SEQ ID NO: 25)]. [Figure 13A-2] Same as above. [Figure 13B-1] Same as above. [Figure 13B-2] Same as above. [Figure 13C] Same as above. [Figure 13D] Same as above. [Figure 13E] Same as above. [Figure 14A]Figures 14A-14D show characterization of CASP8 GGGAGAexp by long-read sequencing. Figure 14A shows an example of the method for long-read sequencing of dCas9-enriched DNA samples. Figure 14B shows mapping of PacBio long-read sequencing reads at the CASP8 GGGAGA repeat locus in polyGR(+) AD groups 1 and 2. Purple lines or bars indicate insertions, and orange / red / green / blue lines indicate single nucleotide polymorphisms (SNPs). Figure 14C shows the percentage of interruptions in CASP8 GGGAGAexp in 18 long-read sequencing reads from two cognitively normal controls and 31 long-read sequencing reads from seven AD cases. Figure 14D shows three representative repeat expansion sequences (RE1, RE2, and RE3) of CASP8 GGGAGAexp detected by long-read sequencing. Interruptions within the repeat region are shown in red and bold, and flanking sequences are bold. [Figure 14B] Same as above. [Figure 14C] Same as above. [Figure 14D] Same as above. [Figure 15A]Figures 15A-15F show analysis of CASP8 RNA transcript levels in CASP8-GGGAGAexp(+) and CASP8-GGGAGAexp(-) AD cases. Figure 15A shows an example of a qRT-PCR strategy for detecting CASP8 exons 7-8 and exon 9. Figure 15B shows the levels of CASP8 exons 7-8 in frontal cortex tissue from CASP8-GGGAGAexp(+) (n=3), CASP8-GGGAGAexp(-) (n=3-4), and normal control cases (n=2). Figure 15C shows the levels of CASP8 exon 9 in frontal cortex tissue from CASP8-GGGAGAexp(+) (n=3), CASP8-GGGAGAexp(-) (n=3-4), and normal control cases (n=2). Data represent mean ± SEM. One-way ANOVA with Holm-Sidak multiple comparison test. ns p>0.05. Figure 15D shows Western blots depicting caspase-8 levels in CASP8-GGGAGAexp(+)AD, CASP8-GGGAGAexp(-)AD, and controls without AD pathology. Figure 15E shows plots depicting full-length caspase-8 levels relative to actin. Data represent mean ± SEM. One-way ANOVA Holm-Sidak multiple comparison test. ns p>0.05, *p<0.05, **p<0.01. Figure 15F shows plots depicting cleaved caspase-8 levels relative to actin. Data represent mean ± SEM. One-way ANOVA Holm-Sidak multiple comparison test. ns p>0.05, *p<0.05, **p<0.01. [Figure 15B] Same as above. [Figure 15C] Same as above. [Figure 15D] Same as above. [Figure 15E] Same as above. [Figure 15F] Same as above. [Figure 16A]Figures 16A-16I show RNA translation from CASP8 GGGAGAexp in cultured cells and characterization with a C-terminal antibody. Figure 16A shows an example construct (6XStop-CASP8-RE-3T: CASP8-hi64-3T, CASP8-i44-3T, CASP8-i64-3T) containing 100 bp of upstream flanking sequence, three CASP8 repeat expansions, and three tag epitopes corresponding to three reading frames. The upstream flanking sequence contained an ATG codon in frame with the HA tag. Figure 16B shows protein blot analysis of mutant proteins from the CASP8 GGGAGAexp minigene detected in frame with FLAG (upper blot) and HA(AUG) (lower blot) in HEK293T cells transfected with the 6XStop-CASP8-RE-3T plasmid. Figure 16C shows protein blot analysis of mutant proteins from the CASP8 GGGAGAexp minigene detected as expressed from the FLAG frame in HEK293T cells transfected with the 6XStop-CASP8-RE-3T plasmid and treated with thapsigargin and / or metformin. Figure 16D shows quantification from protein blot analysis of mutant proteins from the CASP8 GGGAGAexp minigene detected as expressed from the FLAG or HA frame in HEK293T cells transfected with the 6XStop-CASP8-RE-3T plasmid and treated with thapsigargin and / or metformin. Figure 16E shows LDH and MTT assay analysis of T98 cell viability as a result of transfection and expression of disrupted RAN protein from the 6XStop-CASP8-RE-3T plasmid. Figure 16F shows immunofluorescence (IF) analysis of mutant proteins expressed in all three reading frames [FLAG, HA(AUG), and Myc] in HEK293T cells transfected with the 6XStop-CASP8-RE-3T plasmid in all repeat expansion configurations.Figure 16G shows IF analysis using the α-CT-f1S C-terminal antibody, which recognizes a unique C-terminal sequence in the extended protein expressed from the CASP8 GGGAGAexp locus. Figure 16H shows IF analysis using the α-CT-f3S C-terminal antibody, which recognizes a unique C-terminal sequence in the extended protein expressed from the CASP8 GGGAGAexp locus. Figure 16I shows that the α-polyGR antibody recognizes a CASP8 chimeric extension protein containing a polyGR tract. The top panel shows a minigene (AUG-FLAG-hi64 / i44 / i64) containing 100 bp of upstream flanking sequence and three repeat extensions in the CASP8 repeat extension configuration, expressing an AUG-FLAG-tagged CASP8 chimeric polymer protein. The bottom panel shows representative confocal images from a double IF experiment, demonstrating that the α-polyGR antibody signal colocalizes with FLAG tag staining in SH-SY5Y transfected with the AUG-FLAG-hi64 / i44 / i64 plasmid. [Figure 16B] Same as above. [Figure 16C] Same as above. [Figure 16D] Same as above. [Figure 16E] Same as above. [Figure 16F] Same as above. [Figure 16G] Same as above. [Figure 16H] Same as above. [Figure 16I] Same as above. [Figure 17A]Figures 17A-17C show IHC staining analysis using C-terminal CASP8 locus-specific antibodies, α-CT-F3S and α-CT-F1S, in the hippocampus and frontal cortex regions of AD and control cases. Figure 17A shows an example of an IHC image showing staining (red) of CASP8-RAN-Sf3 protein condensates detected in hippocampal tissue from a CASP8 GGGAGAexp(+) AD case but not in CASP8 GGGAGAexp(-) AD or controls (without Alzheimer's disease). Figure 17B shows an example of IHC detection of CASP8-RAN-Sf3 protein condensates in the gray and white matter regions of the frontal cortex from a CASP8 GGGAGAexp(+) AD case. Figure 17C shows IHC detection of CASP8-RAN-Sf1 protein condensates in the gray matter region of the frontal cortex from a CASP8 GGGAGAexp(+) AD case. [Figure 17B] Same as above. [Figure 17C] Same as above. [Figure 18A] Figures 18A-18B show IHC staining analysis using α-CT-f3S antibody in the hippocampus (HC) of AD and control cases. Figure 18A shows IHC analysis of α-CT-f3S staining in the CA, Sub, and DG in AD#2 in addition to staining of the pre-bleed control. Figure 18B shows additional α-CT-f3S IHC analysis of hippocampal sections from a CASP8 GGGAGAexp(+) control (without Alzheimer's disease). [Figure 18B] Same as above. [Figure 19] Figure 19 shows colocalization analysis of poly-GR and α-CT-f3S antibody staining detected by immunofluorescence. Widefield images show that poly-GR staining partially colocalized with α-CT-f3S antibody staining in the frontal cortex of CASP8 GGGAGAexp(+) and poly-GR(+) AD cases, but not in CASP8 GGGAGAexp(-) AD cases. [Figure 20A]Figures 20A-20E show the effects of stress on the translation of CASP8 GGGAGAexp and the toxicity of the CASP8 GGGAGA repeat expansion in cells. Figure 20A shows an example of Western blot analysis of mutant proteins in the HA frame expressed from the 6XStop-CASP8-RE-3T (CASP8-i44-3T and CASP8-i64-3T) plasmid in HEK293T cells with or without thapsigargin (Tg) and with or without metformin (M) treatment (n = 3-4). Figure 20B shows quantification of the data in Figure 20A. Figure 20C shows qRT-PCR analysis of transcript levels of the 6XStop-CASP8-RE-3T plasmid in transfected cells treated with Tg or Tg and metformin (n = 3). Figure 20D shows the toxicity of 6XStop-CASP8-RE-3T minigene in SH-SY5Y cells (n=4). Figure 20E shows the toxicity of 6XStop-CASP8-RE-3T minigene in HEK293T cells (n=4). Cell viability was measured by LDH assay. Data represent mean ± SEM. One-way ANOVA Holm-Sidak multiple comparison test. ns p>0.05, *p<0.05, **p<0.01, ***p<0.001. [Figure 20B] Same as above. [Figure 20C] Same as above. [Figure 20D] Same as above. [Figure 20E] Same as above. [Figure 21A]Figures 21A-21B show analysis of pTau protein in SH-SY5Y cells transfected with FLAG-GR60 or 6XStop-CASP8-RE-3T minigenes. Figure 21A shows IF analysis of endogenous pTau levels at S202 and T205 in polyGR(+) SH-SY5Y cells compared with polyGR(-) SH-SY5Y cells transfected with the FLAG-GR60 minigene containing an alternative codon DNA sequence expressing 3xFLAG-GR60 protein starting at AUG. Figure 21B shows an example of IF analysis of pTau levels at S202 and T205 in SH-SY5Y cells expressing FLAG / Myc CASP8 RAN polymer protein compared with control cells. pTau (S202 and T205) was detected using the AT8 antibody. [Figure 21B] Same as above. [Figure 22] FIG. 22 shows IF co-localization analysis of polyGR and human tau protein in HEK293T cells. [Figure 23A] Figures 23A-23B show long-range PCR of CASP8 GGGAGAexp. Figure 23A shows a schematic diagram showing primers for long-range PCR (LR-PCR) experiments, where the reverse primer is conjugated with a FAM fluorophore to enable downstream fragment analysis. Figure 23B shows fragment analysis of LR-PCR products of CASP8 GGGAGAexp. Case AD#1 showed two extended alleles, with LR-PCR product sizes of 919 bp and 946 bp. Case AD#2 showed a single extended allele, with a product size of 946 bp. Case AD#3 showed no extended allele. BG: background peak. [Figure 23B] Same as above. DETAILED DESCRIPTION OF THE INVENTION
[0037] A "RAN protein (e.g., a repeat-associated non-ATG translated protein)" refers to a polypeptide translated from an mRNA sequence containing a repeat sequence in the absence of an AUG start codon. The repeat sequence in a nucleic acid encoding a RAN protein may be referred to as a "microsatellite repeat" or "extended repeat." Translation of an extended repeat may produce a RAN protein containing a repeat sequence of a single amino acid, a diamino acid, a triamino acid, a quadamino acid, a pentaamino acid, a hexaamino acid, a heptaamino acid, an octaamino acid, a nonaamino acid, or a decaamino acid, which may be referred to as a "polyamino acid repeat."
[0038] Non-limiting examples of expanded repeats found in nucleic acids encoding RAN proteins include those encoding poly(GR)RAN proteins [e.g., poly(GGTCGT), poly(GGCCGT), poly(GGACGT), poly(GGGCGT), poly(GGTCGC), poly(GGCCGC), poly(GGACGC), poly(GGGCGC), poly(GGTCGA), poly(GGCCGA), poly(GGACGA), poly(GGGCGA), poly(GGTCGG), poly(GGCCGG), poly(GGACGG), poly(GGGCGG), poly(GGT AGA), poly(GGCAGA), poly(GGAAGA), poly(GGGAGA), poly(GGTAGG), poly(GGCAGG), poly(GGAAGG), and / or poly(GGGAGG) repeats], those encoding poly(GA)RAN proteins [e.g., (GGTGCT), poly(GGCGCC), poly(GGAGCA), poly(GGGGCG), poly(GGTGCC), poly(GGCGCA), poly(GGAGCG), poly(GGGGCT), poly(GGTGCA), poly(GGCGCG), poly(GGAGCT), poly(GGGGCC) , poly(GGTGCG), poly(GGCGCT), poly(GGAGCC), and / or poly(GGGGCA) repeats], those encoding poly(GP)RAN proteins [e.g., poly(GGTCCT), poly(GGCCCC), poly(GGACCA), poly(GGGCCG), poly(GGTCCC), poly(GGCCCA), poly(GGACCG), poly(GGGCCT), poly(GGTCCA), poly(GGCCCG), poly(GGACCT), poly(GGGCCC), poly(GGTCCG), poly(GGCCCT), poly(GGACCC), and and / or poly(GGGCCA) repeats], and those encoding poly(PR)RAN proteins [e.g., poly(CCTCGT), poly(CCCCGT), poly(CCACGT), poly(CCGCGT), poly(CCTCGC), poly(CCCCGC), poly(CCACGC), poly(CCGCGC), poly(CCTCGA), poly(CCCCGA), poly(CCACGA), poly(CCGCGA), poly(CCTCGG), poly(CCCCGG), poly(CCACGG), poly(CCGCGG), poly(CCTAGA), poly(CCCAGA),poly(CCAAGA), poly(CCGAGA), poly(CCTAGG), poly(CCCAGG), poly(CCAAGG), and / or poly(CCGAGG) repeats.
[0039] RAN proteins may contain multiple poly-amino acid repeats (e.g., at least 10 to 10,000 repeats). RAN proteins containing poly-amino acid repeats may be capable of forming insoluble aggregates within cells and / or tissues. Without wishing to be bound by any particular theory or belief, expression of RAN proteins containing approximately 10 poly-amino acid repeats may not lead to a disease of interest. However, expression of RAN proteins containing more than 10 poly-amino acid repeats (e.g., 40 or more) may result in a RAN protein-associated disease, e.g., a neurological disease such as a neurodegenerative disease.
[0040] The present disclosure relates, at least in part, to the surprising discovery of a novel class of RAN proteins comprising a discontinuous or interrupted poly-amino acid repeat motif. Aspects of the present disclosure further relate to the surprising discovery that the interrupted poly-amino acid repeat motif is expressed in certain RAN protein diseases. In some embodiments, the present disclosure provides methods for identifying a subject as having a RAN protein disease and / or methods for treating a subject having or suspected of having a RAN protein-related disease. In some embodiments, the methods described herein comprise detecting one or more interrupted RAN proteins in a subject. In some embodiments, the methods described herein comprise administering one or more anti-RAN protein agents to the subject.
[0041] Aborted RAN protein In some embodiments, an interrupted RAN protein or discontinuous RAN protein is a RAN protein comprising multiple repeat units, each of which is separated by one or more amino acids. In some embodiments, a repeat unit, repeat motif, or RAN repeat refers to a poly-amino acid repeat [e.g., poly(GA) or poly(GR)] and / or a sequence comprising the same sequence of amino acids as a poly-amino acid repeat but not consecutively [e.g., (GA) or (GR), (LPAC) (SEQ ID NO: 31), etc.]. In some embodiments (e.g., with respect to nucleic acid sequences such as DNA or RNA transcripts), a repeat motif, repeat unit, and RAN repeat may be referred to as a nucleotide stretch in a nucleic acid (e.g., DNA, such as genomic DNA, and RNA, such as mRNA) encoding an interrupted RAN protein.
[0042] In some embodiments, the interrupted RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises repeat units. As a non-limiting example, an interrupted RAN protein comprising GAGAGAFGAGA (SEQ ID NO: 32), GAGAGAFGADGAGA (SEQ ID NO: 33), or GRGRGRFGRVYGRRKDGRGR (SEQ ID NO: 34) would have 90.9% (10 of 11 residues comprised in the repeat unit), 85.7% (12 of 14 residues comprised in the repeat unit), or 30% (6 of 20 residues comprised in the repeat unit), respectively. In some embodiments, the interrupted RAN protein comprises about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 1109%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 0%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequences comprise repeat units. In some embodiments, the interrupted RAN protein comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units. In some embodiments, the interrupted RAN protein comprises an amino acid sequence in which 50-99% (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.) of the amino acid sequence comprises repeat units.However, in some embodiments, the interrupted RAN protein comprises an amino acid sequence in which less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acids comprise repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids are included in repeat units).
[0043] In some embodiments, the interrupted RAN protein comprises at least two repeat units. In some embodiments, the interrupted RAN protein can comprise between about 2 and about 10,000 repeat units. In some embodiments, the interrupted RAN protein can comprise 3-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1100, 1100-1200, 1200-1300, 1300-1400, 1400-1500, 1500-1600, or 1600-1700 repeat units. , 1700-1800, 1800-1900, 2000-2200, 2200-2400, 2400-2600, 2600-2800, 2800-3000, 3000-3500, 3500-4000, 4000-4500, 4500-5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, or 9000-10000 repeat units. In some embodiments, the interrupted RAN protein comprises between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 repeat units.
[0044] In some embodiments, the repeat unit comprises one or more amino acid residues in length. In some embodiments, the repeat unit comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, or at least 200 amino acid residues in length. In some embodiments, the repeat unit is between 2 and 10,000 amino acid residues in length. In some embodiments, the repeat unit comprises between 2 and 500, between 20 and 300, between 30 and 200, between 40 and 100, between 50 and 90, or between 60 and 80 amino acid residues in length. In some embodiments, the repeat unit comprises a length of more than 200 amino acid residues (e.g., 250, 500, 1000, 5000, 10,000, etc.). In some embodiments, the interrupted RAN protein comprises multiple repeat units comprising the same length and / or multiple repeat units comprising different lengths.
[0045] In some embodiments, the repeat unit comprises a sequence of a single amino acid, a di-amino acid, a tri-amino acid, or a quad-amino acid (eg, a tetra-amino acid). In some embodiments, the repeat unit comprises (PR), (GR), (S), (CP), (GP), (G), (A), (GA), (GD), (GE), (GQ), (GT), (L), (LP), (LPAC) (SEQ ID NO: 31), (LS), (P), (PA), (QAGR) (SEQ ID NO: 35), (RE), (SP), (VP), (FP), (GK), (FTPLSLPV) (SEQ ID NO: 36), (LLPSPSRC) (SEQ ID NO: 37), (YSPLPPGV) (SEQ ID NO: 38), (HREGEGSK) (SEQ ID NO: 39), (TGRERGVN) (SEQ ID NO: 40), (PGGRGE) (SEQ ID NO: 41), (GRQRGVNT) (SEQ ID NO: 42), or (GSKHREAE) (SEQ ID NO: 43). In some embodiments, the repeat unit comprises at least two poly-amino acid repeats. In some embodiments, the repeat unit comprises between about 2 and 10,000 poly-amino acid repeats. In some embodiments, the repeat unit comprises between 3 and 100, 100 and 200, 200 and 300, 300 and 400, 400 and 500, 500 and 600, 600 and 700, 700 and 800, 800 and 900, 900 and 1000, 1000 and 1100, 1100 and 1200, 1200 and 1300, 1300 and 1400, 1400 and 1500, 1500 and 1600, 1600 and 1700, 1700 and 1800, 1800 and 2000, 2000 and 2100, 2100 and 2200, 2200 and 2300, 2300 and 2400, 2400 and 2500, 2500 and 2600, 2600 and 2700, 2700 and 2800, 2800 and 2900, 3000 and 3100, 3100 and 3200, 3200 and 3300, 3300 and 3400, 3400 and 3500, 3500 and 3600, 3600 and 3700, 3700 and 3800, 3800 and 3900, 4000 and 4100, 4100 and 4200, 4200 and 4300, 4300 and 4400, 440 and 9000-10000 poly-amino acid repeats. In some embodiments, the repeat unit comprises between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 poly-amino acid repeats.In some embodiments, the repeat unit is selected from the group consisting of poly(proline-arginine) [poly(PR)], poly(glycine-arginine) [poly(GR)], poly(serine) [poly(Ser)], poly(cysteine-proline) [poly(CP)], poly(glycine-proline) [poly(GP)], poly(glycine) [poly(G)], poly(Ala) [polyAla], poly(glycine-alanine) [poly(GA)], poly(glycine-aspartic acid) [poly(GD)], poly(glycine-glutamic acid) [poly(GE)], poly(glycine-glutamine) [poly(GQ)], poly(glycine-threonine) [poly(GT)], poly(leucine) [polyLeu], poly(leucine-proline) [poly(LP)], poly(leucine-proline-alanine-cysteine) [poly(LPAC)] (SEQ ID NO: 31), poly(leucine-serine) [poly(GP)], poly(glycine) [poly(GP)], poly(glycine) [poly(G)], poly(Ala) [polyAla], poly(glycine-alanine) [poly(GA)], poly(glycine-aspartic acid) [poly(GD)], poly(glycine-glutamic acid) [poly(GE)], poly(glycine-glutamine) [poly(GQ)], poly(glycine-threonine) [poly(GT)], poly(leucine) [polyLeu], poly(leucine-proline) [poly(LP)], poly(leucine-proline-alanine-cysteine) [poly(LPAC)] (SEQ ID NO: 31), poly(leucine-serine) [poly(LPAC)] (SEQ ID NO: 32), poly(leucine) [poly(LPAC)] (SEQ ID NO: 33), poly(leucine) [poly(LPAC)] (SEQ ID NO: 34), poly(leucine (LS)], poly(proline) [poly(P)], poly(proline-alanine) [poly(PA)], poly(glutamine-alanine-glycine-arginine) [poly(QAGR)] (SEQ ID NO: 35), poly(arginine-glutamic acid) [poly(RE)], poly(serine-proline) [poly(SP)], poly(valine-proline) [poly(VP)], poly(phenylalanine-proline) [poly(FP)], poly(glycine-lysine) [poly(GK)], poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41), poly(GRQRGVNT) (SEQ ID NO: 42), or poly(GSKHREAE) (SEQ ID NO: 43) repeats.
[0046] In some embodiments, the repeat units are separated by one or more amino acid residues (e.g., non-repeating amino acids). In some embodiments, the repeat units are separated by a plurality of amino acids (e.g., non-repeating amino acids). In some embodiments, the repeat units are separated by about 2 to 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, , 41 pieces, 42 pieces, 43 pieces, 44 pieces, 45 pieces, 46 pieces, 47 pieces, 48 pieces, 49 pieces, 50 pieces, 51 pieces, 52 pieces, 53 pieces, 54 pieces, 55 pieces, 56 pieces, 57 pieces, 58 pieces, 59 pieces, 60 pieces, 61 pieces, 62 pieces, 63 pieces, 6 4 pieces, 65 pieces, 66 pieces, 67 pieces, 68 pieces, 69 pieces, 70 pieces, 71 pieces, 72 pieces, 73 pieces, 74 pieces, 75 pieces, 76 pieces, 77 pieces, 78 pieces, 79 pieces, 80 pieces, 81 pieces, 82 pieces, 83 pieces, 84 pieces, 85 pieces, 86 pieces, 87 pieces In some embodiments, the repeat units are separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, the repeat units are separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.). In some embodiments, the repeat units are separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0047] In some embodiments, the repeat units comprising the poly-amino acid repeats are separated by a plurality of amino acids comprising one or more repeat units. In some embodiments, the plurality of amino acids comprises 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-1000 repeat units. In some embodiments, one or more repeat units in the plurality of amino acids comprise non-contiguous amino acids. In some embodiments, the plurality of amino acids comprises one or more non-contiguous (PR), (GR), (S), (CP), (GP), (G), (A), (GA), (GD), (GE), (GQ), (GT), (L), (LP), (LPAC) (SEQ ID NO: 31), (LS), (P), (PA), (QAGR) (SEQ ID NO: 35), (RE), (SP), (VP), (FP), (GK), (FTPLSLPV) (SEQ ID NO: 36), (LLPSPSRC) (SEQ ID NO: 37), (YSPLPPGV) (SEQ ID NO: 38), (HREGEGSK) (SEQ ID NO: 39), (TGRERGVN) (SEQ ID NO: 40), (PGGRGE) (SEQ ID NO: 41), (GRQRGVNT) (SEQ ID NO: 42), or (GSKHREAE) (SEQ ID NO: 43) repeat units.
[0048] In some embodiments, the interrupted RAN protein comprises one or more of the following RAN repeat units: (PR), (GR), (S), (CP), (GP), (G), (A), (GA), (GD), (GE), (GQ), (GT), (L), (LP), (LPAC) (SEQ ID NO: 31), (LS), (P), (PA), (QAGR) (SEQ ID NO: 35), (RE), (SP), (VP), (FP), (GK), (FTPLSLPV) (SEQ ID NO: 36), (LLPSPSRC) (SEQ ID NO: 37), (YSPLPPGV) (SEQ ID NO: 38), (HREGEGSK) (SEQ ID NO: 39), (TGRERGVN) (SEQ ID NO: 40), (PGGRGE) (SEQ ID NO: 41), (GRQRGVNT) (SEQ ID NO: 42), or (GSKHREAE) (SEQ ID NO: 43). In some embodiments, the interrupted RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises RAN repeat units. In some embodiments, the interrupted RAN comprises less than about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50% of the amino acid sequence comprises RAN repeat units. , 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% contain amino acid sequences containing repeat units. In some embodiments, the interrupted RAN comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units.In some embodiments, the interrupted RAN protein comprises an amino acid sequence in which 50-99% of the amino acid sequence comprises repeat units (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.). However, in some embodiments, the interrupted RAN protein comprises an amino acid sequence in which less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acid sequence comprises repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids comprise repeat units). In some embodiments, the interrupted RAN protein comprises at least two RAN repeats. In some embodiments, the repeat units in the interrupted RAN protein comprise between about 2 and about 10,000 RAN repeats. In some embodiments, the repeat units in the interrupted RAN protein are 3 to 100, 100 to 200, 200 to 300, 300 to 400, 400 to 500, 500 to 600, 600 to 700, 700 to 800, 800 to 900, 900 to 1000, 1000 to 1100, 1100 to 1200, 1200 to 1300, 1300 to 1400, 1400 to 1500, 1500 to 1600, 1600 Contains ~1700, 1700-1800, 1800-1900, 2000-2200, 2200-2400, 2400-2600, 2600-2800, 2800-3000, 3000-3500, 3500-4000, 4000-4500, 4500-5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, or 9000-10000 RAN repeats. In some embodiments, the repeat unit in the interrupted RAN protein comprises between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 RAN repeats.In some embodiments, repeat units comprising RAN repeats are separated by a plurality of amino acids, including one or more RAN repeat units, and the one or more RAN repeat units are not contiguous (e.g., separated by one or more amino acids). In some embodiments, repeat units comprising RAN repeats are separated by a plurality of amino acids, including 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-10000 RAN repeat units. In some embodiments, repeat units comprising RAN repeats are separated by a plurality of amino acids, including 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) RAN repeat units. In some embodiments, each of the one or more RAN repeat units between repeat units comprising RAN repeats is separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, each of the one or more RAN repeat units are separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.).In some embodiments, each of the one or more RAN repeat units is between about 2 and 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 7, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100). In some embodiments, each of the one or more RAN repeat units are separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0049] In some embodiments, the interrupted RAN protein is an interrupted poly(GR) protein. In some embodiments, the interrupted poly(GR) RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(GR)RAN comprises about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(GR)RAN comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(GR)RAN comprises an amino acid sequence in which 50-99% (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.) of the amino acid sequence comprises repeat units. However, in some embodiments, the interrupted poly(GR)RAN protein comprises less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acid sequences comprising repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids are comprised of repeat units). In some embodiments, the interrupted poly(GR)RAN protein comprises at least two poly(GR) repeats.In some embodiments, the repeat unit in the interrupted poly(GR)RAN protein comprises between about 2 and about 10,000 poly(GR) repeats. In some embodiments, the repeat unit in the interrupted poly(GR)RAN protein comprises between 3-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1100, 1100-1200, 1200-1300, 1300-1400, 1400-1500, 1500-1600, 1600-1700, 1800-1900, 1900-2000, 2000-2100, 2100-2200, 2200-2300, 2300-2400, 2400-2500, 2500-2600, 2600-2700, 2700-2800, 2800-2900, 2900-3000, 3000-3100, 3100-3200, 3200-3300, 3300-3400, 3400-3500, 3500-3600, 3600-3700, 3700-3800, 3800-3900, 3900-4000, 4100-4200, 4300-4400, 4400-4500, 4500-4600, 4 and containing 00-1700, 1700-1800, 1800-1900, 2000-2200, 2200-2400, 2400-2600, 2600-2800, 2800-3000, 3000-3500, 3500-4000, 4000-4500, 4500-5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, or 9000-10000 poly(GR) repeats. In some embodiments, the repeat units in the interrupted poly(GR)RAN protein comprise between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 poly(GR) repeats. In some embodiments, the repeat units comprising the poly(GR) repeats are separated by multiple amino acids, including one or more (GR) repeat units, and the one or more (GR) repeat units are not contiguous (e.g., separated by one or more amino acids). In some embodiments, the repeat units comprising the poly(GR) repeat are separated by a plurality of amino acids, including 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-10000 (GR) repeat units.In some embodiments, the repeat units comprising the poly(GR) repeat are separated by a plurality of amino acids, including 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) (GR) repeat units. In some embodiments, each of the one or more (GR) repeat units between the poly(GR)-comprising repeat units is separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, each of the one or more (GR) repeat units between repeat units comprising poly(GR) are separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.). In some embodiments, each of the one or more (GR) repeat units between the repeat units comprising poly(GR) is between about 2 and 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, , 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100).In some embodiments, each of the one or more (GR) repeat units between repeat units comprising poly(GR) are separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0050] In some embodiments, the interrupted RAN protein is an interrupted poly(GA) protein. In some embodiments, the interrupted poly(GA) RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(GA)RAN comprises about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(GA)RAN comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(GA)RAN comprises an amino acid sequence in which 50-99% (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.) of the amino acid sequence comprises repeat units. However, in some embodiments, the interrupted poly(GA)RAN protein comprises less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acid sequences comprising repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids are comprised of repeat units). In some embodiments, the interrupted poly(GA)RAN protein comprises at least two poly(GA) repeats.In some embodiments, the repeat units in the interrupted poly(GA)RAN protein comprise between about 2 and about 10,000 poly(GA) repeats. In some embodiments, the repeat units in the interrupted poly(GA)RAN protein comprise between 3-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1100, 1100-1200, 1200-1300, 1300-1400, 1400-1500, 1500-1600, 1600-1700, 1800-1900, 1900-2000, 2000-2100, 2100-2200, 2200-2300, 2300-2400, 2400-2500, 2500-2600, 2600-2700, 2700-2800, 2800-2900, 2900-3000, 3000-3100, 3100-3200, 3200-3300, 3300-3400, 3400-3500, 3500-3600, 3600-3700, 3700-3800, 3800-3900, 3900-4000, 4100-4200, 4300-4400, 4500-4600, 4600-4700, 4 and 9000-10000 poly(GA) repeats. In some embodiments, the repeat units in the interrupted poly(GA)RAN protein comprise between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 poly(GA) repeats. In some embodiments, the repeat units comprising poly(GA) repeats are separated by a plurality of amino acids, including one or more (GA) repeat units, and the one or more (GA) repeat units are not contiguous (e.g., separated by one or more amino acids). In some embodiments, the repeat units comprising the poly(GA) repeat are separated by a plurality of amino acids, including 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-10000 (GA) repeat units.In some embodiments, repeat units comprising poly(GA) repeats are separated by a plurality of amino acids, including 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) (GA) repeat units. In some embodiments, each of the one or more (GA) repeat units between poly(GA)-comprising repeat units is separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, each of the one or more (GA) repeat units between poly(GA)-containing repeat units is separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.). In some embodiments, each of one or more (GA) repeat units between poly(GA)-containing repeat units is about 2 to 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 1 , 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100).In some embodiments, each of the one or more (GA) repeat units between poly(GA)-containing repeat units is separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0051] In some embodiments, the interrupted RAN protein is an interrupted poly(GP) protein. In some embodiments, the interrupted poly(GP) RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(GP)RAN comprises about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(GP)RAN comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(GP)RAN comprises an amino acid sequence in which 50-99% (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.) of the amino acid sequence comprises repeat units. However, in some embodiments, the interrupted poly(GP)RAN protein comprises less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acid sequences comprising repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids comprise repeat units). In some embodiments, the interrupted poly(GP)RAN protein comprises at least two poly(GP) repeats.In some embodiments, the repeat unit in the interrupted poly(GP)RAN protein comprises between about 2 and about 10,000 poly(GP) repeats. In some embodiments, the repeat unit in the interrupted poly(GP)RAN protein comprises between 3-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1100, 1100-1200, 1200-1300, 1300-1400, 1400-1500, 1500-1600, 1600-1700, 1800-1900, 1900-2000, 2000-2100, 2100-2200, 2200-2300, 2300-2400, 2400-2500, 2500-2600, 2600-2700, 2700-2800, 2800-2900, 2900-3000, 3000-3100, 3100-3200, 3200-3300, 3300-3400, 3400-3500, 3500-3600, 3600-3700, 3700-3800, 3800-3900, 3900-4000, 4100-4200, 4300-4400, 4500-4600, 4600-4700, 4 and containing 00-1700, 1700-1800, 1800-1900, 2000-2200, 2200-2400, 2400-2600, 2600-2800, 2800-3000, 3000-3500, 3500-4000, 4000-4500, 4500-5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, or 9000-10000 poly(GP) repeats. In some embodiments, the repeat units in the interrupted poly(GP)RAN protein comprise between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 poly(GP) repeats. In some embodiments, the repeat units comprising the poly(GP) repeats are separated by multiple amino acids, including one or more (GP) repeat units, and the one or more (GP) repeat units are not contiguous (e.g., separated by one or more amino acids). In some embodiments, the repeat units comprising the poly(GP) repeat are separated by a plurality of amino acids, including 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-10000 (GP) repeat units.In some embodiments, the repeat units comprising the poly(GP) repeat are separated by a plurality of amino acids, including 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) (GP) repeat units. In some embodiments, each of the one or more (GP) repeat units between the poly(GP)-comprising repeat units is separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, each of the one or more (GP) repeat units between poly(GP)-containing repeat units is separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.). In some embodiments, each of one or more (GP) repeat units between repeat units comprising poly(GP) is between about 2 and 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 1 , 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100).In some embodiments, each of the one or more (GP) repeat units between poly(GP)-containing repeat units is separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0052] In some embodiments, the interrupted RAN protein is an interrupted poly(PR) protein. In some embodiments, the interrupted poly(PR) RAN protein comprises an amino acid sequence in which less than 100% of the amino acid sequence comprises repeat units. In some embodiments, the interrupted poly(PR)RAN comprises about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103%, 104%, 105%, 106%, 107%, 108%, 109%, 110%, 111%, 112%, 113%, 114%, 115%, 116%, 117%, 118%, 119%, 120%, 121%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(PR)RAN comprises an amino acid sequence in which between 10% and 50% (e.g., 10-15%, 15-20%, 20-25%, 25-30%, 30-35%, 35-40%, 40-45%, etc.) of the amino acid sequence comprises repeat units. In some embodiments, interrupted poly(PR)RAN comprises an amino acid sequence in which 50-99% (e.g., 51-55%, 55-60%, 60-70%, 70-80%, etc.) of the amino acid sequence comprises repeat units. However, in some embodiments, the interrupted poly(PR)RAN protein comprises less than 1% (e.g., about 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%) of the amino acid sequences comprise repeat units (e.g., a RAN protein comprising 10,000 amino acids, in which fewer than 100 of the 10,000 amino acids are repeat units). In some embodiments, the interrupted poly(PR)RAN protein comprises at least two poly(PR) repeats.In some embodiments, the repeat units in the interrupted poly(PR)RAN protein comprise between about 2 and about 10,000 poly(PR) repeats. In some embodiments, the repeat units in the interrupted poly(PR)RAN protein comprise between 3-100, 100-200, 200-300, 300-400, 400-500, 500-600, 600-700, 700-800, 800-900, 900-1000, 1000-1100, 1100-1200, 1200-1300, 1300-1400, 1400-1500, 1500-1600, 1600-1700, 1800-1900, 1900-2000, 2000-2100, 2100-2200, 2200-2300, 2300-2400, 2400-2500, 2500-2600, 2600-2700, 2700-2800, 2800-2900, 2900-3000, 3000-3100, 3100-3200, 3200-3300, 3300-3400, 3400-3500, 3500-3600, 3600-3700, 3700-3800, 3800-3900, 3900-4000, 4100-4200, 4300-4400, 4500-4600, 4600-4700, 4 and containing 00-1700, 1700-1800, 1800-1900, 2000-2200, 2200-2400, 2400-2600, 2600-2800, 2800-3000, 3000-3500, 3500-4000, 4000-4500, 4500-5000, 5000-6000, 6000-7000, 7000-8000, 8000-9000, or 9000-10000 poly(PR) repeats. In some embodiments, the repeat units in the interrupted poly(PR)RAN protein comprise between 20 and 100, between 50 and 200, between 100 and 500, between 400 and 800, between 700 and 1000, or between 800 and 1500 poly(PR) repeats. In some embodiments, the repeat units comprising the poly(PR) repeats are separated by multiple amino acids, including one or more (PR) repeat units, and the one or more (PR) repeat units are not contiguous (e.g., separated by one or more amino acids). In some embodiments, the repeat units comprising the poly(PR) repeat are separated by a plurality of amino acids, including 1, 2, 3, 4, 5, 6-10, 10-14, 14-18, 18-22, 22-26, 26-30, 30-50, 50-100, 100-250, 250-500, 500-1000, 1000-5000, or 5000-10000 (PR) repeat units.In some embodiments, the repeat units comprising the poly(PR) repeat are separated by a plurality of amino acids, comprising 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) (PR) repeat units. In some embodiments, each of the one or more (PR) repeat units between the repeat units comprising the poly(PR) is separated by about 2 to 20 amino acid residues (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amino acids). In some embodiments, each of the one or more (PR) repeat units between the repeat units comprising the poly(PR) is separated by more than 20 amino acid residues (e.g., 21-25, 25-30, 30-40, 40-50, etc.). In some embodiments, each of one or more (PR) repeat units between repeat units comprising poly(PR) is between about 2 and 100 amino acids (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 1 , 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100).In some embodiments, each of the one or more (PR) repeat units between the repeat units comprising the poly(PR) are separated by more than 100 amino acids (e.g., 101-150, 150-200, 200-300, 300-400, etc.).
[0053] Nucleic acid encoding a disrupted RAN protein Aspects of the present disclosure also relate to nucleic acids encoding the RAN proteins (e.g., interrupted RAN proteins) described herein. In some embodiments, the nucleic acid encoding the interrupted RAN protein comprises DNA. In some embodiments, the DNA encoding the interrupted RAN protein can be transcribed to produce RNA (e.g., RNA comprising an expanded repeat unit). In some embodiments, the nucleic acid encoding the interrupted RAN protein comprises RNA. In some embodiments, the RNA encoding the interrupted RAN protein is mRNA. In some embodiments, the RNA encoding the interrupted RAN protein can be translated (e.g., inside a cell, such as a cell in a subject) to produce the interrupted RAN protein. In some embodiments, the nucleic acid encoding the interrupted RAN protein further comprises one or more non-coding sequences. In some embodiments, the nucleic acid encoding the interrupted RAN protein further comprises one or more regulatory sequences (e.g., an enhancer, a promoter, a transcription start / stop site, a polyA signal, etc.). In some embodiments, the one or more regulatory sequences are operably linked to the sequence encoding the interrupted RAN protein (e.g., linked so as to be able to control the expression level of the interrupted RAN protein).
[0054] In some embodiments, the nucleic acid encoding the interrupted RAN protein is (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x, (GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGAGA)x, (GGTAGG) x, (GGCAGG)x, (GGAAGG)x, (GGGAGG)x, (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GGAGCG)x, (GGGGCT)x, (GGTGCA)x, (GGCGCG)x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, (GGCCCC)x, (GGACCA)x, (GGGCCG)x, (GGTCCC)x, (G GCCCA)x, (GGACCG)x, (GGGCCT)x, (GGTCCA)x, (GGCCCG)x, (GGACCT)x, (GGGCCC)x, (GGTCCG)x, (GGCCCT)x, (GGACCC)x, (GGGCCA)x, (CCTCGT)x, (CCC CGT)x, (CCACGT)x, (CCGCGT)x, (CCTCGC)x, (CCCCGC)x, (CCACGC)x, (CCGCGC)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCGCGA)x, (CCTCGG)x, (CCCCG) and / or (AGC)x, where "x" represents the number of repeat units present in the nucleic acid (e.g., a nucleic acid corresponding to a gene, a chromosomal locus, and / or an RNA such as an mRNA).In some embodiments, "x" comprises an integer between 2 and 200. In some embodiments, "x" comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, "x" is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200.In some embodiments, "x" comprises an integer between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000.
[0055] In some embodiments, the nucleic acid encoding the interrupted RAN protein comprises one or more nucleotides between the extended repeat units. In some embodiments, the one or more nucleotides between the extended repeat units may comprise a non-naturally occurring sequence (e.g., a synthetic sequence, such as a sequence not normally found in a gene). In some embodiments, the one or more nucleotides between the extended repeat units may comprise a sequence corresponding to a protein-coding region of a gene (e.g., an exon region). In some embodiments, the one or more nucleotides between the extended repeat units may comprise a sequence corresponding to a non-coding region of a gene or chromosomal locus (e.g., an intron region, an untranslated region such as a 5' UTR or 3' UTR, etc.). In some embodiments, the one or more nucleotides between the extended repeat units may comprise a sequence corresponding to an intergenic region (e.g., a nucleic acid sequence located between genes on a chromosome).
[0056] In some embodiments, the nucleic acid encoding the interrupted RAN protein comprises one or more nucleotides between the expanded repeat units corresponding to a gene associated with a disease (e.g., a neurological disease). In some embodiments, the gene is selected from the group consisting of amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), and spinocerebellar ataxia type 6 (SCA6). , associated with diseases such as spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), or Fuchs' corneal dystrophy.
[0057] In some embodiments, the "gene associated with amyotrophic lateral sclerosis" (ALS) is C9ORF72. In some embodiments, the "gene associated with frontotemporal dementia" (FTD) is C9ORF72. In some embodiments, the "gene associated with Alzheimer's disease" (AD) is APP, PSEN1, PSEN2, MAPT, or CASP8. In some embodiments, the "gene associated with fragile X syndrome" (FRAXA) is FMR1. In some embodiments, the "gene associated with spinal-bulbar muscular atrophy" (SBMA) is AR. In some embodiments, the "gene associated with dentatorubral-pallidoluysian atrophy" (DRPLA) is ATN1. In some embodiments, the "gene associated with spinocerebellar ataxia type 1" (SCA1) is ATXN1. In some embodiments, the "gene associated with spinocerebellar ataxia type 2" (SCA2) is ATXN2. In some embodiments, the "gene associated with spinocerebellar ataxia type 3" (SCA3) is ATXN3. In some embodiments, the "gene associated with spinocerebellar ataxia type 6" (SCA6) is CACNA1A. In some embodiments, the "gene associated with spinocerebellar ataxia type 7" (SCA7) is ATXN7. In some embodiments, the "gene associated with spinocerebellar ataxia type 8" (SCA8) is ATXN8 or ATXN8OS. In some embodiments, the "gene associated with spinocerebellar ataxia type 12" (SCA12) is PPP2R2B. In some embodiments, the "gene associated with spinocerebellar ataxia type 17" (SCA17) is TBP. In some embodiments, the "gene associated with spinocerebellar ataxia type 36" (SCA36) is NOP56. In some embodiments, the "gene associated with spinocerebellar ataxia type 29" (SCA29) is ITPR1. In some embodiments, the "gene associated with spinocerebellar ataxia type 10" (SCA10) is ATXN10. In some embodiments, the "gene associated with myotonic dystrophy type 1" (DM1) is DMPK. In some embodiments, the "gene associated with myotonic dystrophy type 2" (DM2) is CNBP.In some embodiments, the "gene associated with Fuchs' corneal dystrophy" is TCF4 (eg, a TCF4 gene containing a CTG18.1 repeat expansion).
[0058] In some embodiments, the nucleic acid encoding the interrupted RAN protein is contained in a genome (e.g., the human genome) at one or more loci. In some embodiments, the nucleic acid encoding the interrupted RAN protein is contained in a gene or chromosomal locus set forth in Table 1 or Table 6. In some embodiments, the nucleic acid encoding the interrupted RAN protein is PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8 ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, P FN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDIN2, AIFM2, MGMT, LINC01164, KNDC1, ANK3, MLLT10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX 4L2, AGAP12P, C10orf53, SMPD1, IFITM10, BUD13, TSPAN18, CD82, OTUB1, NADSYN1, MIR4492, CHID1, SMUG1, LINC0093 8, LINC01257, SLC15A4, ASIC1, DCP1B, TMTC2, TNS2, LOC100240735, SOX1, LATS2, RAB20, ANKRD20A9P, FLT1, RCBTB1 , ELK2AP, STON2, FOXN3, TTLL5, BCL11B, BRMS1L, SMAD3, RPLP1, BAHD1, MYO5A, DNM1P46, DNM1P35, TUBGCP4, C16orf95 , OSGIN1, LINC00311, MIR4718, RBFOX1, SBK1, MIR4722, BANP, C16orf78, MIR5189, ADGRG5, NPRL3, ZDHHC1, MIR662, L INC00482, MRPL12, TBC1D3H, WSCD1, TBC1D3B, TBC1D3, TBC1D3C, METRNL, DNAH9, ASGR1, FOXK2, NPEPPS, SARM1, CLUH,TIAF1, LOC440434, PHOSPHO1, TBCD, CYP4F35P, CXADRP3, LINC00668, MEX3C, COX7A1, SCAF1, RFPL4AL1, SIX5, DIRAS1, POLRMT, ZNF554, MED16, S IPA1L3, DOT1L, KMT5C, PDE4C, ZNF480, CEBPA, PTPN18, HAAO, LOC654342, RGPD2, RAB3GAP1, TNS1, FAM95A, NTSR1, FRG1BP, FAM182B, CDH4, PRNP, MIR1257, MIR4758, OGFR, SRC, COL9A3, ZNF512B, PICSAR, SIK1, CYP4F29P, MX1, LARGE1, CRELD2, UPK3A, RRP7A, MIR4762, SHANK3, SHISA8, CCDC1 88, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CACNA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10, TNIP3, ZFYVE2 8, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2, SLC34A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, MEA1 , RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00 051, LRRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8. In some embodiments, the nucleic acid encoding the interrupted RAN protein is selected from the list consisting of a gene selected from the list consisting of a gene encoding an interrupted RAN protein, such as those shown in FIG. 9F [e.g.,As shown in Figure 9F, the nucleic acid contains a RAN repeat unit in the following configuration: a CASP8 sequence, such as one containing an insertion of a RAN repeat unit within a VNTR sequence (SEQ ID NO: 194).
[0059] In some embodiments, the gene or chromosomal locus encoding the disrupted RAN protein comprises at least one mutation compared to the wild-type form of the gene. In some embodiments, the at least one mutation comprises an insertion, substitution, deletion, or a combination thereof. In some embodiments, the at least one mutation comprises one or more repeat units.In some embodiments, at least one mutation is selected from the group consisting of (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x, (GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGAGA)x, (GGTAG G)x, (GGCAGG)x, (GGAAGG)x, (GGGAGG)x, (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GGAGCG)x, (GGGGCT)x, (GG TGCA)x, (GGCGCG)x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, (GGCCCC)x, (GGACCA)x, (GGGCCG)x, (GGTCCC)x, (GGCCCA)x, (GGACCG)x, (GGGCCT)x, (GGTCCA)x, (GGCCCG)x, (GGACCT)x, (GGGCCC)x, (GGTCCG)x, (GGCCCT)x, (GGACCC)x, (GGGCCA )x, (CCTCGT)x, (CCCCGT)x, (CCACGT)x, (CCGCGT)x, (CCTCGC)x, (CCCCGC)x, (CCACGC)x, (CCGCGC)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCG In some embodiments, "x" comprises at least one of (CGA), (CCTCGG), (CCCCGG), (CCACGG), (CCGCGG), (CCTAGA), (CCCAGA), (CCAAGA), (CCGAGA), (CCTAGG), (CCCAGG), (CCAAGG), (CCGAGG), (CCTG), (TCT), (TCC), (TCA), (TCG), (AGT), and / or (AGC), where "x" represents the number of repeat units present in the gene or chromosomal locus. In some embodiments, "x" comprises an integer between 2 and 200.In some embodiments, "x" comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, "x" is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200.In some embodiments, "x" comprises an integer between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000.
[0060] In some embodiments, the nucleic acid comprises a sequence encoding an interrupted poly(GR)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein comprises a plurality of (GGTCGT), (GGCCGT), (GGACGT), (GGGCGT), (GGTCGC), (GGCCGC), GGACGC), (GGGCGC), (GGTCGA), (GGCCGA), (GGACGA), (GGGCGA), (GGTCGG), (GGCCGG), (GGACGG), (GGGCGG), (GGTAGA), (GGCAGA), (GGAAGA), (GGGAGA), (GGTAGG), (GGCAGG), (GGAAGG), and / or (GGGAGG) repeat units. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein comprises one or more repeat units, including poly(GGTCGT), poly(GGCCGT), poly(GGACGT), poly(GGGCGT), poly(GGTCGC), poly(GGCCGC), poly(GGACGC), poly(GGGCGC), poly(GGTCGA), poly(GGCCGA), poly(GGACGA), poly(GGGCGA), poly(GGTCGG), poly(GGCCGG), poly(GGACGG), poly(GGGCGG), poly(GGTAGA), poly(GGCAGA), poly(GGAAGA), poly(GGGAGA), poly(GGTAGG), poly(GGCAGG), poly(GGAAGG), and / or poly(GGGAGG) expanded repeats. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein is contained in a gene associated with a disease (e.g., a neurological disease) described herein. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein is contained in the CASP8 gene. In some embodiments, the CASP8 gene (e.g., mutated CASP8) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GR)RAN protein. In some embodiments, the CASP8 RNA is mRNA that is translated to produce the poly(GR)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein is contained in the ARMCX4 gene.In some embodiments, the ARMCX4 gene (e.g., mutated ARMCX4) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GR)RAN protein. In some embodiments, the ARMCX4 RNA is mRNA that is translated to produce the poly(GR)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GR)RAN protein is contained in the ALK gene. In some embodiments, the ALK gene (e.g., mutated ALK) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GR)RAN protein. In some embodiments, the ALK RNA is mRNA that is translated to produce the poly(GR)RAN protein.
[0061] In some embodiments, the nucleic acid comprises a sequence encoding an interrupted poly(GA)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein comprises a plurality of (GGTGCT), (GGCGCC), (GGAGCA), (GGGGCG), (GGTGCC), (GGCGCA), (GGAGCG), (GGGGCT), (GGTGCA), (GGCGCG), (GGAGCT), (GGGGCC), (GGTGCG), (GGCGCT), (GGAGCC), and / or (GGGGCA) repeat units. In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein comprises one or more repeat units, including poly(GGTGCT), poly(GGCGCC), poly(GGAGCA), poly(GGGGCG), poly(GGTGCC), poly(GGCGCA), poly(GGAGCG), poly(GGGGCT), poly(GGTGCA), poly(GGCGCG), poly(GGAGCT), poly(GGGGCC), poly(GGTGCG), poly(GGCGCT), poly(GGAGCC), and / or poly(GGGGCA) expanded repeats. In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein is comprised in a gene associated with a disease (e.g., a neurological disease) described herein. In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein is comprised in the CASP8 gene. In some embodiments, the CASP8 gene (e.g., a mutated CASP8) can be transcribed to produce an RNA (e.g., mRNA) encoding the interrupted poly(GA)RAN protein. In some embodiments, the CASP8 RNA is an mRNA that is translated to produce poly(GA)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein is contained in the ARMCX4 gene. In some embodiments, the ARMCX4 gene (e.g., mutated ARMCX4) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GA)RAN protein. In some embodiments, the ARMCX4 RNA is an mRNA that is translated to produce poly(GA)RAN protein.In some embodiments, the nucleic acid encoding the interrupted poly(GA)RAN protein is contained in the ALK gene. In some embodiments, the ALK gene (e.g., a mutated ALK) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GA)RAN protein. In some embodiments, the ALK RNA is mRNA that is translated to produce the poly(GA)RAN protein.
[0062] In some embodiments, the nucleic acid comprises a sequence encoding an interrupted poly(GP)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein comprises a plurality of (GGTCCT), (GGCCCC), (GGACCA), (GGGCCG), (GGTCCC), (GGCCCA), (GGACCG), (GGGCCT), (GGTCCA), (GGCCCG), (GGACCT), (GGGCCC), (GGTCCG), (GGCCCT), (GGACCC), and / or (GGGCCA) repeat units. In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein comprises one or more repeat units, including poly(GGTCCT), poly(GGCCCC), poly(GGACCA), poly(GGGCCG), poly(GGTCCC), poly(GGCCCA), poly(GGACCG), poly(GGGCCT), poly(GGTCCA), poly(GGCCCG), poly(GGACCT), poly(GGGCCC), poly(GGTCCG), poly(GGCCCT), poly(GGACCC), and / or poly(GGGCCA) expanded repeats. In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein is comprised in a gene associated with a disease (e.g., a neurological disease) described herein. In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein is comprised in the CASP8 gene. In some embodiments, the CASP8 gene (e.g., a mutated CASP8) can be transcribed to produce an RNA (e.g., mRNA) encoding the interrupted poly(GP)RAN protein. In some embodiments, the CASP8 RNA is an mRNA that is translated to produce poly(GP)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein is contained in the ARMCX4 gene. In some embodiments, the ARMCX4 gene (e.g., mutated ARMCX4) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GP)RAN protein. In some embodiments, the ARMCX4 RNA is an mRNA that is translated to produce poly(GP)RAN protein.In some embodiments, the nucleic acid encoding the interrupted poly(GP)RAN protein is contained in the ALK gene. In some embodiments, the ALK gene (e.g., a mutated ALK) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(GP)RAN protein. In some embodiments, the ALK RNA is mRNA that is translated to produce the poly(GP)RAN protein.
[0063] In some embodiments, the nucleic acid comprises a sequence encoding an interrupted poly(PR)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein comprises a plurality of (CCTCGT), (CCCCGT), (CCACGT), (CCGCGT), (CCTCGC), (CCCCGC), (CCACGC), (CCGCGC), (CCTCGA), (CCCCGA), (CCACGA), (CCGCGA), (CCTCGG), (CCCCGG), (CCACGG), (CCGCGG), (CCTAGA), (CCCAGA), (CCAAGA), (CCGAGA), (CCTAGG), (CCCAGG), (CCAAGG), and / or (CCGAGG) repeat units. In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein comprises one or more repeat units, including poly(CCTCGT), poly(CCCCGT), poly(CCACGT), poly(CCGCGT), poly(CCTCGC), poly(CCCCGC), poly(CCACGC), poly(CCGCGC), poly(CCTCGA), poly(CCCCGA), poly(CCACGA), poly(CCGCGA), poly(CCTCGG), poly(CCCCGG), poly(CCACGG), poly(CCGCGG), poly(CCTAGA), poly(CCCAGA), poly(CCAAGA), poly(CCGAGA), poly(CCTAGG), poly(CCCAGG), poly(CCAAGG), and / or poly(CCGAGG) expanded repeats. In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein is contained in a gene associated with a disease (e.g., a neurological disease) described herein. In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein is contained in the CASP8 gene. In some embodiments, the CASP8 gene (e.g., mutated CASP8) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(PR)RAN protein. In some embodiments, the CASP8 RNA is mRNA that is translated to produce the poly(PR)RAN protein.In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein is contained in the ARMCX4 gene. In some embodiments, the ARMCX4 gene (e.g., mutated ARMCX4) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(PR)RAN protein. In some embodiments, the ARMCX4 RNA is mRNA that is translated to produce the poly(PR)RAN protein. In some embodiments, the nucleic acid encoding the interrupted poly(PR)RAN protein is contained in the ALK gene. In some embodiments, the ALK gene (e.g., mutated ALK) can be transcribed to produce RNA (e.g., mRNA) encoding the interrupted poly(PR)RAN protein. In some embodiments, the ALK RNA is mRNA that is translated to produce the poly(PR)RAN protein.
[0064] RAN protein-related diseases A "subject having or suspected of having a disease (e.g., a neurological disease) associated with RAN protein expression, translation, and / or accumulation" generally refers to a subject who exhibits one or more signs and symptoms of a disease (e.g., a neurological disease such as a neurodegenerative disease). In some embodiments, signs and symptoms include, but are not limited to, memory deficits (e.g., short-term memory loss), confusion, impaired executive function (e.g., attention, planning, flexibility, abstract thinking, etc.), loss of speech, decline or loss of motor skills, etc., or a subject who has, or has been identified as having, one or more genetic mutations associated with RAN protein expression, translation, and / or accumulation. However, in some embodiments, a subject who has or is suspected of having a disease (e.g., a neurological disease) associated with RAN protein expression, translation, and / or accumulation may be a subject with a family history of a disease (e.g., a neurological disease such as a neurodegenerative disease), e.g., a subject suspected, expected, or suspected to be at high risk of developing the disease. In some embodiments, the subject can be a mammal (e.g., a human, mouse, rat, dog, cat, or pig). In some embodiments, the subject is a non-human animal, such as a mouse, rat, guinea pig, cat, dog, horse, camel, etc. In some embodiments, the subject is a human.
[0065] In some embodiments, a subject having or suspected of having a disease associated with RAN protein expression, translation, and / or accumulation (e.g., a neurological disease) is characterized as having a mutation in a gene or chromosomal locus described herein. In some embodiments, the mutation is (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x, (GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGA GA)x, (GGTAGG)x, (GGCAGG)x, (GGAAGG)x, (GGGAGG)x, (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GG AGCG)x, (GGGGCT)x, (GGTGCA)x, (GGCGCG)x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, ( GGCCCC)x, (GGACCA)x, (GGGCCG)x, (GGTCCC)x, (GGCCCA)x, (GGACCG)x, (GGGCCT)x, (GGTCCA)x, (GGCCCG)x, (GGACCT)x, (GGGCCC)x , (GGTCCG)x, (GGCCCT)x, (GGACCC)x, (GGGCCA)x, (CCTCGT)x, (CCCCGT)x, (CCACGT)x, (CCGCGT)x, (CCTCGC)x, (CCCCGC)x, (CCACGC )x, (CCGCGC)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCGCGA)x, (CCTCGG)x, (CCCCGG)x, (CCACGG)x, (CCGCGG)x, (CCTAGA)x, (CCCA GA)x, (CCAAGA)x, (CCGAGA)x, (CCTAGG)x, (CCCAGG)x, (CCAAGG)x, (CCGAGG)x, (CCTG)x, (TCT)x, (TCC)x, (TCA)x, (TCG)x, (AGT)x,or (AGC)x repeat units, where "x" represents the number of repeat units present in the mutated gene and / or chromosomal locus or their RNA (e.g., mRNA) transcript. In some embodiments, "x" comprises an integer between 2 and 200. In some embodiments, "x" comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, "x" is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57 , 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200. In some embodiments, "x" isContains integers between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000. In some embodiments, the mutation is associated with amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA8), spinocerebellar ataxia type 8 (SCA9), spinocerebellar ataxia type 9 (SCA10), spinocerebellar ataxia type 10 (SCA11), spinocerebellar ataxia type 11 (SCA12), spinocerebellar ataxia type 12 (SCA13), spinocerebellar ataxia type 13 (SCA14), spinocerebellar ataxia type 14 (SCA15), spinocerebellar ataxia type 15 (SCA16), spinocerebellar ataxia type 16 (SCA17), spinocerebellar ataxia type 17 (SCA18), spinocerebellar ataxia type 18 (SCA19), spinocerebellar ataxia type 19 ... In some embodiments, the mutation is in a gene associated with a disease (e.g., a neurological disease) such as spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), or Fuchs' corneal dystrophy. In some embodiments, the mutation is in a chromosomal locus or gene set forth in Table 1 or Table 6. In some embodiments, the mutation is in a gene associated with a disease (e.g., a neurological disease) such as PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8, or a chromosomal locus or gene set forth in Table 1 or Table 6. ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, PFN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR197 6, CASZ1, SLC44A3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDIN2, AIFM2, MG MT, LINC01164, KNDC1, ANK3, MLLT10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX4L2, AGAP12P,C10orf53, SMPD1, IFITM10, BUD13, TSPAN18, CD82, OTUB1, NADSYN1, MIR449 2. CHID1, SMUG1, LINC00938, LINC01257, SLC15A4, ASIC1, DCP1B, TMTC2, TN S2, LOC100240735, SOX1, LATS2, RAB20, ANKRD20A9P, FLT1, RCBTB1, ELK2AP STON2, FOXN3, TTLL5, BCL11B, BRMS1L, SMAD3, RPLP1, BAHD1, MYO5A, DNM1P 46. DNM1P35, TUBGCP4, C16orf95, OSGIN1, LINC00311, MIR4718, RBFOX1, SB K1, MIR4722, BANP, C16orf78, MIR5189, ADGRG5, NPRL3, ZDHHC1, MIR662, LI NC00482, MRPL12, TBC1D3H, WSCD1, TBC1D3B, TBC1D3, TBC1D3C, METRNL, DNA H9, ASGR1, FOXK2, NPEPPS, SARM1, CLUH, TIAF1, LOC440434, PHOSPHO1, TBCD. CYP4F35P, CXADRP3, LINC00668, MEX3C, COX7A1, SCAF1, RFPL4AL1, SIX5, DI RAS1, POLRMT, ZNF554, MED16, SIPA1L3, DOT1L, KMT5C, PDE4C, ZNF480, CEBP A, PTPN18, HAAO, LOC654342, RGPD2, RAB3GAP1, TNS1, FAM95A, NTSR1, FRG1B P, FAM182B, CDH4, PRNP, MIR1257, MIR4758, OGFR, SRC, COL9A3, ZNF512B, PI CSAR, SIK1, CYP4F29P, MX1, LARGE1, CRELD2, UPK3A, RRP7A, MIR4762, SHANK 3. SHISA8, CCDC188, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CAC NA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10, TNIP 3. ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2.SLC34A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, MEA1, RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6 , MYL10, GS1-124K5.11, ABCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT, FOXE1, C9or The mutation is contained in a gene selected from the group consisting of f139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8. In some embodiments, the mutation is contained in the ARMCX4, ALK, and / or CASP8 gene. In some embodiments, the mutation comprises a RAN repeat unit in the configuration shown in Figure 9F (e.g., the mutation is contained in a nucleic acid comprising a CASP8 sequence, such as one comprising an insertion of a RAN repeat unit (SEQ ID NO: 194) within a VNTR sequence, as shown in Figure 9F).
[0066] In some embodiments, subjects having or suspected of having a disease (e.g., a neurological disease) associated with RAN protein expression, translation, and / or accumulation are characterized as expressing one or more disrupted RAN proteins (e.g., as a result of a mutation in a gene or chromosomal locus). In some embodiments, the one or more interrupted RAN proteins are selected from the group consisting of poly(proline-arginine) [poly(PR)], poly(glycine-arginine) [poly(GR)], poly(serine) [poly(Ser)], poly(cysteine-proline) [poly(CP)], poly(glycine-proline) [poly(GP)], poly(glycine) [poly(G)], poly(Ala) [polyAla], poly(glycine-alanine) [poly(GA)], poly(glycine-aspartic acid) [poly(GD)], poly(glycine-glutamic acid) [poly(GE)], poly(glycine-glutamine) [poly(GQ)], poly(glycine-threonine) [poly(GT)], poly(leucine) [polyLeu], poly(leucine-proline) [poly(LP)], poly(leucine-proline-alanine-cysteine) [poly(LPAC)] (SEQ ID NO: 31), poly(leucine-serine) [poly(GP)], poly(glycine) [poly(GP)], poly(glycine) [poly(G)], poly(Ala) [polyAla], poly(glycine-alanine) [poly(GA)], poly(glycine-aspartic acid) [poly(GD)], poly(glycine-glutamic acid) [poly(GE)], poly(glycine-glutamine) [poly(GQ)], poly(glycine-threonine) [poly(GT)], poly(leucine) [polyLeu], poly(leucine-proline) [poly(LP)], poly(leucine-proline-alanine-cysteine) [poly(LPAC)] (SEQ ID NO: 31), poly(leucine) [poly(LPAC)] (SEQ ID NO: 32), poly(leucine) [poly(LPAC)] (SEQ ID NO: 33), poly(leucine) [poly(LPAC)] (SEQ ID NO: 34), poly( lysine (LS)], poly(proline) [poly(P)], poly(proline-alanine) [poly(PA)], poly(glutamine-alanine-glycine-arginine) [poly(QAGR)] (SEQ ID NO: 35), poly(arginine-glutamic acid) [poly(RE)], poly(serine-proline) [poly(SP)], poly(valine-proline) [poly(VP)], poly(phenylalanine-proline) [poly(FP)], poly(glycine-lysine) ) [poly(GK)], poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41), poly(GRQRGVNT) (SEQ ID NO: 42), or poly(GSKHREAE) (SEQ ID NO: 43) interrupted RAN proteins.
[0067] In some embodiments, the aggregation pattern of RAN protein (e.g., interrupted RAN protein, such as interrupted poly(GR)RAN protein or interrupted poly(GA)RAN protein) in cells or tissues (e.g., cells or tissues present in a subject) depends on length.For example, RAN protein (e.g., interrupted RAN protein, such as interrupted poly(GR)RAN protein or interrupted poly(GA)RAN protein) with a length of more than 20, more than 48, or more than 80 residues aggregates in a subject (e.g., in brain tissue).In some embodiments, subjects with less than 10 repeats or less than 10 repeat units (e.g., poly-amino acid repeats / repeat units and / or nucleotide expansion repeats / repeat units) do not show signs or symptoms of RAN protein-related diseases characterized by the expression, translation, and / or accumulation of RAN protein. In some embodiments, subjects with between 10 and 40 repeats or repeat units (e.g., poly-amino acid repeats / repeat units and / or nucleotide extension repeats / repeat units) may or may not exhibit one or more signs or symptoms of a RAN protein-associated disease characterized by RAN protein expression, translation, and / or accumulation. In some embodiments, subjects with more than 40 repeats or repeat units (e.g., poly-amino acid repeats / repeat units and / or nucleotide extension repeats / repeat units) exhibit one or more signs or symptoms of a RAN protein-associated disease characterized by RAN protein expression, translation, and / or accumulation. In some embodiments, subjects identified as having a RAN protein-associated disease associated with RAN protein expression, translation, and / or accumulation are characterized by a large (greater than 100) number of repeats or repeat units (e.g., poly-amino acid repeats / repeat units and / or nucleotide extension repeats / repeat units).
[0068] In some embodiments, the disease associated with the expression, translation, and / or accumulation of RAN protein (e.g., truncated RAN protein) is amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA4), spinocerebellar ataxia type 4 (SCA5), spinocerebellar ataxia type 5 (SCA6), spinocerebellar ataxia type 6 (SCA7), spinocerebellar ataxia type 7 (SCA8), spinocerebellar ataxia type 8 (SCA9), spinocerebellar ataxia type 9 (SCA10), spinocerebellar ataxia type 10 (SCA11), spinocerebellar ataxia type 11 (SCA12), spinocerebellar ataxia type 12 (SCA13), spinocerebellar ataxia type 13 (SCA14), spinocerebellar ataxia type 14 (SCA15), spinocerebellar ataxia type 15 (SCA16), spinocerebellar ataxia type 16 (SCA17), spinocerebellar ataxia type 17 (SCA18), spinocerebellar ataxia type 18 (SCA19), spinocerebellar ataxia type 19 ... Spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), or Fuchs' corneal dystrophy.
[0069] In some embodiments, the disease associated with the expression, translation, and / or accumulation of RAN protein (e.g., a truncated RAN protein) is Alzheimer's disease (AD). In some embodiments, a subject expressing one or more RAN proteins is a subject having or suspected of having Alzheimer's disease (AD). A "subject having or suspected of having Alzheimer's disease (AD)" may be a subject exhibiting one or more signs and symptoms of AD, including, but not limited to, memory deficits (e.g., short-term memory loss), confusion, impaired executive function (e.g., attention, planning, flexibility, abstract thinking, etc.), loss of speech, and decline or loss of motor skills. In some embodiments, a subject having or suspected of having AD may be a subject who has, or has been identified as having, one or more genetic mutations associated with AD. In some embodiments, a subject having or suspected of having AD may be a subject who has, or has been identified as having, one or more signs and symptoms associated with one or more mutations associated with AD. Non-limiting examples of mutations associated with AD include mutations in genes such as apolipoprotein (APP), presenilin genes (PSEN1 and PSEN2), tau protein (MAPT), or caspase 8 (CASP8). In some embodiments, subjects with or suspected of having AD are characterized by the accumulation of β-amyloid (Aβ) peptides and hyperphosphorylated tau protein throughout the subject's brain tissue.In some embodiments, the subject has been diagnosed by a health care professional as having AD according to the NINCDS-ADRDA Alzheimer's disease diagnostic criteria as described by McKhann et al. (1984) "Clinical diagnosis of Alzheimer's disease: report of the NINCDS-ADRDA Work Group under the auspices of Department of Health and Human Services Task Force on Alzheimer's Disease". Neurology. 34 (7):939-44.
[0070] In some embodiments, the disease associated with the expression, translation, and / or accumulation of RAN protein (e.g., a disrupted RAN protein) is amyotrophic lateral sclerosis (ALS). In some embodiments, the subject expressing one or more disrupted RAN proteins is a subject who has or is suspected of having amyotrophic lateral sclerosis (ALS). A "subject who has or is suspected of having amyotrophic lateral sclerosis (ALS)" may be a subject who exhibits one or more signs and symptoms of ALS, including, but not limited to, memory deficits (e.g., short-term memory loss), confusion, impaired executive function (e.g., attention, planning, flexibility, abstract thinking, etc.), loss of speech, and decline or loss of motor skills. In some embodiments, a subject who has or is suspected of having ALS may be a subject who has, or has been identified as having, one or more gene mutations associated with ALS. In some embodiments, a subject who has or is suspected of having ALS may be a subject who has, or has been identified as having, one or more signs and symptoms associated with one or more gene mutations associated with ALS. Non-limiting examples of mutations in genes associated with ALS include C9orf72. In some embodiments, the subject has been diagnosed by a medical professional as having ALS.
[0071] Having described embodiments relating to disrupted RAN proteins and diseases and subjects comprising same, the following sections relate to embodiments that may be useful for reducing levels of disrupted RAN protein in cells (e.g., cells ex vivo or cells in vivo as targeted during treatment of a subject) and / or for detecting disrupted RAN protein in a biological sample (e.g., for identifying subjects having or suspected of having a disease associated with RAN protein expression, translation, and / or accumulation). Agents (e.g., therapeutic agents and / or anti-RAN protein agents) and methods previously described (WO2014159247A1, WO2016196324A1, WO2017176813A1, WO2018195110A1, WO2019067587A1, WO2019060918A1, WO2021007110A1, WO2021231887A1, WO2021055880 ... 1061537A1, WO2021072187A2, WO2023077153A1, WO2023102111A1, and WO2023164686A2, the disclosures of which are incorporated herein by reference, relating to detection methods, therapeutic agents, and / or methods of treating RAN protein-associated diseases) may also be useful in addition to or in combination with the embodiments described herein.
[0072] Drugs In some embodiments, the agents (e.g., therapeutic agents and / or anti-RAN protein agents) described herein are useful for reducing RAN protein levels (e.g., disrupted RAN protein levels) in a cell (e.g., a cell in a subject). In some embodiments, a method for reducing disrupted RAN protein levels comprises administering an agent described herein. Non-limiting examples of agents (e.g., therapeutic agents and / or anti-RAN protein agents) include small molecules, nucleic acids (e.g., inhibitory nucleic acids, genes or gene variants, transgenes, recombinant adeno-associated virus (rAAV) genomes, etc.), peptides, proteins (e.g., antibodies or antigen-binding fragments thereof), and rAAV particles.
[0073] In other embodiments, the agents described herein may be useful for detecting RAN proteins (e.g., truncated RAN proteins). For example, embodiments of the present disclosure relating to inhibitory nucleic acids and / or guide RNAs may also be applied to designing nucleic acids that are complementary to target sequences, such as those present in biological samples containing RNA transcripts encoding truncated RAN proteins, that are detected using the methods described herein (e.g., methods involving the use of nucleic acid probes, primers, etc.). Embodiments of the present disclosure relating to antibodies and antigen-binding fragments may also be useful for detecting truncated RAN proteins in biological samples. Thus, those skilled in the art will recognize that the embodiments of agents (e.g., therapeutic agents and / or anti-RAN protein agents) described herein should not be considered limiting and may be useful in the methods, compositions, and kits provided by the present disclosure.
[0074] In some embodiments, an agent may reduce RAN protein levels (e.g., interrupted RAN protein levels) in cells and / or tissues in a subject (e.g., when the agent is administered in an effective amount). In some embodiments, an agent (e.g., a therapeutic agent) may target one or more RAN proteins (e.g., interrupted RAN proteins) and / or modulate a gene or gene product (e.g., a protein) that interacts with one or more RAN proteins (e.g., interrupted RAN proteins). In some embodiments, an agent (e.g., a therapeutic agent) capable of targeting one or more RAN proteins may be capable of reducing the expression, activity, accumulation, and / or aggregation of a RAN protein (e.g., an interrupted RAN protein). In some embodiments, an agent (e.g., a therapeutic agent) may modulate a gene or gene product (e.g., a protein) that genetically and / or physically interacts with one or more RAN proteins (e.g., interrupted RAN proteins) or their genes or chromosomal loci described herein (e.g., genes / chromosomal loci shown in Tables 1 and 6). In some embodiments, genes or gene products (e.g., proteins) that interact (e.g., genetically and / or physically interact) with one or more RAN proteins (e.g., interrupted RAN proteins) or their genes or chromosomal loci can control the transcription and / or translation of the RAN proteins, post-translationally modify the RAN proteins, regulate intracellular trafficking of the RAN proteins, etc.
[0075] In some embodiments, genes and gene products that interact (e.g., genetically and / or physically interact) with one or more RAN proteins (e.g., interrupted RAN proteins) or their genes or chromosomal loci include eukaryotic initiation factor 2 (eIF2), eukaryotic initiation factor 3 (eIF3), protein kinase R (PKR), p62, LC3 I subunit, LC3 II subunit, RISC loading complex subunit, TARBP2, and Toll-like receptor 3 (TLR3).
[0076] In some embodiments, the agent (e.g., a therapeutic agent and / or an anti-RAN protein agent) inhibits eukaryotic initiation factor 2 (eIF2) or protein kinase R (PKR) (e.g., an inhibitor of eIF2 and / or PKR). In some embodiments, the inhibitor of eIF2 is an inhibitor of a serine / threonine kinase. Non-limiting examples of serine / threonine kinases include protein kinase A (PKA), protein kinase C (PKC), Mos / Raf kinase, mitogen-activated protein kinase (MAPK), protein kinase B (AKT kinase), and the like. In some embodiments, the eIF2 inhibitor is a protein kinase R (PKR) inhibitor. Inhibitors of eIF2 and PKR are described, for example, in International Application Publication No. WO 2018 / 195110, the entire contents of which are incorporated herein by reference. In some embodiments, the eIF2 inhibitor can be a direct inhibitor or an indirect inhibitor. In some embodiments, a direct modulator functions by interacting with (e.g., interacting with or binding to) a gene encoding eIF2 (or eIF2α) or the eIF2 protein complex. In some embodiments, an indirect modulator functions by interacting with a gene or protein that regulates the expression or activity of eIF2 or eIF2α (e.g., does not directly interact with a gene or protein encoding eIF2 or eIF2α). In some embodiments, the inhibitor of eIF2 or PKR is a selective inhibitor. A "selective inhibitor" refers to an inhibitor of eIF2 or PKR that preferentially inhibits the activity or expression of one type of eIF2 subunit compared to other types of eIF2 subunits, or that preferentially inhibits the activity or expression of PKR compared to other kinases. In some embodiments, the inhibitor of eIF2 is a selective inhibitor of eIF2α. In some embodiments, the inhibitor of eIF2 is a selective inhibitor of eIF2A.In some embodiments, the inhibitor of eIF2 is a selective inhibitor of protein kinase R (PKR), eg, a selective PKR inhibitor.
[0077] In some embodiments, the agent (e.g., a therapeutic agent) is an inhibitor of eukaryotic initiation factor 3 (eIF3), a multiprotein complex involved in the initiation step of protein translation in eukaryotes. Generally, human eIF3 contains 13 non-identical subunits (e.g., eIF3a-m). Mammalian eIF3 is the largest and most complex initiation factor, containing up to 13 non-identical subunits. Typically, eIF3f is involved in many steps of translation initiation, including stabilizing the ternary complex, mediating mRNA binding to the 40S subunit, and facilitating the dissociation of the 40S and 60S ribosomal subunits. In some embodiments, a therapeutic agent that inhibits the expression or activity of an eIF3 subunit (e.g., eIF3f, eIF3m, eIF3h, or other eIF3 subunit) can be used to reduce or inhibit RAN translation in a cell or subject (e.g., a subject with Alzheimer's disease, which is characterized by RAN protein translation). Inhibitors of eIF3 subunits are further described, for example, in International Application Publication No. WO2017 / 176813, the entire contents of which are incorporated herein by reference. An eIF3 inhibitor can be a direct inhibitor or an indirect inhibitor. Generally, a direct modulator functions by interacting with (e.g., interacting with or binding to) a gene encoding eIF3 (or an eIF3 subunit), or an eIF3 protein complex, or an eIF3 subunit. Generally, an indirect modulator functions by interacting with a gene or protein that regulates the expression or activity of eIF3 or an eIF3 subunit (e.g., does not directly interact with a gene or protein encoding eIF3 or an eIF3 subunit). In some embodiments, the eIF3 inhibitor is a selective inhibitor. A "selective inhibitor" refers to a modulator of eIF3 that preferentially inhibits the activity or expression of one type of eIF3 subunit compared to other types of eIF3 subunits. In some embodiments, the inhibitor of eIF3 is a selective inhibitor of eIF3f. The eIF3 inhibitor can be a protein (e.g., an antibody), a nucleic acid, or a small molecule.Examples of proteins that inhibit eiF3 (eg, eIF3 subunits) include, but are not limited to, polyclonal anti-eIF3 antibodies, monoclonal anti-eIF3 antibodies, measles virus N protein, viral stress-inducible protein p56, and the like.
[0078] As used herein, a "therapeutic agent" may refer to an agent capable of producing a desired result in a subject. In some embodiments, the desired result varies depending on the active agent administered. For example, in some embodiments, an effective amount of rAAV particles may be the amount of particles capable of transferring an expression construct into a host cell, tissue, or organ. In some embodiments, a therapeutically acceptable amount of an anti-RAN protein antibody may be an amount capable of treating a disease (e.g., a neurological disease, such as a neurodegenerative disease) by reducing the expression and / or aggregation of a disrupted RAN protein and / or the appearance or number of RNA foci containing a microsatellite repeat sequence encoding a RAN protein. In certain embodiments, an effective amount is an amount effective to reduce the level of a RAN protein (e.g., a disrupted RAN protein) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98% (e.g., the level of the RAN protein relative to the level of the RAN protein in a cell or subject not administered the therapeutic agent). In certain embodiments, an effective amount is an amount effective to reduce the translation of RAN protein (e.g., a disrupted RAN protein) (e.g., the level of RAN protein relative to the level of RAN protein in a cell or subject not administered the therapeutic agent) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98%.
[0079] As is well known in the medical and veterinary arts, the dosage for a particular subject will vary depending on many factors, including the subject's size, body surface area, age, the particular composition administered, the active ingredients in the composition, the time and route of administration, overall health, and other drugs administered concomitantly. In some embodiments, the methods of the present disclosure involve administering an agent (e.g., a therapeutic agent) in an amount effective to treat a subject (e.g., a subject having or suspected of having a disease associated with RAN protein expression, translation, and / or accumulation).
[0080] As used herein, "treating" a disease refers to reducing the frequency or severity of at least one sign or symptom of a disease or disorder experienced by a subject. In some embodiments, at least one sign or symptom is experienced by a subject with a RAN protein-associated disease. In some embodiments, the RAN protein-associated disease is amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCAA). In some embodiments, the RAN protein-associated disease is characterized by at least one sign or symptom associated with a disease selected from the group consisting of spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), and Fuchs' corneal dystrophy. In some embodiments, the RAN protein-associated disease is characterized by expression of one or more RAN proteins (e.g., disrupted RAN proteins) from a gene or chromosomal locus listed in Table 1 or Table 6. In some embodiments, the RAN protein-associated disease is a RAN protein-associated disease, characterized by the following: PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8, ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, PFN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDIN2, AIFM2, MGMT,LINC01164, KNDC1, ANK3, MLLT10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX4L 2. AGAP12P, C10orf53, SMPD1, IFITM10, BUD13, TSPAN18, CD82, OTUB1, NADS YN1, MIR4492, CHID1, SMUG1, LINC00938, LINC01257, SLC15A4, ASIC1, DCP1 B. TMTC2, TNS2, LOC100240735, SOX1, LATS2, RAB20, ANKRD20A9P, FLT1, RCBT B1, ELK2AP, STON2, FOXN3, TTLL5, BCL11B, BRMS1L, SMAD3, RPLP1, BAHD1, MY O5A, DNM1P46, DNM1P35, TUBGCP4, C16orf95, OSGIN1, LINC00311, MIR4718 BFOX1, SBK1, MIR4722, BANP, C16orf78, MIR5189, ADGRG5, NPRL3, ZDHHC1, M.S IR662, LINC00482, MRPL12, TBC1D3H, WSCD1, TBC1D3B, TBC1D3, TBC1D3C, MET RNL, DNAH9, ASGR1, FOXK2, NPEPPS, SARM1, CLUH, TIAF1, LOC440434, PHOSPH O1, TBCD, CYP4F35P, CXADRP3, LINC00668, MEX3C, COX7A1, SCAF1, RFPL4AL1. SIX5, DIRAS1, POLRMT, ZNF554, MED16, SIPA1L3, DOT1L, KMT5C, PDE4C, ZNF4 80. CEBPA, PTPN18, HAAO, LOC654342, RGPD2, RAB3GAP1, TNS1, FAM95A, NTSR1 FRG1BP, FAM182B, CDH4, PRNP, MIR1257, MIR4758, OGFR, SRC, COL9A3, ZNF5 12B PICSAR SIK1 CYP4F29P MX1 LARGE1 CRELD2 UPK3A RRP7A MIR4762 SHANK3, SHISA8, CCDC188, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP4 3. CACNA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10.TNIP3, ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2, SLC34A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, MEA1, RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT In some embodiments, the RAN protein-associated disease is characterized by the expression of one or more RAN proteins (e.g., a disrupted RAN protein) from a gene selected from the group consisting of: FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8. In some embodiments, the RAN protein-associated disease is characterized by the expression of one or more RAN proteins (e.g., a disrupted RAN protein) from a gene such as ARMCX4, ALK, and / or CASP8. In some embodiments, the subject with a RAN protein-associated disease has been diagnosed with the disease via a method described herein (e.g., a method for detecting a disrupted RAN protein or its RNA transcript or DNA sequence).
[0081] In some embodiments, diseases associated with RAN protein expression, translation, and / or accumulation are associated with the expression, translation, and / or accumulation of one or more disrupted RAN proteins. In some embodiments, a therapeutic agent useful for treating diseases associated with RAN protein expression, translation, and / or accumulation can also be an agent that is a therapeutic agent for treating a subject expressing one or more disrupted RAN proteins. In some embodiments, one or more therapeutic molecules are administered to a subject to treat a disease associated with a RAN protein, such as a disrupted RAN protein. For example, in some embodiments, a subject is administered 2, 3, 4, 5, 6, 7, 8, 9, or 10 therapeutic agents (e.g., proteins, nucleic acids, small molecules, etc., or any combination thereof).
[0082] small molecule In some embodiments, the agent (e.g., a therapeutic agent and / or an anti-RAN protein agent) is a small molecule. In some embodiments, the small molecule inhibits the expression (e.g., RNA and / or protein levels) or activity (e.g., aggregation) of one or more RAN proteins (e.g., disrupted RAN proteins).
[0083] In some embodiments, the agent (e.g., a therapeutic agent) is a small molecule that inhibits a gene or gene product that interacts (e.g., genetically and / or physically interacts) with the RAN protein or its gene or chromosomal locus. In some embodiments, the small molecule is a p62 inhibitor (see, e.g., inhibitors in Leestemaker et al. (2017) Cell Chemical Biology 24, 725-736). In some embodiments, the small molecule is an eIF3 (or eIF3 subunit) inhibitor, such as an mTOR inhibitor (e.g., rapamycin, PP242), an S6 kinase (S6K) inhibitor, or the like. In some embodiments, the small molecule is a TLR3 inhibitor (see, e.g., TLR3 inhibitors in Cheng et al. (2011) J Am Chem Soc 133(11):3764-7). In some embodiments, the small molecule inhibits the expression or activity of eukaryotic initiation factor 2 (eIF2) or a subunit thereof (eg, eIF2A), such as LY364947, salubrinal, Sal003, ISRIB, and the like.In some embodiments, the small molecule is metformin, also known as N,N-dimethylbiguanide (IUPAC name N,N-dimethylimidodicarbonimidic diamide and CAS 657-24-9), 6-amino-3-methyl-2-oxo-N-phenyl-2,3-dihydro-1H-benzo[d]imidazole-1-carboxamide, N-[2-(1H-indol-3-yl)ethyl]-4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, chloroguanide [1-[amino-(4-chloroanilino)methylidene]-2-propan-2-yl-guanidine, CAS 657-24-9, ... 4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, 4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, 4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, 4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, 4-(2-methyl-1H-indol-3-yl)pyrimidin-2-amine, 4-(2-methyl S500-92-5], chlorproguanil [1-[amino-(3,4-dichloroanilino)methylidene]-2-propan-2-ylguanidine, CAS 537-21-3], buformin [N-butylimidodicarbonimidic acid diamide, CAS 692-13-7], or phenformin [2-(N-phenethylcarbamimidoyl)guanidine, CAS 114-86-3], or a pharmaceutically acceptable salt, co-crystal, tautomer, stereoisomer, solvate, hydrate, polymorph, isotopically enriched derivative, or prodrug of any of the biguanides. In some embodiments, the small molecule is buformin or phenformin. In some embodiments, the small molecule is an inhibitor of TARBP2.
[0084] The term "pharmaceutically acceptable salt" refers to a salt that is suitable, within the bounds of sound medical wisdom, for use in contact with the tissues of humans and lower animals without undue toxicity, irritation, allergic reaction, etc., commensurate with a reasonable benefit / risk ratio. Pharmaceutically acceptable salts are well known in the art. For example, Berge et al. describe pharmaceutically acceptable salts in detail in J. Pharmaceutical Sciences, 1977, 66, 1-19, which is incorporated herein by reference. Pharmaceutically acceptable salts of the compounds of the present invention include those derived from suitable inorganic and organic acids and bases. Examples of pharmaceutically acceptable non-toxic acid addition salts are salts of amino groups formed with inorganic acids such as hydrochloric acid, hydrobromic acid, phosphoric acid, sulfuric acid, and perchloric acid, or with organic acids such as acetic acid, oxalic acid, maleic acid, tartaric acid, citric acid, succinic acid, or malonic acid, or by using other methods known in the art, such as ion exchange. Other pharmaceutically acceptable salts include adipate, alginate, ascorbate, aspartate, benzenesulfonate, benzoate, bisulfate, borate, butyrate, camphorate, camphorsulfonate, citrate, cyclopentanepropionate, digluconate, dodecyl sulfate, ethanesulfonate, formate, fumarate, glucoheptonate, glycerophosphate, gluconate, hemisulfate, heptanoate, hexanoate, hydroiodide, 2-hydroxy-ethanesulfonate, lanthanide ... Salts derived from appropriate bases include alkali metal salts, alkaline earth metal salts, ammonium salts, and ammonium salts. + (C 1~4 Alkyl)4 -Representative alkali metal or alkaline earth metal salts include sodium, lithium, potassium, calcium, magnesium, etc. Further pharmaceutically acceptable salts include non-toxic ammonium, quaternary ammonium, and amine cations, formed where appropriate using counterions such as halides, hydroxides, carboxylates, sulfates, phosphates, nitrates, lower alkylsulfonates, and arylsulfonates.
[0085] The term "solvate" refers to a form of a compound that is associated with a solvent, usually through solvolysis. This physical association may involve hydrogen bonding. Conventional solvents include water, methanol, ethanol, acetic acid, DMSO, THF, diethyl ether, etc. Metformin, for example, may be prepared in crystalline form and solvated. Suitable solvates include pharmaceutically acceptable solvates, further including both stoichiometric and non-stoichiometric solvates. In certain cases, for example, when one or more solvent molecules are incorporated into the crystal lattice of a crystalline solid, the solvate is isolable. "Solvate" encompasses both solution-phase and isolable solvates. Representative solvates include hydrates, ethanolates, and methanolates.
[0086] The term "hydrate" refers to a compound associated with water. Typically, the number of water molecules contained in a hydrate of a compound is a fixed ratio to the number of compound molecules in the hydrate. Thus, a hydrate of a compound can be represented, for example, by the general formula R·xH2O, where R is the compound and x is a number greater than 0. A given compound may form more than one hydrate, including, for example, a monohydrate (x is 1), a lower hydrate [where x is a number greater than 0 and less than 1, e.g., a hemihydrate (R·0.5H2O)], and a polyhydrate [where x is a number greater than 1, e.g., a dihydrate (R·2H2O) and a hexahydrate (R·6H2O)].
[0087] The term "tautomer" or "tautomerism" refers to two or more interconvertible compounds resulting from the formal migration of at least one hydrogen atom and at least one change in valence (e.g., from a single bond to a double bond, a triple bond to a single bond, or vice versa). The exact ratio of tautomers varies depending on several factors, including temperature, solvent, and pH. Tautomerization (i.e., the reaction resulting in a pair of tautomers) can be catalyzed by acid or base. Exemplary tautomerizations include keto to enol, amide to imide, lactam to lactim, enamine to imine, and enamine to (different enamine) tautomerization.
[0088] It should also be understood that compounds that have the same molecular formula but that differ in the nature or sequence of bonding of their atoms or in the arrangement of their atoms in space are termed "isomers." Isomers that differ in the arrangement of their atoms in space are termed "stereoisomers."
[0089] Stereoisomers that are not mirror images of one another are called "diastereomers," while those that are non-superimposable mirror images of each other are called "enantiomers." When a compound has an asymmetric center, for example, when it is bonded to four different groups, a pair of enantiomers is possible. Enantiomers can be characterized by the absolute configuration of their asymmetric center and described by the RS ordering rules of Cahn and Prelog or by the way the molecule rotates the plane of polarized light, and are designated as dextrorotatory or levorotatory (i.e., (+)- or (-)-isomer, respectively). Chiral compounds can exist as individual enantiomers or as mixtures thereof. A mixture containing equal proportions of enantiomers is called a "racemic mixture."
[0090] The term "prodrug" refers to a compound having a cleavable group that becomes a pharmaceutically active compound described herein upon solvolysis or under physiological conditions. Examples include, but are not limited to, choline ester derivatives, N-alkylmorpholine esters, and the like. Other derivatives of the compounds described herein are active in both the acid and acid-derivative forms, although the acid-labile forms often offer advantages such as solubility, tissue compatibility, or delayed release in mammalian organisms (see Bundgard, H., Design of Prodrugs, pp. 7-9, 21-24, Elsevier, Amsterdam 1985). Prodrugs include acid derivatives well known to those skilled in the art, such as esters prepared by reacting the parent acid with a suitable alcohol, or amides prepared by reacting the parent acid with a substituted or unsubstituted amine, or anhydride, or mixed anhydride. Simple aliphatic or aromatic esters, amides, and anhydrides derived from acidic groups pendant on the compounds described herein are specific prodrugs. In some cases, it is desirable to prepare double ester prodrugs, such as (acyloxy)alkyl esters or [(alkoxycarbonyl)oxy]alkyl esters. The C1-C8 alkyl, C2-C8 alkenyl, C2-C8 alkynyl, aryl, C7-C8 alkyl esters of the compounds described herein are also suitable. 12 Substituted aryl and C7-C 12 Aryl alkyl esters may be preferred.
[0091] inhibitory nucleic acid In some embodiments, the agent (e.g., a therapeutic agent and / or an anti-RAN protein agent) is an inhibitory nucleic acid. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a nucleic acid sequence (e.g., a target sequence, such as a target sequence in an RNA transcript). In some embodiments, the length and degree of sequence complementarity are sufficient to base pair with the nucleic acid sequence (e.g., the target sequence) in a specific and / or stable manner. In some embodiments, the inhibitory nucleic acid comprises about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, or more nucleotides. In some embodiments, the degree of sequence complementarity required for hybridization with a nucleic acid sequence (e.g., a target sequence) is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%. In some embodiments, the inhibitory nucleic acid comprises 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, or more nucleotides that are complementary to a nucleic acid sequence (e.g., a target sequence).
[0092] In some embodiments, the inhibitory nucleic acid encodes or is capable of hybridizing to a gene product that interacts with a RAN protein (e.g., eIF2 or a subunit thereof, eIF3 or a subunit thereof, PKR, p62, LC3 I subunit, LC3 II subunit, TARBP2, or TLR3). In some embodiments, the eIF2 gene or gene product comprises the sequence set forth in GenBank Accession No. NM_004094.4. In some embodiments, the eIF2A gene or gene product comprises the sequence set forth in GenBank Accession No. NM_032025.4. In some embodiments, the eIF3 gene or gene product is identified by GenBank Accession No. NM_003750.2 (eIF3a), GenBank Accession No. NM_003751.3 (eIF3b), GenBank Accession No. NM_003752.4 (eIF3c), GenBank Accession No. NM_003753.3 (eIF3d), GenBank Accession No. NM_001568.2 (eIF3e), GenBank Accession No. NM_003754.2 (eIF3f), GenBank Accession No. NM_001568.2 (eIF3e), GenBank Accession No. NM_003755.2 (eIF3f), GenBank Accession No. NM_001568 ... In some embodiments, the PKR gene or gene product comprises the sequence set forth in GenBank Accession No. NP_002750.1, such as the sequence set forth in GenBank Accession No. NM_003755.4 (eIF3g), GenBank Accession No. NM_003756.2 (eIF3h), GenBank Accession No. NM_003757.3 (eIF3i), GenBank Accession No. NM_003758.3 (eIF3j), GenBank Accession No. NM_013234.3 (eIF3k), GenBank Accession No. NM_016091.3 (eIF3l), GenBank Accession No. NM_006360.5 (eiF3m), etc. In some embodiments, the PKR gene or gene product comprises the sequence set forth in GenBank Accession No. NP_002750.1.
[0093] In some embodiments, the inhibitory nucleic acid is an antisense oligonucleotide (ASO). In some embodiments, the ASO comprises a short (approximately 15-30 nucleotide) sequence complementary to a nucleic acid sequence (e.g., a target sequence in an RNA transcript encoding an interrupted RAN protein or a gene product that interacts with the RAN protein, such as eIF2 or its subunits, eIF3 or its subunits, PKR, p62, LC3 I subunit, LC3 II subunit, TARBP2, or TLR3). In some embodiments, the ASO comprises naturally occurring nucleotides and / or modified (e.g., chemically modified) nucleotides. In some embodiments, the nucleotides may be modified by replacing ribose with alternative sugar moieties such as 2'-deoxyribose or 2'-O-(2-methoxyethyl)ribose, methylation, and / or replacing internucleotide phosphodiester linkages with phosphorothioate linkages. In some embodiments, modifications of some nucleotides at both the 3' and 5' ends of the ASO inhibit degradation by ubiquitous RNA nucleases active at the termini, thus improving the stability and thus half-life of the antisense oligo. However, in some embodiments, once at least a portion of the ASO is complexed with an mRNA encoding a RAN protein (e.g., a disrupted RAN protein), it may be desirable to promote the activity of RNase H to promote enzymatic degradation of the mRNA complexed with the ASO. In some embodiments, the ASO targets RNA (e.g., mRNA) corresponding to a gene containing a microsatellite repeat sequence. In some embodiments, the antisense oligonucleotide inhibits the translation of one or more disrupted RAN proteins. In some embodiments, the antisense oligonucleotide hybridizes to the mRNA sequence encoding the target protein, thereby blocking translation of the target protein and thereby inhibiting protein synthesis by the ribosomal machinery.
[0094] In some embodiments, the inhibitory nucleic acid is an interfering RNA. In some embodiments, the interfering RNA comprises a sequence that is complementary to between 5 and 50 consecutive nucleotides (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, about 30, about 35, about 40, or about 50 consecutive nucleotides) of a nucleic acid sequence (e.g., a target sequence in an RNA transcript encoding an interrupted RAN protein, or a gene product that interacts with RAN protein, such as eIF2 or a subunit thereof, eIF3 or a subunit thereof, PKR, p62, LC3 I subunit, LC3 II subunit, TARBP2, or TLR3). Non-limiting examples of interfering RNA include dsRNA, siRNA, shRNA, miRNA, and ami-RNA. In some embodiments, the inhibitory nucleic acid is a nucleic acid aptamer (for example, an RNA aptamer or a DNA aptamer). In some embodiments, the inhibitory RNA molecule can be unmodified or modified. Because it is recognized in the art that phosphorothioate modification, 2'-O-methyl modification, etc. improve the stability of oligonucleotides in vivo, in some embodiments, the inhibitory RNA molecule comprises one or more modified oligonucleotides, for example, the oligonucleotides that have undergone the above-mentioned modifications. In some embodiments, the interfering RNA is eIF3f siRNA (e.g., Dharmacon catalog number J-019535-08), eIF3m siRNA (e.g., Dharmacon catalog number J-016219-12), eIF3h siRNA (e.g., Dharmacon catalog number J-003883-07), or a variant thereof, e.g., an shRNA comprising the same sense and / or antisense sequences and further comprising a suitable shRNA loop sequence.
[0095] In some embodiments, the inhibitory nucleic acid is an aptamer. In some embodiments, the aptamer comprises a sequence that is complementary to and / or can bind to a nucleic acid sequence (e.g., a target sequence in an RNA transcript encoding a disrupted RAN protein, or a target sequence in a gene product that interacts with RAN protein, such as eIF2 or its subunit, eIF3 or its subunit, PKR, p62, LC3 I subunit, LC3 II subunit, TARBP2, or TLR3). In some embodiments, the aptamer comprises a sequence that is complementary to and / or can bind to a RAN protein (e.g., a disrupted RAN protein) or a protein that interacts with RAN protein (e.g., eIF2 or its subunit, eIF3 or its subunit, PKR, p62, LC3 I subunit, LC3 II subunit, TARBP2, or TLR3).
[0096] In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a nucleic acid sequence (e.g., a target sequence) present in a gene, chromosome, or RNA transcript (e.g., mRNA) encoding the interrupted RAN protein. In some embodiments, the target sequence comprises an extended repeat unit. In some embodiments, the inhibitory nucleic acid is selected from the group consisting of GGTCGT, GGCCGT, GGACGT, GGGCGT, GGTCGC, GGCCGC, GGACGC, GGGCGC, GGTCGA, GGCCGA, GGACGA, GGGCGA, GGTCGG, GGCCGG, GGACGG, GGGCGG, GGTAGA, GGCAGA, GGAAGA, GGGAGA, GGTAGG, GGCAGG, GGAAGG, GGGAGG GGTGCT, GGCGCC, GGAGCA, GGGGCG, GGTGCC, GGCGCA, GGAGCG, GGGGCT, GGTGCA, GGCGCG, GGAGCT, GGGGCC, GGTGCG, GGCGCT, GGAGCC, GGGGCA GGTCCT, GGCCCC, GGACCA, GGGCCG, GGTCCC, GGCCCA, GGACCG, GGGCCT, GGTCCA, GGCCCG, GGACCT, GGGCCC, GGTCCG, GGCCCT, GGACCC, GGGCCA CCTCGT, CCCCGT, CCACGT, CCGCGT, CCTCGC, CCCCGC, CCACGC, CCGCGC, CCTCGA, CCCCGA, CCACGA, CCGCGA, CCTCGG, CCCCGG, CCACGG, CCGCGG, CCTAGA, CCCAGA, CCAAGA, CCGAGA, CCTAGG, CCCAGG, CCAAGG, CCGAGG TCT, TCC, TCA, TCG, AGT, AGC, CCTG, and / or CAGG repeat units, or a portion thereof.In some embodiments, the target sequence, or a portion thereof, is (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x, (GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGAGA)x, (GGTAGG)x, (G GCAGG)x, (GGAAGG)x, (GGGAGG)x, (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GGAGCG)x, (GGGGCT)x, (GGTGCA)x, (G GCGCG)x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, (GGCCCC)x, (GGACCA)x, (GGGCCG)x, (GGTCCC)x, (GG CCCA)x, (GGACCG)x, (GGGCCT)x, (GGTCCA)x, (GGCCCG)x, (GGACCT)x, (GGGCCC)x, (GGTCCG)x, (GGCCCT)x, (GGACCC)x, (GGGCCA)x, (CCTCGT)x, (CC CCGT)x, (CCACGT)x, (CCGCGT)x, (CCTCGC)x, (CCCCGC)x, (CCACGC)x, (CCGCGC)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCGCGA)x, (CCTCGG)x, (CCC In some embodiments, the nucleic acid encoding an interrupted RAN protein comprises (CGG), (CCACGG), (CCGCGG), (CCTAGA), (CCCAGA), (CCAAGA), (CCGAGA), (CCTAGG), (CCCAGG), (CCAAGG), (CCGAGG), (CCTG), (TCT), (TCC), (TCA), (TCG), (AGT), and / or (AGC) repeat units, where x represents the number of repeat units present in the nucleic acid encoding the interrupted RAN protein. In some embodiments, x comprises an integer between 2 and 200.In some embodiments, x comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, x is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 8, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168 13, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200.In some embodiments, "x" comprises an integer between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000.
[0097] In some embodiments, the inhibitory nucleic acid is a nucleic acid that is associated with amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA9), spinocerebellar ataxia type 9 (SCA10), spinocerebellar ataxia type 10 (SCA11), spinocerebellar ataxia type 11 (SCA12), spinocerebellar ataxia type 12 (SCA13), spinocerebellar ataxia type 13 (SCA14), spinocerebellar ataxia type 14 (SCA15), spinocerebellar ataxia type 15 (SCA16), spinocerebellar ataxia type 16 (SCA17), spinocerebellar ataxia type 17 (SCA18), spinocerebellar ataxia type 18 (SCA19), spinocerebellar ataxia type 19 ... 7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), or Fuchs' corneal dystrophy. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a target sequence corresponding to a gene or chromosomal locus set forth in Table 1 or Table 6. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a target sequence corresponding to a gene or chromosomal locus set forth in Table 1 or Table 6. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a target sequence corresponding to a gene or chromosomal locus set forth in Table 1 or Table 6. ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, PFN1P2, C1orf 61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDI N2, AIFM2, MGMT, LINC01164, KNDC1, ANK3, MLLT10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX4L2, AGAP12P, C10orf53, SMPD1, IFITM 10, BUD13, TSPAN18, CD82, OTUB1, NADSYN1, MIR4492, CHID1, SMUG1, LINC00938, LINC01257, SLC15A4, ASIC1, DCP1B, TMTC2, TNS2,LOC100240735, SOX1, LATS2, RAB20, ANKRD20A9P, FLT1, RCBTB1, ELK2AP, ST ON2, FOXN3, TTLL5, BCL11B, BRMS1L, SMAD3, RPLP1, BAHD1, MYO5A, DNM1P46. DNM1P35, TUBGCP4, C16orf95, OSGIN1, LINC00311, MIR4718, RBFOX1, SBK1. MIR4722, BANP, C16orf78, MIR5189, ADGRG5, NPRL3, ZDHHC1, MIR662, LINC00 482 MRPL12 TBC1D3H WSCD1 TBC1D3B TBC1D3 TBC1D3C METRNL DNAH9A SGR1, FOXK2, NPEPPS, SARM1, CLUH, TIAF1, LOC440434, PHOSPHO1, TBCD, CYP 4F35P, CXADRP3, LINC00668, MEX3C, COX7A1, SCAF1, RFPL4AL1, SIX5, DIRAS 1, POLRMT, ZNF554, MED16, SIPA1L3, DOT1L, KMT5C, PDE4C, ZNF480, CEBPA, PT PN18, HAAO, LOC654342, RGPD2, RAB3GAP1, TNS1, FAM95A, NTSR1, FRG1BP, FA M182B, CDH4, PRNP, MIR1257, MIR4758, OGFR, SRC, COL9A3, ZNF512B, PICSAR SIK1, CYP4F29P, MX1, LARGE1, CRELD2, UPK3A, RRP7A, MIR4762, SHANK3, SH ISA8, CCDC188, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CACNA2D2 TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10, TNIP3, ZFY VE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2, SLC3 4A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, MEA1, RI MS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4.The hybridization target may be capable of hybridizing to a target sequence corresponding to a gene selected from MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8, or a portion thereof. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a target sequence corresponding to ARMCX4, ALK, or CASP8. In some embodiments, the inhibitory nucleic acid is capable of hybridizing to a RAN repeat unit in the configuration shown in Figure 9F (e.g., the target sequence includes a CASP8 sequence, such as one that includes an insertion of a RAN repeat unit (SEQ ID NO: 194) within a VNTR sequence, as shown in Figure 9F).
[0098] Guide nucleic acids and RNA-guided nucleases In some embodiments, the agent (e.g., therapeutic agent) is a guide RNA (gRNA) and / or an RNA-guided nuclease (e.g., a complex comprising a gRNA and a Cas nuclease). The terms "gRNA" and "guide RNA" may be used interchangeably throughout to refer to a nucleic acid comprising a sequence that physically interacts with (e.g., binds to) and localizes an RNA-guided nuclease (e.g., a Cas9 nuclease) to a target sequence. In some embodiments, a gRNA comprises a targeting sequence that is 5' to a scaffold sequence. A "targeting sequence" refers to a sequence of a gRNA used to localize an RNA-guided nuclease (e.g., a Cas9 nuclease) to a target DNA. A "scaffold sequence" refers to a sequence of a gRNA that is responsible for RNA-guided nuclease binding and does not comprise a targeting sequence. An "sgRNA" refers to a gRNA comprising a scaffold sequence that is a fusion of an endogenous bacterial crRNA and a tracrRNA. In some embodiments, an sgRNA comprises a targeting sequence.
[0099] In some embodiments, the targeting sequence, or a portion thereof, hybridizes to (e.g., is partially or fully complementary to) a target sequence in a nucleic acid encoding an interrupted RAN protein. In some embodiments, the targeting sequence is at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or at least 99% complementary to the target sequence in a nucleic acid encoding an interrupted RAN protein. In some embodiments, the targeting sequence is 100% complementary to the target sequence in a nucleic acid encoding an interrupted RAN protein. In some embodiments, the targeting sequence, or a portion thereof, that hybridizes to the target sequence in a nucleic acid encoding an interrupted RAN protein can be 15-25 nucleotides, 18-22 nucleotides, or 19-21 nucleotides in length. In some embodiments, the targeting sequence or a portion thereof that hybridizes to a target sequence in a nucleic acid encoding an interrupted RAN protein is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the targeting sequence or a portion thereof that hybridizes to a target sequence in a nucleic acid encoding an interrupted RAN protein is between 10 and 30, or between 15 and 25 nucleotides in length. In some embodiments, the targeting sequence or a portion thereof that hybridizes to a target sequence in a nucleic acid encoding an interrupted RAN protein is 20 nucleotides in length. In some embodiments, the targeting sequence comprises a sequence set forth in Table 4.
[0100] In some embodiments, the gRNA is selected from the group consisting of GGTCGT, GGCCGT, GGACGT, GGGCGT, GGTCGC, GGCCGC, GGACGC, GGGCGC, GGTCGA, GGCCGA, GGACGA, GGGCGA, GGTCGG, GGCCGG, GGACGG, GGGCGG, GGTAGA, GGCAGA, GGAAGA, GGGAGA, GGTAGG, GGCAGG, GGAAGG, GGGAGG, GGTGCT, GGCGCC, GGAGCA, GGGGCG, GGTGCC, GGCGCA, GGAGCG, GGGGCT, GGTGCA, GGCGCG, GGAGCT, GGGGCC, GGTGCG, GGCGCT, GGAGCC, GGGGCA GGTCCT, GGCCCC, GGACCA, GGGCCG, GGTCCC, GGCCCA, GGACCG, GGGCCT, GGTCCA, GGCCCG, GGACCT, GGGCCC, GGTCCG, GGCCCT, GGACCC, GGGCCA CCTCGT, CCCCGT, CCACGT, CCGCGT, CCTCGC, CCCCGC, CCACGC, CCGCGC, CCTCGA, CCCCGA, CCACGA, CCGCGA, CCTCGG, CCCCGG, CCACGG, CCGCGG, CCTAGA, CCCAGA, CCAAGA, CCGAGA, CCTAGG, CCCAGG, CCAAGG, CCGAGG TCT, TCC, TCA, TCG, AGT, AGC, CCTG, and / or CAGG repeat units, or a portion thereof. In some embodiments, the target sequence is (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x, (GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGAGA)x, (GGTAGG)x, (GGCAGG)x, (GGAAGG)x, (GGGAGG)x, (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GGAGCG)x,(GGGGCT)x, (GGTGCA)x, (GGCGCG)x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, (GGCCCC)x, (GGACCA)x, (GGGCCG)x, (GGTCCC)x, (GGCCCA)x, (GGACC G) x, (GGGCCT) In some embodiments, the nucleic acid encoding an interrupted RAN protein comprises a repeat unit selected from the group consisting of (CACGC)x, (CCGCGC)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCGCGA)x, (CCTCGG)x, (CCCCGG)x, (CCACGG)x, (CCGCGG)x, (CCTAGA)x, (CCCAGA)x, (CCAAGA)x, (CCGAGA)x, (CCTAGG)x, (CCCAGG)x, (CCAAGG)x, (CCGAGG)x, (CCTG)x, (TCT)x, (TCC)x, (TCA)x, (TCG)x, (AGT)x, and / or (AGC)x, wherein x represents the number of repeat units present in the nucleic acid encoding the interrupted RAN protein. In some embodiments, x comprises an integer between 2 and 200. In some embodiments, x comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, x is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79,80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143 , 144, 145, 146, 147, 148, 149, 150, 151, 152, 152, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200. In some embodiments, "x" comprises an integer between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000. In some embodiments, the target sequence, or a portion thereof, is selected from the group consisting of PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8, ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, PFN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1,HYDIN2, AIFM2, MGMT, LINC01164, KNDC1, ANK3, MLLT10, TBC1D12, LRMDA, CC NY, MIR3156-1, DUX4L2, AGAP12P, C10orf53, SMPD1, IFITM10, BUD13, TSPAN 18. CD82, OTUB1, NADSYN1, MIR4492, CHID1, SMUG1, LINC00938, LINC01257. SLC15A4, ASIC1, DCP1B, TMTC2, TNS2, LOC100240735, SOX1, LATS2, RAB20, AN KRD20A9P, FLT1, RCBTB1, ELK2AP, STON2, FOXN3, TTLL5, BCL11B, BRMS1L, SM AD3, RPLP1, BAHD1, MYO5A, DNM1P46, DNM1P35, TUBGCP4, C16orf95, OSGIN1. LINC00311, MIR4718, RBFOX1, SBK1, MIR4722, BANP, C16orf78, MIR5189, AD GRG5, NPRL3, ZDHHC1, MIR662, LINC00482, MRPL12, TBC1D3H, WSCD1, TBC1D3B TBC1D3, TBC1D3C, METRNL, DNAH9, ASGR1, FOXK2, NPEPPS, SARM1, CLUH, and TIA F1, LOC440434, PHOSPHO1, TBCD, CYP4F35P, CXADRP3, LINC00668, MEX3C, CO X7A1, SCAF1, RFPL4AL1, SIX5, DIRAS1, POLRMT, ZNF554, MED16, SIPA1L3, DO T1L, KMT5C, PDE4C, ZNF480, CEBPA, PTPN18, HAAO, LOC654342, RGPD2, RAB3GA P1, TNS1, FAM95A, NTSR1, FRG1BP, FAM182B, CDH4, PRNP, MIR1257, MIR4758. OGFR, SRC, COL9A3, ZNF512B, PICSAR, SIK1, CYP4F29P, MX1, LARGE1, CRELD2 UPK3A, RRP7A, MIR4762, SHANK3, SHISA8, CCDC188, NPTXR, ZNF621, TPRA1, P.S IGZ, LHFPL4, OSTN, GAP43, CACNA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1.DOCK3, DUX4L8, PCDH10, TNIP3, ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2, SLC34A1, CXXC5, PPARGC1B, LOC643201, P4HA 2, THBS2, SEC63, SLC17A5, MEA1, RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, The gene corresponds to a gene selected from DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA) 5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8.
[0101] In some embodiments, the gRNA is a modified gRNA (e.g., a chemically modified gRNA). Methods for designing gRNAs and chemically modified gRNAs will be apparent to those skilled in the art [e.g., Jinek, et al. Science (2012) 337(6096):816-821, Ran, et al. Nature Protocols (2013) 8:2281-2308, International Publication Nos. WO2014 / 093694 and WO2013 / 176772, Vanegas et al., Fungal Biol Biotechnol. 2019; 6: 6, Fu Y et al., Nat Biotechnol 2014 (doi: 10.1038 / nbt.2808), Sternberg SH et al., Nature 2014 (doi: 10.1038 / naturel3011), Rahdar et al. PNAS (2015) 112 (51) E7110-E7117, and Hendel et al., Nat Biotechnol. (2015); 33(9): 985-989, International Publication Nos. WO2017 / 214460, WO2016 / 089433, and WO2016 / 164356].
[0102] As used herein, "RNA-guided nuclease" may be used to refer to a protein comprising a nuclease domain or a variant thereof (e.g., a catalytically inactive nuclease or a partially catalytically inactive nuclease) that physically interacts with an RNA molecule (e.g., a guide RNA) that localizes the nuclease to a site within a target DNA sequence. A "ribonucleoprotein complex" or "RNP complex" may be used to refer to an RNA-guided nuclease (e.g., a Cas nuclease, such as a Cas9 nuclease) bound to a gRNA.
[0103] A variety of RNA-guided nucleases, including Cas nucleases such as Cas9 nuclease, are known in the art [see, e.g., Gill et al. LIPSCOMB 2017. In United States: Inscripta Inc., Price et al. Biotechnol. Bioeng. (2020) 117(60): 1805-1816, International Publication No. WO2015 / 157070; see, e.g., Gao et al. Nat. Biotechnol. (2017) 35(8): 789-792, Liang et al. Nat. Comm. (2022) 13: 3421, Walton et al. Science (2020) 368 (6488): 290-296, Slaymaker et al. Science (2016) 351 (6268): 84-88, Kleinstiver et al. al. Nature (2016) 529: 490-495, Stella et al. Nature Structural & Molecular Biology (2017), Shmakov et al. Mol Cell (2015) 60: 385-397, Strohkendl et al. Mol. Cell (2018) 71: 1-9].
[0104] In some embodiments, the RNA-guided nuclease comprises a Cas nuclease variant (e.g., a Cas9 variant). In some embodiments, the Cas nuclease variant comprises one or more mutations, wherein the one or more mutations comprise one or more amino acid point mutations, substitutions, insertions, and / or deletions. In some embodiments, the Cas nuclease variant is a catalytically inactive or partially catalytically inactive Cas nuclease variant (see, e.g., Yeh et al. Nat. Cell. Biol. (2019) 21: 1468-1478; e.g., Hsu et al. Cell (2014) 157: 1262-1278; Jasin et al. DNA Repair (2016) 44: 6-16; Sfeir et al. Trends Biochem. Sci. (2015) 40: 701-714). In some embodiments, the Cas nuclease variant is Cas9-NRTH, inactive (dead) Cas9 (dCas9), Cas9-NG, or Cas9-NRCH.
[0105] In some embodiments, the RNA-guided nuclease targets and deaminates specific nucleic acid bases. In some embodiments, the RNA-guided nuclease further comprises a deaminase. In some embodiments, the deaminase is fused to the RNA-guided nuclease at one end (e.g., the N-terminus or C-terminus of the RNA-guided nuclease). In some embodiments, the RNA-guided nuclease comprises an adenosine deaminase. In some embodiments, the RNA-guided nuclease comprises a cytidine deaminase. In some embodiments, the RNA-guided nuclease comprises a cytidine deaminase and one or more uracil glycosylase inhibitor domains. In some embodiments, the RNA-guided nuclease is an adenine base editor (ABE) or a cytosine base editor (CBE). Various RNA-guided nucleases and / or deaminases that can be used in conjunction with Cas nucleases for base editing are known in the art [e.g., Komor et al. Nature (2016) 533: 420-424, Rees et al. Nat. Rev. Genet. (2018) 19(12): 770-788, Anzalone et al. Nat. Biotechnol. (2020) 38: 824-844, Eid et al. Biochem. J. (2018) 475(11): 1965-1964, Rees et al. Nature Reviews Genetics (2018) 19:770-788, U.S. Publication No. 2018 / 0312825(A1), U.S. Publication No. 2018 / 0312828(A1), and International Publication No. WO2018 / 165629(A1)].
[0106] In some embodiments, the RNA-guided nuclease is fused to an engineered reverse transcriptase (RT) domain. In some embodiments, the RNA-guided nuclease is a prime editor (see, e.g., Anzalone et al. Nature (2019) 576 (7785):149-157).
[0107] In some embodiments, an RNA-guided nuclease can recognize (e.g., bind to) a PAM sequence. As used herein, the term "protospacer adjacent motif" or "PAM" refers to a DNA sequence that may be required for an RNA-guided nuclease and gRNA (e.g., Cas9 / sgRNA) to probe a specific DNA sequence through Watson-Crick pairing of the guide RNA with the genome. In some embodiments, PAM specificity may depend on the DNA-binding specificity of a "protospacer adjacent motif recognition domain" at the C-terminus of an RNA-guided nuclease, e.g., Cas9. In some embodiments, a PAM comprises a 2-8 base pair DNA sequence (e.g., a 2, 3, 4, 5, 6, 7, or 8 base pair DNA sequence) immediately downstream or upstream of a target sequence that can be directly recognized by an RNA-guided nuclease to facilitate cleavage of the target site or, in the case of a nuclease-deficient Cas, enable binding to DNA at that locus. In some embodiments, recognition of the PAM sequence involves an RNA-guided nuclease and gRNA forming an R-loop.
[0108] In some embodiments, the PAM can be a 5' PAM located upstream of the 5' end of the target sequence. In other embodiments, the PAM can be a 3' PAM located downstream of the 5' end of the target sequence. In some embodiments, the PAM is located on the same strand of the target sequence. In some embodiments, the PAM is 1 to 30 nucleotides upstream or downstream from the target sequence. In some embodiments, the PAM is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream or downstream from the target sequence.
[0109] In some embodiments, the PAM comprises the sequence NGG, NAG, NGCG, NGAG, NGAN, NGNG, NG, GAA, GAT, NNGRRT, NNGRR(N), TTTV, TYCV, TATV, NNNNRYAC, NNNNRYAC, NNNNRYAC, or NAAAAC, where "N" is any nucleotide or base, "R" is A or guanine (G), and "Y" is C or T. In some embodiments, the PAM is selected from the group consisting of GGTCGT, GGCCGT, GGACGT, GGGCGT, GGTCGC, GGCCGC, GGACGC, GGGCGC, GGTCGA, GGCCGA, GGACGA, GGGCGA, GGTCGG, GGCCGG, GGACGG, GGGCGG, GGTAGA, GGCAGA, GGAAGA, GGGAGA, GGTAGG, GGCAGG, GGAAGG, GGGAGG, GGTGCT, GGCGCC, GGAGCA, GGGGCG, GGTGCC, GGCGCA, GGAGCG, GGGGCT, GGTGCA, GGCGCG, GGAGCT, GGGGCC, GGTGCG, GGCGCT, GGAGCC, GGGGCA GGTCCT, GGCCCC, GGACCA, GGGCCG, GGTCCC, GGCCCA, GGACCG, GGGCCT, GGTCCA, GGCCCG, GGACCT, GGGCCC, GGTCCG, GGCCCT, GGACCC, GGGCCA In some embodiments, the PAM is present in a repeat unit comprising a CCTCGT, CCCCGT, CCACGT, CCGCGT, CCTCGC, CCCCGC, CCACGC, CCGCGC, CCTCGA, CCCCGA, CCACGA, CCGCGA, CCTCGG, CCCCGG, CCACGG, CCGCGG, CCTAGA, CCCAGA, CCAAGA, CCGAGA, CCTAGG, CCCAGG, CCAAGG, CCGAGG TCT, TCC, TCA, TCG, AGT, AGC, CCTG, and / or CAGG repeat unit. In some embodiments, the PAM is present in a repeat unit comprising a (GGTCGT)x, (GGCCGT)x, (GGACGT)x, (GGGCGT)x, (GGTCGC)x, (GGCCGC)x, (GGACGC)x, (GGGCGC)x, (GGTCGA)x, (GGCCGA)x, (GGACGA)x, (GGGCGA)x, (GGTCGG)x, (GGCCGG)x,(GGACGG)x, (GGGCGG)x, (GGTAGA)x, (GGCAGA)x, (GGAAGA)x, (GGGAGA)x, (GGTAGG)x, (GGCAGG)x, (GGAAGG)x, (GGGAGG)x , (GGTGCT)x, (GGCGCC)x, (GGAGCA)x, (GGGGCG)x, (GGTGCC)x, (GGCGCA)x, (GGAGCG)x, (GGGGCT)x, (GGTGCA)x, (GGCGCG) x, (GGAGCT)x, (GGGGCC)x, (GGTGCG)x, (GGCGCT)x, (GGAGCC)x, (GGGGCA)x, (GGTCCT)x, (GGCCCC)x, (GGACCA)x, (GGGCCG )x, (GGTCCC)x, (GGCCCA)x, (GGACCG)x, (GGGCCT)x, (GGTCCA)x, (GGCCCG)x, (GGACCT)x, (GGGCCC)x, (GGTCCG)x, (GGCCCT )x, (GGACCC)x, (GGGCCA)x, (CCTCGT)x, (CCCCGT)x, (CCACGT)x, (CCGCGT)x, (CCTCGC)x, (CCCCGC)x, (CCACGC)x, (CCGCG C)x, (CCTCGA)x, (CCCCGA)x, (CCACGA)x, (CCGCGA)x, (CCTCGG)x, (CCCCGG)x, (CCACGG)x, (CCGCGG)x, (CCTAGA)x, (CCCAG In some embodiments, x represents the number of repeat units present in the nucleic acid encoding an interrupted RAN protein. In some embodiments, x comprises an integer between 2 and 200. In some embodiments, x comprises an integer between 2 and 175, between 2 and 150, between 2 and 125, between 2 and 100, between 2 and 75, between 2 and 50, between 2 and 25, between 2 and 10, between 2 and 5, etc. In some embodiments, x is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 14 6, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 22, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 209, 2010, 2011, 2012, 2013, 2014, 2015, 2016, 2017, 2018, 2019, 2020, 2021, 2022, 2023, 2024, 2025, 2026, 2027, 202 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200. In some embodiments, "x" comprises an integer greater than 200. In some embodiments, "x" comprises an integer between 200 and 250, 250 and 300, 300 and 400, 400 and 500, 500 and 750, 750 and 1,000, 1,000 and 1,500, 1,500 and 2,000, 2,000 and 3,000, 3,000 and 4,000, 4,000 and 5,000, 5,000 and 6,000, 6,000 and 7,000, 8,000 and 9,000, or 9,000 and 10,000. In some embodiments, the PAM is selected from the group consisting of PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8 ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP,TCF4、HTT、APP、ARMCX4、PEX14、PTPRF、ACTA1、DNAH14、PFN1P2、C1orf61、WASF2、PGBD2、NBPF15、DDX11L1、BARHL2、MIR1976、CASZ1、SLC44A3、GPR137B、SOX13、CROCC、RNPEP、MIR3121、MPZ、MCL1、HYDIN2、AIFM2、MGMT、LINC01164、KNDC1、ANK3、MLLT10、TBC1D12、LRMDA、CCNY、MIR3156-1、DUX4L2、AGAP12P、C10orf53、SMPD1、IFITM10、BUD13、TSPAN18、CD82、OTUB1、NADSYN1、MIR4492、CHID1、SMUG1、LINC00938、LINC01257、SLC15A4、ASIC1、DCP1B、TMTC2、TNS2、LOC100240735、SOX1、LATS2、RAB20、ANKRD20A9P、FLT1、RCBTB1、ELK2AP、STON2、FOXN3、TTLL5、BCL11B、BRMS1L、SMAD3、RPLP1、BAHD1、MYO5A、DNM1P46、DNM1P35、TUBGCP4、C16orf95、OSGIN1、LINC00311、MIR4718、RBFOX1、SBK1、MIR4722、BANP、C16orf78、MIR5189、ADGRG5、NPRL3、ZDHHC1、MIR662、LINC00482、MRPL12、TBC1D3H、WSCD1、TBC1D3B、TBC1D3、TBC1D3C、METRNL、DNAH9、ASGR1、FOXK2、NPEPPS、SARM1、CLUH、TIAF1、LOC440434、PHOSPHO1、TBCD、CYP4F35P、CXADRP3、LINC00668、MEX3C、COX7A1、SCAF1、RFPL4AL1、SIX5、DIRAS1、POLRMT、ZNF554、MED16、SIPA1L3、DOT1L、KMT5C、PDE4C、ZNF480、CEBPA、PTPN18、HAAO、LOC654342、RGPD2、RAB3GAP1、TNS1、FAM95A、NTSR1、FRG1BP、FAM182B、CDH4、PRNP、MIR1257、MIR4758、OGFR、SRC、COL9A3、ZNF512B、PICSAR、SIK1, CYP4F29P, MX1, LARGE1, CRELD2, UPK3A, RRP7A, MIR4762, SHANK3, SHISA8, CCDC188, NPTXR, Z NF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CACNA2D2, TNK2, IQSEC1, RAD18, PARP14, PLXNA1, DOCK3, DUX4L8, PCDH10, TNIP3, ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX1, ZNF622, SPOCK1, PLEKHG4B, LCP2, S LC34A1, CXXC5, PPARGC1B, LOC643201, P4HA2, THBS2, SEC63, SLC17A5, MEA1, RIMS1, ARID1B, PRKAG2 , EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4, MFSD3, SOX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM27B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8. In some embodiments, the target sequence comprises a RAN repeat unit in the configuration shown in Figure 9F (e.g., the target sequence comprises a CASP8 sequence, such as one comprising an insertion of a RAN repeat unit (SEQ ID NO: 194) within a VNTR sequence, as shown in Figure 9F).
[0110] In some embodiments, the RNA-guided nuclease and gRNA are contacted with a target sequence or one or more cells containing the target sequence. In some embodiments, contacting the RNA-guided nuclease and gRNA with one or more cells comprises contacting the one or more cells with a nucleic acid encoding the gRNA. In some embodiments, the nucleic acid encoding the RNA-guided nuclease is contacted with one or more cells. In some embodiments, the RNA-guided nuclease and gRNA are encoded on the same nucleic acid. In some embodiments, the nucleic acid for delivery of the RNA-guided nuclease and / or gRNA comprises a regulatory sequence (e.g., a promoter sequence for increasing expression of the RNA-guided nuclease and gRNA) operably linked to the sequence encoding the RNA-guided nuclease and / or gRNA. In some embodiments, the RNA-guided nuclease and gRNA are contacted with one or more cells, and the RNA-guided nuclease is in the form of a protein. In some embodiments, the gRNA is bound to the RNA-guided nuclease (e.g., as a ribonucleoprotein complex). In some embodiments, delivery of a complex comprising an RNA-guided nuclease and a gRNA to one or more cells comprises electroporation.
[0111] Genetic mutations In some embodiments, the agent (e.g., a therapeutic agent and / or an anti-RAN protein agent) comprises a protein kinase R (PKR) mutant. As used herein, a "protein kinase R (PKR) mutant" refers to a protein comprising an amino acid sequence that is at least 70%, at least 80%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to wild-type protein kinase R (PKR) (e.g., GenBank Accession No. NP_002750.1), wherein the mutant protein comprises at least one amino acid change (sometimes referred to as a "mutation") compared to the amino acid sequence of wild-type PKR. In some embodiments, the amino acid sequence of the PKR mutant is at least 75%, at least 85%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, or at least 99.9% identical to the amino acid sequence of wild-type PKR. In some embodiments, the amino acid sequence is about 95-99.9% identical to the amino acid sequence of wild-type PKR. In some embodiments, the protein contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, or at least 15 different amino acid sequence mutations compared to the amino acid sequence set forth in the amino acid sequence of wild-type PKR. In some embodiments, the PKR mutant contains a mutation at position 296 (e.g., position 296 of human wild-type PKR). In some embodiments, the mutation at position 296 is K296R. In some embodiments, the PKR mutant is a dominant-negative PKR mutant. In some embodiments, the PKR mutant functions in a dominant-negative manner to inhibit phosphorylation of eIF2α.
[0112] vector In some embodiments, the agent (e.g., a therapeutic agent) is a vector or is expressed from a sequence contained in a vector. In some embodiments, the vector comprises a sequence encoding an inhibitory nucleic acid, gene variant, transgene, gRNA, and / or RNA-guided nuclease described herein. In some embodiments, any of the agents may be encoded by a sequence contained in a vector, provided to a cell (e.g., a cell in a subject), and expressed from the vector.
[0113] In some embodiments, the vector comprises deoxyribonucleotides. In some embodiments, the vector comprises ribonucleotides. In some embodiments, the vector comprises both deoxyribonucleotides and ribonucleotides. In some embodiments, the vector is single-stranded. In some embodiments, the vector is double-stranded. In some embodiments, the vector is circular (e.g., an artificial chromosome, such as an artificial bacterial chromosome, or a plasmid, such as a circular plasmid, nanoplasmid, and minicircle plasmid). In some embodiments, the vector is linear. In some embodiments, the vector is self-complementary. In some embodiments, the vector can be maintained at high levels within the cell using selection methods, such as those involving antibiotic resistance genes. In some embodiments, the vector can include a partitioning sequence that ensures stable inheritance of the vector. In some embodiments, the vector is a high copy number vector. In some embodiments, the vector is about 1 to 60 kb in size, e.g., 1 to 50 kb, 1 to 30 kb, 1 to 20 kb, etc., e.g., 1 to 15 kb, 1 to 10, etc., e.g., 1 to 8 kb, 2 to 7 kb, etc., e.g., 3 to 6 kb, 4 to 5 kb, etc. In some embodiments, the vector is small enough to be efficiently packaged into rAAV viral particles (e.g., less than 5 kb, less than 4 kb, less than 3 kb, less than 2 kb, etc.).
[0114] Recombinant adeno-associated virus (rAAV) nucleic acids and particles In some embodiments, the agent (e.g., therapeutic agent) is a recombinant adeno-associated virus (rAAV) particle. As used herein, the term "adeno-associated virus" or the abbreviation "AAV" refers to the virus itself or its derivatives. This term applies to all AAV subtypes, including both naturally occurring and recombinant forms, unless otherwise specified. The term "recombinant AAV (rAAV)" refers to a recombinant adeno-associated virus, which refers to AAV that contains a nucleic acid sequence (e.g., a heterologous nucleic acid) that is not derived from AAV.
[0115] In some embodiments, the nucleic acid sequence found in an rAAV is the "rAAV genome," which refers to a nucleic acid comprising a heterologous nucleic acid flanked by 5' and 3' AAV inverted terminal repeats (ITRs). In this context, the term "heterologous nucleic acid" can refer to any DNA sequence not normally found between flanking AAV ITRs. In some embodiments, the heterologous nucleic acid comprises at least one transgene. As used herein, "transgene" refers to a DNA sequence encoding at least one RNA to be expressed in a cell. In some embodiments, the heterologous nucleic acid (e.g., transgene) comprises a sequence encoding an agent (e.g., a therapeutic agent) described herein. In some embodiments, the agent is an inhibitory nucleic acid, gene variant, gRNA, and / or RNA-guided nuclease described herein. In some embodiments, the heterologous nucleic acid (e.g., transgene) comprises a sequence encoding an antibody or antigen-binding fragment described herein.
[0116] The term "AAV particle" or "rAAV particle" refers to a particle formed by one or more AAV capsid proteins. In some embodiments, AAV particles and rAAV particles comprise encapsidated nucleic acids (e.g., rAAV particles comprising an rAAV genome). In some embodiments, rAAV particles comprise nucleic acids (e.g., heterologous nucleic acids found between AAV ITRs, such as transgenes) encoding inhibitory nucleic acids, gene variants, gRNAs, and / or RNA-guided nucleases described herein.
[0117] In some embodiments, rAAV particles are packaged using packaging nucleic acid and / or helper nucleic acid. As used herein, "helper nucleic acid" refers to a nucleic acid (e.g., a helper vector or nucleic acid provided in a helper virus) that contains one or more genes (e.g., E1, E2A, E4, and / or VA) that function in trans for productive AAV replication and encapsidation. As used herein, "packaging nucleic acid" refers to a nucleic acid (e.g., a packaging vector) that provides nucleotide sequences (e.g., AAV rep and AAV capsid protein gene sequences) (e.g., accessory functions) on which AAV depends for replication.
[0118] In some embodiments, expression of RNA encoded by a transgene may be under the control of (e.g., operably linked to) one or more regulatory sequences (e.g., enhancers, promoters, transcription initiation sites, translation initiation sites, splicing acceptor / donor sites, transcription termination sites, stop codons, polyA signals, etc.). In some embodiments, regulatory sequences may be found between the AAV ITRs. However, in some embodiments, a nucleic acid comprising a transgene may include one or more regulatory elements that are not operably linked to the transgene.
[0119] In some embodiments, AAV particles and rAAV particles comprise encapsidated nucleic acids (e.g., rAAV particles comprising an rAAV genome). In some embodiments, one or more capsid proteins correspond to an AAV serotype, an AAV serotype derivative, or an AAV pseudotype. Non-limiting examples of AAV particle and rAAV particle serotypes of the present disclosure include mammalian AAV1, mammalian AAV2, mammalian AAV3, mammalian AAV4, mammalian AAV5, mammalian AAV6, mammalian AAV7, mammalian AAV8, mammalian AAV9, and mammalian AAV10. Non-limiting examples of rAAV pseudotypes include mammalian AAV2 / 1, mammalian AAV2 / 5, mammalian AAV2 / 6, mammalian AAV2 / 8, mammalian AAV2 / 9, mammalian AAV3 / 1, mammalian AAV3 / 5, mammalian AAV3 / 8, and mammalian AAV3 / 9, where the slash indicates that an rAAV genome of one serotype is packaged in a capsid from a different serotype (e.g., an rAAV genome containing AAV2 ITRs packaged in an AAV5 capsid would be AAV2 / 5).
[0120] In some embodiments, the rAAV particles can be any of a variety of vectors, including, for example, AAVrh.10, AAVrh.74, AAVhu.14, AAV3a / 3b, AAVrh32.33, AAV-HSC15, AAV-HSC17, AAVhu.37, AAVrh.8, CHt-P6, AAV2.5, AAV6.2, AAV2i8, AAV-HSC15 / 17, AAVM41, AAV9.45, AAV6(Y445F / Y731F), AAV2.5T, AAV-HAE1 / 2, AAV clone 32 / 83, AAVShHIO, AAV2(Y→F), AAV8(Y733F), AAV2.15, AAV2.4, AAVM41 AAVs may be engineered with hybrid or mutant mammalian AAV capsid protein derivatives, such as AAV2(pentaYF), AAV2-BCDG(T491V+K556R), AAV5-M2, AAV5(Y719F), AAV6(T492V+S663V), AAV6(T492V+Y705F+Y731F), AAV6(S551V+S663V), AAV8-C&G(T494V), AAV8-M3, AAV8(Y733F), AAV8(T494V+Y733F), AAV8(Y275F+Y447F+Y733F), AAV9-PHP.B, and AAVr3.45.Such AAV serotypes and derivatives / pseudotypes, as well as methods for producing them, have been previously described [e.g., Mol. Ther. 2012 Apr;20(4):699-708. doi: 10.1038 / mt.2011.287. Epub 2012 Jan 24. The AAV vector toolkit: poised at the clinical crossroads. Asokan Al, Schaffer DV, Samulski RJ, Duan et al, J. Virol., 75:7662-7671, 2001; Halbert et al, J. Virol., 74:1524-1532, 2000; Zolotukhin et al, Methods, 28:158-167, 2002; and Auricchio et al., Hum. Molec. Genet., 10:3075-3081, 2001; see, e.g., U.S. Patent Publication No. 2005 / 0100890(A1), International Publication No. WO 01 / 83692(A2), U.S. Patent Publication No. 2003 / 0103939(A1), and Miller (1996). Proc. Natl. Acad. Sci., 93: 11407-11413].
[0121] In some embodiments, rAAV particles are packaged using packaging and / or helper nucleic acids. Preferably, the AAV helper nucleic acids support efficient AAV vector production without producing detectable wild-type AAV particles (e.g., AAV particles containing functional rep and capsid protein genes). Helper nucleic acids, and methods for producing said nucleic acids, have been previously described and are commercially available [e.g., pDM, pDG, pDP1rs, pDP2rs, pDP3rs, pDP4rs, pDP5rs, pDP6rs, pDG(R484E / R585E), and pDP8.ape plasmids from PlasmidFactory, Bielefeld, Germany; Vector Biolabs, Philadelphia, PA; Cellbiolabs, San Diego, CA; Agilent Technologies, Santa Clara, CA; and other products and services available from Addgene, Cambridge, MA, pxx6, Grimm et al. (1998), Novel Tools for Production and Purification of Recombinant Adenoassociated Virus Vectors, Human Gene Therapy, Vol. 9, 2745-2760; Kern, A. et al. (2003), Identification of a Heparin-Binding Motif on Adeno-Associated Virus Type 2 Capsids, Journal of Virology, Vol. 77, 11072-11081., Grimm et al. (2003), Helper Virus Free, Optically Controllable, and Two-Plasmid-Based Production of Adeno-associated Virus Vectors of Serotypes 1 to 6, Molecular Therapy, Vol. 7, 839-850, Kronenberg et al.(2005), "A Conformational Change in the Adeno-Associated Virus Type 2 Capsid Leads to the Exposure of Hidden VP1 N Termini," Journal of Virology, Vol. 79, 5296-5303, and Moullier, P. and Snyder, RO (2008), "International efforts for recombinant adenoassociated viral vector reference standards," Molecular Therapy, Vol. 16, 1185-1188. In some embodiments, the packaging nucleic acid comprises an AAV rep nucleic acid sequence and an AAV cap nucleic acid sequence. In some embodiments, the AAV rep nucleic acid sequence and / or the AAV cap nucleic acid sequence are of the same AAV serotype as the AAV ITRs flanking the heterologous nucleic acid. In some embodiments, the AAV rep nucleic acid sequence and the AAV ITRs are of the same AAV serotype, but of a different serotype compared to the AAV cap sequence.
[0122] In some embodiments, components cultured within cells to package the rAAV genome into capsids can be provided to the cells in trans. In some embodiments, rAAV particles can be produced using a triple transfection method (described in detail in U.S. Pat. No. 6,001,650). In some embodiments, rAAV particles are produced by transfecting cells with an AAV vector (containing a heterologous nucleic acid flanked by ITR elements) to be packaged in the rAAV particle and at least one AAV helper or packaging nucleic acid. In some embodiments, two nucleic acids are used, including a helper nucleic acid and a packaging nucleic acid.
[0123] Alternatively, in some embodiments, any one or more of the required components (e.g., the heterologous nucleic acid flanked by AAV ITRs, the rep sequence, the cap sequence, and / or the helper nucleic acid) can be provided by a cell engineered to stably contain one or more of the required components (e.g., via genomic integration of the packaging nucleic acid and / or the helper nucleic acid). In some embodiments, the cell contains the required components under the control of either an inducible or constitutive promoter. Examples of suitable inducible and constitutive promoters are provided herein in the description of suitable regulatory sequences for use with heterologous nucleic acids. Methods used to construct engineered nucleic acids or rAAV particles thereof have also been previously described [see, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY]. Similarly, methods for producing rAAV virions are well known, and the selection of a suitable method is not a limitation of the present disclosure. See, e.g., K. Fisher et al., J. Virol., 70:520-532 (1993) and U.S. Patent No. 5,478,745].
[0124] Antibodies and antigen-binding fragments In some embodiments, the agent (e.g., a therapeutic agent and / or an anti-RAN protein agent) may be an anti-RAN protein antibody or antigen-binding fragment thereof. In some embodiments, the anti-RAN antibody may be a polyclonal antibody. In some embodiments, the anti-RAN antibody may be a monoclonal antibody. In some embodiments, the anti-RAN antigen-binding fragment may be derived from a polyclonal antibody. In some embodiments, it may be derived from a monoclonal antibody. In some embodiments, the anti-RAN protein antibody or antigen-binding fragment may bind to extracellular RAN protein, intracellular RAN protein, or both extracellular and intracellular RAN protein.
[0125] In some embodiments, the agent (e.g., a therapeutic agent) is an antibody or antigen-binding fragment capable of recognizing a gene product that interacts with a RAN protein. In some embodiments, the antibody or antigen-binding fragment is anti-eIF2. In some embodiments, the antibody or antigen-binding fragment is anti-eIF2A. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3a. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3b. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3c. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3d. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3e. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3f. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3g. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3h. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3i. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3j. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3k. In some embodiments, the antibody or antigen-binding fragment is anti-eIF3l. In some embodiments, the antibody or antigen-binding fragment is anti-eiF3m. In some embodiments, the antibody or antigen-binding fragment is anti-PKR. In some embodiments, the antibody or antigen-binding fragment is anti-p62. In some embodiments, the antibody or antigen-binding fragment is anti-LC3 I subunit. In some embodiments, the antibody or antigen-binding fragment is anti-LC3 II subunit. In some embodiments, the antibody or antigen-binding fragment is anti-TARBP2.In some embodiments, the antibody or antigen-binding fragment is anti-TLR3).
[0126] "Antibody" broadly refers to an immunoglobulin molecule or a functional mutant, variant, or derivative thereof. It is desirable that the functional mutant, variant, and derivative thereof, as well as antigen-binding fragments, retain the essential epitope-binding characteristics of an Ig molecule. Antibodies are capable of specific binding to a target through at least one antigen recognition site located within the variable region of the immunoglobulin molecule. Generally, an intact or full-length antibody comprises two heavy chains and two light chains. Each heavy chain contains a heavy chain variable region (VH) and first, second, and third constant regions (CH1, CH2, and CH3). Each light chain contains a light chain variable region (VL) and a constant region (CL). The VH and VL regions can be further subdivided into regions of hypervariability called complementarity-determining regions (CDRs), separated by more conserved regions called framework regions (FRs). The CDR components in the heavy chain are referred to as CDRH1, CDRH2, and CDRH3, and the CDR components in the light chain are referred to as CDRL1, CDRL2, and CDRL3.
[0127] CDRs typically refer to Kabat CDRs as described in Sequences of Proteins of Immunological Interest, US Department of Health and Human Services (1991), eds. Kabat et al. Another standard for characterizing antigen-binding sites is to refer to the hypervariable loops described by Chothia. See, e.g., Chothia, D. et al. (1992) J. Mol. Biol. 227:799-817 and Tomlinson et al. (1995) EMBO J. 14:4628-4638. Yet another standard is the AbM definition used by Oxford Molecular's AbM antibody modeling software. See, e.g., the entirety of Protein Sequence and Structure Analysis of Antibody Variable Domains. In: Antibody Engineering Lab Manual (Eds.: Duebel, S., and Kontermann, R., Springer-Verlag, Heidelberg). The embodiments described with respect to Kabat CDRs can alternatively be implemented using similarly described relationships with respect to Chothia hypervariable loops or AbM-defined loops, or any combination of these methods.
[0128] Each VH and VL is composed of three CDRs and four FRs, arranged from amino-terminus to carboxy-terminus in the following order: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. A full-length antibody can be of any class (or subclass thereof), such as IgD, IgE, IgG, IgA, or IgM; the antibody need not be of any particular class. Immunoglobulins can be assigned to different classes depending on the antibody amino acid sequence of the constant domain of their heavy chain. There are five major classes of immunoglobulins: IgA, IgD, IgE, IgG, and IgM, some of which can be further divided into subclasses (isotypes), e.g., IgG1, IgG2, IgG3, IgG4, IgA1, and IgA2. The heavy-chain constant domains corresponding to the different classes of immunoglobulins are called alpha, delta, epsilon, gamma, and mu, respectively. The subunit structures and three-dimensional configurations of different classes of immunoglobulins are well known.
[0129] The term "antigen-binding fragment" refers to any derivative of an antibody that is less than full-length and can specifically bind to a target. Preferably, the antigen-binding fragments provided herein retain the ability to specifically bind to a RAN protein. The antigen-binding fragment may include a heavy chain variable region (VH), a light chain variable region (VL), or both. Each of the VH and VL typically contains three complementarity-determining regions: CDR1, CDR2, and CDR3.
[0130] Examples of antigen-binding fragments include, but are not limited to, Fab, Fab', F(ab')2, scFv, Fv, dsFv, diabody, affibody, and Fd fragments. Antigen-binding fragments may be produced by any suitable means. For example, antigen-binding fragments may be enzymatically or chemically produced by fragmentation of an intact antibody, or recombinantly produced from a gene encoding a partial antibody sequence. Alternatively, antigen-binding fragments may be wholly or partially synthetically produced. Antigen-binding fragments may optionally be single-chain antibody fragments. Alternatively, fragments may contain multiple chains linked together, for example, by disulfide bonds. Antigen-binding fragments may also optionally be multimolecular complexes. Functional antigen-binding fragments typically contain at least about 50 amino acids, more typically at least about 200 amino acids.
[0131] Single-chain Fvs (scFvs) are recombinant antigen-binding fragments consisting of only a variable light chain (VL) and a variable heavy chain (VH) covalently linked to each other by a polypeptide linker. Either the VL or VH can be the NH2-terminal domain. The polypeptide linker can be of various lengths and compositions, as long as it bridges the two variable domains without significant steric interference. Typically, the linker consists primarily of a stretch of glycine and serine residues, interspersed with several glutamic acid or lysine residues for solubility. scFvs are encompassed by the term "antigen-binding fragment."
[0132] Diabodies are dimeric scFvs. Diabody components typically have shorter peptide linkers than most scFvs and prefer to assemble as dimers (see, for example, Holliger, P., et al. (1993) Proc. Natl. Acad. Sci. USA 90:6444-6448; Poljak, RJ, et al. (1994) Structure 2: 1121-1123). Diabodies are also encompassed by the term "antigen-binding fragment."
[0133] An Fv fragment is an antigen-binding fragment consisting of one VH and one VL domain held together by non-covalent interactions. The two domains of an Fv fragment, VL and VH, can be encoded by separate genes, but can be linked using recombinant techniques with a synthetic linker that allows them to be combined into a single protein chain (in which the VL and VH domains pair to form a monovalent molecule) (known as single-chain Fv (scFv); see, e.g., Bird et al. (1988) Science 242:423-426 and Huston et al. (1988) Proc. Natl. Acad. Sci. USA 85:5879-5883). Such single-chain antibodies are also intended to be encompassed by the term "antigen-binding fragment" of an antibody. The term dsFv is used herein to refer to an Fv with an engineered intermolecular disulfide bond to stabilize the VH-VL pair. dsFvs are also encompassed by the term "antigen-binding fragment."
[0134] An F(ab')2 fragment is an antigen-binding fragment essentially equivalent to that obtained from an immunoglobulin (typically an IgG) by digestion with the enzyme pepsin at pH 4.0 to 4.5. This fragment can be produced recombinantly. F(ab')2 is also encompassed by the term "antigen-binding fragment."
[0135] A Fab fragment is an antigen-binding fragment essentially equivalent to that obtained by reduction of one or more disulfide bridges connecting the two heavy chain pieces in a F(ab')2 fragment. Fab' fragments can be recombinantly produced. Fab' is also encompassed by the term "antigen-binding fragment."
[0136] A Fab fragment is an antigen-binding fragment essentially equivalent to that obtained by digestion of immunoglobulins (typically IgG) with the enzyme papain. Fab fragments can be produced recombinantly. The heavy chain segment of a Fab fragment is the Fd fragment. Fab fragments are also encompassed by the term "antigen-binding fragment."
[0137] Affibodies are small proteins containing a three-helix bundle that function as antigen-binding molecules (e.g., antibody mimetics). Generally, affibodies are approximately 58 amino acids long and have a molar mass of approximately 6 kDa. Affibody molecules with unique binding properties are obtained by randomizing 13 amino acids located within the two alpha helices involved in the binding activity of the parent protein domain. Specific affibody molecules that bind to a desired target protein can be isolated from a pool (library) containing hundreds of millions of different variants using methods such as phage display. Affibodies are also encompassed by the term "antigen-binding fragment."
[0138] The term "human antibody" refers to an antibody obtained from a human subject, e.g., having variable and constant regions substantially corresponding to or derived from an antibody encoded by human germline immunoglobulin sequences or variants thereof. A human antibody may contain one or more amino acid residues not encoded by human germline immunoglobulin sequences (e.g., mutations introduced by random or site-specific mutagenesis in vitro or by somatic mutation in vivo). Such mutations may occur in one or more CDRs, particularly CDR3, or in one or more framework regions. In some embodiments, a human antibody may have at least one, two, three, four, five, or more positions replaced with an amino acid residue not encoded by human germline immunoglobulin sequences. However, the term "human antibody," as used herein, is not intended to include antibodies in which CDR sequences derived from the germline of another mammalian species, such as a mouse, have been grafted onto human framework sequences.
[0139] As used herein, the term "recombinant human antibody" refers to any human antibody that is prepared, expressed, created, or isolated by recombinant means, e.g., an antibody expressed using a recombinant expression vector transfected into a host cell, an antibody isolated from a recombinant combinatorial human antibody library [Hoogenboom HR, (1997) TIB Tech. 15:62-70, Azzazy H., and Highsmith WE, (2002) Clin. Biochem. 35:425-445, Gavilondo JV, and Larrick JW (2002) BioTechniques 29: 128-145, Hoogenboom H., and Chames P. (2000) Immunology Today 21:371-378], or an antibody isolated from an animal (e.g., a mouse) that is transgenic for human immunoglobulin genes [e.g., Taylor, LD, et al. (1992) Nucl. Acids Res. 20:6287-6295; Kellermann SA., and Green LL (2002) Current Opinion in Biotechnology 13:593-597; Little M. et al (2000) Immunology Today 21:364-370], or other means involving splicing human immunoglobulin gene sequences into other DNA sequences. Such recombinant human antibodies have variable and constant regions as defined above. However, in certain embodiments, such recombinant human antibodies may be subjected to in vitro mutagenesis (or, when animals transgenic for human Ig sequences are used, in vivo somatic mutagenesis), such that the amino acid sequences of the VH and VL regions of the recombinant antibodies, while derived from and related to human germline VH and VL sequences, may not naturally occur in the in vivo human antibody germline repertoire.
[0140] In some embodiments, the anti-RAN protein antibody or antigen-binding fragment is selected from the group consisting of anti-poly-serine, anti-poly(GR), anti-poly(PR), anti-poly(CP), anti-poly(GP), anti-poly(G), anti-poly(A), anti-poly(GA), anti-poly(GD), anti-poly(GE), anti-poly(GQ), anti-poly(GT), anti-poly(L), anti-poly(LP), anti-poly(LPAC) (SEQ ID NO: 31), anti-poly(LS), anti-poly(P), anti-poly(PA), anti-poly(QAGR) (SEQ ID NO: 35), anti-poly(RE), anti-poly(SP), anti-poly(VP), anti-poly(FP), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 36), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 37), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 38), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 39), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 40), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 41), anti-poly(L), anti-poly(LPAC) (SEQ ID NO: 42), anti-poly(LPAC) (SEQ ID NO: 43), anti-poly(LPAC) (SEQ ID NO: 44), anti-poly(LPAC) (SEQ ID NO: 45), anti-poly(LPAC) (SEQ ID NO: 46), anti-poly(LPAC) (SEQ ID NO: 47), anti-poly(LPAC) (SEQ ID NO: 48), anti-poly(LPAC) (SEQ ID NO: 49), anti-poly(LPAC) (SEQ ID NO: 50), anti-poly(LPAC) (SEQ ID NO: 51), anti-poly(LPAC) (SEQ ID anti-poly(GK), anti-poly(FTPLSLPV) (SEQ ID NO: 36), anti-poly(LLPSPSRC) (SEQ ID NO: 37), anti-poly(YSPLPPGV) (SEQ ID NO: 38), anti-poly(HREGEGSK) (SEQ ID NO: 39), anti-poly(TGRERGVN) (SEQ ID NO: 40), anti-poly(PGGRGE) (SEQ ID NO: 41), anti-poly(GRQRGVNT) (SEQ ID NO: 42), and / or anti-poly(GSKHREAE) (SEQ ID NO: 43) antibodies [also referred to as α-poly(Ser), α-poly(PR), α-poly(GR), etc.], or antigen-binding fragments thereof.
[0141] In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GA). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(Ser). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(PR). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GR). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(Leu). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(Ala). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(LPAC) (SEQ ID NO: 31). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(QAGR) (SEQ ID NO: 35). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(CP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(G). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GD). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GE). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GQ). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GT). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(LP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(LS). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(P). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(PA). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(RE).In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(SP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(VP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(FP). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GK). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(FTPLSLPV) (SEQ ID NO: 36). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(LLPSPSRC) (SEQ ID NO: 37). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(YSPLPPGV) (SEQ ID NO: 38). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(HREGEGSK) (SEQ ID NO: 39). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(TGRERGVN) (SEQ ID NO: 40). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(PGGRGE) (SEQ ID NO: 41). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GRQRGVNT) (SEQ ID NO: 42). In some embodiments, the antibody or antigen-binding fragment specifically binds to poly(GSKHREAE) (SEQ ID NO: 43).
[0142] In some embodiments, the anti-RAN antibody or antigen-binding fragment is directed against a portion of a truncated RAN protein that does not contain poly-amino acid repeats, e.g., the C-terminus of a truncated RAN protein [e.g., poly(CP), poly(GP), poly(G), poly(A), poly(GA), poly(GD), poly(GE), poly(GQ), poly(GR), poly(GT), poly(L), poly(LP), poly(LPAC) (SEQ ID NO: 31), poly(LS), poly(P), poly(PA), poly(PR), poly(QAGR) (SEQ ID NO: 35), poly(RE)] , poly(Ser), poly(SP), poly(VP), poly(FP), poly(GK), poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41) protein, poly(GRQRGVNT) (SEQ ID NO: 42), and / or poly(GSKHREAE) (SEQ ID NO: 43). Examples of anti-RAN antibodies targeting the C-terminus of the RAN protein are disclosed, for example, in U.S. Publication No. 2013 / 0115603, the entire contents of which are incorporated herein by reference.
[0143] In some embodiments, a set (or combination) of anti-RAN antibodies or antigen-binding fragments [e.g., poly(CP), poly(GP), poly(G), poly(A), poly(GA), poly(GD), poly(GE), poly(GQ), poly(GR), poly(GT), poly(L), poly(LP), poly(LPAC) (SEQ ID NO: 31), poly(LS), poly(P), poly(PA), poly(PR), poly(QAGR) (SEQ ID NO: 32), poly(RE), poly(Ser), poly(SP), poly(VP), poly(FP), poly(GK), poly(F)], poly(Ser ... A combination of two or more anti-RAN antibodies selected from poly(TPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41), poly(GRQRGVNT) (SEQ ID NO: 42), and poly(GSKHREAE) (SEQ ID NO: 43) is administered to a subject for the purpose of treating a disease associated with a RAN protein (e.g., a truncated RAN protein).
[0144] It should be understood that in some embodiments, the present disclosure contemplates variants (e.g., homologs) of the amino acid and nucleic acid sequences of the heavy and light chain variable regions of an antibody. "Homology" refers to the percent identity between two polynucleotides or two polypeptide moieties. The term "substantial homology," when referring to a nucleic acid or fragment thereof, indicates that when optimally aligned with another nucleic acid (or its complementary strand), with appropriate nucleotide insertions or deletions, the aligned sequences share about 90-100% nucleotide sequence identity. For example, in some embodiments, nucleic acid sequences that share substantial homology have at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity. When referring to a polypeptide or fragment thereof, the term "substantial homology" indicates that when optimally aligned with another polypeptide, with appropriate gaps, insertions, or deletions, the aligned sequences share about 90-100% nucleotide sequence identity. The term "highly conserved" refers to at least 80% identity, preferably at least 90% identity, and more preferably greater than 97% identity. For example, in some embodiments, highly conserved proteins share at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity. In some cases, highly conserved can refer to 100% identity. Identity is readily determined by one of skill in the art, for example, by use of algorithms and computer programs known to those of skill in the art.
[0145] In some embodiments, the RAN antibodies of the present disclosure bind to the truncated RAN protein with high affinity, e.g., 10 -7 Under M, 10 -8 Under M, 10 -9 Under M, 10 -10 Under M, 10 -11The antibody or antigen-binding fragment may bind to a truncated RAN protein with an affinity of less than 5 pM or a lower Kd. For example, the anti-RAN antibody or antigen-binding fragment may bind to a truncated RAN protein with an affinity of between 5 pM and 500 nM, e.g., between 50 pM and 100 nM, e.g., between 500 pM and 50 nM. The present disclosure also includes antibodies or antigen-binding fragments that compete for binding to a truncated RAN protein with any of the antibodies described herein and have an affinity of 50 nM or less (e.g., 20 nM or less, 10 nM or less, 500 pM or less, 50 pM or less, or 5 pM or less). The affinity and binding rate of an anti-RAN protein antibody can be tested using any method known in the art, including, but not limited to, biosensor technology (e.g., OCTET or BIACORE).
[0146] In some embodiments, an anti-RAN antibody of the present disclosure may comprise one or more of the VH, VL, and CDR amino acid sequences shown in the table below: In some embodiments, an anti-RAN protein antibody may be produced using one or more of the nucleic acids shown in the table below.
[0147] [Table 1]
[0148] [Table 2-1] [Table 2-2]
[0149] [Table 3-1] [Table 3-2]
[0150] [Table 4-1] [Table 4-2]
[0151] In some embodiments, antibody clone 27B11.A7 binds to polyGA. In some embodiments, clone 27B11.A7 is an IgG1 antibody. In some embodiments, antibody clone 23H2.D1.B5 binds to polyGA. In some embodiments, antibody clone 23H2.D1.B5 is an IgG3 antibody. In some embodiments, antibody clone 16A3.C8 binds to poly(Ser). In some embodiments, antibody clone 16A3.C8 is an IgG1 antibody. In some embodiments, antibody clone HL2362-2G4 binds to polyPR. In some embodiments, antibody clone HL2362-2G4 is an IgG2A kappa antibody.
[0152] Production of anti-RAN antibodies Typically, polyclonal antibodies are produced by inoculation into a suitable mammal, such as a mouse, rabbit, or goat. By injecting an antigen into the mammal, B lymphocytes are induced to produce IgG immunoglobulins specific to the antigen. The polyclonal IgG is purified from the serum of the mammal. Monoclonal antibodies are generally produced from a single cell line (e.g., a hybridoma cell line). In some embodiments, anti-RAN antibodies are purified (e.g., isolated from serum).
[0153] Exemplary anti-RAN antibodies disclosed herein were produced using the antigens shown in Table 13. In some embodiments, the antigen is selected from the group consisting of poly(proline-arginine) [poly(PR)]; poly(glycine-arginine) [poly(GR)]; poly(serine) [poly(Ser)]; poly(cysteine-proline) [poly(CP)]; poly(glycine-proline) [poly(GP)]; poly(glycine) [poly(G)]; poly(Ala) [polyAla]; poly(glycine-alanine) [poly(GA)]; poly(glycine-asparagine) [poly(Ala)]; poly(glycine-glutamic acid) [poly(GE)]; poly(glycine-glutamine) [poly(GQ)]; poly(glycine-threonine) [poly(GT)]; poly(leucine) [poly(Leu)]; poly(leucine-proline) [poly(LP)]; poly(leucine-proline-alanine-cysteine) [poly(LPAC)] (SEQ ID NO: 31); poly(leucine-serine) [poly(LS)]; poly(proline) ) [poly(P)]; poly(proline-alanine) [poly(PA)]; poly(glutamine-alanine-glycine-arginine) [poly(QAGR)] (SEQ ID NO: 35); poly(arginine-glutamic acid) [poly(RE)]; poly(serine-proline) [poly(SP)], poly(valine-proline) [poly(VP)], poly(phenylalanine-proline) [poly(FP)], poly(glycine-lysine) [poly(GK)], poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41), poly(GRQRGVNT) (SEQ ID NO: 42), and poly(GSKHREAE) (SEQ ID NO: 43).
[0154] [Table 5]
[0155] Numerous methods can be used to obtain anti-RAN antibodies. For example, antibodies can be produced using recombinant DNA methods. Monoclonal antibodies can also be produced by generating hybridomas according to known methods [see, e.g., Kohler and Milstein (1975) Nature, 256:495-499]. The hybridomas thus formed are then screened using standard methods, such as enzyme-linked immunosorbent assays (ELISAs; e.g., RCA-based ELISAs or rtPCR-based ELISAs) and surface plasmon resonance (e.g., OCTET or BIACORE) analysis, to identify one or more hybridomas that produce antibodies that specifically bind to the antigen. Any form of the antigen (e.g., truncated RAN protein), such as recombinant antigens, naturally occurring forms, or any mutant or fragment thereof, can be used as an immunogen. One exemplary method for generating antibodies involves screening a protein expression library, such as a phage display library or ribosome display library, that expresses antibodies or fragments thereof (e.g., scFvs). Phage display is described, for example, in U.S. Pat. No. 5,223,409 to Ladner et al., Smith (1985) Science 228:1315-1317, Clackson et al. (1991) Nature, 352:624-628, Marks et al. (1991) J. Mol. Biol., 222:581-597, WO92 / 18619, WO91 / 17271, WO92 / 20791, WO92 / 15679, WO93 / 01288, WO92 / 01047, WO92 / 09690, and WO90 / 02809.
[0156] In another embodiment, monoclonal antibodies are obtained from non-human animals and then modified, e.g., chimerized, using recombinant DNA techniques known in the art. Various techniques for producing chimeric antibodies have been described. See, for example, Morrison et al., Proc. Natl. Acad. Sci. USA 81:6851, 1985; Takeda et al., Nature 314:452, 1985; U.S. Patent No. 4,816,567 to Cabilly et al.; U.S. Patent No. 4,816,397 to Boss et al.; European Patent Publication Nos. EP 171496 and 0173494 to Tanaguchi et al.; and British Patent No. GB 2177096B.
[0157] Antibodies can be humanized by methods known in the art. For example, monoclonal antibodies with desired binding specificity can be commercially humanized (Scotgene, Scotland and Oxford Molecular, Palo Alto, Calif.). Fully humanized antibodies, such as those expressed in transgenic animals, are within the scope of the present invention (see, for example, Green et al. (1994) Nature Genetics 7, 13, and U.S. Patent Nos. 5,545,806 and 5,569,825). For further antibody production techniques, see Antibodies: A Laboratory Manual, Second Edition. Edited by Edward A. Greenfield, Dana-Farber Cancer Institute, (C)2014. The present disclosure is not necessarily limited to a particular source, production method, or other particular characteristics of the antibody.
[0158] Some aspects of the present disclosure relate to isolated cells (e.g., host cells) transformed with a polynucleotide or vector. The host cell can be a prokaryotic or eukaryotic cell. The polynucleotide or vector present in the host cell may be integrated into the genome of the host cell or maintained extrachromosomally. The host cell can be any prokaryotic or eukaryotic cell, such as a bacterial, insect, fungal, plant, animal, or human cell. In some embodiments, the fungal cell is, for example, of the genus Saccharomyces, particularly the species S. cerevisiae. The term "prokaryotic" includes all bacteria that can be transformed or transfected with DNA or RNA molecules for expression of antibodies or corresponding immunoglobulin chains. Prokaryotic hosts can include Gram-negative and Gram-positive bacteria, such as E. coli, Salmonella typhimurium, Serratia marcescens, and Bacillus subtilis. The term "eukaryotic" includes yeast, higher plants, insects, and vertebrate cells, e.g., mammalian cells such as NSO and CHO cells. Depending on the host used in a recombinant production procedure, antibodies or immunoglobulin chains encoded by the polynucleotides can be glycosylated or non-glycosylated. Antibodies or corresponding immunoglobulin chains can include an initial methionine amino acid residue.
[0159] In some embodiments, once the vector has been incorporated into a suitable host, the host may be maintained under conditions suitable for high-level expression of the nucleotide sequence, followed by collection and purification of immunoglobulin light chains, heavy chains, light / heavy chain dimers or intact antibodies, antigen-binding fragments, or other immunoglobulin forms, as desired. See Beychok, Cells of Immunoglobulin Synthesis, Academic Press, NY, (1979). Thus, the polynucleotide or vector is introduced into cells, thereby producing antibodies or antigen-binding fragments. Furthermore, transgenic animals, preferably mammals, containing the aforementioned host cells can be used for large-scale production of antibodies or antibody fragments.
[0160] Transformed host cells can be grown and cultured in fermentors according to techniques known in the art to achieve optimal cell growth. Once expressed, whole antibodies, dimers thereof, individual light and heavy chains, other immunoglobulin forms, or antigen-binding fragments can be purified according to standard procedures in the art, including ammonium sulfate precipitation, affinity columns, column chromatography, gel electrophoresis, and the like. See Scopes, "Protein Purification," Springer Verlag, NY (1982). The antibody or antigen-binding fragment can then be isolated from the growth medium, cell lysates, or cell membrane fractions. Isolation and purification of antibodies or antigen-binding fragments expressed, for example, in microorganisms, can be by any conventional means, such as preparative chromatographic separation and immunological separation, for example, those involving the use of monoclonal or polyclonal antibodies against antibody constant regions, etc.
[0161] Aspects of the present disclosure relate to hybridomas, which are an indefinitely prolonged source of monoclonal antibodies. As used herein, "hybridoma cells" refers to immortalized cells derived from the fusion of a B lymphoblast with a myeloma fusion partner. To prepare monoclonal antibody-producing cells (e.g., hybridoma cells), individual animals (e.g., mice) with confirmed antibody titers are selected, and their spleens or lymph nodes are harvested two to five days after the final immunization. The antibody-producing cells contained therein are fused with myeloma cells to prepare hybridomas producing the desired monoclonal antibodies. Antibody titers in antisera can be measured, for example, by reacting the antisera with a labeled protein, as described below, and then measuring the activity of the labeling agent bound to the antibody. Cell fusion can be performed according to known methods, such as those described by Kochler and Milstein [Nature 256:495 (1975)]. As a fusion promoter, for example, polyethylene glycol (PEG) or Sendai virus (HVJ) is used.
[0162] Examples of myeloma cells include NS-1, P3U1, SP2 / 0, and AP-1. The ratio of the number of antibody-producing cells (spleen cells) to the number of myeloma cells used is preferably about 1:1 to about 20:1. PEG (preferably PEG1000 to PEG6000) is preferably added at a concentration of about 10% to about 80%. Cell fusion can be efficiently carried out by incubating a mixture of both cells at about 20°C to about 40°C, preferably about 30°C to about 37°C, for about 1 to 10 minutes.
[0163] Various methods can be used to screen hybridomas that produce antibodies (e.g., antibodies against the tumor antigens or autoantibodies of the present invention). For example, the supernatant of a hybridoma is added to a solid phase (e.g., a microplate) to adsorb the antibody directly or with a carrier, and then a radioactively or enzyme-labeled anti-immunoglobulin antibody (when mouse cells are used in cell fusion, an anti-mouse immunoglobulin antibody is used) or protein A is added to detect monoclonal antibodies against the protein bound to the solid phase. Alternatively, the supernatant of a hybridoma is added to a solid phase to adsorb the anti-immunoglobulin antibody or protein A, and then a radioactively or enzyme-labeled protein is added to detect monoclonal antibodies against the protein bound to the solid phase.
[0164] Monoclonal antibody selection can be performed by any known method or its modifications. Typically, animal cell culture media supplemented with HAT (hypoxanthine, aminopterin, thymidine) are used. Any selection and growth medium can be used as long as hybridomas can grow. For example, RPMI 1640 medium containing 1% to 20%, preferably 10% to 20%, fetal bovine serum (GIT medium) containing 1% to 10% fetal bovine serum, or serum-free medium for culturing hybridomas (SFM-101, Nissui Seiyaku) can be used. Culture is typically performed at 20°C to 40°C, preferably 37°C, for approximately 5 days to 3 weeks, preferably 1 to 2 weeks, under approximately 5% CO2 gas. The antibody titer of the supernatant of the hybridoma culture can be measured in the same manner as described above for the antibody titer of the antiprotein in the antiserum.
[0165] Instead of obtaining immunoglobulins directly from hybridoma cultures, immortalized hybridoma cells may be used as a source of rearranged heavy and light chain loci for subsequent expression and / or genetic manipulation. Rearranged antibody genes can be reverse transcribed from the appropriate mRNA to produce cDNA. If desired, the heavy chain constant region can be replaced with one of a different isotype or removed entirely. Variable regions may be linked to encode a single-chain Fv region. Multiple Fv regions may be linked to confer binding capabilities to more than one target, or chimeric heavy and light chain combinations may be used. Any suitable method may be used to clone antibody variable regions and generate recombinant antibodies.
[0166] In some embodiments, appropriate nucleic acids encoding the heavy and / or light chain variable regions are obtained and inserted into expression vectors that can be transfected into standard recombinant host cells. A variety of such host cells can be used. In some embodiments, mammalian host cells can be advantageous for efficient processing and production. Exemplary mammalian cell lines useful for this purpose include CHO cells, 293 cells, or NSO cells. Antibodies or antigen-binding fragments can be produced by culturing the modified recombinant host under culture conditions appropriate for host cell growth and expression of the coding sequences. Antibodies or antigen-binding fragments can be recovered by isolating them from the culture. Expression systems can be designed to include a signal peptide so that the resulting antibody is secreted into the culture medium. However, intracellular production is also possible.
[0167] The present disclosure also includes polynucleotides encoding at least the variable regions of the immunoglobulin chains of the antibodies described herein. In some embodiments, the variable regions encoded by the polynucleotides comprise at least one complementarity-determining region (CDR) of the VH and / or VL variable regions of the antibodies produced by any one of the above-mentioned hybridomas.
[0168] The polynucleotide encoding the antibody or antigen-binding fragment can be, for example, DNA, cDNA, RNA, or synthetically produced DNA or RNA, or a recombinantly produced chimeric nucleic acid molecule comprising any of these polynucleotides, alone or in combination. In some embodiments, the polynucleotide is part of a vector. Such vectors can contain additional genes, such as marker genes, that allow for the selection of the vector in a suitable host cell and under suitable conditions.
[0169] In some embodiments, the polynucleotide is operably linked to an expression control sequence that allows expression in prokaryotic or eukaryotic cells. Expression of the polynucleotide involves transcribing the polynucleotide into translatable mRNA. Regulatory elements ensuring expression in eukaryotic cells, preferably mammalian cells, are well known to those skilled in the art. These may include regulatory sequences that facilitate transcription initiation and, optionally, a poly-A signal that facilitates transcription termination and transcript stabilization. Additional regulatory elements may include transcriptional and translational enhancers and / or naturally associated or heterologous promoter regions. Possible regulatory elements that allow expression in prokaryotic host cells include, for example, the PL, Lac, Trp, or Tac promoters in E. coli. Examples of regulatory elements that allow expression in eukaryotic host cells are the AOX1 or GAL1 promoter in yeast, or the CMV promoter, SV40 promoter, RSV promoter (Rous sarcoma virus), CMV enhancer, SV40 enhancer, or globin intron in mammalian and other animal cells.
[0170] In addition to elements responsible for transcription initiation, such regulatory elements may include transcription termination signals, such as an SV40 poly-A site or a tk poly-A site, downstream of the polynucleotide. Furthermore, depending on the expression system used, leader sequences capable of directing the polypeptide into a cellular compartment or secreting it into the medium may be added to the coding sequence of the polynucleotide, and are well known in the art. The leader sequence is assembled with translation initiation and termination sequences at the appropriate stage, and preferably, the leader sequence is capable of directing the secretion of the translated protein or a portion thereof, for example, into the extracellular medium. Where appropriate, heterologous polynucleotide sequences may be used that encode fusion proteins containing specific C- or N-terminal peptides that confer desired characteristics, such as stabilization or simplified purification of the expressed recombinant product.
[0171] In some embodiments, the polynucleotides encoding at least the variable domains of the light and / or heavy chains may encode the variable domains of both immunoglobulin chains or only one of them. Similarly, the polynucleotides may be under the control of the same promoter or may be separately controlled for expression. Furthermore, some aspects relate to vectors, particularly plasmids, cosmids, viruses, and bacteriophages, conventionally used in genetic engineering, that contain polynucleotides encoding the variable domains of an immunoglobulin chain of an antibody or antigen-binding fragment, optionally in combination with polynucleotides encoding the variable domains of other immunoglobulin chains of the antibody.
[0172] In some embodiments, expression control sequences are provided as eukaryotic promoter systems in vectors capable of transforming or transfecting eukaryotic host cells, although control sequences for prokaryotic hosts may also be used. Expression vectors derived from viruses such as retroviruses, vaccinia virus, adeno-associated virus, herpesvirus, or bovine papillomavirus may be used to deliver polynucleotides or vectors to target cell populations (e.g., to engineer cells to express antibodies or antigen-binding fragments). A variety of suitable methods can be used to construct recombinant viral vectors. In some embodiments, polynucleotides and vectors can be reconstituted into liposomes for delivery to target cells. Vectors containing polynucleotides (e.g., heavy and / or light chain variable domains of immunoglobulin chain coding sequences and expression control sequences) can be transferred into host cells by suitable methods, which vary depending on the type of cellular host.
[0173] qualification Some aspects of the present disclosure relate to antibody-drug conjugates that target one or more disrupted RAN proteins. As used herein, "antibody-drug conjugate" refers to a molecule comprising an antibody or its antigen-binding fragment linked to a targeted molecule (e.g., a bioactive molecule such as a therapeutic molecule, and / or a detectable label). Thus, in some embodiments, the antibodies or antigen-binding fragments of the present disclosure can be modified with a detectable label, including but not limited to, an enzyme, a prosthetic group, a fluorescent material, a luminescent material, a bioluminescent material, a radioactive material, a positron-emitting metal, a non-radioactive paramagnetic metal ion, and an affinity label, for detection and isolation of one or more disrupted RAN proteins. Detectable substances can be coupled or conjugated to the polypeptides of the present disclosure directly or indirectly through an intermediary (e.g., a linker known in the art) using techniques known in the art. Non-limiting examples of suitable enzymes include horseradish peroxidase, alkaline phosphatase, β-galactosidase, glucose oxidase, or acetylcholinesterase; non-limiting examples of suitable prosthetic group complexes include streptavidin / biotin and avidin / biotin; non-limiting examples of suitable fluorescent materials include biotin, umbelliferone, fluorescein, fluorescein isothiocyanate, rhodamine, dichlorotriazinylamine fluorescein, dansyl chloride, or phycoerythrin; an example of a luminescent material includes luminol; non-limiting examples of bioluminescent materials include luciferase, luciferin, and aequorin; and examples of suitable radioactive materials include radioactive metal ions, e.g., alpha emitters, or e.g., iodine ( 131 I, 125 I, 123 I, 121 I), carbon ( 14 C), sulfur ( 35 S), tritium ( 3 H), indium ( 115 mIn, 113 mIn, 112 In, 111In), and technetium ( 99 Tc, 99 mTc), thallium ( 201 Ti), Gallium ( 68 Ga, 67 Ga), palladium ( 103 Pd), molybdenum ( 99 Mo), xenon ( 133 Xe), fluorine ( 18 F), 153 Sm, Lu, 159 Gd, 149 Pm, 140 La, 175 Yb, 166 Ho, 90 Y, 47 Sc, 86 R, 188 Re, 142 Pr, 105 Rh, 97 Ru, 68 Ge, 57 Co, 65 Zn, 85 Sr, 32 P, 153 Gd, 169 Yb, 51 Cr, 54 Mn, 75 Se, and tin ( 113 Sn, 117 Detectable substances can be coupled or conjugated directly to the anti-RAN antibodies or antigen-binding fragments of the present disclosure using techniques known in the art, or indirectly through an intermediary (e.g., a linker known in the art). Anti-RAN antibodies conjugated to detectable substances can be used for the diagnostic assays described herein.
[0174] In some embodiments, the antibody or antigen-binding fragment of the present disclosure may be modified with a therapeutic moiety (e.g., a therapeutic agent). In some embodiments, the antibody is coupled to the targeted agent via a linker. As used herein, the term "linker" refers to a sequence, such as a molecule or amino acid sequence, that attaches one molecule or sequence to another, like a bridge. "Linked," "conjugated," or "coupled" means attached or attached by a covalent bond, a non-covalent bond, or other bond such as van der Waals forces. The antibody described by the present disclosure can be linked to the targeted agent (e.g., a therapeutic moiety or a detectable moiety) directly, for example, as a fusion protein with a protein or peptide detectable moiety (with or without an appropriate linking sequence, e.g., a flexible linker sequence), or via a chemical coupling moiety. Several such coupling moieties are known in the art, such as peptide linkers or chemical linkers described in International Patent Application Publication No. WO 2009 / 036092. In some embodiments, the linker is a flexible amino acid sequence. Examples of flexible amino acid sequences include glycine- and serine-rich linkers containing a stretch of two or more glycine residues. In some embodiments, the linker is a photolinker. Examples of photolinkers include ketyl-reactive benzophenone (BP), anthraquinone (AQ), nitrene-reactive nitrophenyl azide (NPA), and carbene-reactive phenyl-(trifluoromethyl)diazirine (PTD).
[0175] composition In some embodiments, the composition (e.g., pharmaceutical composition) comprises one or more agents (e.g., therapeutic agents and / or anti-RAN proteins) described herein. In some embodiments, the composition (e.g., pharmaceutical composition) comprises a small molecule described herein. In some embodiments, the composition (e.g., pharmaceutical composition) comprises a nucleic acid described herein. In some embodiments, the composition (e.g., pharmaceutical composition) comprises an inhibitory nucleic acid, genetic variant, gRNA, and / or RNA-guided nuclease described herein. In some embodiments, the composition (e.g., pharmaceutical composition) comprises an antibody (e.g., an anti-RAN protein antibody) or antigen-binding fragment.
[0176] In some embodiments, the composition comprises an agent (e.g., a therapeutic agent) described herein and a pharmaceutically acceptable carrier. In some embodiments, the composition comprises an anti-RAN antibody and a pharmaceutically acceptable carrier. A composition is said to be a "pharmaceutically acceptable carrier" if its administration can be tolerated by a recipient patient. Sterile phosphate-buffered saline is one example of a pharmaceutically acceptable carrier. Other suitable carriers are well known in the art. See, for example, REMINGTON'S PHARMACEUTICAL SCIENCES, 18th Ed. (1990). As used herein, the term "pharmaceutically acceptable carrier" is intended to include any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like, that are compatible with pharmaceutical administration. The use of such media and agents for pharmaceutically active substances is well known in the art. Except insofar as a conventional media or agent is incompatible with the active compound, its use in the composition is contemplated. Supplementary active compounds can also be incorporated into the compositions. Pharmaceutical compositions can be prepared as described below. The active ingredient may be mixed or blended with any conventional pharmaceutically acceptable carrier or excipient. The composition may be sterile.
[0177] As will be understood by those skilled in the art, any conventionally used administration form, vehicle, or carrier that is inert to the active substance can be used to prepare and administer the pharmaceutical compositions of the present disclosure. Examples of such methods, vehicles, and carriers are described, for example, in Remington's Pharmaceutical Sciences, 4th ed. (1970), the disclosure of which is incorporated herein by reference. A person skilled in the art, armed with the principles of the present disclosure, will have no difficulty in determining suitable and appropriate vehicles, excipients, and carriers, or in combining active ingredients therewith to form the pharmaceutical compositions of the present disclosure.
[0178] Typically, a composition (e.g., a pharmaceutical composition) is formulated to deliver an effective amount of an agent (e.g., an anti-RAN antibody). Generally, an "effective amount" of an active agent refers to an amount sufficient to induce a desired biological response (e.g., ameliorate one or more symptoms of AD or ALS). The effective amount of an agent may vary depending on factors such as the desired biological endpoint, the pharmacokinetics of the compound, the disease being treated (e.g., AD or ALS, repeat expansion disease, etc.), the mode of administration, and the patient. In certain embodiments, an effective amount is an amount effective to reduce the level of RAN protein (e.g., a truncated RAN protein) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98% (e.g., the level of a truncated RAN protein relative to the level of a truncated RAN protein in a cell or subject not administered the therapeutic agent). In certain embodiments, an effective amount is an amount effective to reduce translation of RAN protein (e.g., a disrupted RAN protein) (e.g., the level of the disrupted RAN protein relative to the level of RAN protein in a cell or subject not administered the therapeutic agent) by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 98%.
[0179] In some embodiments, an effective amount (also referred to as a therapeutically effective amount) of an agent (e.g., a therapeutic agent such as an anti-RAN antibody) is an amount sufficient to ameliorate at least one adverse effect associated with a disease associated with a RAN protein (e.g., a disrupted RAN protein), such as, for example, memory loss, cognitive impairment, loss of coordination, speech disorders, etc. In some embodiments, the neurological disease associated with a RAN protein (e.g., a disrupted RAN protein) is amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 4 (SCA5), spinocerebellar ataxia type 5 (SCA6), spinocerebellar ataxia type 6 (SCA7), spinocerebellar ataxia type 7 (SCA8), spinocerebellar ataxia type 8 (SCA9), spinocerebellar ataxia type 9 (SCA10), spinocerebellar ataxia type 10 (SCA11), spinocerebellar ataxia type 11 (SCA12), spinocerebellar ataxia type 12 (SCA13), spinocerebellar ataxia type 13 (SCA14), spinocerebellar ataxia type 14 (SCA15), spinocerebellar ataxia type 15 (SCA16), spinocerebellar ataxia type 16 (SCA17), spinocerebellar ataxia type 17 The neurological disease associated with disrupted RAN protein is selected from the group consisting of spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), and Fuchs' corneal dystrophy. In some embodiments, the neurological disease associated with disrupted RAN protein is ALS or AD. The therapeutically effective amount contained in the pharmaceutical composition varies from case to case depending on several factors, such as the type, size, and condition of the patient being treated, the intended administration method, and the patient's adaptability to the intended dosage form. Generally, each dosage form contains an amount of active agent to provide from about 0.1 to about 250 mg / kg, preferably from about 0.1 to about 100 mg / kg. Those skilled in the art will be able to empirically determine the appropriate therapeutically effective amount.
[0180] By selecting from various active compounds and considering factors such as efficacy, relative bioavailability, patient weight, severity of adverse side effects, and the selected administration format, in conjunction with the teachings provided herein, an effective preventive or therapeutic treatment regimen can be designed that is sufficiently effective to treat a particular subject without causing substantial toxicity.The effective amount for a particular application may vary depending on factors such as the disease or condition being treated, the specific therapeutic agent being administered, the size of the subject, or the severity of the disease or condition.Those skilled in the art can empirically determine the effective amount of a particular nucleic acid and / or other therapeutic agent without undue experimentation.
[0181] The pharmaceutical compositions may comprise suitable solid- or gel-phase carriers or excipients. Examples of such carriers or excipients include, but are not limited to, calcium carbonate, calcium phosphate, various sugars, starches, cellulose derivatives, gelatin, and polymers such as polyethylene glycols.
[0182] Suitable liquid or solid pharmaceutical preparation forms include, for example, aqueous or saline solutions for inhalation, microencapsulated, encochleated, coated on fine gold particles, contained in liposomes, atomized, aerosolized, pelleted for implantation into the skin, or dried on a sharp object for penetration into the skin through a scratch. Pharmaceutical compositions also include granules, powders, tablets, coated tablets, (micro)capsules, suppositories, syrups, emulsions, suspensions, creams, drops, or preparations for delayed release of active compounds, in which excipients and additives, and / or auxiliary agents such as disintegrants, binders, coating agents, swelling agents, lubricants, flavorings, sweeteners, or solubilizers are conventionally used, as described above. The pharmaceutical compositions are suitable for use in a variety of drug delivery systems. For a brief review of methods for drug delivery, see Langer R (1990) Science 249:1527-1533, which is incorporated herein by reference.
[0183] The compounds may be administered neat or in the form of a pharmaceutically acceptable salt. When used in medicine, salts should be pharmaceutically acceptable, but non-pharmaceutically acceptable salts may be conveniently used to prepare pharmaceutically acceptable salts. Such salts include, but are not limited to, those prepared from hydrochloric acid, hydrobromic acid, sulfuric acid, nitric acid, phosphoric acid, maleic acid, acetic acid, salicylic acid, p-toluenesulfonic acid, tartaric acid, citric acid, methanesulfonic acid, formic acid, malonic acid, succinic acid, naphthalene-2-sulfonic acid, and benzenesulfonic acid. Such salts may also be prepared as alkali metal or alkaline earth salts, such as sodium, potassium, or calcium salts of the carboxylic acid group.
[0184] The composition can be conveniently presented in unit dosage form and can be prepared by any method well known in the art of pharmaceuticals.All methods include the step of associating the compound with a carrier that constitutes one or more accessory ingredients.Generally, the composition is prepared by uniformly and intimately associating the compound with a liquid carrier, a finely divided solid carrier, or both, and then, if necessary, shaping the product.Liquid dosage units are vials or ampoules.Solid dosage units are tablets, capsules, and suppositories.
[0185] Detection Method As used herein, a "biological sample" may refer to any specimen derived from or obtained from a subject having or suspected of having a disease (e.g., a neurological disease) associated with the expression, translation, and / or accumulation of RAN protein. In some embodiments, the biological sample is blood, serum (e.g., plasma with clotting proteins removed), or cerebrospinal fluid (CSF). In some embodiments, the biological sample is a tissue sample, e.g., central nervous system (CNS) tissue such as brain tissue or spinal cord tissue. Those skilled in the art will recognize other biological samples, such as cells (e.g., brain cells, nerve cells, skin cells, etc.), that are suitable for the methods described by the present disclosure.
[0186] In some embodiments, methods for detecting one or more disrupted RAN proteins in a biological sample are useful for monitoring the progression of diseases associated with RAN protein expression, translation, and / or accumulation. In some embodiments, the biological sample is selected from the group consisting of amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), and spinocerebellar ataxia type 7 (SCA7). In some embodiments, the biological sample is obtained from a subject having or suspected of having a disease selected from the group consisting of PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8, and / or PSEN1. In some embodiments, the biological sample is obtained from a subject expressing one or more disrupted RAN proteins from a gene or chromosomal locus listed in Table 1 or Table 6. In some embodiments, the biological sample is obtained from a subject having or suspected of having a disease selected from the group consisting of PSEN1, PSEN2, MAPT, FMR1, AR, ATN1, ATXN1, ATXN2, ATXN3, CACNA1A, ATXN7, ATXN8. ATXN8OS, PPP2R2B, TBP, NOP56, ITPR1, ATXN10, DMPK, CNBP, TCF4, HTT, APP, ARMCX4, PEX14, PTPRF, ACTA1, DNAH14, PFN1P2, C1orf61, WASF2, PGBD2, NBPF15, DDX11L1, BARHL2, MIR1976, CASZ1, SLC44A 3, GPR137B, SOX13, CROCC, RNPEP, MIR3121, MPZ, MCL1, HYDIN2, AIFM2, MGMT, LINC01164, KNDC1, ANK 3, MLLT10, TBC1D12, LRMDA, CCNY, MIR3156-1, DUX4L2, AGAP12P, C10orf53, SMPD1, IFITM10, BUD13,TSPAN18、CD82、OTUB1、NADSYN1、MIR4492、CHID1、SMUG1、LINC00938、LINC0 1257、SLC15A4、ASIC1、DCP1B、TMTC2、TNS2、LOC100240735、SOX1、LATS2、RA B20、ANKRD20A9P、FLT1、RCBTB1、ELK2AP、STON2、FOXN3、TTLL5、BCL11B、BRM S1L、SMAD3、RPLP1、BAHD1、MYO5A、DNM1P46、DNM1P35、TUBGCP4、C16orf95、OS GIN1、LINC00311、MIR4718、RBFOX1、SBK1、MIR4722、BANP、C16orf78、MIR51 89、ADGRG5、NPRL3、ZDHHC1、MIR662、LINC00482、MRPL12、TBC1D3H、WSCD1、TB C1D3B、TBC1D3、TBC1D3C、METRNL、DNAH9、ASGR1、FOXK2、NPEPPS、SARM1、CLU H、TIAF1、LOC440434、PHOSPHO1、TBCD、CYP4F35P、CXADRP3、LINC00668、MEX3 C、COX7A1、SCAF1、RFPL4AL1、SIX5、DIRAS1、POLRMT、ZNF554、MED16、SIPA1L 3、DOT1L、KMT5C、PDE4C、ZNF480、CEBPA、PTPN18、HAAO、LOC654342、RGPD2、R AB3GAP1、TNS1、FAM95A、NTSR1、FRG1BP、FAM182B、CDH4、PRNP、MIR1257、MIR 4758、OGFR、SRC、COL9A3、ZNF512B、PICSAR、SIK1、CYP4F29P、MX1、LARGE1、CR ELD2, UPK3A, RRP7A, MIR4762, SHANK3, SHISA8, CCDC188, NPTXR, ZNF621, TPRA1, PIGZ, LHFPL4, OSTN, GAP43, CACNA2D2, TNK2, IQSEC1, RAD18, PARP14, PL XNA1, DOCK3, DUX4L8, PCDH10, TNIP3, ZFYVE28, MSMO1, ANKRD50, FGFR4, IRX 1、ZNF622、SPOCK1、PLEKHG4B、LCP2、SLC34A1、CXXC5、PPARGC1B、LOC643201、P4HA2, THBS2, SEC63, SLC17A5, MEA1, RIMS1, ARID1B, PRKAG2, EN2, NXPH1, NUB1, DPP6, MYL10, GS1-124K5.11, ABCB4, MFSD3, S OX17, MTDH, RRS1-AS1, SDCBP, DOCK5, SHARPIN, LINC00051, LRRC6, NAPRT, FOXE1, C9orf139, FAM27C, AQP7P1, TLE4, NCS1, FAM2 In some embodiments, the biological sample is obtained from a subject expressing one or more disrupted RAN proteins from a gene selected from the group consisting of: 7B, C9orf50, TOR1A, PNPLA7, MIR4473, PRRX2, DAB2IP, C9orf72, GPSM1, FAM230C, RNA (e.g., mRNA)5-8SN5, SUPT20HL2, SUPT20HL1, FAM236A, RPL10, AVPR2, SHROOM2, FAM226A, ALK, and CASP8. In some embodiments, the biological sample is obtained from a subject expressing one or more disrupted RAN proteins from a gene such as ARMCX4, ALK, and / or CASP8.
[0187] In some embodiments, the biological sample has been subjected to one or more processing steps before being used in the detection method. In some embodiments, the biological sample has been subjected to one or more of enzymatic digestion (e.g., by nucleases and / or proteases), contact with chemicals (e.g., for the purposes of cell permeabilization, cell lysis, and / or improving sample stability), or storage for a predetermined period of time (e.g., about 6, 5, 4, 3, 2, or 1 week, or 6, 5, 4, 3, 2, or 1 day) and / or at a predetermined temperature (e.g., at 25°C, 4°C, -20°C, or lower). In some embodiments, the biological sample has been subjected to one or more steps to remove and / or enrich one or more cell types in the biological sample. In some embodiments, the blood sample has been subjected to one or more steps to isolate and / or enrich leukocytes and / or lymphocytes present in the blood sample.
[0188] In some embodiments, the methods described herein involve detecting RAN protein (e.g., a disrupted RAN protein) or an RNA transcript or DNA sequence encoding a RAN protein in a biological sample. In some embodiments, the methods described herein involve subjecting the biological sample to one or more detection methods to determine the presence, absence, or level (e.g., relative to a control sample) of RAN protein (e.g., a disrupted RAN protein). In some embodiments, the methods described herein may involve obtaining, or having obtained, a biological sample from a subject having, or suspected of having, a disease associated with RAN protein expression, translation, and / or accumulation. In some embodiments, the methods described herein may involve performing, or having performed, one or more detection methods on a biological sample obtained from a subject having, or suspected of having, a disease associated with RAN protein expression, translation, and / or accumulation.
[0189] In some embodiments, the differential aggregation properties of RAN proteins having different lengths (e.g., interrupted RAN proteins such as interrupted poly(GR)RAN protein or interrupted poly(GA)RAN protein) can be used to detect RAN proteins (e.g., interrupted RAN proteins) in biological samples. In some embodiments, longer RAN proteins (e.g., interrupted RAN proteins such as interrupted poly(GR)RAN protein or interrupted poly(GA)RAN protein) are found at higher levels in biological samples such as blood, serum, or CSF. In some embodiments, RAN proteins having poly-amino acid repeats greater than 40, greater than 50, greater than 60, greater than 70, or greater than 80 amino acid residues in length (e.g., interrupted RAN proteins such as interrupted poly(GR)RAN protein or interrupted poly(GA)RAN protein) can be detected in biological samples.
[0190] In some embodiments, the methods described herein include detecting a RAN protein (e.g., a disrupted RAN protein) or an RNA transcript or DNA sequence encoding the RAN protein in a biological sample and a control sample. In some embodiments, detecting the presence, absence, or level (e.g., relative level) of a RAN protein (e.g., a disrupted RAN protein) or an RNA transcript or DNA sequence encoding the RAN protein in a biological sample further includes subjecting the control sample to the same detection conditions (e.g., the same assay) as those to which the biological sample was subjected. In some embodiments, the control sample can be a sample lacking a RAN protein (e.g., a disrupted RAN protein) or an RNA transcript or DNA sequence encoding the RAN protein. In some embodiments, the control sample can be a sample containing a normal amount (e.g., a non-pathogenic amount or an amount not associated with disease) of a nucleic acid containing an expanded repeat.
[0191] In some embodiments, the methods described herein involve detecting changes in the level of a RAN protein (e.g., a disrupted RAN protein) or its RNA transcript or DNA sequence. In some embodiments, a RAN protein (e.g., a disrupted RAN protein) or its RNA transcript or DNA sequence is detected (e.g., at elevated levels) when the signal corresponding to the presence of the RAN protein or its RNA transcript or DNA sequence in the biological sample is relatively high (e.g., at least 1.1-10.0-fold or at least 10%-1000% increased) compared to the signal corresponding to the presence of the RAN protein or its RNA transcript or DNA sequence in the control sample. In some embodiments, a RAN protein (e.g., a disrupted RAN protein) or its RNA transcript or DNA sequence is detected (e.g., at reduced levels) when the signal corresponding to the presence of the RAN protein or its RNA transcript or DNA sequence in the biological sample is relatively low (e.g., at least 1.1-10.0-fold or at least 10%-1000% decreased) compared to the signal corresponding to the presence of the RAN protein or its RNA transcript or DNA sequence in the control sample.
[0192] In some embodiments, the methods described herein include obtaining or having obtained a first biological sample from a subject at a first time point and obtaining or having obtained a second biological sample from the subject at a second time point. In some embodiments, the first biological sample or the second biological sample is a control sample. In some embodiments, the first time point occurs before a therapeutic agent is administered to the subject, and the second time point occurs after the therapeutic agent is administered to the subject. In some embodiments, the first time point occurs during the course of treatment of the subject with a therapeutic agent, and the second time point occurs after the subject has completed the course of treatment with the therapeutic agent. However, in other embodiments, the subject may be receiving treatment with a therapeutic agent during both the first and second time points, or may not be receiving treatment during either the first or second time points. In some embodiments, the methods described herein include performing or having performed a first assay on the first biological sample and performing or having performed a second assay on the second biological sample. In some embodiments, the first assay and the second assay are performed at the same time or at different times. In some embodiments, the first assay and the second assay comprise the detection method described herein. In some embodiments, the first assay and the second assay comprise the same assay or different assays.
[0193] In some embodiments, methods for detecting one or more disrupted RAN proteins in a biological sample are useful for monitoring the progression of diseases associated with RAN protein expression, translation, and / or accumulation. In some embodiments, the diseases associated with disrupted RAN proteins include amyotrophic lateral sclerosis (ALS), frontotemporal dementia (FTD), Huntington's disease (HD), Alzheimer's disease (AD), Fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), and spinocerebellar ataxia type 6. In some embodiments, the neurological disease associated with the disrupted RAN protein is selected from the group consisting of spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), and Fuchs' corneal dystrophy. In some embodiments, the neurological disease associated with the disrupted RAN protein is Alzheimer's disease (AD) or amyotrophic lateral sclerosis (ALS).
[0194] For example, in some embodiments, biological samples are obtained from the subject before and after the start of the treatment regimen (e.g., 1 week, 2 weeks, 1 month, 6 months, or 1 year), and the amounts of disrupted RAN protein detected in the samples are compared. In some embodiments, if the level (e.g., amount) of disrupted RAN protein in the post-treatment sample is reduced compared to the pre-treatment level (e.g., amount) of disrupted RAN protein, the treatment regimen is successful. In some embodiments, the level of disrupted RAN protein in the subject's biological samples (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more samples) is continuously monitored during the treatment regimen (e.g., measured on 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more separate occasions).
[0195] In some embodiments, if RAN protein (e.g., a truncated RAN protein) or its RNA transcript or DNA sequence is detected at elevated levels relative to a control sample, an agent (e.g., a therapeutic agent) may be administered to the subject. However, in some embodiments, the methods described herein may include administering an agent (e.g., a therapeutic agent) to a subject without performing a detection method prior to administering the agent. In some embodiments, the methods described herein may include subjecting a biological sample to one or more detection methods after administering an agent (e.g., a therapeutic agent) to a subject.
[0196] Protein detection In some embodiments of the methods described by the present disclosure, a sample (e.g., a biological sample) is processed by an antibody-based capture process to isolate one or more disrupted RAN proteins in the sample. Typically, antibody-based capture methods involve contacting the sample with one or more (e.g., two, three, four, five, or more) anti-RAN protein antibodies. In some embodiments, the one or more anti-RAN antibodies are conjugated to a solid support (e.g., a scaffold, resin beads, etc.). In some embodiments, antibody-based capture methods involve physically separating and / or isolating the disrupted RAN proteins bound by the anti-RAN antibodies, for example, by eluting the disrupted RAN proteins by a chromatographic method such as affinity chromatography or ion exchange chromatography.
[0197] In some embodiments, the biological sample may be subjected to an antigen retrieval procedure before contacting with an anti-RAN antibody. As used herein, "antigen retrieval" (also referred to as epitope retrieval or antigen unmasking) refers to a process of treating a biological sample (e.g., blood, serum, CSF, etc.) under conditions that expose antigens (e.g., epitopes) that were previously inaccessible to detection agents (e.g., antibodies, aptamers, and other binding molecules) prior to the process. Generally, antigen retrieval methods include steps including, but not limited to, heating, pressure treatment, enzymatic digestion, treatment with a reducing agent, treatment with an oxidizing agent, treatment with a crosslinking agent, treatment with a denaturing agent (e.g., detergent, ethanol, acid), or a change in pH, or any combination of the foregoing. Several antigen retrieval methods are known in the art, including, but not limited to, protease-induced epitope retrieval (PIER) and heat-induced epitope retrieval (HIER). In some embodiments, antigen retrieval procedures reduce background and increase the sensitivity of detection techniques (e.g., immunohistochemistry (IHC), immunoblot (such as Western blot), ELISA, etc.).
[0198] In some embodiments, detection of RAN proteins (e.g., truncated RAN proteins) in biological samples can be performed by immunoassays, including the use of detection agents or probes to identify the presence of proteins or peptides (e.g., truncated RAN proteins). In some embodiments, detection of one or more truncated RAN proteins is performed by immunoblots (e.g., dot blots, 2-D gel electrophoresis, Western blots, etc.), electrochemiluminescence immunoassays (e.g., Meso-Scale Detection (MSD)), immunohistochemistry (IHC), ELISA (e.g., RCA-based ELISA or RT-PCR-based ELISA), label-free immunoassays such as surface plasmon resonance biolayer interferometry, immunoquantitative PCR, bead-based immunoassays, immunoprecipitation, immunostaining, or immunoelectrophoresis. However, in some embodiments, methods for detecting RAN proteins can also include mass spectrometry, such as GC-MS, LC-MS, and MALDI-TOF-MS.
[0199] In some embodiments, the detection agent is an antibody or antigen-binding fragment. In some embodiments, the antibody is an anti-RAN protein antibody or an antigen-binding fragment thereof, such as anti-poly(Ser), anti-poly(GR), anti-poly(PR), anti-poly(CP), anti-poly(GP), anti-poly(G), anti-poly(A), anti-poly(GA), anti-poly(GD), anti-poly(GE), anti-poly(GQ), anti-poly(GT), anti-poly(L), anti-poly(LP), anti-poly(LPAC) (SEQ ID NO: 31), anti-poly(LS), anti-poly(P), anti-poly(PA), anti-poly(QAGR) (SEQ ID NO: 35), anti-poly(RE), anti-poly(SP), anti-poly( Anti-poly(VP), anti-poly(FP), anti-poly(GK), anti-poly(FTPLSLPV) (SEQ ID NO: 36), anti-poly(LLPSPSRC) (SEQ ID NO: 37), anti-poly(YSPLPPGV) (SEQ ID NO: 38), anti-poly(HREGEGSK) (SEQ ID NO: 39), anti-poly(TGRERGVN) (SEQ ID NO: 40), anti-poly(PGGRGE) (SEQ ID NO: 41), anti-poly(GRQRGVNT) (SEQ ID NO: 42), and / or anti-poly(GSKHREAE) (SEQ ID NO: 43) (also referred to as α-poly(Ser), α-poly(PR), α-poly(GR), etc.). In some embodiments, the anti-RAN protein antibody or antigen-binding fragment thereof targets (e.g., specifically binds to) the amino acid repeat region of the interrupted RAN protein (e.g., PRPRPRPRPR (SEQ ID NO: 130), GRGRGRGRGR (SEQ ID NO: 131), SSSSSSSSS (SEQ ID NO: 132), etc.). In some embodiments, the anti-RAN protein antibody or antigen-binding fragment thereof targets (e.g., specifically binds) an epitope comprising amino acids at the characteristic reading frame-specific C-terminus translated 3' to the repeated amino acids. In some embodiments, the anti-RAN protein antibody or antigen-binding fragment thereof targets (e.g., specifically binds) an epitope comprising amino acids bridging the C-terminus of the amino acid repeat region and the N-terminus of the characteristic reading frame-specific C-terminus translated 3' to the repeated amino acids.
[0200] In some embodiments, the anti-RAN antibody or antigen-binding fragment thereof is directed against a portion of a truncated RAN protein that does not contain poly-amino acid repeats, e.g., the C-terminus of a truncated RAN protein [e.g., poly(GR), poly(PR), poly(Ser), poly(CP), poly(GP), poly(G), poly(A), poly(GA), poly(GD), poly(GE), poly(GQ), poly(GT), poly(L), poly(LP), poly(LPAC) (SEQ ID NO: 31), poly(LS), poly(P), poly(PA), poly(QAGR) (SEQ ID NO: 32)]. 5), poly(RE), poly(SP), poly(VP), poly(FP), poly(GK), poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41) to the C-terminus of the protein], poly(GRQRGVNT) (SEQ ID NO: 42), and / or poly(GSKHREAE) (SEQ ID NO: 43).
[0201] In some embodiments, a set (or combination) of anti-RAN antibodies or antigen-binding fragments thereof [e.g., anti-poly(Ser), anti-poly(GR), anti-poly(PR), anti-poly(CP), anti-poly(GP), anti-poly(G), anti-poly(A), anti-poly(GA), anti-poly(GD), anti-poly(GE), anti-poly(GQ), anti-poly(GT), anti-poly(L), anti-poly(LP), anti-poly(LPAC) (SEQ ID NO: 31), anti-poly(LS), anti-poly(P), anti-poly(PA), anti-poly(QAGR) (SEQ ID NO: 35), anti-poly(RE), anti-poly(SP), anti-poly(VP), anti-poly(FP), anti-poly(GK)] , anti-poly(FTPLSLPV) (SEQ ID NO: 36), anti-poly(LLPSPSRC) (SEQ ID NO: 37), anti-poly(YSPLPPGV) (SEQ ID NO: 38), anti-poly(HREGEGSK) (SEQ ID NO: 39), anti-poly(TGRERGVN) (SEQ ID NO: 40), anti-poly(PGGRGE) (SEQ ID NO: 41), anti-poly(GRQRGVNT) (SEQ ID NO: 42), and / or anti-poly(GSKHREAE) (SEQ ID NO: 43), or a combination of two or more anti-RAN antibodies or antigen-binding fragments thereof selected from are used to detect one or more disrupted RAN proteins in a biological sample.
[0202] In some embodiments, the detection agent is an aptamer (e.g., an RNA aptamer, a DNA aptamer, or a peptide aptamer). In some embodiments, the aptamer is an aptamer that targets a disrupted RAN protein [e.g., poly(Ser), poly(PR), poly(GR), poly(CP), poly(GP), poly(G), poly(A), poly(GA), poly(GD), poly(GE), poly(GQ), poly(GT), poly(L), poly(LP), poly(LPAC) (SEQ ID NO: 31), poly(LS), poly(P), poly(PA), poly(QAGR) (SEQ ID NO: 35), poly(RE), poly(SP), poly(QAGR) (SEQ ID NO: 36), poly(RE), poly(SP), poly(QAGR) (SEQ ID NO: 37), poly(Ser), poly(PR), poly(GR), poly(CP), poly(GP), poly(G), poly(A), poly(GA), poly(GD), poly(GE), poly(GQ), poly(GT), poly(L), poly(LP), poly(LPAC) (SEQ ID NO: 31), poly(LS), poly(P), poly(PA), poly(QAGR) (SEQ ID NO: 35), poly(RE), poly(SP), poly(QAGR) (SEQ ID NO: 37 ...SP), poly(QAGR) (SEQ ID NO poly(VP), poly(FP), poly(GK), poly(FTPLSLPV) (SEQ ID NO: 36), poly(LLPSPSRC) (SEQ ID NO: 37), poly(YSPLPPGV) (SEQ ID NO: 38), poly(HREGEGSK) (SEQ ID NO: 39), poly(TGRERGVN) (SEQ ID NO: 40), poly(PGGRGE) (SEQ ID NO: 41), poly(GRQRGVNT) (SEQ ID NO: 42), and / or poly(GSKHREAE) (SEQ ID NO: 43).
[0203] Nucleic acid detection In some embodiments, nucleic acid hybridization-based methods are used to identify the presence of interrupted RAN proteins or microsatellite repeat sequences encoding interrupted RAN proteins in a biological sample (e.g., a biological sample obtained from a subject). In some embodiments, nucleic acid hybridization-based methods use nucleic acids capable of hybridizing to a nucleic acid sequence (e.g., a target sequence) (e.g., a nucleotide extension repeat or repeat unit thereof). As used herein, a nucleic acid "capable of hybridizing to" or "capable of detecting" a nucleic acid sequence (e.g., a target sequence) refers to a nucleic acid that comprises a length and a degree of sequence complementarity sufficient to base pair with a nucleic acid sequence (e.g., a target sequence) in a specific and / or stable manner. In some embodiments, the length required for hybridization is about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 20-30, 30-40, 40-50, 50-75, 75-100, or more than 100 nucleotides in length. In some embodiments, the degree of sequence complementarity required for hybridization is at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%. Non-limiting examples of nucleic acids capable of hybridizing to a nucleic acid sequence (e.g., a target sequence) include nucleic acid probes (e.g., detectable probes), guide RNAs (gRNAs), primers, aptamers (e.g., RNA or DNA aptamers), and other forms of antisense oligonucleotides. Non-limiting examples of sequences that may be useful for detecting nucleic acids encoding interrupted RAN proteins are found in Tables 4 and 8.
[0204] In some embodiments, detecting a nucleic acid sequence encoding an interrupted RAN protein uses a detectable nucleic acid probe (e.g., a fluorophore-conjugated DNA probe). Generally, a "detectable nucleic acid probe" refers to a nucleic acid sequence that specifically binds (e.g., hybridizes) to a target sequence and includes a detectable moiety, such as a fluorescent moiety, a radioactive moiety, a chemiluminescent moiety, an electroluminescent moiety, biotin, a peptide tag (e.g., a poly-His tag, a FLAG tag, etc.), or the like. In some embodiments, the detectable nucleic acid probe is a DNA probe or an RNA probe. In some embodiments, the DNA probe or RNA probe is conjugated to a fluorophore. In some embodiments, the detectable nucleic acid probe is chemically modified. In some embodiments, the detectable nucleic acid probe is useful for localizing RAN protein translation by fluorescence in situ hybridization (FISH).
[0205] In some embodiments, the biological sample may be contacted with a plurality of detectable nucleic acid probes. The number of nucleic acid probes in the plurality may vary. In some embodiments, the plurality of nucleic acid probes may be between 2 and 100 (2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, , 45, 46, 47, 48, 49, 50, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleic acid probes. In some embodiments, a plurality comprises more than 100 probes. The nucleic acid probes may be the same or different sequences. In some embodiments, the plurality of detectable nucleic acid probes comprises probes that hybridize to a target sequence encoding an interrupted RAN protein.
[0206] The method for detecting the nucleic acid encoding the interrupted RAN protein may include a concentration step. "Enrichment" refers to the process of increasing the amount and / or concentration of target nucleic acid in a sample relative to other nucleic acids in the sample. Generally, enrichment can occur by increasing the number of target nucleic acid sequences in a sample (for example, by amplifying the target sequence, for example, by polymerase chain reaction (PCR)), or by reducing the amount or concentration of non-target nucleic acid sequences in a sample (for example, by separating or isolating the target nucleic acid sequence from the non-target sequence).
[0207] In some embodiments, the methods described herein include enriching a biological sample for nucleic acid sequences (e.g., microsatellite repeat sequences) encoding interrupted RAN proteins. In some embodiments, enrichment involves contacting the biological sample with 1) a labeled (e.g., biotinylated) dCas9 protein and 2) one or more single-stranded guide RNAs (sgRNAs) that specifically bind to the nucleic acid repeat sequences encoding the interrupted RAN proteins. In some embodiments, the labeled dCas9 protein and one or more sgRNAs are provided together as a single molecule (e.g., a dCas9-sgRNA complex). In some embodiments, after the biological sample contains the labeled dCas9 protein and one or more sgRNAs, the nucleic acid sequences encoding the one or more interrupted RAN proteins are isolated from the labeled dCas9 protein and sgRNAs by affinity chromatography, for example, as described by Liu et al. (2017) Cell 170:1028-1043.
[0208] In some embodiments, detecting one or more interrupted RAN proteins involves next-generation sequencing (NGS). In some embodiments, the enrichment step (e.g., dCas9-based enrichment) is performed on the sample using a guide RNA. In some embodiments, the guide RNA used in enrichment targets NGG protospacer adjacent motif (PAM)-containing repeats. In other embodiments, the guide RNA used in enrichment targets non-NGG PAM-containing repeats. In some embodiments, the non-NGG PAM-containing repeats include extended repeats. In some embodiments, the guide RNA used in enrichment enriches non-NGG PAM-containing repeat extensions that are longer than the corresponding normal allele (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100 repeats longer). In some embodiments, the guide RNAs used in the enrichment simultaneously identify multiple repeat expansions, in some embodiments, including sequences with non-NGG PAMs. In some embodiments, the gRNAs comprise the sequences shown in Table 4.
[0209] In some embodiments, the nucleic acid used for the detection methods described herein is capable of hybridizing to a nucleic acid sequence (e.g., a target sequence) present in a gene, chromosome, or RNA transcript (e.g., mRNA) encoding an interrupted RAN protein. In some embodiments, the target sequence comprises one or more extended repeats or repeat units. In some embodiments, the target sequence, or a portion thereof, is selected from the group consisting of GGTCGT, GGCCGT, GGACGT, GGGCGT, GGTCGC, GGCCGC, GGACGC, GGGCGC, GGTCGA, GGCCGA, GGACGA, GGGCGA, GGTCGG, GGCCGG, GGACGG, GGGCGG, GGTAGA, GGCAGA, GGAAGA, GGGAGA, GGTAGG, GGCAGG, GGAAGG, GGGAGG, GGTGCT, GGCGCC, GGAGCA, GGGGCG, GGTGCC, GGCGCA, GGAGCG, GGGGCT, GGTGCA, GGCGCG, GGAGCT, GGGGCC, GGTGCG, GGCGCT, GGAGCC, GGGGCA GGTCCT, GGCCCC, GGACCA, GGGCCG, GGTCCC, GGCCCA, GGACCG, GGGCCT, GGTCCA, GGCCCG, GGACCT, GGGCCC, GGTCCG, GGCCCT, GGACCC, GGGCCA CCTCGT, CCCCGT, CCACGT, CCGCGT, CCTCGC, CCCCGC, CCACGC, CCGCGC, CCTCGA, CCCCGA, CCACGA, CCGCGA, CCTCGG, CCCCGG, CCACGG, CCGCGG, CCTAGA, CCCAGA, CCAAGA, CCGAGA, CCTAGG, CCCAGG, CCAAGG, CCGAGG TCT, TCC, TCA, TCG, AGT, AGC, CCTG, or CAGG repeat units.In some embodiments, the target sequence...
Claims
1. A method for identifying a subject as having a RAN protein disease, the method comprising detecting one or more interrupted RAN proteins in a biological sample obtained from the subject, each of the one or more interrupted RAN proteins comprising multiple RAN repeat units.
2. The method of claim 1, wherein the one or more disrupted RAN proteins comprise a poly-GA-disrupted RAN protein.
3. At least one of the interrupted RAN proteins is (GGGGCT) x Expanded repeat or (GAAGGA) x 3. The method of claim 1 or 2, wherein the extended repeat is translated from the expanded repeat, and wherein x comprises an integer between 2 and 200.
4. The method of any one of claims 1 to 3, wherein the one or more interrupted RAN proteins comprises a poly-GR interrupted RAN protein.
5. At least one of the interrupted RAN proteins is (GGGAGA) x 5. The method of claim 1 , wherein the extended repeat is translated from the sequence X.
6. 6. The method of any one of claims 1 to 5, wherein at least one of the interrupted RNA proteins is transcribed from a gene or chromosomal locus shown in Table 1, and wherein at least one of the interrupted RNA proteins may be transcribed from ARMCX4, ALK, and / or CASP8.
7. The method of any one of claims 1 to 6, wherein at least one of the interrupted RAN proteins comprises at least one amino acid residue between each RAN repeat unit.
8. The method of claim 7, wherein at least one of the interrupted RAN proteins contains between 2 and 20 amino acid residues between each RAN repeat unit.
9. The method of any one of claims 1 to 8, wherein the subject is a human.
10. The method of claim 1 , wherein the detecting comprises performing an assay on the biological sample.
11. 11. The method of claim 10, wherein the assay comprises an antibody-based capture assay, a binding assay, a hybridization assay, immunoblot analysis, Western blot analysis, immunohistochemistry, dCas9-based enrichment, label-free immunoassay, immunoquantitative PCR, mass spectrometry, bead-based immunoassay, immunoprecipitation, immunostaining, immunoelectrophoresis, and / or ELISA.
12. The RAN protein disease is selected from the group consisting of amyotrophic lateral sclerosis (ALS), Huntington's disease (HD), Alzheimer's disease (AD), fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), and the like. 8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), Alzheimer's disease (AD), or Fuchs' corneal dystrophy (e.g., CTG181).
13. The method of claim 12, wherein the RAN protein disease is Alzheimer's disease (AD) or amyotrophic lateral sclerosis (ALS).
14. 14. The method of any one of claims 1 to 13, further comprising administering to the subject one or more anti-RAN protein agents.
15. 15. The method of claim 14, wherein the one or more anti-RAN protein agents comprise a protein, peptide, nucleic acid, or small molecule.
16. The method of claim 15, wherein the protein comprises an antibody, and the antibody may be an anti-poly-GA antibody or an anti-poly-GR antibody, and the anti-poly-GA antibody may specifically bind to the poly-GA repeat region of the subject's RAN protein, and the anti-poly-GR antibody may specifically bind to the poly-GR repeat region of the subject's RAN protein, and the anti-poly-GA antibody or anti-poly-GR antibody may be a monoclonal antibody.
17. 15. The method of claim 14, wherein the one or more anti-RNA protein agents comprise a nucleic acid, which may be double-stranded RNA (dsRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), artificial microRNA (amiRNA), aptamer, or antisense oligonucleotide (ASO), and which may comprise a region of complementarity to a nucleic acid sequence encoding a poly-GA or poly-GR repeat expansion in the subject, and which may comprise a region of complementarity to a nucleic acid sequence present at a chromosomal locus or gene set forth in Table 1.
18. 15. The method of claim 14, wherein the one or more anti-RAN protein agents comprise a small molecule, and the small molecule may be metformin.
19. A method described in any one of claims 14 to 18, wherein the administration results in a reduction in the transcription, translation, expression, accumulation, or aggregation of RAN protein in the subject relative to the level of transcription, translation, expression, aggregation, or accumulation of RAN protein in the subject prior to the administration.
20. A method of treating a subject having or suspected of having a RAN protein-associated disease or disorder, comprising administering to said subject one or more anti-RAN protein agents.
21. The method of claim 20, wherein the one or more anti-RAN protein agents target one or more interrupted RAN proteins, each of the one or more interrupted RAN proteins comprising multiple RAN repeat units.
22. 22. The method of claim 21, wherein the one or more interrupted RAN proteins comprise a poly-GA interrupted RAN protein.
23. At least one of the interrupted RAN proteins is (GGGGCT) x Expanded repeat or (GAAGGA) x 23. The method of claim 22, wherein the extended repeat is translated from the expanded repeat, wherein x comprises an integer between 2 and 200.
24. 22. The method of claim 21, wherein the one or more disrupted RAN proteins comprise a poly-GR-disrupted RAN protein.
25. At least one of the interrupted RAN proteins is (GGGAGA) x 25. The method of claim 24, wherein the extended repeat is translated from the expanded repeat, wherein x comprises an integer between 2 and 200.
26. 26. The method of any one of claims 21 to 25, wherein at least one of the interrupted RNA proteins is transcribed from a gene or chromosomal locus set forth in Table 1, and wherein at least one of the interrupted RNA proteins may be transcribed from ARMCX4, ALK, and / or CASP8.
27. 27. The method of any one of claims 21 to 26, wherein at least one of the interrupted RAN proteins comprises at least one amino acid residue between each RAN repeat unit.
28. 28. The method of claim 27, wherein at least one of the interrupted RAN proteins comprises between 2 and 20 amino acid residues between each RAN repeat unit.
29. 29. The method of any one of claims 20 to 28, wherein the subject is a human.
30. The RAN protein disease is selected from the group consisting of amyotrophic lateral sclerosis (ALS), Huntington's disease (HD), Alzheimer's disease (AD), fragile X syndrome (FRAXA), spinal-bulbar muscular atrophy (SBMA), dentatorubral-pallidoluysian atrophy (DRPLA), spinocerebellar ataxia type 1 (SCA1), spinocerebellar ataxia type 2 (SCA2), spinocerebellar ataxia type 3 (SCA3), spinocerebellar ataxia type 6 (SCA6), spinocerebellar ataxia type 7 (SCA7), spinocerebellar ataxia type 8 (SCA8), and the like. 8), spinocerebellar ataxia type 12 (SCA12), spinocerebellar ataxia type 17 (SCA17), spinocerebellar ataxia type 36 (SCA36), spinocerebellar ataxia type 29 (SCA29), spinocerebellar ataxia type 10 (SCA10), myotonic dystrophy type 1 (DM1), myotonic dystrophy type 2 (DM2), Alzheimer's disease (AD), or Fuchs' corneal dystrophy (e.g., CTG181).
31. The method of claim 30, wherein the RAN protein disease is Alzheimer's disease (AD) or amyotrophic lateral sclerosis (ALS).
32. 32. The method of any one of claims 20 to 31, wherein the one or more anti-RAN protein agents comprise a protein, peptide, nucleic acid, or small molecule.
33. 33. The method of claim 32, wherein the protein comprises an antibody, and the antibody may be an anti-poly-GA antibody or an anti-poly-GR antibody, wherein the anti-poly-GA antibody may specifically bind to the poly-GA repeat region of the subject's RAN protein, and wherein the anti-poly-GR antibody may specifically bind to the poly-GR repeat region of the subject's RAN protein, and wherein the anti-poly-GA antibody or anti-poly-GR antibody may be a monoclonal antibody.
34. 33. The method of claim 32, wherein the one or more anti-RNA protein agents comprise a nucleic acid, which may be double-stranded RNA (dsRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), artificial microRNA (amiRNA), aptamer, or antisense oligonucleotide (ASO), and which may comprise a region of complementarity to a nucleic acid sequence encoding a poly-GA or poly-GR repeat expansion in the subject, and which may comprise a region of complementarity to a nucleic acid sequence present at a chromosomal locus or gene set forth in Table 1.
35. 33. The method of claim 32, wherein the one or more anti-RAN protein agents comprise a small molecule, and the small molecule may be metformin.
36. A method described in any one of claims 20 to 35, wherein the administration results in a reduction in the transcription, translation, expression, accumulation, or aggregation of RAN protein in the subject relative to the level of transcription, translation, expression, aggregation, or accumulation of RAN protein in the subject prior to the administration.
37. 37. The method of any one of claims 20 to 36, wherein the subject is identified as having a RAN protein disorder according to the method of any one of claims 1 to 19.
38. 38. The method of any one of claims 1 to 19 or 37, wherein the biological sample is tissue, blood, serum, or cerebrospinal fluid (CSF), and the tissue may be brain tissue or spinal cord tissue.