Compositions and methods for treating Huntington's disease by editing the mutant huntingtin gene

RNA-guided nuclease systems targeting mutant huntingtin alleles using CRISPR-Cas protein complexes address the incomplete removal of mutant huntingtin protein in Huntington's disease, achieving significant protein level reductions.

JP2026516656APending Publication Date: 2026-05-26LIFEEDIT THERAPEUTICS INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
LIFEEDIT THERAPEUTICS INC
Filing Date
2024-04-12
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Current RNA interference therapies do not completely eliminate mutant huntingtin protein expression in Huntington's disease, necessitating the development of more effective genome editing systems for targeted genomic modifications.

Method used

Utilization of RNA-guided nuclease (RGN) systems, including CRISPR-Cas protein complexes with guide RNAs, to specifically target and cleave mutant huntingtin (mutHTT) alleles, reducing mutHTT protein levels.

Benefits of technology

The RGN systems effectively reduce mutHTT protein levels in both cellular and animal models, demonstrating therapeutic potential for Huntington's disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026516656000001_ABST
    Figure 2026516656000001_ABST
Patent Text Reader

Abstract

Compositions and methods for cleaving mutant huntingtin (mutHTT) alleles are provided. The compositions comprise CRISPR RNA, a guide RNA, and an encoding nucleic acid molecule. Vectors and host cells comprising the nucleic acid molecule are also provided. Further provided are RNA-induced nuclease (RGN) systems for cleaving mutHTT alleles, wherein the RGN system comprises an RNA-induced nuclease and a guide RNA. The compositions are useful for cleaving or modifying mutHTT alleles and / or modifying the expression of mutHTT alleles. The compositions are even more useful for the treatment of Huntington's disease (HD), particularly in an allele-specific manner.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Application No. 63 / 495,725 filed on 12 April 2023, U.S. Provisional Application No. 63 / 497,904 filed on 24 April 2023, U.S. Provisional Application No. 63 / 518,231 filed on 8 August 2023, U.S. Provisional Application No. 63 / 593,881 filed on 27 October 2023, and U.S. Provisional Application No. 63 / 555,290 filed on 19 February 2024, each of which is incorporated herein by reference in its entirety.

[0002] References to sequence listings submitted electronically as XML files. This application includes a sequence listing submitted in XML format via the USPTO Patent Centre, which is incorporated herein by reference in its entirety. The XML copy was created on 10 April 2024, is named L103438_1300WO_0257_6_SL, and is 251,299 bytes in size.

[0003] This invention relates to the fields of molecular biology and gene editing. [Background technology]

[0004] Huntington's disease (HD) is a hereditary neurodegenerative disorder caused by an expansion of the cytosine-adenine-guanine (CAG) trinucleotide in the huntingtin (HTT) gene (Huntington's Disease Collaborative Research Group, 1993, Cell 72:971-983). The resulting polyglutamine (PolyQ)-containing mutant HTT disrupts the function of wild-type HTT (wtHTT), causing neurological stress and dysfunction (Kaemmerer et al, 2019, Degenerative Neurological and Neuromuscular Disease 9:3-17). HD patients develop striatal atrophy along with cognitive impairment, followed by progressive psychiatric and motor disorders (Ross et al., 2014, Nat. Rev. Neurol. 10:204-216). Mouse models expressing either the full-length human HTT containing an expanded CAG repeat or exon 1 reproduce the HD pathology (Southwell et al., 2016, Human Molecular Genetics 25(17):3654-3675; Slow et al., 2003, Human Molecular Genetics 12(13):1555-1567; Raamsdonk et al, 2007, Neurobiology of Disease 26:189-200; Southwell et al., 2017, Human Molecular Genetics 26(6):1115-1132). The reduction in mutant HTT levels in HD animal models, which leads to improvements in motor and neuropathological abnormalities, supports the therapeutic approach of HTT reduction (Miniarikova et al., 2016, Mol Ther Nucleic Acids 5(3):e297; Caron et al., 2020, Nucleic Acids Research 48(1):36-54; Spronck et al., 2019, Mol Ther Methods Clin Dev 13:334-343; Stanek et al., 2014, Human Gene Therapy 25(5):461-474).

[0005] While RNA interference has been developed to reduce mutant huntingtin protein levels, RNAi therapy does not completely eliminate mutant huntingtin expression. Targeted genome editing offers an opportunity to introduce changes at the genomic level. Editing or modifying the targeted genome is rapidly becoming an important tool for basic and applied research because it allows for genomic modifications such as nucleic acid cleavage, nucleic acid deletion, nucleic acid insertion, nucleotide substitution within nucleic acids, and regulation of gene expression at specific locations within the genome, as well as many other possible modifications. The initial efforts in genome editing involved designing nucleases, i.e., proteins that can edit nucleic acids in order to recognize the target nucleic acid sequence to be edited and specifically bind to it. However, manipulating nucleases requires considerable time and experimentation to obtain one that is effective for editing a particular sequence. Genome editing systems using RNA-inducible nucleases, such as the clustered regular-spacing short palindromic sequence repeat (CRISPR)-associated (Cas) protein of the CRISPR-Cas bacterial system, function by complexing the nuclease with a guide RNA. Hybridization of guide RNA to a specific target sequence enables editing at a specific location in the genome. Therefore, genome editing systems using RNA-induced nucleases may be less expensive and more efficient for editing genome sequences because nucleic acids are typically easier to design and redesign compared to nucleases.

[0006] Therefore, patients suffering from diseases such as Huntington's disease, which are associated with specific deletions in the genome, would benefit from the development of RNA-induced nuclease systems that can edit genomic deletions for therapeutic purposes. [Overview of the project]

[0007] Compositions and methods for cleaving mutant huntingtin (mutHTT) alleles are provided. The compositions include CRISPR RNAs, guide RNAs, and nucleic acid molecules encoding them. Vectors and host cells containing the nucleic acid molecules are also provided. A RNA-guided nuclease (RGN) system for cleaving the mutHTT allele is further provided, wherein the RGN system includes a RNA-guided nuclease and a guide RNA. The compositions are useful for cleaving or modifying the mutHTT allele and / or modifying the expression of the mutHTT allele. The compositions are further useful for the treatment of Huntington's disease (HD), particularly in an allele-specific manner.

[0008] A method for cleaving the mutHTT allele intracellularly includes introducing an RGN or a nucleic acid molecule encoding the RGN, and a guide RNA or a nucleic acid molecule encoding the guide RNA, wherein the mutHTT allele includes a single nucleotide polymorphism (SNP) allele at exon 50, the SNP allele generates a protospacer adjacent motif (PAM), and the RGN can recognize the PAM and cleave the mutHTT allele.

[0009] In a subject that needs to alleviate or delay the onset of one or more symptoms of HD, a method of doing so includes administering an RGN system to the subject, wherein the subject's mutHTT allele includes a SNP allele at exon 50 that generates a protospacer adjacent motif (PAM) recognized by the RGN. The RGN then cleaves and edits the mutHTT allele, and the level of the mutHTT protein encoded by the mutHTT allele is reduced compared to a control subject or a wild-type HTT protein. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] [Figure 1]This shows the percentage of insertions and / or deletions (indels) in patient fibroblasts nucleofected with APG07433.1 nuclease and SGN002908 or SGN002911 guide RNA, or with APG05586 nuclease and SGN004282 guide RNA. [Figure 2A] These figures represent the percentage of indels and the percentage of edited reads in patient fibroblasts nucleofected with APG05586 nuclease and SGN004282, SGN008949, or SGN007707 guide RNAs, respectively. [Figure 2B] Figures 2B show the percentage of indels and the percentage of edited reads in patient fibroblasts nucleofected with APG05586 nuclease and SGN004282, SGN008949, or SGN007707 guide RNA, respectively. No edited reads for the C allele were detected in any of the patient fibroblasts tested with APG05586 nuclease or SGN004282 guide RNA. [Figure 3A] This study provides immunofluorescence analysis of neuronal marker genes in forebrain neurons derived from induced pluripotent stem cells (iPSCs). (Left panel) shows that iPSC-derived forebrain neurons express the neuronal markers Tuj1 (β-tubulin III, green) and gamma-aminobutyric acid (GABA) (red) (10x magnification). (Center panel) shows that they also express the neuronal marker Tuj1 (green) and the mature neuron marker microtubule-associated protein 2 (MAP2) (red) (10x magnification). (Right panel) shows that iPSC-derived cells express Tuj1 (green) and the forebrain neuron-specific marker FoxG1 (red) (20x magnification). The nuclei are labeled with DAPI (blue). [Figure 3B]This paper provides immunofluorescence analysis of neuronal marker genes in forebrain neurons derived from induced pluripotent stem cells (iPSCs). FACS analysis of iPSC-derived forebrain neurons is shown. Negative controls are shown in the leftmost panel. The panels from the center left to the far right display staining results for the following antibody combinations: Ki67 and nerve filament heavy chain (NEFH) (center left), GABA and MAP2 (center right), and GFAP and Tuj1 (far right). [Figure 4] This provides immunofluorescence analysis of the neuronal marker gene cAMP regulatory phosphoprotein with an apparent molecular weight of 32 kDa (DARPP-32), GABA, MAP2, or Ctip2 in medium-sized spiny neurons derived from iPSCs. [Figure 5] As shown in Figure 1, a graph is provided showing the percentage of indels in induced pluripotent stem cells (iPSCs), neural progenitor cells (NPCs), forebrain neuron progenitor cells (FBPs), or forebrain neurons (FBNs), using the indicated guide RNA (and appropriate nuclease). [Figure 6] This study demonstrates AAV5-mediated intrastriatal delivery of APG07433.1, resulting in substantial levels of AAV5 vector DNA in clinically relevant brain regions: the striatum and cortex. Dose-response was observed at both 4 weeks and 3 months post-administration. Each point represents an individual mouse with the indicated mean ± SE. Naive and vehicle-treated animals are not shown because they are below the lower limit of quantification (LLOQ). vg = viral genome. [Figure 7] AAV5-mediated intrastriatal delivery of APG07433.1 is demonstrated, resulting in strong APG07433.1 transgene expression four weeks after administration. dPCR analysis from the right striatum. Each point represents an individual mouse with the indicated mean ± SE. [n=3-6 mice per group]. [Figure 8]This shows dose-dependent reduction of mutant HTT protein in the striatum at 4 weeks and 3 months after administration of AAV5-JeT-APG07433.1-SGN002908. Each point represents an individual mouse with the indicated mean ± SE. [n=2-6 mice per group at 4 weeks, n=2-10 mice per group at 3 months]. **P<0.01, ****P<0.0001. [Figure 9] This shows the dose-dependent reduction of mutant HTT mRNA evaluated 3 months after intrastriatal administration of AAV5-JeT-APG07433.1-SGN002908. Each point represents an individual mouse with the indicated mean ± SE. [n=2-10 mice per group]. Compared to naive animals at the same time, *P<0.05, **P<0.01. [Figure 10] This document provides the results of indel next-generation sequencing (NGS) analysis showing editing in the striatum of animals treated with 6.4E10vg and 3.6E11vg of AAV5-JeT-APG07433.1-SGN002908 at 4 weeks and 3 months post-administration. Each point represents an individual mouse with the indicated mean ± SE. [2–4 mice per group for 4 weeks, 2–10 mice per group for 3 months]. [Figure 11] This shows that intrastriatal injection of 1.72E11vg of AAV5-hU6-SGN004282-JeT-APG05586 resulted in robust AAV5 vector levels in the striatum, leading to APG05586 nuclease editing and a 30% reduction in mutHTT protein. Each point represents an individual animal with mean ± SE. The naive and vehicle cohorts were below the quantification limit of the vector's in vivo distribution. *P<0.05, unpaired t-test against concurrent naive control cohort. [Figure 12] This graph shows the in vivo distribution of the vector genome in the striatum and cortex after intrastriatal delivery accompanied by subsequent APG05586 nuclease expression. Each point represents an individual animal with mean ± SE. The naive cohort was below the quantification limit for vector in vivo distribution. [Figure 13]This shows that intrastriatal administration of AAV5-hU6-SGN004282-JeT-APG05586 and AAV5-hU6-SGN004282-hSyn-APG05586 resulted in a reduction of mutHTT protein, confirming editing. Each point represents an individual animal with mean ± SE. Compared to concurrently naive animals using unpaired Student's t-test, **P<0.01. [Figure 14] This shows that intrastriatal delivery of 2.84E11vg of AAV5-hU6-SGN004282-hSyn-APG05586 resulted in widespread vector genome distribution within the striatum, leading to APG05586 nuclease expression. Each point represents an individual animal with mean ± SE. The naive cohort was below the quantification limit of vector in vivo distribution. [Figure 15] This study shows that intrastriatal administration of AAV5-hU6-SGN004282-hSyn-APG05586 resulted in a reduction of mutHTT mRNA and protein, confirming genome editing in BACHD mice. Each point represents an individual animal with mean ± SE. Compared to concurrently naive animals using unpaired Student's t-test, *P<0.05, ****P<0.0001. [Figure 16A] This report shows the in vivo distribution and nuclease expression of AAV in BACHD mice. It demonstrates AAV5-mediated intrastriatal delivery of the codon-optimized construct APG05586 in BACHD mice, resulting in substantial levels of AAV5 vector DNA in clinically relevant brain regions: the striatum and cortex. Placement was evaluated 6 weeks post-administration. Each point represents an individual mouse with the indicated mean ± SE. Naive and vehicle-treated animals are not shown due to being below LLOQ. [Figure 16B]This report shows the in vivo distribution and nuclease expression of AAV in BACHD mice. It also shows nuclease expression from AAV5-mediated intrastriatal delivery of the codon-optimized construct APG05586 in BACHD mice. Distribution was evaluated 6 weeks post-administration. Each point represents an individual mouse with the indicated mean + SE. Naive and vehicle-treated animals are not shown because they are below LLOQ. [Figure 17] This study shows that an AAV5 cassette expressing SGN04282, driven by the hU6 (249-318bp) promoter and mammalian codon-optimized APG05586, directed by various promoters (Jet, hSyn, CMVeb, EFS), reduced mutant huntingtin protein after intrastriatal administration in BACHD mice. Each point represents an individual mouse with the indicated mean ± SE. *P<0.05, **P<0.01, ****P<0.0001, one-way ANOVA with Dunnett's assessment of post-hoc tests. [Figure 18] This shows confirmation of editing by NGS indel analysis. Genome editing was confirmed by the percentage of indel events in the striatum and cortex following intrastriatal administration of an AAV5 cassette expressing SGN004282, driven by the hU6 (249-318 bp) promoter and mammalian codon-optimized APG05586, directed by various promoters (Jet, hSyn, CMVeb, EFS) in BACHD mice. Each point represents an individual mouse with the indicated mean ± SE. [Figure 19A] This report provides the results of introducing an increased amount of AAV5-hU6-SGN004282-hSyn-APG05586mco-SV40pA containing mammalian codon-optimized APG05586 into BACHD mice after 6 weeks. It also provides the in vivo distribution of the vector in the striatum and cortex as viral dose increases. [Figure 19B]We present the results of introducing an increased amount of AAV5-hU6-SGN004282-hSyn-APG05586mco-SV40pA, which contains mammalian codon-optimized APG05586, into BACHD mice after 6 weeks. This shows a reduction in mutHTT protein in the striatum with increasing viral dose. [Figure 19C] We provide the results of introducing an increased amount of AAV5-hU6-SGN004282-hSyn-APG05586mco-SV40pA containing mammalian codon-optimized APG05586 into BACHD mice after 6 weeks. Figure 19B shows the dose escalation study as a percentage of mutHTT reduction. In addition, the comparison of the dose of 2.7e11vg of the viral genome (vg) on ​​the left side of the graph in Figure 19C with the same dose in optimization study number 2 demonstrates the effect of codon optimization of APG05586. Finally, administration of 7.32e10vg of the codon-optimized construct demonstrates the reproducibility of its effect on mutHTT protein levels. [Figure 20A] This demonstrates the specificity of capillary electrophoresis (CE) immunoassays for mutHTT and wild-type HTT. Electrophoretic diagrams of SDS sample blanks are provided, showing the full scale. [Figure 20B] This document demonstrates the specificity of capillary electrophoresis (CE) immunoassays for mutHTT and wild-type HTT. Electrophoretic diagrams of SDS sample blanks are provided, along with magnified views. [Figure 20C] This demonstrates the specificity of capillary electrophoresis (CE) immunoassays for mutHTT and wild-type HTT. Electrophoretic graphs of Q73HTT (mutHTT) or Q7HTT (wild-type HTT) at various concentrations (0.12 ng / ml to 30 ng / ml) are provided. [Figure 20D] This demonstrates the specificity of capillary electrophoresis (CE) immunoassays for mutHTT and wild-type HTT. Electrophoretic graphs of Q73HTT (mutHTT) or Q7HTT (wild-type HTT) at various concentrations (0.12 ng / ml to 30 ng / ml) are provided. [Figure 20E]This shows the specificity of capillary electrophoresis (CE) immunoassays for mutHTT and wild-type HTT. The electrophoretic maps of Q73HTT and Q7HTT from Figures 20C and 20D are shown as Western blots, respectively. [Figure 21A] This shows the specificity of CE immunoassays for mutHTT (Q73HTT) and wild-type HTT (Q7HTT) proteins. A graph demonstrating the linearity of Q7HTT is provided. [Figure 21B] This shows the specificity of CE immunoassays for mutHTT (Q73HTT) and wild-type HTT (Q7HTT) proteins. A graph demonstrating the linearity of Q73HTT is provided. [Figure 22A] The results of CE immunoassays of brain samples from BACHD mice treated with an AAV5 construct containing SGN004282 and codon-optimized APG05586 are shown, compared to untreated animals. Western blots are provided. [Figure 22B] The results of CE immunoassays of brain samples from BACHD mice treated with an AAV5 construct containing SGN004282 and codon-optimized APG05586 are shown, compared to untreated animals. Electrophoretic graphs are provided. [Figure 23A] The results of CE immunoassays on brain samples from BACHD mice treated with CMV (treatment 1 and cohort A), EFS (treatment 2 and cohort B), or untreated mice (naive or cohort G) are shown. Compared to naive mice, the treated mice show a reduction in mutHTT. [Figure 23B] The results of CE immunoassays on brain samples from BACHD mice treated with CMV (treatment 1 and cohort A), EFS (treatment 2 and cohort B), or untreated mice (naive or cohort G) are shown. Electrophoretic images from these tests are provided. [Figure 24A]This paper provides results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. It provides the in vivo distribution of the respective vectors, guide RNA, APG05586 mRNA, and APG05586 protein in the striatum with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 24B] This paper provides results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. It provides the in vivo distribution of the respective vectors, guide RNA, APG05586 mRNA, and APG05586 protein in the striatum with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 24C]This paper provides results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. It provides the in vivo distribution of the respective vectors, guide RNA, APG05586 mRNA, and APG05586 protein in the striatum with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 24D] This paper provides results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. It provides the in vivo distribution of the respective vectors, guide RNA, APG05586 mRNA, and APG05586 protein in the striatum with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 24E]We provide results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each having a mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. The percentage of indel formation in the striatum with increasing viral dose is shown. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 24F] We provide results in the striatum after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each having a mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. The low, medium, and high doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA were 2.05E10vg, 7.28E10vg, and 2.94E11vg, respectively. The low, medium, and high doses of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA were 2.05E10vg, 7.28E10vg, and 2.05E11vg, respectively. This shows a decrease in mutHTT protein in the striatum with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25A]We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. Each provides the in vivo distribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the cortex with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25B] We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. Each provides the in vivo distribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the cortex with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25C]We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. Each provides the in vivo distribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the cortex with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25D] We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. Each provides the in vivo distribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the cortex with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25E]We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each containing a mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. This shows a reduction in mutHTT protein in the cortex with increasing viral dose. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 25F] We provide cortical results after intrastriatal introduction of increased amounts of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA, each having a mammalian codon-optimized APG05586, into BACHD mice after 12 weeks. The low, medium, and high doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA were 2.05E10vg, 7.28E10vg, and 2.94E11vg, respectively. The low, medium, and high doses of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA were 2.05E10vg, 7.28E10vg, and 2.05E11vg, respectively. The percentage of indel formation in the cortex with increasing viral dose is shown. The results for AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are shown on the left side of the graph, and the results for AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are shown on the right side. [Figure 26A] Results in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 123 are provided 6 weeks after intrastriatal administration of the test material. The in vivo distribution of the vector, APG05586mco mRNA, and guide RNA are provided, respectively. [Figure 26B] Results in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 123 are provided 6 weeks after intrastriatal administration of the test material. The in vivo distribution of the vector, APG05586mco mRNA, and guide RNA are provided, respectively. [Figure 26C] Results in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 123 are provided 6 weeks after intrastriatal administration of the test material. The in vivo distribution of the vector, APG05586mco mRNA, and guide RNA are provided, respectively. [Figure 26D] Results in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 123 are provided 6 weeks after intrastriatal administration of the test material. The percentage of indel formation is shown. [Figure 26E] Results in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122, or SEQ ID NO: 123 are provided 6 weeks after intrastriatal administration of the test material, showing a reduction in mutHTT protein. [Figure 27A] This study demonstrates improved activity of AAV constructs containing c-MYC NLS and NLS linker proteins in generating indels in HEK293t cells and iPSC-derived astrocytes with AAV5 or AAV6 serotypes. Immunofluorescence images of iPSC-derived astrocytes are provided. [Figure 27B] This shows improved activity of AAV constructs containing c-MYC NLS and NLS linker proteins in generating indels in HEK293t cells and iPSC-derived astrocytes with AAV5 or AAV6 serotypes. Flow cytometry analysis of iPSC-derived astrocytes stained with anti-glial fibrillary acidic protein (GFAP)-488 antibody is also shown. [Figure 27C]This shows improved activity of AAV constructs containing c-MYC NLS and NLS linker proteins in generating indels in HEK293t cells and iPSC-derived astrocytes with AAV5 or AAV6 serotypes. The indel rates after AAV6 (Figures 27C and 27D) or AAV5 (Figure 27E) transduction in iPSC-derived astrocytes (Figure 27C) or HEK293t cells (Figures 27D and 27E) using Sequence IDs 36, 121, 122, or 123 are shown. [Figure 27D] This shows improved activity of AAV constructs containing c-MYC NLS and NLS linker proteins in generating indels in HEK293t cells and iPSC-derived astrocytes with AAV5 or AAV6 serotypes. The indel rates after AAV6 (Figures 27C and 27D) or AAV5 (Figure 27E) transduction in iPSC-derived astrocytes (Figure 27C) or HEK293t cells (Figures 27D and 27E) using Sequence IDs 36, 121, 122, or 123 are shown. [Figure 27E] This shows improved activity of AAV constructs containing c-MYC NLS and NLS linker proteins in generating indels in HEK293t cells and iPSC-derived astrocytes with AAV5 or AAV6 serotypes. The indel rates after AAV6 (Figures 27C and 27D) or AAV5 (Figure 27E) transduction in iPSC-derived astrocytes (Figure 27C) or HEK293t cells (Figures 27D and 27E) using Sequence IDs 36, 121, 122, or 123 are shown. [Figure 28] This shows the in vivo distribution of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) after bilateral intrastriatal administration in adult cynomolgus monkeys. Animals were administered 225 μl / animal (75 μl / caudate nucleus + 150 μl / putamen). Low dose (N=2), high dose (N=3). The vector genome was determined by qPCR using a primer-probe set targeting the APG05586mco sequence. [Figure 29] This study shows the expression of APG05586mco in the brain of adult cynomolgus monkeys after bilateral intrastriatal administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). Animals were administered 225 μl / animal (75 μl / caudate nucleus + 150 μl / putamen). Low dose (N=2), high dose (N=3). mRNA transcripts were quantified by qPCR using a primer-probe set targeting the APG05586mco sequence. [Figure 30] This study shows the expression of SGN004282 guide RNA in the brain of adult cynomolgus monkeys after bilateral intrastriatal administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123). Animals were administered 225 μl / animal (75 μl / caudate nucleus + 150 μl / putamen). Low dose (N=2), high dose (N=3). mRNA transcripts were quantified by qPCR using a primer-probe set targeting the SGN004282 sequence. [Figure 31] This shows the in vivo distribution of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) in peripheral tissues after bilateral intrastriatal administration in adult cynomolgus monkeys. Animals were administered 225 μl / animal (75 μl / caudate nucleus + 150 μl / putamen). Low dose (N=2), high dose (N=3). The sample panel was used as suggested by the ICH S12 guideline: Nonclinical Biodistribution Considerations for Gene Therapy Products. After low-dose administration, 3 out of 12 tissues contained vector DNA, and after high-dose administration, 6 out of 12 tissues contained vector DNA. There is no evidence of vector DNA presence in the testes / ovaries. *: No evidence of vector DNA presence. [Figure 32A] This diagram shows a scheme for immunogenicity testing to assay samples obtained from cynomolgus monkeys before and after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). CSF = cerebrospinal fluid, PBMC = peripheral blood mononuclear cells, DC = dendritic cells, MHCII = major histocompatibility complex II. IAV = influenza A virus-derived peptide pool. R10 = negative control, medium only. PHA / SEB = phytohemagglutinin-positive control - nonspecific TCR independent stimulation. [Figure 32B] This diagram shows a scheme for immunogenicity testing to assay samples obtained from cynomolgus monkeys before and after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). CSF = cerebrospinal fluid, PBMC = peripheral blood mononuclear cells, DC = dendritic cells, MHCII = major histocompatibility complex II. IAV = influenza A virus-derived peptide pool. R10 = negative control, medium only. PHA / SEB = phytohemagglutinin-positive control - nonspecific TCR independent stimulation. [Figure 32C] This diagram shows a scheme for immunogenicity testing to assay samples obtained from cynomolgus monkeys before and after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). CSF = cerebrospinal fluid, PBMC = peripheral blood mononuclear cells, DC = dendritic cells, MHCII = major histocompatibility complex II. IAV = influenza A virus-derived peptide pool. R10 = negative control, medium only. PHA / SEB = phytohemagglutinin-positive control - nonspecific TCR independent stimulation. [Figure 32D]This diagram shows a scheme for immunogenicity testing to assay samples obtained from cynomolgus monkeys before and after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). CSF = cerebrospinal fluid, PBMC = peripheral blood mononuclear cells, DC = dendritic cells, MHCII = major histocompatibility complex II. IAV = influenza A virus-derived peptide pool. R10 = negative control, medium only. PHA / SEB = phytohemagglutinin-positive control - nonspecific TCR independent stimulation. [Figure 33] This shows pre-existing and post-treatment (day 29) anti-APG05586mco nuclease antibodies measured in the serum of cynomolgus monkeys. No increase in serum antibodies against APG05586mco was observed in cynomolgus monkeys (N=6) after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123). [Figure 34] This shows the level of total anti-APG05586mco antibody in the cerebrospinal fluid of cynomolgus monkeys after administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). [Figure 35A]This study demonstrates dose-dependent distribution, nuclease transgene expression, and muHTT protein reduction in clinically relevant HD mouse models. Four weeks after intrastriatal administration of vehicle- or AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) to BACHD mice, striatal tissue was collected, and bulk lysate tissue samples were evaluated for AAV vector, nuclease transgene expression (mRNA and protein), and muHTT protein reduction. Each score represents the mean ± SE for 4–6 animals per dose evaluation. Vehicle-treated animals showed zero percentage reduction in muHTT protein, with lower LLOQ scores for both vector and transgene assays. [Figure 35B] This study demonstrates dose-dependent distribution, nuclease transgene expression, and muHTT protein reduction in clinically relevant HD mouse models. Four weeks after intrastriatal administration of vehicle- or AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) to BACHD mice, striatal tissue was collected, and bulk lysate tissue samples were evaluated for AAV vector, nuclease transgene expression (mRNA and protein), and muHTT protein reduction. Each score represents the mean ± SE for 4–6 animals per dose evaluation. Vehicle-treated animals showed zero percentage reduction in muHTT protein, with lower LLOQ scores for both vector and transgene assays. [Figure 35C]This study demonstrates dose-dependent distribution, nuclease transgene expression, and muHTT protein reduction in clinically relevant HD mouse models. Four weeks after intrastriatal administration of vehicle- or AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) to BACHD mice, striatal tissue was collected, and bulk lysate tissue samples were evaluated for AAV vector, nuclease transgene expression (mRNA and protein), and muHTT protein reduction. Each score represents the mean ± SE for 4–6 animals per dose evaluation. Vehicle-treated animals showed zero percentage reduction in muHTT protein, with lower LLOQ scores for both vector and transgene assays. [Figure 35D] This study demonstrates dose-dependent distribution, nuclease transgene expression, and muHTT protein reduction in clinically relevant HD mouse models. Four weeks after intrastriatal administration of vehicle- or AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp)(SEQ ID NO: 123) to BACHD mice, striatal tissue was collected, and bulk lysate tissue samples were evaluated for AAV vector, nuclease transgene expression (mRNA and protein), and muHTT protein reduction. Each score represents the mean ± SE for 4–6 animals per dose evaluation. Vehicle-treated animals showed zero percentage reduction in muHTT protein, with lower LLOQ scores for both vector and transgene assays. [Modes for carrying out the invention]

[0011] Many modifications and other embodiments of the invention described herein will be apparent to those skilled in the art, who have benefit from the teachings presented in the foregoing description and the accompanying drawings. Therefore, it should be understood that the invention is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be within the scope of the appended embodiments. Specific terms are used herein, but they are used in a general and descriptive sense only and not for limiting purposes.

[0012] I. Overview Huntington's disease (HD) is an autosomal dominant disorder that causes progressive degeneration of nerve tissue in the brain. Huntington's disease is the result of an expanded trinucleotide repeat in the HTT gene, where a significant increase in the three-nucleotide (cytosine-adenine-guanine, CAG) motif repeat in this gene produces a polyglutamine (poly-Q) tract in the mutant HTT protein that disrupts the function of the wild-type huntingtin protein.

[0013] The compositions and methods of this disclosure utilize single nucleotide polymorphisms (SNPs) in mutant huntingtin (mutHTT) alleles that generate protospacer flanking motifs enabling cleavage of the mutHTT allele by RNA-induced nucleases (RGNs). Although not constrained by theory, editing of the mutHTT allele can introduce premature stop codons that cause degradation by nonsense mutation-dependent degradation, thereby generating frameshift indels that knock out the full-length mutHTT gene. Since wild-type huntingtin has been shown to support important cellular and neuronal functions, selective strategies targeting only disease-associated mutant HTT are preferred and are being explored in preclinical and clinical settings (O'Regan et al., 2020, Sci.Rep.10:17269; Tabrizi et al., 2019, N Engl J Med 380:2307-2316). Accordingly, in some embodiments, the compositions and methods of this disclosure provide allelic-specific approaches that target only the mutHTT allele, rather than the wtHTT allele. The allele-specific approach described herein targets cells and patients heterozygous for a SNP, where the SNP allele, linked to a CAG expansion on the mutHTT allele, generates a PAM for the RGN. When the PAM-recognizing RGN is introduced into cells or patients along with a guide RNA targeting a sequence adjacent to the PAM, the mutHTT allele is cleaved at or near the SNP, an indel (insertion or deletion) is introduced, and mutHTT mRNA and protein levels are reduced. Due to heterozygosity of the cell or patient at the SNP, only the mutHTT allele is cleaved, only the mutHTT protein level is reduced, while the wtHTT allele and protein levels remain unchanged.

[0014] While other studies have demonstrated editing of mutHTT in HD models, none have taken a SNP-derived PAM-dependent approach that enables allele-specific reduction of mutHTT levels within exon 50. This disclosure is the first to provide a single AAV delivery construct containing both an RNA-induced nuclease and gRNA targeting exon 50 of the HTT mutant allele in a SNP-derived PAM-dependent approach, the construct being delivered in vivo to both the striatum and cortex, two regions known to be crucial for the pathogenesis of Huntington's disease. Importantly, this construct demonstrates allele-specific reduction of mutant HTT mRNA and protein in vitro and in vivo. In particular, the approach of this disclosure enables at least a 40% reduction in mutHTT mRNA and protein levels for at least 12 weeks post-treatment.

[0015] Accordingly, the compositions and methods disclosed herein can be used for the treatment of HD in subjects requiring it by reducing HTT levels, and in some embodiments, this reduction is allele-specific, with only mutHTT levels being reduced and wild-type HTT (wtHTT) expression remaining unchanged. In some embodiments, the compositions and methods disclosed herein can reduce mutHTT mRNA and protein levels by at least 40% in at least 50% of striatal neurons.

[0016] II. Huntintin (HTT) gene Huntington's disease (HD) is a genetic autosomal dominant disorder characterized by progressive degeneration of neurons in the brain, caused by an expansion of the CAG repeat in the first exon of the huntingtin gene on chromosome 4 (Huntington's Disease Collaborative Research Group, 1993, Cell 72:971-983). Disruption of the wild-type HTT protein by a polyglutamine-containing mutant HTT protein causes neuronal stress and dysfunction, ultimately leading to striatal atrophy, cognitive impairment, progressive mental and motor disorders (Kaemmerer et al, 2019, Degenerative Neurological and Neuromuscular Disease 9:3-17; Ross et al., 2014, Nat. Rev. Neurol. 10:204-216).

[0017] The huntingtin gene is large, spanning 180kb and consisting of 67 exons. A non-restrictive example of the HTT gene is the human HTT gene, shown as NCBI gene ID No. 3064, and a non-restrictive example of the HTT protein is the human huntingtin protein, shown as NCBI reference sequence ID No. NP_001375421.1, and shown herein as Sequence ID 85 (both incorporated herein by reference), which contains 21 glutamines within a polyQ tract and matches the GRCh38 reference genome.

[0018] The CAG triplet repeat region is located within exon 1 of the HTT gene. Individuals with more than 26 CAG repeats are likely to pass on the expanded CAG repeats to their offspring. HTT genes with 27-35 CAG repeats are considered to be intermediate alleles with approximately 0% probability of developing the disease phenotype, but individuals with these intermediate alleles may pass on the expanded repeats to their descendants. Individuals with 36-39 CAG repeats are likely to develop disease symptoms, but their alleles are thought to have incomplete penetrance. Huntington's disease patients have 40 or more CAG repeats and have approximately 100% probability of developing disease symptoms. These HD patients with 56 or more CAG repeats typically have an early onset of the disease in childhood or adolescence, which is classified as juvenile Huntington's disease or juvenile-onset Huntington's disease (JHD) (Tabrizi et al., 2022, Lancet Neurol 21:632-644). Therefore, in some embodiments, cells modified by the compositions and methods of this disclosure have an HTT gene having at least 27 CAG repeats in a CAG repeat region within exon 1, which is referred herein to as a mutant HTT gene or allele or mutHTT gene or allele. A CAG repeat region having at least 27 CAG repeats is also referred herein to as a CAG repeat expansion. The wild-type HTT gene or allele or wtHTT gene or allele has fewer than 27 CAG repeats in exon 1, typically 15 to 20 CAG repeats. Subjects treated by the compositions and methods of this disclosure have an HTT gene having at least 36 CAG repeats, and in some embodiments, at least 40 CAG repeats. Most HD patients are heterozygous for the expanded CAG repeats and therefore have one mutHTT allele having at least 36 or at least 40 CAG repeats and one wtHTT allele having fewer than 27 CAG repeats.

[0019] According to the present invention, a cell or subject further contains a single nucleotide polymorphism (SNP) allele within the mutHTT allele, which may be either a major allele (present in the majority of the human population) or a minor allele (present in a small number of populations). It should be noted that a particular genomic location may have minor alleles for multiple SNPs. As used herein, “single nucleotide polymorphism allele” or “SNP allele” refers to a single nucleotide difference between members of a population at a particular site in the genome, the difference being a single nucleotide substitution for another nucleotide. SNPs present on the mutHTT allele generate or are contained within a PAM that can be recognized by RGN. Thus, the SNP allele that generates a PAM is linked to an expansion of the CAG repeat of mutHTT, or in other words, the SNP allele resides on the same copy of the HTT gene as the expansion of the CAG repeat (mutHTT allele).

[0020] Non-limiting examples of SNPs that can be targeted for allele-specific approaches to treat HD are found in Tables 1 and 2 herein and include NCBI dbSNP No. rs362331. In some embodiments, the SNP allele generates a PAM having the nucleotide sequences NNNNCC, NNRYA, NNGRR, and / or NNGG. In some embodiments, a method for cleaving an intracellular mutant huntingtin (mutHTT) allele, wherein the mutHTT allele contains a first single nucleotide polymorphism (SNP) allele in exon 1, the first SNP allele generates a protospacer adjacent motif (PAM) selected from NNNNCC, NNRYA, NNGRR, and / or NNGG, the method comprising introducing an RNA-induced nuclease (RGN) or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA, the RGN being able to recognize the PAM and cleave the mutHTT allele. In some of these embodiments, the RGN has at least 80%, 85%, 90%, 95%, or greater sequence identity with any one of sequence numbers 7, 11, 13, and 15. The SNP allele may be located in exon 1 on either side of the CAG repeat. In these embodiments, the RGN that recognizes the SNP allele on one side (5' or 3') of the CAG repeat can be used in combination with another nuclease that cleaves the opposite end of the CAG repeat to generate an in-frame excision of the CAG repeat region from the mutHTT allele.

[0021] In some embodiments, the SNP allele that is linked to or “homophase” with the expansion of the CAG repeat of mutHTT is located in exon 50. The rs362331 SNP results in a nucleotide difference at position 151 in exon 50 of the HTT gene, as shown herein as SEQ ID NOs: 1 and 2. The major allele of the rs362331 SNP contains T or thymine at position 151 in exon 50 of the HTT gene (SEQ ID NO: 1), and its minor allele contains C or cytosine at that position (SEQ ID NO: 2). In genomes in which the HTT gene contains the “T” allele (major allele) of the rs362331 SNP, the presence of T at position 151 in exon 50 generates a PAM motif (NNRYA) of RGN APG05586 (shown as SEQ ID NO: 7) or its active variant or fragment. The RGN APG05586 PAM motif lies on the reverse complement of the DNA strand containing the T allele, and the "A" in NNRYA PAM is the complementary base of the T allele. Therefore, in individuals containing at least one "T" allele of the rs362331 SNP, APG05586 can bind to and cleave the "T" allele of the rs362331 SNP, along with a corresponding guide RNA complementary to the upstream (5') target sequence of the PAM motif.

[0022] Alternatively, in genomes containing the "C" allele (minor allele) of the rs362331 SNP, the presence of C at position 151 in exon 50 (SEQ ID NO: 2) generates the PAM motif of RGN APG07433.1 (shown as SEQ ID NO: 3, PAM of NNNNCC), the PAM motif of APG01604 (shown as SEQ ID NO: 11, PAM of NNGRR), and the PAM motif of LPG10145 (shown as SEQ ID NO: 15, PAM of NNGG), or any of these active variants or fragments. Thus, in individuals containing at least one "C" allele of the rs362331 SNP, APG07433.1, APG01604, or LPG10145 can bind to and cleave the "C" allele of the rs362331 SNP, along with a corresponding guide RNA complementary to the upstream (5') target sequence of the PAM motif.

[0023] By cleaving the mutHTT allele near the SNP site, a frameshift mutation can be generated within the HTT gene, resulting in early termination and reduced levels of mutHTT mRNA and protein compared to cells or subjects in the absence of RGN and its related guide RNA.

[0024] In most cases, subjects containing the mutHTT allele are heterozygous for that allele and include both the mutHTT allele and the wild-type HTT (wtHTT) allele. In some embodiments, the PAM produced by the SNP allele within mutHTT is not found in the wild-type HTT (wtHTT) allele, and therefore the method is allele-specific because the introduced or administered RGN, along with its corresponding guide RNA, cleaves and edits only the mutHTT allele and not the wtHTT allele, reducing only the levels of mutHTT mRNA and protein. An RGN polypeptide or RGN system that cannot cleave the wild-type HTT allele means that the RGN polypeptide or RGN system can either not cleave the wild-type HTT allele at all or only cleave it to a negligible level, resulting in no significant reduction in the levels of wtHTT mRNA and / or wtHTT protein. For example, if wtHTT can maintain support for important cellular and neuronal functions, and / or if symptoms of Huntington's disease are not present in an in vivo environment (i.e., in subjects heterozygous for the mutHTT allele and administered the RGN polypeptide or RGN system), then a non-significant reduction exists. In some embodiments, the RGN polypeptide or RGN system is cleaved at a negligible level such that the levels of wtHTT mRNA and / or wtHTT protein are reduced to ≤5%, ≤4%, ≤3%, ≤2%, ≤1%, ≤0.9%, ≤0.8%, ≤0.7%, ≤0.6%, ≤0.5%, ≤0.4%, ≤0.3%, ≤0.2%, or ≤0.1% compared to the levels of wtHTT mRNA and / or wtHTT protein in vitro or in vivo without the introduction of the RGN polypeptide or RGN system of this disclosure.

[0025] III. Guide RNA This disclosure provides a guide RNA and a polynucleotide encoding it, which target an RNA-induced nuclease (RGN) associated with a target nucleotide sequence within a mutant HTT allele. The term “guide RNA” includes a nucleotide sequence (e.g., a spacer) that is sufficiently complementary to the target nucleotide sequence within the mutant HTT allele, hybridizing with the target sequence and directing the sequence-specific binding of the associated RGN to the target nucleotide sequence. In some embodiments, if the target nucleotide sequence is double-stranded as DNA, the target nucleotide sequence includes a non-target strand (including a PAM sequence) and a target strand that hybridizes with the spacer of the guide RNA. In these embodiments, the guide RNA is sufficiently complementary to the target strand of the double-stranded target sequence (e.g., the target DNA sequence within the mutant HTT allele) so that the guide RNA hybridizes with the target strand and directs the sequence-specific binding of the associated RGN to the target sequence (e.g., the target DNA sequence within the mutant HTT allele). Therefore, in some embodiments, the guide RNA contains a spacer identical to the non-target strand sequence, except that uracil (U) is replaced with thymine (T) in the guide RNA.

[0026] Each guide RNA of an RGN is one or more RNA molecules (generally one or two) that can bind to the RGN and induce the RGN to bind to a specific target sequence, and in these embodiments where the RGN has nickase or nuclease activity, also cleaves the target and / or non-target strands. Generally, guide RNAs include CRISPR RNA (crRNA) and trans-activated CRISPR RNA (tracrRNA), although some RGNs do not require tracrRNA. Natural guide RNAs containing both crRNA and tracrRNA generally consist of two distinct RNA molecules that hybridize with each other via the repeat sequence of crRNA and the anti-repeat sequence of tracrRNA. In certain embodiments, crRNA and tracrRNA are linked together by a multinucleotide linker (e.g., a 4-nucleotide linker) to form a single guide RNA molecule, and crRNA and tracrRNA hybridize with each other via the repeat sequence of crRNA and the anti-repeat sequence of tracrRNA. Therefore, guide RNA encompasses a single guide RNA (sgRNA), and the crRNA and tracrRNA segments are located on the same RNA molecule or strand. Guide RNA may contain unnatural guide RNA not found in nature, or may be chemically modified, contain unnatural crRNA and / or tracrRNA molecules, and / or contain sequences not found in the corresponding natural molecules.

[0027] The present invention provides a CRISPR RNA (crRNA) or a polynucleotide encoding a CRISPR RNA that targets an RGN associated with a target sequence within a mutant HTT allele. As used herein, the term "crRNA" refers to an RNA molecule or a portion thereof that includes a spacer, which is a nucleotide sequence that directly hybridizes with the target strand of the target sequence, and a CRISPR repeat that includes a nucleotide sequence that is recognized by the RGN molecule and forms a structure either by itself or in cooperation with a hybridized tracrRNA. As used herein, the terms "tracrRNA" or "transactivating crRNA" refer to an RNA molecule that includes an anti-repeat sequence that is sufficiently complementary to hybridize with at least a portion of the CRISPR repeat of the crRNA to form a structure recognized by the RGN molecule. In some embodiments, additional secondary structures (e.g., stem-loops) within the tracrRNA molecule are required to bind to the RGN.

[0028] crRNA includes spacers and CRISPR repeats. The “spacer” has a nucleotide sequence that directly hybridizes with the target strand of the target sequence of interest (e.g., the target DNA sequence in a mutant HTT allele). The spacer is manipulated to have complete or partial complementarity with the target strand of the target sequence of interest (e.g., the target DNA sequence in a mutant HTT allele). In some embodiments, the spacer may consist of about 8 to about 30 nucleotides or more. For example, the spacer may have a nucleotide length of about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more. In some embodiments, the spacer has a nucleotide length of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more. In some embodiments, the spacer has a nucleotide length of approximately 10 to approximately 26, or approximately 12 to approximately 30. In some embodiments, the spacer has a nucleotide length of approximately 30. In some embodiments, the spacer has a nucleotide length of 30.In some embodiments, the degree of complementarity between the spacer and the target strand of the target sequence (e.g., the target DNA sequence within a mutant HTT allele) is approximately 50% or greater, approximately 60% or greater, approximately 70% or greater, approximately 75% or greater, approximately 80% or greater, approximately 81% or greater, approximately 82% or greater, approximately 83% or greater, approximately 84% or greater, approximately 85% or greater, and approximately 80% or greater when optimally aligned using a suitable alignment algorithm. 50% to 99% or more, including but not limited to 6% or more, approximately 87% or more, approximately 88% or more, approximately 89% or more, approximately 90% or more, approximately 91% or more, approximately 92% or more, approximately 93% or more, approximately 94% or more, approximately 95% or more, approximately 96% or more, approximately 97% or more, approximately 98% or more, approximately 99% or more, or more. In some embodiments, the degree of complementarity between the spacer and the target strand of the target sequence (e.g., the target DNA sequence within a mutant HTT allele) is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher, when optimally aligned using a suitable alignment algorithm. In some embodiments, the spacer may be sequence-identical to the non-target strand of the target sequence. In some of these embodiments where the target sequence is the target DNA sequence, the spacer may be sequence-identical to the non-target strand of the target DNA sequence, with the exception that thymine (T) in the non-target strand is replaced by uracil (U) in the spacer.In embodiments, the spacer does not include a secondary structure and can be predicted using any suitable polynucleotide folding algorithm known in the art, including but not limited to mFold (see, e.g., Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, e.g., Gruber et al. (2008) Cell 106(1):23-24).

[0029] In some embodiments, the spacer of the Disclosure has a nucleotide sequence represented as SEQ ID NO: 80, or a nucleotide sequence that differs from SEQ ID NO: 80 by one or two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 80 by two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 80 by one nucleotide. In some embodiments, the spacer has a nucleotide sequence represented as SEQ ID NO: 80. In some embodiments, the spacer of the Disclosure has a nucleotide sequence represented as SEQ ID NO: 81, or a nucleotide sequence that differs from SEQ ID NO: 81 by one or two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 81 by two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 81 by one nucleotide. In some embodiments, the spacer has a nucleotide sequence represented as SEQ ID NO: 81. In some embodiments, the spacer of the Disclosure has a nucleotide sequence represented as SEQ ID NO: 82, or a nucleotide sequence that differs from SEQ ID NO: 82 by one or two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 82 by two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 82 by one nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 82. In some embodiments, the spacer of the present disclosure has a nucleotide sequence shown as SEQ ID NO: 83, or a nucleotide sequence that differs from SEQ ID NO: 83 by one or two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 83 by two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 83 by one nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 83.In some embodiments, the spacer of the present disclosure has a nucleotide sequence represented as SEQ ID NO: 84, or a nucleotide sequence that differs from SEQ ID NO: 84 by one or two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 84 by two nucleotides. In some embodiments, the spacer has a nucleotide sequence that differs from SEQ ID NO: 84 by one nucleotide. In some embodiments, the spacer has a nucleotide sequence represented as SEQ ID NO: 84.

[0030] Along with the spacer, the crRNA further comprises a CRISPR RNA (crRNA) repeat. The CRISPR RNA repeat comprises a nucleotide sequence that forms a structure recognized by the RGN molecule, either by itself or in cooperation with hybridized tracrRNA. In some embodiments, the CRISPR RNA repeat may contain about 8 to about 30 nucleotides, or more. For example, the CRISPR repeat may have a nucleotide length of about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more. In some embodiments, the CRISPR repeat has a nucleotide length of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA antirepeat is approximately 50% or greater, approximately 60% or greater, approximately 70% or greater, approximately 75% or greater, approximately 80% or greater, approximately 81% or greater, approximately 82% or greater, approximately 83% or greater, approximately 84% or greater, approximately 85% or greater, approximately 86% or greater, approximately 87% or greater, approximately 88% or greater, approximately 89% or greater, approximately 90% or greater, approximately 91% or greater, approximately 92% or greater, approximately 93% or greater, approximately 94% or greater, approximately 95% or greater, approximately 96% or greater, approximately 97% or greater, approximately 98% or greater, approximately 99% or greater, or more, when optimally aligned using a preferred alignment algorithm.In certain embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA antirepeat is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher, when optimally aligned using a preferred alignment algorithm.

[0031] In some embodiments, the CRISPR repeat comprises one nucleotide sequence from SEQ ID NOs: 4, 8, 12, or 16, or an active variant or fragment thereof, which, when contained within a guide RNA, can direct the sequence-specific binding of the relevant RNA-induced nuclease provided herein to the target DNA sequence of the present disclosure within the mutant HTT allele. In some embodiments, the active CRISPR repeat variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with the nucleotide sequence represented as one of SEQ ID NOs: 4, 8, 12, 16, and 106. In some embodiments, the active CRISPR repeat fragment contains at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 consecutive nucleotides of the nucleotide sequence shown as one of SEQ ID NOs: 4, 8, 12, 16, and 106. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 4 by one or two nucleotides. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 4 by two nucleotides. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 4 by one nucleotide. In some embodiments, the CRISPR repeat contains a nucleotide sequence shown as SEQ ID NOs: 4. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 8 by one or two nucleotides. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 8 by two nucleotides. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NOs: 8 by one nucleotide. In some embodiments, the CRISPR repeat includes a nucleotide sequence shown as SEQ ID NO: 8.In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 12 by one or two nucleotides. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 12 by two nucleotides. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 12 by one nucleotide. In some embodiments, the CRISPR repeat includes the nucleotide sequence shown as SEQ ID NO: 12. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 16 by one or two nucleotides. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 16 by two nucleotides. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 16 by one nucleotide. In some embodiments, the CRISPR repeat includes the nucleotide sequence shown as SEQ ID NO: 16. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 106 by one or two nucleotides. In some embodiments, the CRISPR repeat includes a nucleotide sequence that differs from SEQ ID NO: 106 by two nucleotides. In some embodiments, the CRISPR repeat contains a nucleotide sequence that differs from SEQ ID NO: 106 by one nucleotide. In some embodiments, the CRISPR repeat contains a nucleotide sequence shown as SEQ ID NO: 106.

[0032] In some embodiments, the crRNA is an engineered sequence that does not exist in nature. In some embodiments, the specific CRISPR repeat is not ligated to a naturally engineered spacer, and the CRISPR repeat is considered heterogeneous to the spacer. In some embodiments, the spacer is an engineered sequence that does not exist in nature.

[0033] The guide RNAs of this disclosure include crRNA and trans-activated CRISPR RNA (tracrRNA), however, some compositions and methods of this disclosure utilize RGN polypeptides that do not require tracrRNA. The tracrRNA molecule comprises a nucleotide sequence containing a region having sufficient complementarity to hybridize to the CRISPR repeat of the crRNA, which is referred to herein as an anti-repeat. In some embodiments, the tracrRNA molecule further comprises a region having a secondary structure (e.g., a stem-loop) or forms a secondary structure when hybridized with the corresponding crRNA. In embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat is located at the 5' end of the molecule, and the 3' end of the tracrRNA contains the secondary structure. This region of the secondary structure generally contains several hairpin structures, including a nexus hairpin found adjacent to the anti-repeat. The nexus forms the core of the interaction between the guide RNA and the RGN and is located at the intersection between the guide RNA, the RGN, and the target sequence. Nexus hairpins often have conserved nucleotide sequences in the bases of their hairpin stems, and the motif UNANNC is found in many nexus hairpins in tracrRNA. In embodiments, the guide RNA or RGN system of the present disclosure uses tracrRNAs that contain non-canonical sequences in the bases of the hairpin stems of those nexus hairpins, including UNANNG and CNANNC. In some embodiments, the guide RNA or RGN system of the present disclosure uses tracrRNAs that contain non-canonical sequences of UNANNG or CNANNC in the bases of their nexus hairpin stems. Often, a terminal hairpin is present at the 3' end of the tracrRNA, which can vary in structure and number, but often contains a string of U at the 3' end following a GC-rich Rho-independent transcriptional terminator hairpin.For example, see Briner et al. (2014) Molecular Cell 56:333-339, Briner and Barrangou (2016) Cold Spring Harb Protoc, doi:10.1101 / pdb.top090902, and U.S. Publication No. 2017 / 0275648, each of which is incorporated herein by reference in its entirety.

[0034] In some embodiments, the antirepeat of tracrRNA, which is fully or partially complementary to the CRISPR repeat, contains about 8 to about 30 nucleotides, or more. For example, the base-pairing region between the tracrRNA antirepeat and the CRISPR repeat can have a nucleotide length of about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more. In some embodiments, the base-pairing region between the tracrRNA antirepeat and the CRISPR repeat has a nucleotide length of 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or more. In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA antirepeat is approximately 50% or greater, approximately 60% or greater, approximately 70% or greater, approximately 75% or greater, approximately 80% or greater, approximately 81% or greater, approximately 82% or greater, approximately 83% or greater, approximately 84% or greater, approximately 85% or greater, approximately 86% or greater, approximately 87% or greater, approximately 88% or greater, approximately 89% or greater, approximately 90% or greater, approximately 91% or greater, approximately 92% or greater, approximately 93% or greater, approximately 94% or greater, approximately 95% or greater, approximately 96% or greater, approximately 97% or greater, approximately 98% or greater, approximately 99% or greater, or more, when optimally aligned using a preferred alignment algorithm.In some embodiments, the degree of complementarity between a CRISPR repeat and its corresponding tracrRNA antirepeat is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or higher, when optimally aligned using a preferred alignment algorithm.

[0035] In some embodiments, the entire tracrRNA can contain approximately 60 to more than 210 nucleotides. For example, the tracrRNA can have nucleotide lengths of approximately 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210, or more. In some embodiments, the tracrRNA has a nucleotide length of 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more. In some embodiments, the tracrRNA has a nucleotide length of approximately 70 to 105 nucleotides, including approximately 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, and 105 nucleotide lengths. In this embodiment, the tracrRNA has a nucleotide length of 70 to 105 nucleotides, including lengths of 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, and 105 nucleotides.

[0036] In some embodiments, the tracrRNA comprises one of the nucleotide sequences of SEQ ID NOs. 5, 9, 13, 17, 107, and 120, or an active variant or fragment thereof, which, when contained within a guide RNA, can direct the sequence-specific binding of the relevant RNA-induced nuclease provided herein to a target sequence within a mutant HTT allele. In some embodiments, the active tracrRNA sequence variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with one of the nucleotide sequences shown as SEQ ID NOs. 5, 9, 13, 17, 107, and 120. In some embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides from any one of the nucleotide sequences shown as SEQ ID NOs. 5, 9, 13, 17, 107, and 120. In some embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides from the nucleotide sequence shown as SEQ ID NO. 5. In a particular embodiment, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides from the nucleotide sequence shown as SEQ ID NO. 9. In certain embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 13. In certain embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 17.In certain embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 107. In certain embodiments, the active tracrRNA sequence fragment contains at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 120. In some embodiments, the active tracrRNA sequence fragment contains the nucleotide sequence shown as SEQ ID NO: 5. In some embodiments, the active tracrRNA sequence fragment contains the nucleotide sequence shown as SEQ ID NO: 9. In some embodiments, the active tracrRNA sequence fragment contains the nucleotide sequence shown as SEQ ID NO: 13. In some embodiments, the active tracrRNA sequence fragment contains the nucleotide sequence shown as SEQ ID NO: 17. In some embodiments, the active tracrRNA sequence fragment includes a nucleotide sequence shown as SEQ ID NO: 107. In some embodiments, the active tracrRNA sequence fragment includes a nucleotide sequence shown as SEQ ID NO: 120.

[0037] Two polynucleotide sequences can be considered substantially complementary when they hybridize to each other under stringent conditions. Similarly, if a guide RNA bound to an RGN binds to a target sequence under stringent conditions, the RGN is considered to bind to a specific target sequence in a sequence-specific manner. "Stringent conditions" or "stringent hybridization conditions" refer to conditions under which two polynucleotide sequences hybridize to each other to a detectably higher degree than other sequences (e.g., at least twice the background). Stringent conditions are sequence-dependent and will vary in different contexts. Typically, stringent conditions would be salt concentrations of less than approximately 1.5 M Na+ ions at pH 7.0–8.3, typically about 0.01–1.0 M Na+ ion concentration (or other salts), and temperatures of at least approximately 30°C for short sequences (e.g., 10–50 nucleotides) and at least approximately 60°C for long sequences (e.g., over 50 nucleotides). Stringent conditions can also be achieved by adding destabilizers such as formamide. Exemplary low-stringency conditions include hybridization with a buffer of 30–35% formamide, 1M NaCl, and 1% SDS (sodium dodecyl sulfate) at 37°C, and washing with 1–2 times SSC (20 times SSC = 3.0M NaCl / 0.3M trisodium citrate) at 50–55°C. Exemplary moderate-stringency conditions include hybridization with 40–45% formamide, 1.0M NaCl, and 1% SDS at 37°C, and washing with 0.5–1 times SSC at 55–60°C. Exemplary high-stringency conditions include hybridization with 50% formamide, 1M NaCl, and 1% SDS at 37°C, and washing with 0.1 times SSC at 60–65°C. Optionally, the washing buffer may contain approximately 0.1% to 1% SDS. The duration of hybridization is generally less than approximately 24 hours, typically about 4 to 12 hours. The washing duration should be long enough to reach equilibrium.

[0038] Tm is the temperature at which 50% of the complementary target sequence hybridizes to a perfectly matched sequence (under defined ionic strength and pH). For DNA-DNA hybrids, Tm can be approximated by the equation from Meinkoth and Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6(log M) + 0.41(%GC) - 0.61(%form) - 500 / L (wherein M is the molar concentration of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, %form is the percentage of formamide in the hybridization solution, and L is the length of the hybrid in the base pair). Generally, stringent conditions are selected so that, at defined ionic strength and pH, they are about 5°C lower than the thermal melting point (Tm) of a particular sequence and its complementary strand. However, strictly stringent conditions can utilize hybridization and / or washing at temperatures 1, 2, 3, or 4°C below the melting point (Tm). Moderate stringent conditions can utilize hybridization and / or washing at temperatures 6, 7, 8, 9, or 10°C below the melting point (Tm), and low stringency conditions can utilize hybridization and / or washing at temperatures 11, 12, 13, 14, 15, or 20°C below the melting point (Tm). Using this formula, hybridization and washing compositions, and desired Tm, those skilled in the art will understand that variations in the stringency of hybridization and / or washing solutions are essentially described.Extensive guides on nucleic acid hybridization can be found in Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2 (Elsevier, New York) and Ausubel et al., eds. (1995) Current Protocols in Molecular Biology, Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See also Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual (2nd ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0039] The term "sequence-specific" can also refer to the binding of RGN polypeptides to a target sequence with higher affinity than binding to a randomized background sequence.

[0040] Guide RNA can be either single guide RNA (sgRNA) or dual guide RNA (dgRNA). Single guide RNA includes crRNA and tracrRNA on a single RNA molecule, while a dual guide RNA system includes crRNA and tracrRNA residing on two different RNA molecules that are hybridized to each other via at least a portion of the CRISPR repeats of crRNA and at least a portion of the tracrRNA (i.e., an anti-repeat) that may be fully or partially complementary to the CRISPR repeats of crRNA. In embodiments where the guide RNA is single guide RNA, crRNA and tracrRNA are separated by a linker nucleotide sequence. The crRNA repeats and tracrRNA linked by the nucleotide linker can be referred to as the sgRNA backbone. The backbone can also refer to the crRNA repeats and tracrRNA of the dgRNA.

[0041] The guide RNA backbone may include one of the nucleotide sequences of SEQ ID NOs: 140, 141, and 142, or an active variant or fragment thereof, which, when contained within the guide RNA, can direct the sequence-specific binding of the relevant RNA-induced nuclease provided herein to a target sequence within the mutant HTT allele. The sgRNA or dgRNA backbone may be manipulated to be shorter or longer than its native length while still retaining its function. In some embodiments, the manipulated sgRNA or dgRNA backbone is about 2 to 30 nucleotides shorter than the same backbone before manipulation. In some embodiments, the manipulated sgRNA or dgRNA backbone is about 2 nucleotides shorter, about 4 nucleotides shorter, about 6 nucleotides shorter, about 8 nucleotides shorter, about 10 nucleotides shorter, about 12 nucleotides shorter, about 14 nucleotides shorter, about 16 nucleotides shorter, about 18 nucleotides shorter, about 20 nucleotides shorter, about 22 nucleotides shorter, about 24 nucleotides shorter, about 26 nucleotides shorter, about 28 nucleotides shorter, about 30 nucleotides shorter, or more nucleotides shorter than the same backbone before manipulation. In some embodiments, the backbone of the manipulated sgRNA or dgRNA is about 2 to 18 nucleotides shorter than the same backbone before manipulation. In some embodiments, the backbone of the manipulated sgRNA or dgRNA is about 2 nucleotides shorter, about 4 nucleotides shorter, about 6 nucleotides shorter, about 8 nucleotides shorter, about 10 nucleotides shorter, about 12 nucleotides shorter, about 14 nucleotides shorter, about 16 nucleotides shorter, or about 18 nucleotides shorter than the same backbone before manipulation. In some embodiments, the backbone of the manipulated sgRNA or dgRNA is about 14 nucleotides shorter than the same backbone before manipulation. In some embodiments, the backbone of the manipulated sgRNA or dgRNA is about 16 nucleotides shorter than the same backbone before manipulation. In some embodiments, the backbone of the manipulated sgRNA or dgRNA is about 20 nucleotides shorter than the same backbone before manipulation.

[0042] In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain approximately 60 to more than approximately 120 nucleotides. For example, the backbone may contain approximately 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, and approximately The nucleotide length can be 94, approximately 95, approximately 96, approximately 97, approximately 98, approximately 99, approximately 100, approximately 101, approximately 102, approximately 103, approximately 104, approximately 105, approximately 106, approximately 107, approximately 108, approximately 109, approximately 110, approximately 111, approximately 112, approximately 113, approximately 114, approximately 115, approximately 116, approximately 117, approximately 118, approximately 119, approximately 120, or more. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain a nucleotide length of approximately 66 to more than approximately 110. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain a nucleotide length of approximately 66. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain a nucleotide length of approximately 70. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain about 76 nucleotides in length. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain about 90 nucleotides in length. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain about 94 nucleotides in length. In some embodiments, the active skeletal variant of the guide RNA of this disclosure may contain about 110 nucleotides in length.

[0043] The active scaffold fragment of the guide RNA of this disclosure is at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 of any one of the nucleotide sequences shown as SEQ ID NOs: 140, 141, or 142. It may contain 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, or more consecutive nucleotides. In some embodiments, the active backbone fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 140. In some embodiments, the active scaffold fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more consecutive nucleotides of the nucleotide sequence shown as SEQ ID NO: 141.In some embodiments, the active scaffold fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68 of the nucleotide sequence shown as SEQ ID NO: 142, Contains 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, or more consecutive nucleotides.

[0044] The active skeletal variants of the sgRNAs of this disclosure may have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with any one of SEQ ID NOs. In some embodiments, the active skeletal variants of the sgRNAs of this disclosure have a nucleotide sequence having at least 80% sequence identity with any one of SEQ ID NOs. In some embodiments, the active skeletal variants of the sgRNAs of this disclosure have a nucleotide sequence having at least 85% sequence identity with any one of SEQ ID NOs. In some embodiments, the active skeletal variants of the sgRNAs of this disclosure have a nucleotide sequence having at least 90% sequence identity with any one of SEQ ID NOs. In some embodiments, the active scaffold variant of the sgRNA of the Disclosure has a nucleotide sequence having at least 95% sequence identity with any one of SEQ ID NOs: 140-142. In some embodiments, the active scaffold variant of the sgRNA of the Disclosure has a nucleotide sequence shown as any one of SEQ ID NOs: 140-142.

[0045] Generally, the linker nucleotide sequence connecting crRNA and tracrRNA is a sequence that does not contain complementary bases to avoid the formation of a secondary structure containing or having nucleotides within the linker nucleotide sequence. In some embodiments, the linker nucleotide sequence between crRNA and tracrRNA has a nucleotide length of at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, or more. In some embodiments, the linker nucleotide sequence of a single guide RNA has a nucleotide length of at least 4. In some particular embodiments, the linker nucleotide sequence of a single guide RNA has a nucleotide length of 4. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence represented as AAAG, GAAA, ACUU, and CAAAGG. In some particular embodiments, the linker nucleotide sequence includes a nucleotide sequence represented as AAAG. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence represented as GAAA. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence represented as ACUU. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence shown as CAAAGG.

[0046] In some embodiments, the sgRNA has a nucleotide sequence shown as one of sequence numbers 6, 10, 14, 18, and 25-29.

[0047] Single or dual guide RNAs can be synthesized chemically or via in vitro transcription. Assays for determining sequence-specific binding between RGNs and guide RNAs are known in the art and include, but are not limited to, in vitro binding assays between expressed RGNs and guide RNAs, which can be used in pull-down detection assays in which the guide RNA:RGN complex is tagged with a detectable label (e.g., biotin) and captured via a detectable label (e.g., streptavidin beads). A control guide RNA having a sequence or structure unrelated to the guide RNA can be used as a negative control for non-specific binding of RGNs to RNA.

[0048] In some embodiments, the guide RNA can be introduced into target cells as an RNA molecule. The guide RNA can be transcribed in vitro or chemically synthesized. In some embodiments, a nucleic acid molecule encoding the guide RNA is introduced into the target cells. In some embodiments, the nucleic acid molecule encoding the guide RNA is operably ligated to a promoter (e.g., an RNA polymerase III promoter). The promoter can be a native promoter or heterologous to the guide RNA-coding nucleic acid molecule.

[0049] In some embodiments, the guide RNA can be introduced into target cells as a ribonucleoprotein complex, as described herein, and the guide RNA binds to an RGN polypeptide.

[0050] Guide RNA directs the relevant RGN to the desired specific target nucleotide sequence through hybridization of the guide RNA to the target sequence of interest. The target sequence can be bound (or cleaved in some embodiments) by an RNA-inducible nuclease in vitro or intracellularly. The target sequence can include DNA, RNA, or a combination of both, and can be single-stranded or double-stranded. In some embodiments, the target sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, episomal DNA, or RNA molecules (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). In those embodiments where the target sequence is a chromosomal sequence, the chromosomal sequence can be a nuclear or mitochondrial chromosome sequence. In the compositions and methods of this disclosure, the target sequence is located within a double-stranded target nucleic acid molecule (e.g., a target DNA sequence). More specifically, the target sequence is located within a mutant HTT allele. In some embodiments, the target sequence is unique to the target genome. In some embodiments, the target sequence comprises a target chain and a non-target chain, and the target sequence has a nucleotide sequence shown as one of sequence numbers 75-79 and 130.

[0051] The target sequence is adjacent to a protospacer facilitation motif (PAM), and the non-target strand of the target sequence is the strand containing the PAM. The PAM is directly adjacent to the target sequence and often contains N, where N represents any nucleotide. In some embodiments, the PAM contains about 1 to about 10 N (including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 N). In certain embodiments, the PAM contains 1 to 10 N, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 N. The PAM can be the 5' or 3' of the target sequence on its non-target strand. In some embodiments, the PAM is the 3' of the target sequence on its non-target strand for the guide RNA and RGN systems of this disclosure. Generally, PAM is a consensus sequence of about 3 to 4 nucleotides, but in certain embodiments, it can be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides long.

[0052] In some embodiments, the PAM sequence adjacent to the target sequence of the present disclosure on the non-target strand includes a consensus sequence represented as one of NNNNCC, NNRYA, NNGRR, and NNGG. In some embodiments, the PAM sequence adjacent to the target sequence on the non-target strand includes a consensus sequence represented as NNNNCC. In some embodiments, the PAM sequence adjacent to the target sequence on the non-target strand includes a consensus sequence represented as NNRYA. In some embodiments, the PAM sequence adjacent to the target sequence on the non-target strand includes a consensus sequence represented as NNGRR. In some embodiments, the PAM sequence adjacent to the target sequence on the non-target strand includes a consensus sequence represented as NNGG. In some embodiments, the PAM sequence is 3' of the target sequence on the non-target strand.

[0053] It is well known in the art that the specificity of a PAM sequence to a given nuclease enzyme is influenced by the promoter used to express the RGN, or by the enzyme concentration which can be modified by altering the amount of ribonucleoprotein complex delivered to the cell (see, for example, Karvelis et al. (2015) Genome Biol 16:253).

[0054] Upon recognizing a corresponding PAM sequence, the RGN can cleave one or both strands of the target sequence at a specific cleavage site. As used herein, the cleavage site consists of two specific nucleotides in the target sequence, and the target and / or non-target strands of the target sequence are cleaved by the RGN. The cleavage site can include the first and second, second and third, third and fourth, fourth and fifth, fifth and sixth, seventh and eighth, or eighth and ninth nucleotides from the PAM in either the 5' or 3' direction. In some embodiments, the cleavage site can be 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or more than 20 nucleotides from the PAM in either the 5' or 3' direction. Because RGNs can cleave target sequences and shift their ends, in certain embodiments, cleavage sites are defined based on the distance of two nucleotides from the PAM on the non-target strand of the target sequence, and, for the target strand, the distance of two nucleotides from the complementary strand of the PAM.

[0055] IV. RNA-induced nucleases and other nucleases In some embodiments of a method for treating Huntington's disease by cleaving the mutHTT allele, RGN is a type II CRISPR-Cas polypeptide. In some embodiments, RGN is a type V CRISPR-Cas polypeptide. In some embodiments, RGN is a Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Casl2a, Casl2b, Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, SpCas9-NG, LbCasl2a, AsCasl2a, Cas9-KKH, cyclic substitution Cas9, Argonaut (Ago), SmacCas9, or Spy-macCas9 domain.

[0056] RNA-induced nuclease systems, including the guide RNA of this disclosure, are provided herein. The term RNA-induced nuclease (RGN) refers to a polypeptide that is sequence-directed to a specific target sequence (e.g., a target DNA sequence in a mutant HTT allele) by binding to a guide RNA molecule that hybridizes with the target strand of the target sequence (e.g., a target DNA sequence in a mutant HTT allele). Active fragments or variants of naturally occurring RGNs maintain binding to the target nucleotide sequence in an RNA-induced sequence-specific manner. Cleavage of the target strand of the target sequence by an RGN can result in single-strand or double-strand breaks, but generally produces double-strand breaks.

[0057] The RGN system of this disclosure includes an RGN that binds to a target sequence disclosed herein. In some embodiments, the RGN recognizes a consensus nucleotide sequence on its non-target strand (wherein N is A, C, T, or G, R is G or A, and Y is C or T) that includes NNNNCC, NNRYA, NNGRR, and NNGG at 3' of the target sequence, as well as a PAM having its active fragment or variant. In some embodiments, the RGN recognizes a consensus nucleotide sequence on its non-target strand (wherein N is A, C, T, or G) that includes NNNNCC at 3' of the target sequence, as well as a PAM having its active fragment or variant. In some embodiments, the RGN recognizes a consensus nucleotide sequence on its non-target strand (wherein N is A, C, T, or G, R is G or A, and Y is C or T) that includes NNRYA at 3' of the target sequence, as well as a PAM having its active fragment or variant. In some embodiments, the RGN recognizes a consensus nucleotide sequence containing NNGRR at 3' of the target sequence on its non-target strand (where N is A, C, T, or G, and R is G or A), as well as a PAM having an active fragment or variant thereof. In some embodiments, the RGN recognizes a consensus nucleotide sequence containing NNGG at 3' of the target sequence on its non-target strand (where N is A, C, T, or G), as well as a PAM having an active fragment or variant thereof. In some embodiments, the active fragment or variant of the RGN that recognizes such a PAM sequence can bind, and in some embodiments, can cleave or nick the target sequence.

[0058] The RGN polypeptides of this disclosure may include a linker domain 1 (L1), a linker domain 2 (L2), a wedge (WED) domain, a RuvC nuclease domain, an HNH nuclease domain, a bridge helix (BH) domain, a Rec domain, or a PAM interaction (PI) domain. In some embodiments, the RuvC domain is a RuvCIII domain. The Rec or recognition lobe mediates nucleic acid binding via multiple Rec domains (e.g., Rec1-3) by sensing nucleic acids, modulates HNH conformational transitions, and locks the catalytic HNH domain at the cleavage site. The wedge domain is involved in the recognition of the guide RNA scaffold. The arginine-rich bridge helix (BH) domain connects the nuclease lobe and the recognition lobe.

[0059] Non-limiting examples of domains within the APG07433.1 RGN polypeptide, shown as Sequence ID No. 3, include, all with reference to Sequence ID No. 3, RuvC-I from amino acid residues 1-54, BH from amino acid residues 55-83, REC1 from amino acid residues 84-244, REC2 from amino acid residues 245-462, RuvC-II from amino acid residues 463-521, L1 from amino acid residues 522-552, HNH from amino acid residues 553-672, L2 from amino acid residues 673-685, RuvC-III from amino acid residues 686-833, WED from amino acid residues 834-938, and PI from amino acid residues 939-1071.

[0060] Non-limiting examples of domains within the APG05586 RGN polypeptide, shown as Sequence ID No. 7, include the following domains: RuvC-I from amino acid residues 1-33, BH from amino acid residues 34-71, REC1 from amino acid residues 72-232, REC2 from amino acid residues 233-468, RuvC-II from amino acid residues 469-517, L1 from amino acid residues 518-552, HNH from amino acid residues 553-672, L2 from amino acid residues 673-687, RuvC-III from amino acid residues 688-837, WED from amino acid residues 838-998, and PI from amino acid residues 999-1150, all of which refer to Sequence ID No. 7.

[0061] Non-exclusive examples of domains within the APG01604 RGN polypeptide, as shown in Sequence ID No. 11, include the following domains: RuvC-I from amino acid residues 1-40, BH from amino acid residues 41-74, REC1 from amino acid residues 75-223, REC2 from amino acid residues 224-430, RuvC-II from amino acid residues 431-483, L1 from amino acid residues 484-516, HNH from amino acid residues 517-631, L2 from amino acid residues 632-651, RuvC-III from amino acid residues 652-775, WED from amino acid residues 776-909, and PI from amino acid residues 910-1052, all of which refer to Sequence ID No. 11.

[0062] Non-exclusive examples of domains within the LPG10145 RGN polypeptide, as shown in Sequence ID No. 15, include the following domains: RuvC-I from amino acid residues 1-42, BH from amino acid residues 43-79, REC1 from amino acid residues 80-236, REC2 from amino acid residues 237-476, RuvC-II from amino acid residues 477-524, L1 from amino acid residues 525-560, HNH from amino acid residues 561-676, L2 from amino acid residues 677-690, RuvC-III from amino acid residues 691-828, WED from amino acid residues 829-976, and PI from amino acid residues 977-1130, all of which refer to Sequence ID No. 15.

[0063] The RGN systems of this disclosure may include RGNs that include a PAM interaction domain that contributes to the recognition and binding of a PAM site. In certain embodiments, the PAM interaction domain of the RGN has a sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a PAM interaction domain having an amino acid sequence that has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, and recognizes the PAM sequence NNNNCC.

[0064] In some embodiments, the PAM interaction domain of the RGN has a sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA.

[0065] In some embodiments, the PAM interaction domain of the RGN has a sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0066] In some embodiments, the PAM interaction domain of the RGN has a sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. The PAM interaction domains of APG07433.1 nuclease (shown as SEQ ID NO: 3), APG05586 nuclease (shown as SEQ ID NO: 7), APG01604 nuclease (shown as SEQ ID NO: 11), and LPG10145 nuclease (shown as SEQ ID NO: 15) were determined by aligning the nuclease sequences to known RNA-induced nucleases with elucidated structures, including Staphylococcus aureus (PDB: 5CZZ-Chain-A), Neisseria menigitidis 1 (PDB: 6JDV_1|Chain), and Streptococcus thermophilus (6M0W_4|Chain), and identifying regions at similar positions in the protein alignment.

[0067] The RGN systems of this disclosure may include RGN polypeptides each comprising at least one nuclease domain, each involved in cleaving a single strand of a nucleic acid molecule. The nuclease domain may include a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0068] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence represented as one of SEQ ID NOs: 147, 148, 149, and 150.

[0069] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 151, 152, 153, and 154.

[0070] An RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 155, 156, 157, and 158.

[0071] The RGN system of this disclosure may include an RGN polypeptide comprising a PAM interaction domain that contributes to the recognition and binding of a PAM site, and further comprising at least one nuclease domain, each of which is involved in the cleavage of a single strand of nucleic acid molecule. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNNNCC and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, recognizing the PAM sequence NNNNCC, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0072] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNRYA and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, recognizing the PAM sequence NNRYA, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150.

[0073] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNGRR and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, recognizing the PAM sequence NNGRR, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154.

[0074] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136, recognizes the PAM sequence NNGG, and may further include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, recognizing the PAM sequence NNGG, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158.

[0075] In some embodiments, the compositions and methods of the present disclosure use RGNs, or active variants or fragments thereof, that can bind to a target sequence adjacent to (i.e., capable of recognizing) a PAM consensus sequence, represented as one of NNNNCC, NNRYA, NNGRR, and NNGG. In some embodiments, the PAM sequence is the 3' of the target sequence on its non-target strand. In some embodiments, the RGN binds to a guide RNA having a sequence represented as one of SEQ ID NOs: 6, 10, 14, 18, and 25-29. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence represented as SEQ ID NO: 3 can recognize the PAM sequence of NNNNCC and binds to a guide RNA containing a CRISPR repeat, or an active variant or fragment thereof, represented as SEQ ID NO: 4, and a tracrRNA, or an active variant or fragment thereof, represented as SEQ ID NO: 5. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 7 can recognize the PAM sequence of NNRYA and bind to a guide RNA containing a CRISPR repeat, or its active variant or fragment, shown as SEQ ID NO: 8 or 106, and a tracrRNA, or its active variant or fragment, shown as SEQ ID NO: 9 or 107. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 11 can recognize the PAM sequence of NNGRR and bind to a guide RNA containing a CRISPR repeat, or its active variant or fragment, shown as SEQ ID NO: 12, and a tracrRNA, or its active variant or fragment, shown as SEQ ID NO: 13 or 120.In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 15 can recognize the PAM sequence of NNGG and bind to a guide RNA containing the CRISPR repeat shown as SEQ ID NO: 16, or its active variant or fragment, and the tracrRNA shown as SEQ ID NO: 17, or its active variant or fragment.

[0076] Non-limiting examples of RGNs useful in the methods and compositions of this disclosure include APG07433.1, APG05586, APG01604, and LPG10145 RNA-inducible nucleases (their amino acid sequences are shown as SEQ ID NOs: 3, 7, 11, and 15, respectively), as well as active fragments or variants thereof that retain the ability to bind to target sequences in an RNA-inducible sequence-specific manner. In some embodiments, active variants of the RGNs disclosed herein include amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% or more sequence identity with the amino acid sequence shown as SEQ ID NO: 3. In some embodiments, the active variants of RGN disclosed herein include an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with the amino acid sequence shown as SEQ ID NO: 7. In some embodiments, the active variants of RGN disclosed herein include an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% or more sequence identity with the amino acid sequence shown as SEQ ID NO: 11. In some embodiments, the active variants of RGN disclosed herein include amino acid sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the amino acid sequence shown as SEQ ID NO: 15.In some embodiments, the active fragment of APG07433.1RGN contains at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or 1050 or more consecutive amino acid residues of the amino acid sequence shown as SEQ ID NO: 3. In some embodiments, the active fragment of APG05586RGN contains at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or 1050 or more consecutive amino acid residues of the amino acid sequence shown as SEQ ID NO: 7. In some embodiments, the active fragment of APG01604RGN contains at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or 1050 consecutive amino acid residues of the amino acid sequence shown as SEQ ID NO: 11. In some embodiments, the active fragment of LPG10145RGN contains at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, or 1050 consecutive amino acid residues of the amino acid sequence shown as SEQ ID NO: 15.

[0077] The compositions and methods of the present disclosure may include an RGN capable of binding to a target sequence of the present disclosure, or an RGN having an amino acid sequence shown as SEQ ID NO: 3, or an active variant or fragment thereof, wherein the RGN can bind to a target sequence adjacent to a PAM consensus sequence shown as NNNNCC. In some embodiments, the PAM sequence is the 3' of the target sequence on its non-target strand. In some embodiments, the RGN binds to a guide RNA having a sequence shown as SEQ ID NO: 27 or 28. In some embodiments, the RGN binds to the guide RNA and includes a CRISPR repeat shown as SEQ ID NO: 4, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 5, or an active variant or fragment thereof.

[0078] The compositions and methods of the present disclosure may include an RGN capable of binding to a target sequence of the present disclosure, or an RGN having an amino acid sequence shown as SEQ ID NO: 7, or an active variant or fragment thereof, wherein the RGN can bind to a target sequence adjacent to a PAM consensus sequence shown as NNRYA. In some embodiments, the PAM sequence is 3' of the target sequence on its non-target strand. In some embodiments, the RGN binds to a guide RNA having a sequence shown as SEQ ID NO: 25 or 26. In some embodiments, the RGN binds to a guide RNA comprising a CRISPR repeat shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof.

[0079] The compositions and methods of the present disclosure may include RGNs capable of binding to target sequences of the present disclosure, or RGNs having an amino acid sequence shown as SEQ ID NO: 11, or an active variant or fragment thereof, wherein the RGNs may bind to target sequences adjacent to a PAM consensus sequence shown as NNGRR. In some embodiments, the PAM sequence is 3' of the target sequence on its non-target strand. In some embodiments, the RGNs bind to guide RNA having a sequence shown as SEQ ID NO: 29. In some embodiments, the RGNs bind to guide RNA comprising a CRISPR repeat shown as SEQ ID NO: 12, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof.

[0080] The compositions and methods of the present disclosure may include RGNs capable of binding to target sequences of the present disclosure, or RGNs having an amino acid sequence shown as SEQ ID NO: 15, or an active variant or fragment thereof, wherein the RGNs may bind to target sequences adjacent to a PAM consensus sequence shown as NNGG. In some embodiments, the PAM sequence is 3' of the target sequence on its non-target strand. In some embodiments, the RGNs bind to guide RNA having a sequence shown as SEQ ID NO: 18. In some embodiments, the RGNs bind to guide RNA comprising a CRISPR repeat shown as SEQ ID NO: 16, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 17, or an active variant or fragment thereof.

[0081] In some embodiments, nucleases other than RGNs are used in the compositions and methods of the present disclosure. These nucleases bind to the opposite end or inside of the expanded trinucleotide repeat of the HTT gene from the target sequence of the present disclosure. As used herein, the term “nuclease” refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides within a nucleic acid molecule. Generally, a nuclease is an endonuclease that can cleave phosphodiester bonds between nucleotides within a nucleic acid molecule. In some embodiments, the sequence-specific nuclease is selected from the group consisting of meganucleases, zinc finger nucleases, TAL effector DNA-binding domain nuclease fusion proteins (TALENs), and RNA-inducible nucleases (RGNs) or variants thereof, and the nuclease activity is reduced or inhibited.

[0082] As used herein, the terms “meganucleases” or “homing endonucleases” refer to endonucleases that bind to recognition sites in double-stranded DNA of a length of 12–40 bp. Non-exclusive examples of meganucleases belong to the LAGLIDADG family, which includes the conserved amino acid motif LAGLIDADG (SEQ ID NO: 139). The term “meganucleases” may refer to dimeric or single-stranded meganucleases.

[0083] As used herein, the terms "zinc finger nuclease" or "ZFN" refer to a chimeric protein comprising a zinc finger DNA-binding domain and a nuclease domain.

[0084] As used herein, the terms "TAL effector DNA-binding domain nuclease fusion protein" or "TALEN" refer to a chimeric protein comprising a TAL effector DNA-binding domain and a nuclease domain.

[0085] According to the present invention, the target sequence of this disclosure within a mutant HTT allele is bound by an RGN. The target strand of the target sequence hybridizes with a guide RNA associated with the RGN. If the polypeptide then has nuclease activity, the target strand and / or non-target strand of the target sequence (e.g., the target DNA sequence) can subsequently be cleaved by the RGN. The terms “cleaving” or “cutting” refer to the hydrolysis of at least one phosphodiester bond in the backbone of one or both strands of the double-stranded target sequence (target DNA sequence), which may result in either a single-strand break or a double-strand break in the target DNA sequence. The cleavage of the target sequence of this disclosure may result in a shifted break or a blunt end.

[0086] The compositions and methods of this disclosure may utilize RGN or other nucleases containing at least one nuclear localization signal (NLS) to enhance the transport of RGN to the cell nucleus. Nuclear localization signals are known in the art and generally contain stretches of basic amino acids (see, e.g., Lange et al., J. Biol. Chem. (2007) 282:5101-5105). In some embodiments, the RGN contains two, three, four, five, or six or more nuclear localization signals. The nuclear localization signals may be heterologous NLS. Non-limiting examples of nuclear localization signals useful for the RGN of this disclosure are the nuclear localization signals of SV40 large T antigen, nucleoplasmin, and c-Myc (see, e.g., Ray et al. (2015) Bioconjug Chem 26(6):1004-7). In some embodiments, the RGN includes an NLS sequence, indicated as SEQ ID NO: 86, 87, or 125. The RGN or other nuclease may contain one or more NLS sequences at its N-terminus, C-terminus, or both. For example, the RGN may contain two NLS sequences in its N-terminal region and four NLS sequences in its C-terminal region. In some embodiments, the RGN or other nuclease includes an SV40 NLS (such as the sequence indicated as SEQ ID NO: 86) at its N-terminus and a nucleoplasmin NLS (such as the sequence indicated as SEQ ID NO: 87) at its C-terminus. In some embodiments, the RGN or other nuclease includes a c-Myc NLS (such as the sequence indicated as SEQ ID NO: 125) at both its N-terminus and C-terminus. If the NLS is bound to the N-terminus, C-terminus, or both of the RGN or other nuclease, an NLS linker protein may be present to separate the RGN or other nuclease from the NLS. In some embodiments, the NLS linker protein ligates the RGN polypeptide or other nuclease to the NLS. Such NLS linker proteins may have an amino acid length of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more.In some embodiments, the NLS linker protein has an amino acid length of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, or at least 8. In some embodiments, the NLS linker protein between or connecting to the NLS and the RGN or other nuclease has the sequence shown as SEQ ID NO: 127. In some embodiments, the RGN or other nuclease includes a c-Myc NLS at its N-terminus (such as the sequence shown as SEQ ID NO: 125) isolated from the nuclease protein by the NLS linker protein having the sequence shown as SEQ ID NO: 127, and a c-Myc NLS at its C-terminus (such as the sequence shown as SEQ ID NO: 125) isolated from the nuclease protein by the NLS linker protein having the sequence shown as SEQ ID NO: 127. RGN polypeptides or other nucleases can link to c-Myc NLS (such as the sequence shown as SEQ ID NO: 125) at their N-terminus and to c-Myc NLS (such as the sequence shown as SEQ ID NO: 125) at their C-terminus, and RGN polypeptides or other nucleases can link to the N-terminal c-Myc NLS and the C-terminal c-Myc NLS, respectively, by an NLS linker protein having the sequence shown as SEQ ID NO: 127.

[0087] In some embodiments, the compositions and methods of the present disclosure utilize RGN or other nucleases comprising at least one cell-permeable domain that facilitates the cellular uptake of RGN. Cell-permeable domains are known in the art and generally include stretches of positively charged amino acid residues (i.e., polycationic cell-permeable domains), alternating polar and nonpolar amino acid residues (i.e., amphipathic cell-permeable domains), or hydrophobic amino acid residues (i.e., hydrophobic cell-permeable domains) (see, for example, Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-permeable domain is trans-activated transcription activator (TAT) derived from human immunodeficiency virus 1.

[0088] Nuclear localization signals and / or cell permeability domains may be located at the N-terminus, C-terminus, and / or within the RGN.

[0089] V. Nucleic acid molecules encoding RNA-induced nucleases, single guide RNAs, CRISPR RNAs, and / or tracrRNAs. This disclosure provides nucleic acid molecules comprising or encoding RGN, crRNA, tracrRNA, and / or sgRNA as disclosed herein.

[0090] The use of the terms “polynucleotide” or “nucleic acid molecule” is not intended to limit this disclosure to polynucleotides containing DNA. Those skilled in the art will recognize that polynucleotides may include ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNA), PNA-DNA chimeras, locked nucleic acids (LNA), and phosphothiolate-binding sequences. The polynucleotides disclosed herein also encompass all forms of sequences, including, but not limited to, single-stranded, double-stranded, DNA-RNA hybrids, triple structures, stem and loop structures.

[0091] In some embodiments of the compositions and methods of this disclosure that include a nucleic acid molecule encoding an RGN, the nucleic acid molecule is an mRNA (messenger RNA) molecule. mRNA refers to any polynucleotide encoding a polypeptide of interest, which can be translated to produce the encoded polypeptide of interest in vitro, in vivo, in situ, or ex vivo. In some embodiments, the basic components of the mRNA molecule include at least a coding region, a 5'UTR, a 3'UTR, a 5' cap, and a poly-A tail. In some embodiments, the mRNA encoding an RGN useful for the methods and compositions of this disclosure may include one or more structural and / or chemical modifications or alterations that confer useful properties to the polynucleotide. For example, useful properties of the mRNA include the absence of substantial induction of the innate immune response of the cell into which the mRNA is introduced. A “structural” feature or modification is one in which two or more linked nucleotides are inserted, deleted, replicated, inverted, or randomized in the mRNA without significant chemical modification of the nucleotides themselves. Chemical bonds are inevitably broken and reconfigured to result in structural modifications; therefore, structural modifications have chemical properties and are thus chemical modifications. However, structural modifications will result in different sequences of nucleotides. Chemical modifications of mRNA can include 5-methylcytosine, N1-methyl-pseudridine, pseudouridine, 2-thiouridine, 4-thiouridine, 5-methoxyuridine, 2'-fluoroguanosine, 2'-fluorouridine, 5-bromolidine, 5-(2-carbomethoxyvinyl)uridine, 5-[3(1-E-propenylamino)]uridine, α-thiocytidine, N6-methyladenosine, 5-methylcytidine, N4-acetylcytidine, 5-formylcytidine, or combinations thereof in mRNA.

[0092] Nucleic acid molecules encoding RGNs can be codon-optimized for expression in the organism of interest (e.g., mammals). A “codon-optimized” coding sequence is a polynucleotide coding sequence with codon usage frequencies designed to mimic the preferred codon usage or transcriptional conditions of a particular host cell. Expression in a particular host cell or organism is enhanced as a result of changes in one or more codons at the nucleic acid level, without altering the translated amino acid sequence. Nucleic acid molecules can be codon-optimized whole or partially. Codon tables and other reference literature providing a wide range of biological preference information are available in the art (see, for example, Gaspar et al. (2012) Bioinformatics 28(20):2683-2684, Komar et al. (1998) Biol.Chem. 379(10):1295-1300, and Inouye et al. (2015) Protein Expr.Purif. 109:47-54). A non-limiting example of a codon-optimized code sequence for RGN useful in the compositions and methods of this disclosure is shown as Sequence ID No. 88.

[0093] The polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA provided herein may be provided in expression cassettes for in vitro expression or expression in cells, embryos, or organisms of interest. The cassette comprises 5' and 3' regulatory sequences operably ligated to the polynucleotides encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein, enabling the expression of the polynucleotides. The cassette may further include at least one additional gene or genetic element to be co-transformed into an organism. If an additional gene or element is included, the component is operably ligated. The term “operably ligated” is intended to mean a functional linkage between two or more elements. For example, an operable linkage between a promoter and a coding region of interest (e.g., a region encoding RGN, crRNA, tracrRNA, and / or sgRNA) is a functional linkage that enables the expression of the coding region of interest. Operatively ligated elements may be continuous or discontinuous. When used to refer to the joining of two protein-coding regions, "operably linked" or "operably fused" means that the coding regions are intended to be within the same reading frame. For example, "operably fused" polypeptides may mean that the structure and / or biological activity of each individual peptide is also present in the fusion. Alternatively, additional genes or elements may be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the RGN of this disclosure may reside on one expression cassette, while the nucleotide sequence encoding crRNA, tracrRNA, or complete guide RNA may reside on a separate expression cassette. Such an expression cassette is provided with multiple restriction sites and / or recombination sites for inserting polynucleotides so that they are under transcriptional regulation of the regulatory region. The expression cassette may further include selectable marker genes.

[0094] An expression cassette comprises, in the 5'-3' direction of transcription, a transcription (and, in some embodiments, translation) initiation region (i.e., promoter), a polynucleotide encoding RGN-, crRNA-, tracrRNA-, and / or sgRNA- as disclosed herein, and a transcription (and, in some embodiments, translation) termination region (i.e., termination region) that functions in the organism of interest. The promoter as disclosed herein can lead or drive the expression of the coding sequence in a host cell. Regulatory regions (e.g., promoter, transcription regulatory region, and translation termination region) may be heterogeneous to or from the host cell, whether endogenous or heterogeneous. As used herein, “heterogeneous” with respect to a sequence means a sequence of foreign origin, or, if of the same species, a sequence whose composition and / or genomic locus has been substantially modified from its native form by intentional artificial intervention. As used herein, a chimeric gene comprises a coding sequence operably ligated to a transcription initiation region that is heterogeneous with respect to the coding sequence.

[0095] Convenient termination regions include those derived from simian virus (SV40), human growth hormone (hGH), bovine growth hormone (BGH), and rabbit β-globin (rbGlob). See also Proudfoot (1991) Cell 64:671-674, Munroe et al. (1990) Gene 91:151-158, Schek et al. (1992) Molecular and Cellular Biology 12(12):5386-5393, Gil and Proudfoot (1987) Cell 49(3):399-406, Goodwin and Rottman (1992) The Journal of Biological Chemistry 267(23):16330-16334, and Lanois and Acheson (1988) EMBO J.7(8):2515-2522.

[0096] Additional regulatory signals include, but are not limited to, transcription initiation sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, and termination signals. For example, see Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), hereafter referred to as "Sambrook11", Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and the references cited therein.

[0097] In the preparation of expression cassettes, various DNA fragments may be manipulated to provide DNA sequences in appropriate orientation and, if necessary, in appropriate reading frames. For this purpose, adapters or linkers may be used to join the DNA fragments, or other operations may be involved to provide convenient restriction sites, removal of unwanted DNA, removal of restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, and resubstitution (e.g., transposition and base change) may be involved.

[0098] Many promoters can be used in the practice of the present invention. Promoters can be selected based on the desired results. Generally, RGN expression is under the control of the RNA polymerase II promoter, and therefore, the RGN coding sequence can be operably ligated to the RNA polymerase II promoter. The expression of crRNA, tracrRNA, or sgRNA is generally under the control of the RNA polymerase III promoter, and therefore, the coding sequences of these elements can be operably ligated to the RNA polymerase III promoter. Non-limiting examples of RNA polymerase III promoters useful for the expression of crRNA, tracrRNA, and sgRNA include the mammalian U6, U3, H1, and 7SL RNA promoters, and rice U6 and U3 promoters, e.g., the human U6 micronucleus promoter or cleaved forms thereof, e.g., the sequence shown as SEQ ID NO: 89 or 128, and the promoters shown herein as SEQ ID NOs: 96-105, as disclosed in U.S. Provisional Application No. 63 / 209,660 filed June 11, 2021, and International Application PCT / US2022 / 032940 filed June 10, 2022 (each of which is incorporated herein by reference in its entirety).

[0099] Nucleic acids can be combined with constitutive, inducible, growth-stage specific, cell-type specific, tissue-preferential, tissue-specific, or other promoters for expression in the organism of interest.

[0100] Exemplary constitutive promoters for expression in cells as disclosed herein include the SV40 early promoter, the mouse mammary tumor virus long-term repeat (LTR) promoter, and the adenovirus major late promoter (Ad Examples include the MLP), herpes simplex virus (HSV) promoter, cytomegalovirus (CMV) promoters such as the CMV earliest promoter region (CMVIE), Roussarcoma virus (RSV) promoter, human ubiquitin C promoter (UBC), human U6 micronucleus promoter (U6), cleaved U6 promoter, enhanced U6 promoter, human H1 promoter derived from RNA polymerase III (H1), human elongation factor 1α promoter (EF1A), human beta-actin promoter (ACTB), human or mouse phosphoglycerate kinase 1 promoter (PGK), chicken β-actin promoter conjugated with CMV early enhancer (CAGG), yeast transcription elongation factor promoter (TEF1), elongation factor 1α short (EFS) promoter, and JeT promoter (see, for example, U.S. Patent Publication 2002 / 0098547, which is incorporated herein by reference in its entirety).For example, Miyagishi et al. (2002) Nature Biotechnology 20:497-500, Xia et al. (2003) Nucleic Acids Res.31(17):e100-e100, Pasleau et al. (1985) Gene 38:227-232, Martin-Gallardo et al. (1988) Gene 70:51-56, Oellig and Seliger (1990) J Neurosci Res 26:390-396, Manthorpe et al. (1993) Hum Gene Ther 4:419-431, Yew et al. (1997) Hum Gene Ther 8:575-584, Xu et al. (2001) Gene 272:149-156, Nguyen et al. (2008)J See Surg Res 148:60-66, Costa et al. (2005) Nat Meth. 2:259-260, and Lam and Truong (2020) ACS Synth. Biol. 9(10):2625-2631. In some embodiments, the RGN coding sequence is operably ligated to a constitutive promoter, which may be a cytomegalovirus (CMV) promoter, a cleaved CMV promoter such as the CMVeb promoter shown as SEQ ID NO: 90, an elongation factor 1α short (EFS) promoter shown as SEQ ID NO: 91, or a JeT promoter shown as SEQ ID NO: 92.

[0101] Examples of inductive promoters include stress-modulating promoters such as the Hsp70 and Hsp90 promoters (Wurm et al. (1986) Proc. Natl. Acad. Sci. USA. 83:5414-5418, Nover L. Heat Shock Response. CRC Press; Boca Raton, FL, USA: 1991), metal-modulating promoters (Mayo et al. (1982) Cell. 29:99-108, Searle et al. (1985) Mol. Cell. Biol. 5:1480-1489), and hormone-responsive promoters including glucocorticoid-responsive promoters (Hynes et al. (1981) Proc. Natl. Acad. Sci. USA. 78:2038-2042, Klock et al. (1987) Nature. 329:734-736). Examples of chemically modified prokaryotic promoters used include the isopropyl-β-D-thiogalactopyranoside (IPTG) regulatory promoter, the lactose regulatory promoter, and the tetracycline regulatory promoter (see, for example, Gossen et al. (1993) Trends Biochem Sci. 18:471-475, Gossen and Bujard (1992) Proc. Natl Acad. Sci. USA 89:5547-5551, and Zhou et al. (2006) Gene Ther. 13:1382-1390).Inducible expression can be obtained using an operator system containing AlcR / acetaldehyde, ArgR / L-arginine, BirA / biotinyl-AMP, CymR / chmate, EthR / 2-phenylethyl butyrate, HdnoR / 6-hydroxynicotine, HucR / uric acid, MphR(A) / macrolide, PIP / streptogramin, Rex / NADH, RheA / heat, ScbR / SCB1, TraR / 3-oxo-C8-HSL, and TtgR / phloretin, for example, U.S. Patent No. 8,728,759B2, U.S. Patent No. 7,745,592B2, Weber and Fussenegger (2004) Methods Mol. Biol. 267:451-466, Hartenbach et al. (2007) Nucleic Acids Res. 35:e136, Weber et al. al.(2009) Metab.Eng.11:117-124, Weber et al.(2008) Proc.Natl.Acad.Sci.USA.105:9994-9998, Malphettes et al.(2005) Nucleic Acids Res.33:e107, Kemmer et al. al. (2010) Nat.Biotechnol.28:355-360, Weber et al. (2002) Nat.Biotechnol.20:901-907, Fussenegger et al. (2000) Nat.Biotechnol.18:1203-1208, Weber et al. al. (2006) Metab. Eng. 8:273-280, Weber et al. (2003) Nucleic Acids See Res.31:e69, Weber et al. (2003) Nucleic Acids Res.31:e71, Neddermann et al. (2003) EMBO Rep.4:159-165, and Gitzinger et al. (2009) Proc.Natl.Acad.Sci.USA.106:10638-10643.Inducible expression includes the rapamycin-inducible interaction system between FKBP12 (FK506-binding protein 12) and mTOR (Rivera et al. (1996) Nat. Med. 2: 1028-1032, Belshaw et al. (1996) Proc. Natl. Acad. Sci. USA. 93: 4604-46077), the abscisic acid (ABA) regulatory interaction between PYL1 (abscisic acid receptor) and ABI1 (protein phosphatase 2C56) (Liang et al. (2011) Sci. Signal. 4(164): rs2-rs2), and the photo-inducible protein-protein interaction system (Wang et al. (2012) Nat. Methods. 9: 266-269, Yamada et al. It can be obtained using a protein-protein-protein interaction system, including al. (2018) Cell. Rep. 25: 487-500).

[0102] Tissue-specific or tissue-preferential promoters can be used to target the expression of expression constructs within specific tissues. In some embodiments, the tissue-specific or tissue-preferential promoter is active in mammalian tissues. Examples of tissue-specific or tissue-preferential promoters include promoters that preferentially initiate transcription in specific tissues, such as the brain. A “tissue-specific” promoter is a promoter that initiates transcription only in specific tissues. Unlike constitutive gene expression, tissue-specific expression is the result of several interacting levels of gene regulation. Therefore, promoters from homologous or closely related species may be preferable to achieve efficient and reliable expression of transgenes in specific tissues. In some embodiments, expression includes a tissue-preferential promoter. A “tissue-preferential” promoter is a promoter that preferentially initiates transcription in specific tissues, such as the brain, but not necessarily completely or solely.

[0103] In embodiments, nucleic acid molecules encoding RGN, crRNA, tracrRNA, and / or sgRNA include cell type-specific promoters. A “cell type-specific” promoter is a promoter that primarily drives expression in a particular cell type in one or more organs. Some examples of cells in which a cell type-specific promoter may be primarily active include, for example, neurons. Nucleic acid molecules may also include cell type-preferential promoters. A “cell type-preferential” promoter is a promoter that primarily drives, but not necessarily entirely or solely, expression in a particular cell type in one or more organs. Some examples of cells in which a cell type-preferential promoter may be preferentially active include, for example, neurons. Neurons can include neural progenitor cells, forebrain progenitor cells, striatal neurons, medium spiny neurons, and cortical neurons. A cell type-preferential promoter may be preferentially active in the brain of non-neuronal cells such as glial cells. Glial cells can include microglia, astrocytes, and oligodendrocytes. In some embodiments, the cell type preferential promoter may be preferentially activated in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or a combination thereof.

[0104] RGN coding sequences can be operably ligated to brain or neuron-specific promoters such as the human synapsin I (Syn) promoter, 65kDa or 67kDa glutamate carboxylase (GAD65 or GAD67, respectively) promoter, homeobox Dlx5 / 6 promoter, preprotachykinin 1 (Tac1) promoter, neuron-specific enolase (NSE), dopaminergic receptor 1 (Drd1a) promoter or dopaminergic receptor 2 (DRD2) promoter, glial fibrillary acidic protein (GFAP) promoter, or the 32kDa dopamine and cyclic AMP-regulated phosphoprotein (DARP32) promoter (see, for example, Delzor et al., 2012, Hum Gene Ther Methods 23(4):242-254, the whole of which is incorporated by reference). A non-limiting example of a promoter that can be used to drive RGN expression for use in the compositions and methods of this disclosure, which are neuron-specific, is the human synapsin I (Syn) promoter. The Syn promoter may have the nucleotide sequence shown as Sequence ID No. 93.

[0105] Nucleic acid sequences encoding RGN, crRNA, tracrRNA, and / or sgRNA can be operably ligated to a promoter sequence recognized by phage RNA polymerase, for example, for in vitro mRNA synthesis. In embodiments, in vitro transcription RNA can be purified for use in the method described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence, or a variation thereof. In embodiments, expressed proteins and / or RNA can be purified for use in the genome modification method described herein.

[0106] In some embodiments, polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA can also be ligated to a polyadenylation (polyA) signal and / or at least one transcription termination sequence. In some embodiments, the coding sequence (e.g., a nucleic acid molecule encoding RGN, crRNA, tracrRNA, and / or sgRNA) is ligated to a monkey virus (SV40) polyA tail, such as shown as SEQ ID NO: 94, or a bovine growth hormone polyadenylation (bGHpolyA) tail, such as shown as SEQ ID NO: 95. For example, see Proudfoot (1991) Cell 64:671-674; Munroe et al. (1990) Gene 91:151-158; Schek et al. (1992) Molecular and Cellular Biology 12(12):5386-5393; Gil and Proudfoot (1987) Cell 49(3):399-406; Goodwin and Rottman (1992) The Journal of Biological Chemistry 267(23):16330-16334, and Lanois and Acheson (1988) EMBO J.7(8):2515-2522.

[0107] In addition, the RGN encoding sequence may also be ligated to a sequence(s) encoding at least one nuclear localization signal, at least one cell permeability domain, and / or at least one signal peptide capable of transporting a protein to a specific subcellular location, as described elsewhere herein.

[0108] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA can be present in a vector or multiple vectors. “Vector” refers to a polynucleotide composition for transferring, delivering, or introducing nucleic acids into host cells. Suitable vectors include plasmid vectors, phagemids, cosmids, artificial / minichromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated virus vectors, baculovirus vectors). Vectors may include additional expression regulatory sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), selectable marker sequences (e.g., antibiotic resistance genes), origins of replication, etc. Further information can be found in “Current Protocols in Molecular Biology” Ausubel et al., John Wiley & Sons, New York, 2003 or “Molecular Cloning: A Laboratory Manual” Sambrook & Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0109] The vector may also include selectable marker genes for selecting transformed cells. These selectable marker genes are used to select transformed cells or tissues. Marker genes include those encoding antibiotic resistance, such as neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT). Marker genes include those for specific nutrients or substances, such as dihydrofolate reductase (DHFR; Simonsen and Levinson (1983) Proc. Natl. Acad. Sci. USA 80:2495-2499), histidinol dehydrogenase (hisD; Hartman and Mulligan (1988) Proc. Natl. Acad. Sci. USA 85:8047-8051), puromycin-N-acetyltransferase (PAC or puro; de la Luna et al. (1988) Gene 62:121-126), thymine kinase (TK; Littlefield (1964) Science 145:709-710), and xanthine-guanine phosphoribosyltransferase (XGPRT or gpt; Mulligan and It may be possible to include genes that enable selective growth on Berg (1981) Proc. Natl. Acad. Sci. USA 78:2072-2076).

[0110] As shown, an organism of interest can be transformed using an expression construct comprising nucleotide sequences encoding RGN, crRNA, tracrRNA, and / or sgRNA. The method for transformation involves introducing the nucleotide construct into the organism of interest. "Introducing" means introducing the nucleotide construct into a host cell so that the construct can access the interior of the host cell. The methods of this disclosure do not require a specific method for introducing the nucleotide construct into the host organism, as long as the nucleotide construct can access the interior of at least one cell of the host organism. The host cell can be a eukaryotic cell or a prokaryotic cell. In some embodiments, the eukaryotic host cell is a mammalian cell, an avian cell, or an insect cell. In some embodiments, the eukaryotic cell comprising, expressing, or being modified by the RGN system of this disclosure is a human cell. In some embodiments, eukaryotic cells containing, expressing, or being modified by the RGN system of the Disclosure, the crRNA, tracrRNA, sgRNA, and / or RGN, are stem cells, including induced pluripotent stem cells. In some embodiments, mammalian or human cells containing, expressing, or being modified by the RGN system of the Disclosure, the crRNA, tracrRNA, sgRNA, and / or RGN, are cardiomyocytes, neurons, glial cells, or retinal ganglion cells.

[0111] Methods for introducing nucleotide constructs into host cells are known in the art and include, but are not limited to, stable transformation methods, transient transformation methods, and virus-mediated methods.

[0112] The methods disclosed herein can result in transformed organisms or cell lines derived from these transformed cells.

[0113] A “transgenic organism,” “transformed organism,” or “stable transformed” organism, cell, or tissue refers to an organism that incorporates or has incorporated polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA as defined herein. Other exogenous or endogenous nucleic acid sequences or DNA fragments are also recognized to be incorporated into host cells. Transformation of host cells can be carried out by infection, conjugation, transfection, microinjection, electroporation, microprojection, gene gun or particle collision, electroporation, silica / carbon fiber, sonication-mediated, PEG-mediated, calcium phosphate coprecipitation, polycationic DMSO technology, DEAE dextran procedure, and viral-mediated, liposome-mediated, etc. Viral-mediated introduction of polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA includes retrovirus, lentivirus, adenovirus, and adeno-associated virus-mediated introduction and expression.

[0114] Transformation may result in the stable or transient integration of nucleic acids into cells. “Stable transformation” is intended to mean that a nucleotide construct introduced into a host cell is integrated into the host cell’s genome and can be inherited by its offspring. “Transient transformation” is intended to mean that a polynucleotide introduced into a host cell is not integrated into the host cell’s genome.

[0115] In some embodiments, the transformed cells may be introduced into an organism. These cells may originate from the organism, and the cells are transformed using an ex vivo approach. These cells can be autogenic (returning to the same origin and subject) or allogeneic (the donor and recipient subjects are of the same species). Generally, the donor and recipient of allogeneic cells are a complete or partial HLA match.

[0116] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA, or containing crRNA, tracrRNA, and / or sgRNA, can also be used to transform any prokaryotic species, including but not limited to archaea and bacteria (e.g., Bacillus sp., Klebsiella sp., Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.).

[0117] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA, or containing crRNA, tracrRNA, and / or sgRNA, can be used to transform any eukaryotic species, including but not limited to animals (e.g., mammals, humans, mice, rats, non-human primates, insects, fish, birds, and reptiles), fungi, amoebas, algae, and yeast.

[0118] Nucleic acids can be introduced into mammalian, insect, or avian cells or target tissues using conventional viral and nonviral-based gene transfer methods. Using such methods, nucleic acids encoding components of the RGN system can be administered to cells in culture or into a host organism. Nonviral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery vehicles such as liposomes. Viral vector delivery systems include DNA and RNA viruses, which, after delivery to cells, have either an episomal genome or an integrated genome. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel & Feigner, TIBTECH 11:211-217 (1993); Mitani & Caskey, TIBTECH 11:162-166 (1993); Dillon, TIBTECH 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer & Perricaudet, British Medical Bulletin 51(1):31-44 (1995); Haddada et al., in Current Topics in Microbiology and See Immunology, Doerfler and Bohm (eds) (1995) and Yu et al., Gene Therapy 1:13-26 (1994).

[0119] Nonviral delivery methods of nucleic acids include lipofection, nucleofection, microinjection, gene guns, virosomes, liposomes, immunoliposomes, polycational or lipid: nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced uptake of DNA. Lipofection is described, for example, in U.S. Patents 5,049,386, 4,946,787, and 4,897,355, and lipofection reagents are commercially available (e.g., Transfectam® and Lipofectin®). Cationic and neutral lipids suitable for efficient receptor-recognition lipofection of polynucleotides include those described in Feigner, WO91 / 17424, and WO91 / 16024. Delivery may be to cells (in vitro or ex vivo administration) or target tissues (e.g., in vivo administration). The preparation of lipid:nucleic acid complexes, including target liposomes such as immunolipid complexes, is well known to those skilled in the art (e.g., Crystal, Science 270:404-410 (1995), Blaese et al., Cancer Gene Ther. 2:291-297 (1995), Behr et al., Bioconjugate Chem. 5:382-389 (1994), Remy et al., Bioconjugate Chem. 5:647-654 (1994), Gao et al., Gene Therapy 2:710-722 (1995), Ahmad et al., Cancer See Res.52:4817-4820 (1992), U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028, and 4,946,787).

[0120] The use of RNA or DNA virus-based systems for nucleic acid delivery utilizes highly evolved processes to target viruses to specific cells in the body and transport the viral payload to the nucleus. Viral vectors can be administered directly to patients (in vivo) or used to treat cells in vitro, and modified cells can be administered to patients selectively (ex vivo). Conventional virus-based systems can include retroviral, lentiviral, adenovirus, adeno-associated virus, and herpes simplex virus vectors for gene transfer. Integration into the host genome is possible with retroviral, lentiviral, and adeno-associated virus gene transfer methods, often resulting in long-term expression of the inserted transgene. In addition, high transduction efficiencies have been observed in many different cell types and target tissues.

[0121] Retroviral tropism can be modified by incorporating exogenous envelope proteins and expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically produce high viral titers. Therefore, the selection of a retroviral gene transduction system will depend on the target tissue. Retroviral vectors consist of cis-acting long terminal repeats with the ability to package exogenous sequences up to 6–10 kb. The smallest cis-acting LTRs are sufficient for vector replication and packaging and are then used to integrate therapeutic genes into target cells to provide persistent transgene expression. Widely used retroviral vectors include those based on mouse leukemia virus (MuLV), Divonian primate leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., J.Viral.66:2731-2739 (1992), Johann et al., J.Viral.66:1635-1640 (1992), Sommnerfelt et al., Viral.176:58-59 (1990), Wilson et al., J.Viral.63:2374-2378 (1989), Miller et al., J.Viral.65:2220-2224 (1991), and PCT / US94 / 05700).

[0122] For applications where transient expression is preferred, adenovirus-based systems may be used. Adenovirus-based vectors enable very high transduction efficiency in many cell types and do not require cell division. High titers and expression levels have been obtained using such vectors. These vectors can be produced in large quantities using relatively simple systems. Adeno-associated virus ("AAV") vectors may also be used to transduce target nucleic acids into cells, for example, in the in vitro production of nucleic acids and peptides, as well as in in vivo and ex vivo gene therapy procedures (see, e.g., West et al., Virology 160:38-47 (1987), U.S. Patent No. 4,797,368, WO93 / 24641, Katin, Human Gene Therapy 5:793-801 (1994), Muzyczka, 1. Clin. Invest. 94:1351 (1994). For the construction of recombinant AAV vectors, see U.S. Patent No. 5,173,414, Tratschin et al., Mol. Cell. Biol. 5:3251-3260 (1985), Tratschin, et al. This has been described in numerous publications, including al., Mol.Cell.Biol.4:2072-2081(1984); Hermonat & Muzyczka, PNAS 81:6466-6470(1984), and Samulski et al., J.Viral.63:03822-3828(1989). Packaging cells are typically used to form viral particles capable of infecting host cells. Such cells include 293 cells that package adenoviruses, and ψJ2 cells or PA317 cells that package retroviruses.

[0123] Viral vectors used in gene therapy are typically produced by creating cell lines that package nucleic acid vectors into viral particles. The vectors typically contain the minimum viral sequences necessary for packaging and subsequent integration into the host, with other viral sequences replaced by expression cassettes for the polynucleotide(s) to be expressed. Missing viral functions are typically supplied trans by the packaging cell line. For example, AAV vectors used in gene therapy typically contain only the ITR sequences derived from the AAV genome necessary for packaging and integration into the host genome. The viral DNA is packaged into cell lines containing helper plasmids that encode other AAV genes, namely rep and cap, but lack the ITR sequences.

[0124] Cell lines can be infected with adenovirus as a helper virus. The helper virus promotes replication of the AAV vector and expression of the AAV gene from the helper plasmid. The helper plasmid is not packaged in significant quantities due to the lack of an ITR sequence. Adenovirus contamination can be reduced, for example, by heat treatment in which the adenovirus is more susceptible than AAV. Additional methods for delivering nucleic acids to cells are known to those skilled in the art. See, for example, US2003 / 0087817 (incorporated herein by reference).

[0125] Non-limiting examples of AAV vectors useful in the compositions and methods of this disclosure are the AAV2, AAV3, AAV5, AAV6, and AAV9 vectors (see, for example, Pupo et al., 2022, Molecular Therapy 30(12):P3515-3541, which is incorporated by reference in its entirety). In some embodiments, the vector is AAV5 or AAV6. In some embodiments, the AAV vector has a sequence indicated as one of SEQ ID NOs. 30-39 or 121-123, and the AAV packages the vector sequence.

[0126] In some embodiments, host cells are transiently or nontransiently transfected with one or more nucleic acid molecules or vectors described herein. In some embodiments, cells are transfected to resemble their naturally occurring state in the subject. In some embodiments, the cells to be transfected are taken from a subject, e.g., a Huntington's disease patient. In embodiments, the cells are derived from a subject, e.g., cells taken from a cell line. In some embodiments, the cell line may be mammalian, insect, or avian cells. A wide variety of cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFL, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182 , A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, T IB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, HeLa T4.COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelium, BALB / 3T3 mouse embryonic fibroblasts, 3T3 SWISS, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B 16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA 2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-l0A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO -MAC6, MTD-lA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW Examples include -145, OPCN / OPCT cell lines, Peer, PNT-1A / PNT2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THPl cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and their transgenic varieties. Cell lines are available from various sources known to those skilled in the art (see, for example, the American Type Culture Collection (ATCC) (Manassas, Va.)).

[0127] In some embodiments, a novel cell line containing one or more vector-derived sequences is established using cells transfected with one or more nucleic acid molecules or vectors described herein. In some embodiments, a novel cell line containing cells that are transiently transfected with components of the RGN system described herein (e.g., by transient transfection of one or more vectors or by transfection with RNA) and modified by the activity of the RGN system is established using cells that contain the modification but lack any other exogenous sequences.

[0128] In some embodiments, one or more nucleic acid molecules or vectors described herein are used to produce non-human transgenic animals. In some embodiments, the transgenic animal is a mammal such as a mouse, rat, hamster, rabbit, cow, or pig.

[0129] VI. Variants and Fragments of Polypeptides and Polynucleotides This disclosure provides active variants and fragments of crRNA repeats, crRNAs, tracrRNAs, sgRNAs, and RGNs. Active variants or fragments of naturally occurring (i.e., wild-type) RGNs bind to the target sequences described herein within mutant HTT alleles in an RNA-inducible sequence-specific manner. In some embodiments, the target sequences described herein include nucleotide sequences shown as any one of SEQ ID NOs. 75-79 and 130. In some embodiments, the Disclosure provides active variants and fragments of RGNs having an amino acid sequence shown as any one of SEQ ID NOs: 3, 7, 11, and 15; active variants and fragments of naturally occurring CRISPR repeats including a sequence shown as any one of SEQ ID NOs: 4, 8, 12, 16, and 106; active variants and fragments of naturally occurring tracrRNAs such as one of a sequence shown as any one of SEQ ID NOs: 5, 9, 13, 17, 107, and 120; and active variants and fragments of sgRNAs such as a sequence shown as any one of SEQ ID NOs: 25-29, as well as polynucleotides encoding them. In some embodiments, the sgRNAs of the Disclosure include sgRNAs shown as SEQ ID NOs: 6, 10, 14, or 18, the sgRNAs including any spacer useful for targeting a target sequence in the mHTT allele, and a backbone that can be bound by the RGN polypeptide of SEQ ID NOs: 3, 7, 11, or 15, or its active variant or fragment, respectively.

[0130] The activity of a variant or fragment may be altered compared to the polynucleotide or polypeptide of interest, but the variant and fragment must retain the function of the polynucleotide or polypeptide of interest. For example, the variant or fragment may have increased activity, decreased activity, a different spectrum of activity, or any other alteration of activity compared to the polynucleotide or polypeptide of interest.

[0131] Naturally occurring RGN polypeptide fragments and variants, such as those disclosed herein, retain sequence-specific RNA-induced DNA-binding activity. In embodiments, naturally occurring RGN polypeptide fragments and variants, such as those disclosed herein, retain nuclease activity (single-stranded or double-stranded).

[0132] Naturally occurring CRISPR repeat fragments and variants, such as those disclosed herein, retain the ability to bind to RNA-inducible nucleases (complexed with guide RNA) in a sequence-specific manner and guide the RNA-inducible nucleases to a target sequence when they are part of a guide RNA (including tracrRNA).

[0133] Naturally occurring tracrRNA fragments and variants, such as those disclosed herein, retain the ability to guide RNA-induced nucleases (complexed with guide RNA) to target sequences in a sequence-specific manner when they are part of a guide RNA (including CRISPR RNA).

[0134] The sgRNA fragments and variants disclosed herein, among others, retain the ability to induce RNA-induced nucleases (complexed with sgRNA) to target sequences in a sequence-specific manner.

[0135] The term “fragment” refers to a portion of a polynucleotide or polypeptide sequence of the present disclosure. A “fragment” or “biologically active portion” comprises a polynucleotide containing a sufficient number of consecutive nucleotides to maintain biological activity (i.e., when contained within a guide RNA, it binds to an RGN and directs the RGN to a target sequence in a sequence-specific manner). A “fragment” or “biologically active portion” comprises a polypeptide containing a sufficient number of consecutive amino acid residues to maintain biological activity (i.e., when complexed with a guide RNA, it binds to a target sequence in a sequence-specific manner). Fragments of RGN proteins include fragments shorter than the full-length sequence by using an alternative downstream start site. A biologically active portion of an RGN protein is, for example, an RGN that binds to a target nucleotide sequence disclosed herein, or an RGN having an amino acid sequence shown as one of SEQ ID NOs: 3, 7, 11, and 15, containing 10, 25, 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550 The polypeptide can contain 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700, or more consecutive amino acid residues. Such biologically active portions can be prepared by recombinant techniques and evaluated for sequence-specific RNA-induced DNA binding activity. A biologically active fragment of a CRISPR repeat sequence may contain at least eight consecutive nucleotides from any one of SEQ ID NOs: 4, 8, 12, 16, and 106. The biologically active portion of the CRISPR repeat sequence can be a polynucleotide containing, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides from any one of sequence numbers 4, 8, 12, 16, and 106.The biologically active portion of tracrRNA can be a polynucleotide containing, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, or more consecutive nucleotides from any one of sequence numbers 5, 9, 13, 17, 107, and 120. The biologically active portion of sgRNA can be a polynucleotide containing, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, or more consecutive nucleotides from any one of sequence numbers 6, 10, 14, 18, and 25-29.

[0136] In general, “variant” is intended to mean a substantially similar sequence. With respect to polynucleotides, a variant includes the deletion and / or addition of one or more nucleotides at one or more internal sites within a native polynucleotide, and / or the substitution of one or more nucleotides at one or more sites within a native polynucleotide. As used herein, “native” or “wild-type” polynucleotides or polypeptides include, respectively, naturally occurring nucleotide sequences or amino acid sequences. With respect to polynucleotides, conserved variants include sequences that encode the natural amino acid sequence of the gene of interest due to the degeneracy of the gene code. These naturally occurring allelic variants can be identified using well-known molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques, as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those produced using site-directed mutagenesis but still encoding the polypeptide or polynucleotide of interest. Generally, the variants of specific polynucleotides disclosed herein have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% or more of sequence identity with the specific polynucleotide, as determined by the sequence alignment programs and parameters described elsewhere in this Spec.

[0137] Variants of specific polynucleotides (i.e., reference polynucleotides) disclosed herein can also be evaluated by comparing the sequence identity percentage between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The sequence identity percentage between any two polypeptides can be calculated using the sequence alignment programs and parameters described elsewhere herein. When any given pair of polynucleotides disclosed herein is evaluated by comparing the sequence identity percentage shared by the two polypeptides they encode, the sequence identity percentage between the two encoded polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, and 99% or greater.

[0138] In certain embodiments, the polynucleotides of the present disclosure encode an RNA-inducible nuclease polypeptide comprising an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with an amino acid sequence that encodes an RGN that binds to a target sequence disclosed herein or an amino acid sequence shown as any one of SEQ ID NOs.

[0139] The biologically active variants of the RGN polypeptides of this disclosure may differ by only about 1 to 15 amino acid residues, only about 1 to 10 amino acid residues, for example, only about 6 to 10 amino acid residues, only 5 amino acid residues, only 4 amino acid residues, only 3 amino acid residues, only 2 amino acid residues, or only 1 amino acid residue. In some embodiments, the polypeptide may include cleavage at the N-terminus or C-terminus, which is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 35 It can contain 0, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, or 1700 or more amino acid deletions.

[0140] In some embodiments, the polynucleotides of the present disclosure include or encode crRNA repeats that have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with a nucleotide sequence represented as any one of SEQ ID NOs: 4, 8, 12, 16, and 106.

[0141] The polynucleotides of this disclosure include or can encode tracrRNAs that have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with any one of the nucleotide sequences shown as any one of SEQ ID NOs.

[0142] The polynucleotides of this disclosure include or can encode sgRNAs that have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more identity with any one of the nucleotide sequences shown as SEQ ID NOs: 6, 10, 14, 18, and 25-29.

[0143] The biologically active variants of the CRISPR repeats, crRNAs, tracrRNAs, or sgRNAs of this disclosure may differ by only about 1 to 15 nucleotides, only about 1 to 10 nucleotides, for example, only about 6 to 10 nucleotides, only 5 nucleotides, only 4 nucleotides, only 3 nucleotides, only 2 nucleotides, or only 1 nucleotide. In some embodiments, the polynucleotide may include a 5' or 3' cleavage, which may include deletions of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 95, 100, 105, 110 or more nucleotides from either the 5' or 3' end of the polynucleotide.

[0144] As used herein, a sequence that is “different” from a parent sequence by a certain number of amino acids or nucleotides may differ due to amino acid or nucleotide substitutions, amino acid or nucleotide additions, and / or amino acid or nucleotide deletions. For example, a nucleotide sequence that differs from a parent nucleotide sequence by only one nucleotide may have one nucleotide substitution and may be one nucleotide longer or one nucleotide shorter than the parent sequence.

[0145] Modifications can be performed on the RGN polypeptides, CRISPR repeats, crRNAs, tracrRNAs, and sgRNAs provided herein, and are recognized as producing variant proteins and polynucleotides. Human-designed alterations may be introduced through the application of site-directed mutagenesis techniques. Alternatively, native, previously unknown, or unconfirmed polynucleotides and / or polypeptides structurally and / or functionally related to the sequences disclosed herein may also be identified as being within the scope of this disclosure. Conservative amino acid substitutions may be performed in non-conserved regions that do not alter the function of the RGN protein. Alternatively, modifications that improve the activity of RGN may be performed.

[0146] Variant polynucleotides and proteins also include sequences and proteins derived from mutagenic and recombinogenic procedures such as DNA shuffling. Such procedures involve manipulating one or more different RGN proteins disclosed herein (e.g., SEQ ID NOs: 3, 7, 11, or 15) to create novel RGN proteins with desired properties. Thus, a library of recombinant polynucleotides is generated from a population of related sequence polynucleotides containing sequence regions that have substantial sequence identity and can undergo homologous recombination in vitro or in vivo. For example, this approach can be used to shuffle sequence motifs encoding a domain of interest between RGN sequences provided herein and other known RGN genes, resulting in increased K in the case of enzymes. mNovel genes encoding proteins with improved desired properties, such as those mentioned above, may be obtained. Strategies for such DNA shuffling are known in the art. For example, see Stemmer (1994) Proc. Natl. Acad. Sci. USA 91:10747-10751; Stemmer (1994) Nature 370:389-391; Crameri et al. (1997) Nature Biotech. 15:436-438; Moore et al. (1997) J. Mol. Biol. 272:336-347; Zhang et al. (1997) Proc. Natl. Acad. Sci. USA 94:4504-4509; Crameri et al. (1998) Nature 391:288-291, as well as U.S. patents No. 5,605,793 and 5,837,458. A “shuffled” nucleic acid is a nucleic acid produced by a shuffling procedure, such as any shuffling procedure described herein. A shuffled nucleic acid is produced by recombining two or more nucleic acids (or strings) (physically or virtually), for example, artificially and in an arbitrarily recursive manner. Generally, one or more screening steps are used in the shuffling process to identify the nucleic acid of interest, and this screening step can be performed before or after any recombination step. In some (but not all) embodiments of shuffling, it is desirable to perform multiple rounds of recombination before selection in order to increase the diversity of the pool being screened. The overall process of recombination and selection is arbitrarily repeated recursively. Depending on the context, shuffling can refer to the overall process of recombination and selection, or alternatively, simply to the recombination portion of the overall process.

[0147] As used herein, “sequence identity” or “identity” in the context of two polynucleotide or polypeptide sequences refers to residues in two sequences that are identical when aligned to the maximum extent possible within a particular comparison window. It is recognized that non-identical residue positions are often distinguished by conserved amino acid substitutions (where an amino acid residue is replaced by another amino acid residue that has similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule). Protein sequences that differ by such conserved substitutions are said to have “sequence similarity” or “similarity.” Means for measuring sequence similarity are well known to those skilled in the art. Typically, this involves scoring conserved substitutions as partial mismatches rather than complete mismatches. Thus, for example, if identical amino acids are given a score of 1 and non-conservative substitutions are given a score of 0, then conserved substitutions are given a score of 0 to 1. Scoring of conserved substitutions is calculated, for example, to be performed by the program PC / GENE (Intelligenetics, Mountain View, California).

[0148] As used herein, “percentage of sequence identity” is determined by comparing two optimally aligned sequences across a comparison window, where portions of the polynucleotide sequence within the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (without additions or deletions) for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions in which identical nucleic acid bases or amino acid residues occur in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions within the comparison window, and multiplying the result by 100 to obtain the percentage of sequence identity.

[0149] Unless otherwise stated, the sequence identity / similarity values ​​provided herein refer to values ​​obtained using GAP version 10 with the following parameters: GAP weight 50 and length weight 3, and nucleotide sequence identity % and similarity % using the nwsgapdna.cmp scoring matrix; GAP weight 8 and length weight 2, and amino acid sequence identity % and similarity % using the BLOSUM62 scoring matrix; or any equivalent program. "Equivalent program" refers to any sequence comparison program that, for any two sequences in question, produces alignments having identical nucleotide or amino acid residue matches and identical sequence identity percentages when compared to the corresponding alignment produced by GAP version 10.

[0150] Two sequences are "optimally aligned" when aligned for similarity scoring using a defined amino acid substitution matrix (e.g., BLOSUM62), gap presence penalty, and gap elongation penalty, to achieve the highest possible score for that pair of sequences. Amino acid substitution matrices and their use in quantifying similarity between two sequences are well known in the art and are described, for example, Dayhoff et al. (1978) “A model of evolutionary change in proteins.” In “Atlas of Protein Sequence and Structure,” Vol. 5, Suppl. 3 (ed. M Dayhoff), pp. 345-352. Natl. Biomed. Res. Found., Washington, DC, and Henikoff et al. (1992) Proc. Natl. Acad. Sci. USA 89:10915-10919. The BLOSUM62 matrix is ​​often used as the default scoring substitution matrix in sequence alignment protocols. A gap presence penalty is imposed for introducing a single amino acid gap into one of the aligned sequences, and a gap elongation penalty is imposed for each additional empty amino acid position inserted into an already open gap. Alignment is defined by the amino acid positions of each sequence in which the alignment begins and ends, and is optionally defined by inserting one or more gaps into one or both sequences to achieve the highest possible score. Optimal alignment and scoring can be achieved manually, but this process is facilitated by the use of computer-implemented alignment algorithms, such as GapBLAST2.0, described in Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402 and made publicly available on the National Center for Biotechnology Information Website (www.ncbi.nlm.nih.gov).Optimal alignments, including multiple alignments, are available, for example, through www.ncbi.nlm.nih.gov and can be prepared using PSI-BLAST as described by Altschul et al. (1997) Nucleic Acids Res. 25:3389-3402.

[0151] With respect to a nucleotide or amino acid sequence that is optimally aligned with a reference sequence, each nucleotide or amino acid residue "corresponds" to a position in the reference sequence where the nucleotide or residue is paired in the alignment. The "position" is indicated by a number that sequentially identifies each nucleotide in the reference nucleotide sequence based on its position relative to the 5' end, or each amino acid in the reference amino acid sequence based on its position relative to the N-terminus. Due to deletions, insertions, cleavages, fusions, etc., which must be considered when determining optimal alignment, the number of nucleotide positions or amino acid residues in a test sequence, determined simply by counting from the 5' or N-terminus, is not necessarily the same as the number of corresponding positions in the reference sequence. For example, if the aligned test sequence has a deletion, there is no nucleotide or amino acid corresponding to the position of the deletion site in the reference sequence. If an insertion exists in the aligned reference sequence, that insertion does not correspond to any nucleotide or amino acid position in the reference sequence. In the case of cleavage or fusion, there may be stretches of nucleotides or amino acids in the reference or aligned sequence that do not correspond to any nucleotide or amino acid in the corresponding sequence.

[0152] VII. RGN systems and ribonucleoprotein complexes for binding to target sequences, and methods for producing them. This disclosure provides an RGN system for binding to a target sequence in a mutant HTT allele. As used herein, the RGN system comprises at least one RGN polypeptide, or a polynucleotide comprising a nucleotide sequence encoding an RGN polypeptide, and one or more guide RNAs capable of forming a complex (ribonucleoprotein (RNP) complex) with the RGN polypeptide. The RGN system comprises a) one or more guide RNAs, or one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs, and b) an RGN polypeptide, or a polynucleotide comprising a nucleotide sequence encoding an RGN polypeptide, wherein one or more guide RNAs can form a complex with the RGN polypeptide so as to direct the RGN polypeptide to bind to a target sequence in the mutant HTT allele. The guide RNAs hybridize to the target strand of the target sequence in the mutant HTT allele and also form a complex with the RGN polypeptide, thereby directing the RGN polypeptide to bind to the target sequence. In some embodiments, the target sequence comprises a nucleotide sequence shown as any one of SEQ ID NOs. 75-79 and 130. In some embodiments, the RGN can recognize consensus PAM sequences represented as NNNNCC, NNRYA, NNGRR, or NNGG. In some embodiments, the RGN comprises an amino acid sequence represented as SEQ ID NOs: 3, 7, 11, and 15, or an active variant or fragment thereof. In some embodiments, the RGN comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, or at least 95%, or more sequence identity with any one of SEQ ID NOs: 3, 7, 11, and 15. In some embodiments, the guide RNA comprises a CRISPR repeat containing a nucleotide sequence represented as any one of SEQ ID NOs: 4, 8, 12, 16, and 106, or an active variant or fragment thereof.In some embodiments, the guide RNA comprises a tracrRNA containing one of the nucleotide sequences shown as one of SEQ ID NOs: 5, 9, 13, 17, 107, and 120, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises an sgRNA containing one of the nucleotide sequences shown as one of SEQ ID NOs: 6, 10, 14, 18, and 25-29, or an active variant or fragment thereof. The guide RNA in the system can be a single guide RNA or a dual guide RNA. In some embodiments, the system comprises an RNA-inducible nuclease heterologously to the guide RNA, and the RGN and guide RNA are not found to be naturally complexed with each other (i.e., bound to each other).

[0153] In some embodiments, the PAM interaction domain of the RGN polypeptide has a sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC.

[0154] In some embodiments, the PAM interaction domain of the RGN polypeptide has a sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA.

[0155] In some embodiments, the PAM interaction domain of the RGN polypeptide has a sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0156] In some embodiments, the PAM interaction domain of the RGN polypeptide has a sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG.

[0157] The RGN systems of this disclosure may include RGN polypeptides each comprising at least one nuclease domain, each involved in cleaving a single strand of a nucleic acid molecule. The nuclease domain may include a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0158] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence represented as one of SEQ ID NOs: 147, 148, 149, and 150.

[0159] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 151, 152, 153, and 154.

[0160] An RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 155, 156, 157, and 158.

[0161] The RGN system of this disclosure may include an RGN polypeptide comprising a PAM interaction domain that contributes to the recognition and binding of a PAM site, and further comprising at least one nuclease domain, each of which is involved in the cleavage of a single strand of nucleic acid molecule. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNNNCC and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, recognizing the PAM sequence NNNNCC, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0162] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNRYA and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, recognizing the PAM sequence NNRYA, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150.

[0163] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNGRR and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, recognizing the PAM sequence NNGRR, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154.

[0164] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136, recognizes the PAM sequence NNGG, and may further include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, recognizing the PAM sequence NNGG, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158.

[0165] The systems for binding to target sequences of interest provided herein may be ribonucleoprotein complexes, which are at least one molecule of RNA bound to at least one protein. The ribonucleoprotein complexes provided herein comprise at least one guide RNA as an RNA component and an RNA-inducible nuclease as a protein component. Such ribonucleoprotein complexes can be purified from cells or organisms that naturally express RGN polypeptides and have been engineered to express a specific guide RNA that is specific to the target sequence of interest (the target sequence of the mutant HTT allele). Alternatively, ribonucleoprotein complexes can be purified from cells or organisms transformed with polynucleotides (e.g., mRNA) encoding RGN polypeptides and guide RNAs, and cultured under conditions that allow for the expression of RGN polypeptides and guide RNAs. In some embodiments, ribonucleoprotein complexes are purified from cells or organisms transformed with polynucleotides (e.g., mRNA) encoding RGN polypeptides, and synthetically induced gRNAs are introduced. Thus, methods for producing RGN polypeptides or RGN ribonucleoprotein complexes are provided. Such a method involves culturing cells containing a nucleotide sequence encoding an RGN polypeptide (and, in some embodiments, a guide RNA) under conditions in which the RGN polypeptide (and, in some embodiments, a guide RNA) is expressed. The RGN polypeptide or RGN ribonucleoprotein can then be purified from the lysates of the cultured cells. In some embodiments, the nucleotide sequence encoding the RGN polypeptide includes mRNA (messenger RNA). In some embodiments, a method for assembling an RNP complex involves combining one or more of the guide RNAs of the disclosure with one or more of the RGN polypeptides of the disclosure under conditions suitable for the formation of an RNP complex.

[0166] Methods for purifying RGN polypeptides or RGN ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reversed-phase chromatography, immunoprecipitation). In certain methods, RGN polypeptides are recombinantly produced and include purification tags to aid in their purification, which include, but are not limited to, glutathione-S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tags, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG (e.g., 3X FLAG tag), HA, nus, Softag1, Softag3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotin carboxyl carrier protein (BCCP), and calmodulin. Generally, tagged RGN polypeptides or RGN ribonucleoprotein complexes are purified using immobilized metal affinity chromatography. It will be understood that other forms of chromatography or other similar methods known in the art, including, for example, immunoprecipitation, may be used either alone or in combination.

[0167] An isolated or purified polypeptide, or its biologically active portion, substantially or essentially contains components that would normally accompany or interact with the polypeptide as it would be found in its naturally occurring environment. Therefore, an isolated or purified polypeptide, if produced by recombinant technology, substantially contains no other cellular material or culture medium, or, if chemically synthesized, substantially contains no chemical precursors or other chemicals. A protein substantially free of cellular material includes protein preparations containing approximately 30%, 20%, 10%, 5%, or less than 1% (by dry weight) of contaminating proteins. When a protein or its biologically active portion of the disclosed material is produced recombinantly, the culture medium, optimally, contains approximately 30%, 20%, 10%, 5%, or less than 1% (by dry weight) of chemical precursors or non-target protein chemicals. Similarly, an isolated polynucleotide or nucleic acid molecule is removed from its naturally occurring environment. Isolated polynucleotides, if chemically synthesized or removed from gene loci via phosphodiester bond cleavage, are substantially free of chemical precursors or other chemicals. Isolated polynucleotides can be vectors, part of the composition of substances, or incorporated into cells unless the cell is the original environment of the polynucleotide.

[0168] The specific methods provided herein for binding to and / or cleaving a target sequence of interest involve the use of an in vitro assembled RGN ribonucleoprotein complex. In vitro assembly of the RGN ribonucleoprotein complex can be carried out using any method known in the art to contact the RGN polypeptide with the guide RNA under conditions that allow the RGN polypeptide to bind to the guide RNA. As used herein, “contact,” “bring into contact,” and “contacted” mean placing together the components of a desired reaction under conditions suitable for carrying out the desired reaction. The RGN polypeptide can be purified from a biological sample, cell lysate, or culture medium, produced via in vitro translation, or chemically synthesized. The guide RNA can be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGN polypeptide and guide RNA can be contacted in solution (e.g., buffered saline) to enable in vitro assembly of the RGN ribonucleoprotein complex.

[0169] Some embodiments of this disclosure provide kits comprising one or more elements of the RGN system described herein, including guide RNA (i.e., crRNA, tracrRNA, and / or sgRNA), RGN, and / or polynucleotides encoding it, cells, and a complete RGN system, and in some embodiments, another type of nuclease. In some embodiments, the kit includes suitable reagents, buffers, and / or instructions for using one or more elements of the RGN system for, for example, in vitro or in vivo nucleic acid editing. The reagents may be provided in any suitable container, such as vials, bottles, or tubes. The reagents may be used in a process that utilizes one or more elements of the RGN system. For example, restriction enzymes may be included for cloning polynucleotides encoding RGN or guide RNA into a vector. In some embodiments, the kit includes instructions for designing and using guide RNA (i.e., crRNA, tracrRNA, and / or sgRNA) suitable for targeted editing of nucleic acid sequences. The reagents may be provided in a form usable in a particular assay, or in a form that requires the addition of one or more other components before use (e.g., concentrate or lyophilized form). The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10.

[0170] Kits comprising one or more elements of the RGN system of this disclosure are useful in a wide variety of applications, including modification (e.g., deletion, insertion, transposition, inactivation, activation) of target polynucleotides in multiple cell types. Accordingly, kits comprising one or more elements of the RGN system of this disclosure may be useful, for example, in gene therapy, drug screening, disease diagnosis, and prognosis.

[0171] In some embodiments, the kit of the Disclosure includes a pharmaceutical kit comprising the pharmaceutical composition described herein. In some embodiments, the pharmaceutical kit includes (a) a container comprising the composition of the Disclosure in lyophilized form, and (b) a second container comprising a pharmaceutically acceptable diluent for injection (e.g., sterile water). The pharmaceutically acceptable diluent can be used to reconstitute or dilute the lyophilized compound of the Disclosure. Optionally, a notice in the form prescribed by a government agency regulating the manufacture, use, or sale of a pharmaceutical or biological product may accompany such container(s), the notice reflecting approval by the agency for manufacture, use, or sale for human administration.

[0172] VIII. Methods for binding to, cleaving, or modifying a target sequence This disclosure provides methods for binding to, cleaving, and / or modifying (i.e., editing) a target sequence within a mutant HTT allele. The methods include introducing an RGN system, comprising at least one guide RNA or polynucleotide encoding it, and at least one RGN polypeptide or polynucleotide encoding it, into a cell containing the target sequence. In some embodiments, delivery is ex vivo, and the cell containing the target sequence may be a stem cell, zygote, embryonic cell, or gamete. In some embodiments, the stem cell is an induced pluripotent stem cell (iPSC) or a mesenchymal stem cell (MSC). In some embodiments, the methods include delivering the RGN system in vivo, and the cell containing the target sequence is in vivo. In some embodiments, the target sequence within the mutant HTT allele has a nucleotide sequence represented as one of sequence numbers 75-79 and 130. In some embodiments, the RGN can recognize a consensus PAM sequence represented as one of NNNNCC, NNRYA, NNGRR, or NNGG. In some embodiments, the RGN comprises the amino acid sequence shown as SEQ ID NOs: 3, 7, 11, and 15, or an active variant or fragment thereof. The guide RNA of the system can be a single guide RNA or a dual guide RNA.

[0173] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 3. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 5, or an active variant or fragment thereof. In some embodiments, the guide RNA includes an sgRNA containing the nucleotide sequence shown as SEQ ID NO: 27 or 28, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC.

[0174] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 7. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof. In some embodiments, the guide RNA includes an sgRNA containing the nucleotide sequence shown as SEQ ID NO: 25 or 26, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA.

[0175] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 11. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof. In some embodiments, the guide RNA includes an sgRNA containing the nucleotide sequence shown as SEQ ID NO: 29, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0176] The RGN may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 15. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 17, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, and recognizes the PAM sequence NNGG.

[0177] The RGN systems of this disclosure may include RGN polypeptides each comprising at least one nuclease domain, each involved in cleaving a single strand of a nucleic acid molecule. The nuclease domain may include a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0178] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence represented as one of SEQ ID NOs: 147, 148, 149, and 150.

[0179] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 151, 152, 153, and 154.

[0180] An RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 155, 156, 157, and 158.

[0181] The RGN system of this disclosure may include an RGN polypeptide comprising a PAM interaction domain that contributes to the recognition and binding of a PAM site, and further comprising at least one nuclease domain, each of which is involved in the cleavage of a single strand of nucleic acid molecule. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNNNCC and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, recognizing the PAM sequence NNNNCC, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0182] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNRYA and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, recognizing the PAM sequence NNRYA, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150.

[0183] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNGRR and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, recognizing the PAM sequence NNGRR, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154.

[0184] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136, recognizes the PAM sequence NNGG, and may further include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, recognizing the PAM sequence NNGG, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158.

[0185] In certain embodiments, the RGN and / or guide RNA are heterogeneous to the cell into which the RGN and / or guide RNA (or polynucleotides encoding at least one of the RGN and / or guide RNA) are introduced.

[0186] In embodiments of the method, which includes the delivery of a polynucleotide encoding a guide RNA and / or RGN polypeptide, cells can then be cultured under conditions in which the guide RNA and / or RGN polypeptide is expressed. In some embodiments, the method includes contacting a target nucleic acid molecule with an RGN ribonucleoprotein complex. In some embodiments, the method includes introducing the RGN ribonucleoprotein complex into cells containing the target nucleic acid molecule. The RGN ribonucleoprotein complex can be a complex purified from a biological sample, a recombinantly produced and subsequently purified complex, or a complex assembled in vitro, as described herein. In embodiments in which the RGN ribonucleoprotein complex that contacts the target nucleic acid molecule or cells containing the target nucleic acid molecule is assembled in vitro, the method may further include in vitro assembly of the complex before contact with the target nucleic acid molecule or cells containing the target nucleic acid molecule.

[0187] Purified or in vitro assembled RGN ribonucleoprotein complexes can be introduced into cells using any method known in the art, including but not limited to electroporation. Alternatively, RGN polypeptides and / or polynucleotides encoding or containing guide RNA can be introduced into cells using any method known in the art (e.g., electroporation).

[0188] Upon delivery to or contact with a target nucleic acid molecule or a cell containing the target nucleic acid molecule, the guide RNA directs the RGN to bind to a target sequence within the target nucleic acid molecule in a sequence-specific manner. In those embodiments where the RGN has nuclease activity, the RGN polypeptide cleaves the target sequence upon binding. The target sequence can then be modified (i.e., edited) via endogenous repair mechanisms such as non-homologous end joining (NHEJ).

[0189] Methods for measuring the binding of RGN polypeptides to target sequences are known in the art and include chromatin immunoprecipitation assays, gel mobility shift assays, DNA pull-down assays, reporter assays, microplate capture, and detection assays. Similarly, methods for measuring the cleavage or modification of target nucleic acid molecules containing target sequences are known in the art and include in vitro or in vivo cleavage assays in which cleavage is confirmed by PCR, sequencing, or gel electrophoresis, with or without appropriate labeling (e.g., radioisotopes, fluorescent substances) attached to the target sequence to facilitate the detection of degradation products. Alternatively, Nicking-induced exponential amplification (NTEXPAR) assays can be used (see, e.g., Zhang et al. (2016) Chem. Sci. 7:4951-4957). In vivo cleavage can be evaluated using the Surveyor assay (Guschin et al. (2010) Methods Mol Biol 649:247-256).

[0190] This method may include the use of only one RGN and only one guide RNA. A single double-strand break within or near an expanded trinucleotide repeat has been shown to result in the loss or reduction of the length of the repeat region, possibly due to destabilization of the repeat tract (Richard et al., PLoS ONE (2014), 9(4):e95611; Mittelman et al., Proc Natl Acad Sci USA (2009), 106(24):9607-12; van Agtmaal et al. Mol Ther. 2017 Jan 4;25(1):24-43). In some embodiments, the RGN and its associated guide RNA may recognize the PAM generated by the SNP within the mutant HTT allele, resulting in the mutant HTT allele being cleaved by the RGN system, leading to a reduction in the levels of mutant HTT protein and / or mutant HTT mRNA.

[0191] This method may involve the use of a single type of RGN complexed with two or more guide RNAs. In some embodiments, the method involves the use of two types of RGNs, each complexed with a guide RNA. The two or more guide RNAs may target different regions of the mutant HTT allele. For example, the first guide RNA may target the 5' proximal to the expanded trinucleotide repeat within the HTT gene, and the second guide RNA may target the 3' proximal to the expanded trinucleotide repeat, enabling the excision of the expanded trinucleotide repeat.

[0192] Double-strand breaks introduced by RGN polypeptides can be repaired by the non-homologous end-joining (NHEJ) repair process. Due to the error-prone nature of NHEJs, repair of double-strand breaks can result in mutations in the target sequence. In certain embodiments, a “mutation” in a nucleic acid molecule refers to a change in the nucleotide sequence of the nucleic acid molecule, which can be the deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. In some embodiments, cleavage of the mutHTT allele results in the introduction and premature termination of indels (insertions and / or deletions), leading to a reduction in mutHTT mRNA and / or protein levels. In some embodiments, cleavage of the mutHTT allele results in the introduction of premature stop codons, leading to a reduction in mutHTT protein levels.

[0193] In some embodiments, a cell into which an RGN and / or guide RNA (or a polynucleotide(s) encoding at least one of the RGN and / or guide RNA) is introduced has a mutant HTT allele containing at least 27 CAG repeats in exon 1 (e.g., 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, or more than 36 CAG repeats). The cell may also contain a mutant HTT allele containing at least 36 CAG repeats in exon 1 (e.g., 36, 37, 38, 39, 40, or more than 40 CAG repeats). In other embodiments, the cell contains a mutant HTT allele containing at least 40 CAG repeats in exon 1 (e.g., 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, or more than 56 CAG repeats). The cell may also contain a mutant HTT allele containing at least 56 CAG repeats in exon 1 (e.g., 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or more than 70 CAG repeats).

[0194] This method may include the use of RGN polypeptides or RGN systems that cannot cleave the wild-type HTT allele. An RGN polypeptide or RGN system that cannot cleave the wild-type HTT allele means that it can either not cleave the wild-type HTT allele at all, or only cleaves it to a negligible level, resulting in no significant reduction in wtHTT mRNA and / or wtHTT protein levels, for example, wtHTT being able to maintain support for important cellular and neuronal functions, and / or the absence of Huntington's disease symptoms in an in vivo environment (i.e., in subjects heterozygous for the mutHTT allele and administered the RGN polypeptide or RGN system). In some embodiments, the RGN polypeptide or RGN system is cleaved to a negligible level such that the level of wtHTT mRNA and / or wtHTT protein is reduced to ≤5%, ≤4%, ≤3%, ≤2%, ≤1%, ≤0.9%, ≤0.8%, ≤0.7%, ≤0.6%, ≤0.5%, ≤0.4%, ≤0.3%, ≤0.2%, or ≤0.1% compared to the level of wtHTT mRNA and / or wtHTT protein in vitro or in vivo without the introduction of the RGN polypeptide or RGN system of the Disclosure.

[0195] IX. Cells containing polynucleotide gene modifications Provided herein are cells and organisms comprising a target sequence within a mutant HTT allele modified using processes mediated by RGN, crRNA, tracrRNA, and / or sgRNA as described herein. Cells comprising the modified target sequence within a mutant HTT allele may be Huntington's disease patient cells, stem cells, zygotes, embryonic cells, or gametes. In some embodiments, the stem cells are induced pluripotent stem cells (iPSCs) or mesenchymal stem cells (MSCs). In some embodiments, the cells are derived from induced pluripotent stem cells (iPSCs) or mesenchymal stem cells (MSCs). In some embodiments, the cells comprising the modified target sequence within a mutant HTT allele are in vitro or ex vivo. In some embodiments, the cells comprising the modified target sequence within a mutant HTT allele are in vitro. The modified cells (e.g., embryonic cells, zygotes, gametes) may be developed in an organism under appropriate conditions. RGN introduced into cells to modify the target sequence within the mutant HTT allele can recognize a consensus PAM sequence containing one of NNNNCC, NNRYA, NNGRR, and NNGG. In some embodiments, the target sequence within the mutant HTT allele has a nucleotide sequence shown as one of SEQ ID NOs. 75-79 and 130. The guide RNA in the system can be a single guide RNA or a dual guide RNA.

[0196] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 3. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 5, or an active variant or fragment thereof. In some embodiments, the guide RNA includes an sgRNA containing the nucleotide sequence shown as SEQ ID NO: 27 or 28, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC.

[0197] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 7. In some embodiments, the guide RNA contains a CRISPR repeat comprising the nucleotide sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof. In some embodiments, the guide RNA contains a tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof. In some embodiments, the guide RNA contains an sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 25 or 26, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA.

[0198] The RGN may contain an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 11. In some embodiments, the guide RNA contains a CRISPR repeat comprising the nucleotide sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof. In some embodiments, the guide RNA contains a tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof. In some embodiments, the guide RNA contains an sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 29, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM interaction domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0199] The RGN may include an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or more sequence identity with SEQ ID NO: 15. In some embodiments, the guide RNA includes a CRISPR repeat containing the nucleotide sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof. In some embodiments, the guide RNA includes a tracrRNA containing the nucleotide sequence shown as SEQ ID NO: 17, or an active variant or fragment thereof. In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, and recognizes the PAM sequence NNGG.

[0200] The RGN systems of this disclosure may include RGN polypeptides each comprising at least one nuclease domain, each involved in cleaving a single strand of a nucleic acid molecule. The nuclease domain may include a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0201] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence represented as one of SEQ ID NOs: 147, 148, 149, and 150.

[0202] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 151, 152, 153, and 154.

[0203] An RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence represented as any one of SEQ ID NOs: 155, 156, 157, and 158.

[0204] The RGN system of this disclosure may include an RGN polypeptide comprising a PAM interaction domain that contributes to the recognition and binding of a PAM site, and further comprising at least one nuclease domain, each of which is involved in the cleavage of a single strand of nucleic acid molecule. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNNNCC and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133, recognizing the PAM sequence NNNNCC, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

[0205] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNRYA and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134, recognizing the PAM sequence NNRYA, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150.

[0206] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135, and may further include a nuclease domain having an amino acid sequence that recognizes the PAM sequence NNGRR and has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135, recognizing the PAM sequence NNGRR, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 151, 152, 153, and 154.

[0207] An RGN polypeptide of the present disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may include a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136, recognizes the PAM sequence NNGG, and may further include a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may further comprise a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136, recognizing the PAM sequence NNGG, and having a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 155, 156, 157, and 158.

[0208] The modified cells may be eukaryotes (e.g., mammals). In some embodiments, cells modified by the methods of the present disclosure include stem cells (e.g., induced pluripotent stem cells, mesenchymal stem cells), nerve cells, and glial cells. Stem cells are totipotent, pluripotent, or multipotent cells that can differentiate into one or more different cell types. The term "totipotent" refers to the ability of a cell to differentiate into any type of cell in a differentiated organism, as well as into cells of extraembryonic material such as the placenta. "Pluripotent" refers to a cell line that can differentiate into any terminal cell type. The term "multipotent" refers to a cell line that can differentiate into at least two terminal cell types. The term “induced pluripotent stem cells” or “iPSCs” refers to a type of pluripotent stem cell that is similar to embryonic stem cells, formed by introducing certain embryonic genes (e.g., OCT4, SOX2, and KLF4 transgenes) (see, for example, Takahashi and Yamanaka Cell 126, 663-676 (2006), incorporated herein by reference) into somatic (e.g., adult) cells. Examples of somatic cells include, but are not limited to, bone marrow cells, epithelial cells, fibroblasts, hematopoietic cells, hepatocytes, intestinal cells, mesenchymal cells, bone marrow progenitor cells, nerve cells, glial cells, and spleen cells. Alternatively, iPSCs can be produced by reprogramming somatic cells to enter an embryonic stem cell-like state by forcing them to express factors important for maintaining the “stem cell nature” of embryonic stem cells (ESCs). Reprogramming factors may be expressed from an expression cassette contained in one or more vectors, such as an embedded vector, a chromosomally non-embedded RNA virus vector, or an episomal vector such as the EBV element system (Yu et al. (2009) Science, 324(5928):797-801).In some embodiments, reprogramming proteins or RNA (such as mRNA or miRNA) can be directly introduced into somatic cells by protein or RNA transfection (Yakubov et al. (2010) Biochemical and biophysical research communications, 394(1):189-193).

[0209] Mesenchymal stem cells (MSCs) can give rise to connective tissue, bone, cartilage, and cells in the circulatory and lymphatic systems. MSCs are found in the mesenchyme, which is a portion of the embryonic mesoderm containing loosely packed, spindle-shaped or star-shaped unspecific cells. In some embodiments, MSCs are CD34 - MSCs include stem cells. MSCs can be isolated from a variety of sources, including bone marrow, umbilical cord blood, (mobilized) peripheral blood, and adipose tissue (Horwitz et al. Clarification of the nomenclature for MSC: the International Society for Cellular Therapy position statement. Cytotherapy (2005) 7:393-395).

[0210] Stem cells (e.g., iPSCs or MSCs) containing target sequences within mutated HTT alleles modified by the described RGN system can be differentiated into neurons or glial cells. The term "differentiating cells" refers to changing a default cell type (genotype and / or phenotype) to a non-default cell type (genotype and / or phenotype). For example, differentiating stem cells (e.g., iPSCs or MSCs) refers to inducing the stem cells to divide into progeny cells that have different characteristics from the stem cells (e.g., iPSCs or MSCs), such as genotype (i.e., changes in gene expression determined by genetic analysis such as microarrays) and / or phenotype (i.e., changes in protein expression). One or more small molecules, growth factor proteins, and other growth conditions can be used to facilitate the transition from a non-specification state, e.g., a stem cell state, to a more specific cell fate (e.g., a neuron). In some embodiments, stem cell differentiation directs the stem cells towards cellular pathways leading to somatic cells. For example, factors that differentiate stem cells into nerve cells may include Wnt activator, SMAD inhibitors (e.g., Noggin peptide, SB-431542), nerve growth factors (e.g., brain-derived neurotrophic factor, nerve growth factor, glial-derived neurotrophic factor), and / or the introduction of polynucleotides to express nerve genes (e.g., neurogenin-2, NeuroD1). Differentiation may involve culturing pluripotent stem cells and / or their progeny cells in adherent or suspension culture.

[0211] Neurons that can be modified by processes utilizing RGN polypeptides, crRNAs, tracrRNAs, and / or guide RNAs as described herein may include neural progenitor cells, forebrain progenitor cells, striatal neurons, medium spiny neurons, and cortical neurons. Non-neuronal brain cells that can be modified by processes utilizing RGN polypeptides, crRNAs, tracrRNAs, and / or guide RNAs as described herein include glial cells. Glial cells may include microglia, astrocytes, and oligodendrocytes. In some embodiments, cells useful in this disclosure include mammalian or human cells present in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or combinations thereof.

[0212] Also provided herein are embryonic cells, zygotes, or gametes containing a mutant HTT allele modified by a process utilizing RGN, crRNA, tracrRNA, and / or sgRNA as described herein. The modified cells and organisms may be heterozygous or homozygous with respect to the modified mutant HTT allele. In some embodiments, the modified cells and organisms are heterozygous with respect to the modified mutant HTT allele.

[0213] Chromosomal modifications in cells containing a target sequence within a mutant HTT allele having the RGN system of this disclosure can result in downregulation of the expression of mutant HTT protein and / or mutant HTT mRNA. In some embodiments, chromosomal modifications result in a reduction or removal of mutant HTT mRNA compared to the level of mutant HTT mRNA in cells not chromosomally modified with the RGN system. In some embodiments, chromosomal modifications result in a reduction or removal of mutant HTT protein compared to the level of mutant HTT protein in cells not chromosomally modified with the RGN system. Mutant HTT protein levels can be measured by assays including immunoassays using antibodies capable of distinguishing between wild-type HTT protein and mutant HTT protein (e.g., Western blotting, ELISA, single-molecule count immunoassay, immunoprecipitation assay combined with flow cytometry, time-resolved fluorescence energy transfer (TR-FRET)), and examples include the JESS capillary Western blotting assay described herein. Mutant HTT mRNA levels can be measured, for example, by RT-qPCR or array-based methods.

[0214] X. Pharmaceutical Compositions A pharmaceutical composition is provided comprising: crRNA and its active variants and fragments or encoding polynucleotides as disclosed herein; tracrRNA and its active variants and fragments or encoding polynucleotides as disclosed herein; sgRNA and its active variants and fragments or encoding polynucleotides as disclosed herein; RGN polypeptide and its active variants and fragments or encoding polynucleotides as disclosed herein; RGN system as disclosed herein; or RNP complex as disclosed herein comprising RGN polypeptide and gRNA; or vector as disclosed herein (e.g., viral vector); and a pharmaceutically acceptable carrier.

[0215] A pharmaceutical composition is a composition used to prevent, reduce the intensity of, cure, or otherwise treat a target condition or disease, comprising an active ingredient (i.e., an RGN polypeptide, an RGN-coding polynucleotide, a gRNA, a gRNA-coding polynucleotide, an RGN system, an RNP complex, or a vector) and a pharmaceutically acceptable carrier.

[0216] As used herein, “pharmaceutically acceptable carrier” means a material that does not cause significant irritation to the organism and does not inhibit the activity and properties of the active ingredient (i.e., RGN polypeptide, RGN-coding polynucleotide, gRNA, gRNA-coding polynucleotide, RGN system, RNP complex, or vector). The carrier must be sufficiently pure and sufficiently low in toxicity to be suitable for administration to the subject being treated. The carrier may be inactive or may retain pharmaceutically beneficial properties. In some embodiments, the pharmaceutically acceptable carrier comprises one or more suitable solid or liquid fillers, diluents, or encapsulating materials suitable for administration to humans or other vertebrates. In some embodiments, the pharmaceutically acceptable carrier does not exist in nature. In some embodiments, the pharmaceutically acceptable carrier and the active ingredient are not found together in nature.

[0217] The pharmaceutical compositions used in the methods of this disclosure can be formulated with suitable carriers, excipients, and other agents that provide suitable transport, delivery, tolerance, etc. Numerous suitable formulations are known to those skilled in the art. See, for example, Remington, The Science and Practice of Pharmacy (21st ed. 2005). Suitable formulations include, for example, powders, pastes, ointments, jellies, waxes, oils, lipids, lipid (cationic or anionic)-containing vesicles (such as LIPOFECTIN vesicles), lipid nanoparticles, DNA complexes, anhydrous absorbent pastes, oil-in-water emulsions and water-in-oil emulsions, emulsion carbowaxes (polyethylene glycol of various molecular weights), semi-solid gels, and semi-solid mixtures containing carbowaxes. Pharmaceutical compositions for oral or parenteral use may be prepared into dosage forms in unit doses suitable to accommodate the dose of the active ingredient. Such dosage forms in unit doses include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc.

[0218] A pharmaceutical composition containing an active ingredient (i.e., an RGN polypeptide, an RGN-coding polynucleotide, a gRNA, a gRNA-coding polynucleotide, an RGN system, or an RNP complex, or a vector) can be emulsified or presented as a liposome composition, provided that the emulsification procedure does not adversely affect the active ingredient or the patient.

[0219] Additional agents included in a pharmaceutical composition may include pharmaceutically acceptable salts. Examples of pharmaceutically acceptable salts include acid addition salts (formed using the free amino group of a polypeptide) formed using inorganic acids such as hydrochloric acid or phosphoric acid, or organic acids such as acetic acid, tartaric acid, or mandelic acid. Salts formed using the free carboxyl group may also be derived from inorganic bases such as sodium, potassium, ammonium, calcium, or ferric hydroxide, and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, or procaine.

[0220] Physiologically and pharmaceutically acceptable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions containing no materials in addition to the active ingredient and water, or containing a buffer such as sodium phosphate, physiological saline, or both at a physiological pH value, e.g., phosphate-buffered saline. Furthermore, aqueous carriers may contain two or more buffer salts, as well as salts such as sodium chloride and potassium chloride, dextrose, sucrose, mannose, polyethylene glycol, and other solutes. Liquid compositions may also contain liquid phases, both in addition to and excluding water. Examples of such additional liquid phases include glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of active compound used in a cell composition effective for treating a particular disorder or condition may depend on the nature of the disorder or condition and can be determined by standard clinical techniques.

[0221] In some embodiments, the pharmaceutical composition comprises one or more molecules having surfactant properties that allow them to interact with biological membranes, e.g., PLURONIC® or poloxamer, e.g., PLURONIC® F68 (poloxamer 188, P188). In some embodiments, the pharmaceutical composition comprises 0.001% to 0.1% poloxamer. In some embodiments, the pharmaceutical composition comprises 0.001% poloxamer.

[0222] The RGN polypeptides, guide RNAs, RGN systems, polynucleotides encoding them, RNP complexes, or vectors of this disclosure can be formulated with pharmaceutically acceptable excipients such as carriers, solvents, stabilizers, adjuvants, and diluents, depending on the specific mode of administration and dosage form. In some embodiments, these pharmaceutical compositions are formulated to achieve a physiologically compatible pH, ranging from about pH 3 to about pH 11 and from about pH 3 to about pH 7, depending on the formulation and route of administration. In some embodiments, the buffer can be adjusted to a pH range of about pH 5.0 to about pH 8. In some embodiments, the composition may contain a therapeutically effective amount of at least one active ingredient described herein (i.e., RGN polypeptide, RGN-coding polynucleotide, gRNA, gRNA-coding polynucleotide, RGN system, RNP complex, or vector) together with one or more pharmaceutically acceptable excipients. In some embodiments, the composition includes a combination of the active ingredients described herein, or a second active ingredient useful for treating or preventing bacterial growth (e.g., an antimicrobial or antibacterial agent), or a combination of the reagents of the Disclosure.

[0223] Suitable excipients include, for example, carrier molecules containing large, slowly metabolized polymers such as proteins, polysaccharides, polylactic acid, polyglycolic acid, high molecular weight amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients include antioxidants (e.g., but not limited to ascorbic acid), chelating agents (e.g., but not limited to EDTA), carbohydrates (e.g., but not limited to dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (e.g., but not limited to oil, water, saline, glycerol, and ethanol), wetting agents or emulsifiers, and pH buffering agents.

[0224] In some embodiments, the formulation is provided in a unit-dose or multi-dose container, such as a sealed ampoule and vial, and stored under freeze-dried (lyophilized) conditions that require the addition of a sterile liquid carrier, such as saline, water for injection, semi-liquid form, or gel, immediately prior to use. The temporary injection solutions and suspensions may be prepared from sterile powders, granules and tablets of the aforementioned types. In some embodiments, the active ingredient is frozen in a unit-dose or multi-dose container and later thawed for injection or dissolved in a buffered liquid solution that is held / stabilized under refrigeration until use.

[0225] The therapeutic agent(s) may be included in a controlled release system. It is often desirable to delay the absorption of a drug from a subcutaneous, intrathecal or intramuscular injection in order to extend the effect of the drug. This can be achieved by using a crystalline or amorphous liquid suspension that is poorly water-soluble. The absorption rate of the drug then depends on its dissolution rate, which in turn can depend on the crystal size and crystal form. Alternatively, the delayed absorption of a parenterally administered drug form is achieved by dissolving or suspending the drug in an oily vehicle. In some embodiments, the use of a long-term sustained release implant may be particularly suitable for the treatment of chronic diseases. Long-term sustained release implants are well known to those skilled in the art.

[0226] Therapeutic agents (plural), such as the crRNA and its active variants and fragments, or encoding polynucleotides, as disclosed herein; the tracrRNA and its active variants and fragments, or encoding polynucleotides, as disclosed herein; the sgRNA and its active variants and fragments, or encoding polynucleotides, as disclosed herein; the RGN polypeptide and its active variants and fragments, or encoding polynucleotides, as disclosed herein; the RGN system, as disclosed herein; or the RNP complex, as disclosed herein, comprising the RGN polypeptide and gRNA; or the vector, as disclosed herein (e.g., viral vector), may be isolated, or substantially or essentially absent, from components that normally accompany or interact with the therapeutic agent in its naturally occurring environment, or from chemical precursors or other chemicals used to synthesize the therapeutic agent, or by recombinant technology, or from media produced by endotoxins and / or associated pyrogens. Endotoxins include toxins trapped inside microorganisms and released only upon degradation or death of the microorganism. Pyromogens also include pyrogen-induced thermostable substances (glycoproteins) from the outer membranes of bacteria and other microorganisms. Both substances can cause fever, hypotension, and shock when administered to humans. Because of the potential for adverse effects, even small amounts of endotoxin must be removed from intravenously administered drug solutions. The U.S. Food and Drug Administration ("FDA") has set an upper limit of 5 endotoxin units (EU) / dose / kg body weight per hour for intravenous drug use (The United States Pharmaceutical Convention, pharmaceutical form 26(1):223(2000)). In certain specific embodiments, the levels of endotoxin and pyrogens in the composition are less than about 1 EU / mg, or less than about 0.1 EU / mg, or less than about 0.01 EU / mg, or less than about 0.001 EU / mg. In some embodiments, the levels of endotoxin and pyrogens in the composition are 0.0138 EU / mg or less.

[0227] The therapeutic agent or pharmaceutical composition can have a purity of at least 80%, 85%, 90%, 95%, or more. The therapeutic agent or pharmaceutical composition can have low levels or undetectable levels of endotoxins or other impurities.

[0228] The pharmaceutical composition may be stored frozen, refrigerated, or at room temperature. The storage conditions may be below freezing, for example, less than about -10°C, or less than about -20°C, or less than about -40°C, or less than about -70°C. The storage conditions are generally less than room temperature, for example, less than about 32°C, or less than about 30°C, or less than about 27°C, or less than about 25°C, or less than about 20°C, or less than about 15°C. In some embodiments, the formulation is stored at 2°C to 8°C. For example, the formulation may be isotonic with blood or may have an ionic strength that mimics physiological conditions.

[0229] In some embodiments, the pharmaceutical composition is stable under the storage conditions. The stability can be measured using any suitable means in the art. Generally, a stable formulation shows an increase of less than 5% of degradation products or impurities. In some embodiments, the formulation is stable under storage conditions for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, at least about 6 months, at least about 1 year, or at least about 2 years or more. In some embodiments, the formulation is stable at 25°C for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, at least about 6 months, or at least about 1 year or more.

[0230] When used for in vivo administration, the pharmaceutical compositions of this disclosure should be sterile. The formulations of this disclosure may be sterilized by various sterilization methods, including sterile filtration and irradiation. In one embodiment, the formulation is sterilized by filtration through a pre-sterilized 0.22-micron filter. The sterile compositions for injection may be formulated in accordance with conventional pharmaceutical practices, such as those described in "Remington: The Science & Practice of Pharmacy", 21st edition, Lippincott Williams & Wilkins, (2005).

[0231] XI. How to treat Huntington's disease A method for treating Huntington's disease (HD) in subjects requiring treatment is provided herein. The method comprises administering to a subject requiring treatment a pharmaceutical composition comprising the crRNA or polynucleotide encoding

[0232] In some embodiments, the treatment includes in vivo gene editing by administering the RGN system of the Disclosure, the polynucleotide(s) encoding it, the RNP complex of the Disclosure, or the vector of the Disclosure. In some embodiments, the treatment includes ex vivo gene editing of zygotes, embryonic cells, or gametes to correct a genetic error at an early stage of life. In some embodiments, a therapeutic composition comprising the RGN system, the polynucleotide(s) encoding it, the RNP complex, or the vector of the Disclosure targets target cells in vivo. In some embodiments, cells targeted for gene editing of mutated HTT alleles include stem cells, neurons (e.g., medium spiny neurons, cortical neurons), and glial cells (e.g., astrocytes, oligodendrocytes, and microglia).

[0233] Huntington's disease (HD) is a fatal, monogenic neurodegenerative disorder characterized by progressive chorea (involuntary movements), neuropsychiatric dysfunction, and cognitive impairment. Symptoms typically appear between the ages of 35 and 44, and the average life expectancy after onset is 10 to 25 years.

[0234] HD is known to be caused by an expansion of the triplet cytosine-adenine-guanine (CAG) repeat at the end of exon 1 of the huntingtin (HTT) gene. The CAG repeat encodes polyglutamine at the N-terminus of the huntingtin (HTT) protein. Normal HTT alleles contain 15–20 CAG repeats, while alleles containing 27–35 CAG repeats are considered intermediate alleles with little likelihood of developing the disease phenotype. HTT alleles containing 35 or more CAG repeats can be considered potentially HD-causing alleles, posing a risk of developing the disease. Alleles containing 36–39 CAG repeats are considered to have incomplete penetrance, and individuals carrying these alleles may or may not develop the disease (or may develop symptoms later in life), while alleles containing 40 or more CAG repeats are considered to have complete penetrance. A Huntington's disease integrated staging system (HD-ISS) has been developed to define the disease stages from birth to death and classify individuals with HD considering clinical criteria, clinical biomarkers, and functional assessments (Tabrizi et al. Lancet Neurol 2022;21:632-644).

[0235] Juvenile-onset hemorrhagic disorder (JHD) is a form of hemorrhagic disorder (HD) that affects children and teenagers. Individuals with JHD (under 21 years of age) are often found to have 60 or more CAG repeats. Symptoms of JHD include changes in personality, coordination, behavior, speech, or cognitive abilities. Physical changes also occur, including rigidity, leg stiffness, clumsiness, slow movement, tremors, or myoclonus. In contrast to adult HD, seizures and rigidity are common, while chorea is rare. JHD progresses more rapidly than adult HD, and death can occur within 10 years of onset.

[0236] The mutated HTT allele is usually inherited from one parent as a dominant trait. Children born to HD patients have a 50% chance of developing the disease if the other parent is not affected. In some cases, parents may have an intermediate HD allele and be asymptomatic, while in others, due to repeat expansion, the child may develop the disease. In addition, the HD allele can also exhibit a phenomenon known as prediction, where increased severity or a decrease in age of onset is observed over several generations, due to the unstable nature of the repeat region during spermatogenesis.

[0237] The expansion of this repeat leads to the formation of mutated HTT proteins that can form aggregates within cells, disrupt normal cellular function, and / or have pathological interactions with other molecules. Ultimately, the presence of mutated HTT proteins causes striatal neurodegeneration that progresses to widespread brain atrophy.

[0238] Trinucleotide expansion in the HTT gene causes neuronal loss in medium spiny gamma-aminobutyric acid (GABA) process neurons in the striatum, and also in the neocortex. In some embodiments, medium spiny neurons (MSNs) containing enkephalin and projecting into the lateral globus pallidus, and / or MSNs containing substance P and projecting into the internal globus pallidus are affected. Other brain regions significantly affected in individuals with HD include the substantia nigra, cortical laminas 3, 5, and 6, the CA1 region of the hippocampus, the angular gyrus of the parietal lobe, Purkinje cells of the cerebellum, the lateral nucleus prominens of the hypothalamus, and the paracentral fasciculus complex of the thalamus (Walker (2007) Lancet 369:218-228). Currently, there is no therapeutic treatment for HD, but experimental approaches based on drugs, cell therapy, and gene therapy are being explored.

[0239] RGN systems, polynucleotides encoding components of RGN systems, RNP complexes, vectors, or compositions comprising any of these are useful for modifying mutant HTT alleles in vivo in HD patient cells. In some embodiments, modifying a mutant HTT allele involves modifying a target sequence within exon 50 of the HTT gene. For example, the APG07433.1 RGN is used with a suitable guide RNA selected from SEQ ID NOs. 27 and 28 to modify the target sequence within exon 50 of the mutant HTT allele. Another example is the APG05586 RGN, used with a suitable guide RNA selected from either SEQ ID NOs. 25 and 26 to modify the target sequence within exon 50 of the mutant HTT allele. Yet another example is the APG01604 RGN, used with a suitable guide RNA having the nucleotide sequence shown as SEQ ID NO. 29 to modify the target sequence within exon 50 of the mutant HTT allele.

[0240] Modification of the target sequence in the mutant HTT allele involves cleaving the mutant HTT allele in vivo in HD patient cells. In some embodiments, the cleavage is within exon 50 of the mutant HTT allele. In some embodiments, after cleavage of the target sequence, non-homologous end joining (NHEJ) occurs, leading to nucleotide insertions and / or deletions (indels) at the cleavage site, as well as disruption of the mutant HTT coding sequence, resulting in a reduction in the levels of mutant HTT mRNA and / or protein. In some embodiments, cleavage of the mutHTT allele results in the introduction of an early stop codon, leading to a reduction in mutHTT protein levels.

[0241] As used herein, the term “subject” refers to any individual for whom diagnosis, treatment, or therapy is desired. In some embodiments, the subject is an animal. In some embodiments, the subject is a mammal. In some embodiments, the subject is a human.

[0242] The methods for treating HD in this disclosure utilize SNPs occurring in the human genome. SNP alleles that generate PAM sites recognized by the RGN system and RNP complex of this disclosure have been identified. If a mutant HTT allele contains a SNP allele that generates a PAM recognized by the RGN system of this disclosure, these RGN systems can be used to cleave the mutant HTT allele, resulting in a reduction in the levels of mutant HTT mRNA and / or protein. In some embodiments, the mutant HTT allele contains a PAM containing a SNP allele. Therefore, subjects that are suitable for treatment with the compositions described herein (e.g., RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN system, or RNP complex) can be assayed before treatment to determine whether the mutant HTT allele contains a SNP allele that generates a PAM recognized by the RGN system described herein. In some embodiments, subjects that are suitable for treatment with the RGN system described herein further have a wild-type HTT allele that does not contain a SNP allele that generates a PAM and is therefore heterozygous for the SNP. Since it is desirable to have a wild-type HTT allele that does not contain the SNP allele that produces PAM, the administered composition cleaves only the mutant HTT allele containing the SNP allele that produces PAM, and does not cleave the wild-type HTT allele. Therefore, in some embodiments, PAM is present only on the mutant HTT allele and not on the wild-type HTT allele, and the treatment is considered allele-specific. The method may include the use of an RGN system, a polynucleotide coding component of the RGN system, an RNP complex, a vector, or a composition containing any of these, which cannot cleave the wild-type HTT allele.In some embodiments, RGN systems, polynucleotide coding components of RGN systems, RNP complexes, vectors, or compositions comprising any of these, which are unable to cleave the wild-type HTT allele, are unable to cleave the wild-type HTT allele at all, or only negligibly so, as a result, the levels of wtHTT mRNA and / or wtHTT protein are not significantly reduced, for example, wtHTT can maintain support for important cellular and neuronal functions, and / or symptoms of Huntington's disease are absent in an in vivo environment (i.e., in subjects heterozygous for the mutHTT allele and administered the RGN system).

[0243] Determining whether a subject has a mutant HTT allele containing a SNP allele that produces a PAM recognized by the RGN system described herein, and / or whether the SNP is heterozygous, can be achieved by sequencing, array-based hybridization, and / or PCR-based methods (e.g., long-range PCR) performed on a biological sample obtained from the subject. In some embodiments, the biological sample may include blood, cells, and / or cerebrospinal fluid. In some embodiments, the mutant HTT allele contains a PAM containing a SNP allele.

[0244] The composition administered to a subject requiring it includes an RGN capable of recognizing PAMs produced (i.e., including) by a SNP allele (e.g., within exon 50 of the mutHTT allele) that may contain the NNNNCC, NNRYA, NNGRR, and NNGG PAM sequences. The SNP that produces the PAMs may be located within exon 50 of the mutHTT allele. In some embodiments, the SNP is a thymine at the position corresponding to position 151 of SEQ ID NO: 1, which produces the PAM sequence NNRYA. In some embodiments, the NNRYA PAM sequence includes a SNP that is a thymine at the position corresponding to position 151 of SEQ ID NO: 1. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that includes an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 7. In some embodiments, the NNRYA PAM sequence is recognized by an RGN containing the amino acid sequence of SEQ ID NO: 7. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA containing the nucleotide sequence of SEQ ID NO: 8 or 106, or a crRNA repeat having a nucleotide sequence that differs from SEQ ID NO: 8 or 106 by one or two nucleotides. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA containing a tracrRNA having a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 9 or 107. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA containing a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA containing a spacer having a nucleotide sequence complementary to the target sequence having the nucleotide sequence of SEQ ID NO: 75 or 76. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 80 or 81, or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by one or two nucleotides. In some embodiments, the guide RNA is a single guide RNA. In some embodiments, the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

[0245] The SNP that generates the PAM may be located in exon 50 of the mutant HTT allele. In some embodiments, the SNP is a cytosine at position 151 of SEQ ID NO: 2, which generates the PAM sequences NNNNCC, NNGRR, and NNGG. In some embodiments, the PAM sequences NNNNCC, NNGRR, and NNGG include a SNP that is a cytosine at position 151 of SEQ ID NO: 2. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that includes an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 3. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that includes the amino acid sequence of SEQ ID NO: 3. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA containing a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 4. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA containing a tracrRNA having a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 5. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA containing a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 5. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA containing a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 77 or 78. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 82 or 83, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 82 or 83. In some embodiments, the guide RNA is a single guide RNA.In some embodiments, the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

[0246] In some embodiments, the NNGRR PAM sequence is recognized by an RGN containing an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 11. In some embodiments, the NNGRR PAM sequence is recognized by an RGN containing an amino acid sequence of SEQ ID NO: 11. In some embodiments, the NNGRR PAM sequence is recognized by an RGN that binds to a guide RNA containing the nucleotide sequence of SEQ ID NO: 12, or a crRNA repeat having a nucleotide sequence that differs from SEQ ID NO: 12 by one or two nucleotides. In some embodiments, the NNGRR PAM sequence is recognized by an RGN that binds to a guide RNA containing a tracrRNA having a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 13 or 120. In some embodiments, the NNGRR PAM sequence is recognized by an RGN that binds to a guide RNA containing a crRNA repeat having the nucleotide sequence of SEQ ID NO: 12 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 13 or 120. In some embodiments, the NNGRR PAM sequence is recognized by an RGN that binds to a guide RNA containing a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 79. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 84, or a nucleotide sequence that differs from SEQ ID NO: 84 by one or two nucleotides. In some embodiments, the guide RNA is a single guide RNA. In some embodiments, the single guide RNA has the nucleotide sequence of SEQ ID NO: 29.

[0247] In some embodiments, the NNGG PAM sequence is recognized by an RGN containing an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 15. In some embodiments, the NNGG PAM sequence is recognized by an RGN containing an amino acid sequence of SEQ ID NO: 15. In some embodiments, the NNGG PAM sequence is recognized by an RGN that binds to a guide RNA containing the nucleotide sequence of SEQ ID NO: 16, or a crRNA repeat having a nucleotide sequence that differs from SEQ ID NO: 16 by one or two nucleotides. In some embodiments, the NNGG PAM sequence is recognized by an RGN that binds to a guide RNA containing a tracrRNA having a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more sequence identity with SEQ ID NO: 17. In some embodiments, the NNGG PAM sequence is recognized by an RGN that binds to a guide RNA containing a crRNA repeat having the nucleotide sequence of SEQ ID NO: 16 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 17. In some embodiments, the guide RNA is a single guide RNA.

[0248] As used herein, “treatment” or “treating” refers to an approach to obtain beneficial or desired outcomes, including but not limited to therapeutic and / or preventive benefits. Therapeutic benefit means any therapeutically relevant improvement or effect in one or more diseases, conditions, or symptoms during treatment. For preventive benefits, the composition may be administered to subjects at risk of developing certain diseases, conditions, or symptoms, or subjects reporting one or more physiological symptoms of the disease or biomarkers of the disease (e.g., changes in neuronal phenotype), even if the disease, condition, or symptoms have not yet manifested. In some embodiments, subjects at risk of developing HD are subjects containing a mutant HTT allele as defined herein. In some embodiments, subjects at risk of developing HD are defined as subjects having a mutant HTT allele containing more than 35 CAG repeats. In some embodiments, subjects at risk of developing HD are defined as subjects having a mutant HTT allele containing at least 40 CAG repeats. In some embodiments, subjects at risk of developing HD are defined as subjects having a mutant HTT allele containing more than 56 CAG repeats. In some embodiments, treatment may be administered after the onset of one or more symptoms and / or after the disease has been diagnosed. In some embodiments, treatment may be administered asymptomatically, for example, to prevent or delay the onset of symptoms, or to inhibit the onset or progression of the disease. For example, treatment may be administered to susceptible individuals before the onset of symptoms (for example, in light of a history of symptoms and / or genetic or other susceptibility factors). Treatment may also be continued after the recovery of symptoms, for example, to prevent or delay their recurrence. In some embodiments, the number of CAG repeats influences the decision of when to treat a subject. In some embodiments, subjects having a mutant HTT allele containing more than 56 CAG repeats are treated as adolescents or young adults before the onset of any HD symptoms or the presentation of physiological biomarkers.

[0249] In some embodiments, the compositions and methods of the present disclosure are used to effect (i.e., reduce) or delay the onset of one or more symptoms of Huntington's disease in a subject that requires doing so. The symptoms may be reduced by about 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95% and 100%, or at least 20 - 30%, 20 - 40%, 20 - 50%, 20 - 60%, 20 - 70%, 20 - 80%, 20 - 90%, 20 - 95%, 20 - 100%, 30 - 40%, 30 - 50%, 30 - 60%, 30 - 70%, 30 - 80%, 30 - 90%, 30 - 95%, 30 - 100%, 40 - 50%, 40 - 60%, 40 - 70%, 40 - 80%, 40 - 90%, 40 - 95%, 40 - 100%, 50 - 60%, 50 - 70%, 50 - 80%, 50 - 90%, 50 - 95%, 50 - 100%, 60 - 70%, 60 - 80%, 60 - 90%, 60 - 95%, 60 - 100%, 70 - 80%, 70 - 90%, 70 - 95%, 70 - 100%, 80 - 90%, 80 - 95%, 80 - 100%, 90 - 95%, 90 - 100% or 95 - 100% compared to a control value or a control subject (e.g., one not receiving a therapeutic agent). The onset of one or more symptoms of Huntington's disease may be delayed by about 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 12 years, 15 years, 20 years, 25 years, 30 years or more compared to a control value or a control subject (e.g., one not receiving a therapeutic agent). In some embodiments, the compositions and methods of the present disclosure can prevent the occurrence of one or more symptoms of Huntington's disease in a subject.

[0250] In some embodiments, the compositions and methods of the present disclosure are used to improve (i.e., reduce) or delay the onset of one or more biomarkers of Huntington's disease in subjects that require such improvement. The biomarkers are increased by approximately 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, and 100% compared to a control value or a control subject (e.g., one not treated with the therapy), or by at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, and 30-1 It may be reduced by 00%, 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100%, or 95-100%. The onset of one or more Huntington's disease biomarkers may be delayed by approximately 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 12 years, 15 years, 20 years, 25 years, 30 years or more compared to a control value or a control subject (e.g., one not treated with the drug). In some embodiments, the compositions and methods of the present disclosure can prevent the onset of one or more Huntington's disease biomarkers in a subject.

[0251] Subjects at risk of developing HD or who have HD may be identified by a variety of means, including cognitive assessments and / or neurological or neuropsychiatric examinations, motor tests, sensory tests, psychiatric assessments, brain imaging, family history, and / or genetic testing. Subjects requiring this may have symptoms of HD, be diagnosed with HD, and / or be asymptomatic for HD.

[0252] The compositions of this disclosure (e.g., including RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects staging using the Huntington's Disease Integrated Staging System (HD-ISS) (Tabrizi et al. Lancet Neurol 2022;21:632-644, the contents of which are incorporated in their entirety by reference). In this system, HD subjects in stage 0 have 40 or more CAG repeats. HD subjects in stage 1 have 40 or more CAG repeats and possess pathogenic biomarkers (e.g., putamen volume and / or caudate nucleus volume). HD subjects in stage 2 have 40 or more CAG repeats and possess pathogenic biomarkers, as well as signs or symptoms (e.g., as measured by the Total Motor Score (TMS) and / or Symbol Digit Modality Test (SDMT)). Stage 3 HD subjects have 40 or more CAG repeats and present with etiological biomarkers, signs or symptoms, and functional changes (e.g., as measured by independence scales and / or total functional capacity (TFC)).

[0253] The compositions of this disclosure (e.g., including RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects diagnosed using a prognostic index for Huntington's disease or its derivatives (Long JD et al., Movement Disorders, 2017, 32(2), 256-263, the contents of which are incorporated in their entirety by reference). This prognostic index predicts the probability of motor diagnosis using the following four components: (1) Total Motor Score (TMS) from the Unified Huntington's Disease Rating Scale (UHDRS), (2) Symbolic Digit Modality Test (SDMT), (3) Baseline Age, and (4) CAG Expansion. In some embodiments, the prognostic index for HD is given by the following formula PI HD = 51 × TMS + (-34) × SDMT + 7 × age × (CAG - 34), and in the formula, PI HDA higher value indicates a greater risk of diagnosis or symptom onset. In some embodiments, the prognostic index for HD is given by the following normalized formula PIN, which gives the units of standard deviation to be interpreted in the context of a 50% 10-year survival rate. HD =( PI HD It is calculated as -883) / 1044, and in the formula, PI HD <0 indicates a 10-year survival rate of over 50%, and PIN HD A score of >0 suggests a 10-year survival rate of less than 50%. In some embodiments, prognostic indicators may be used to identify subjects who will develop symptoms of HD within a few years but who do not yet have clinically diagnosable symptoms. Furthermore, these asymptomatic patients may be selected for and treated with the compositions of the present disclosure during the asymptomatic period.

[0254] The compositions of this disclosure (including, for example, RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects undergoing biomarker evaluation. Potential biomarkers in the blood for HD include, but are not limited to, 8-hydroxy-2-deoxyguanosine (8-OhdG) oxidative stress markers, metabolic markers (e.g., creatine kinase, branched-chain amino acids), cholesterol metabolites (e.g., 24-OH cholesterol), immune and inflammatory proteins (e.g., clatherin, complement components, interleukin 6 and 8), gene expression changes (e.g., transcriptome markers), endocrine markers (e.g., cortisol, ghrelin and leptin), brain-derived neurotrophic factor (BDNF), and adenosine 2A receptor. Potential biomarkers for brain imaging of HD include, but are not limited to, striatal volume, putamen volume, caudate nucleus volume, subcortical white matter volume, cortical thickness, whole brain volume, and ventricular volume. Brain imaging can be performed using functional imaging (e.g., functional MRI), positron emission tomography (PET) (e.g., with fluorodeoxyglucose), and magnetic resonance spectroscopy (e.g., lactate). Potential biomarkers for quantitative clinical tools for HD include, but are not limited to, quantitative motor assessments, exercise physiological assessments (e.g., transcranial magnetic stimulation), and quantitative eye movement measurements. Non-limited examples of quantitative clinical biomarker assessments include tongue force variability, metronome-guided tapping, grip strength, eye movement assessments, and cognitive tests. Non-limited examples of multicenter observational studies include PREDICT-HD and TRACK-HD. In some embodiments, the biomarker for HD is the level of wild-type huntingtin (HTT) mRNA and / or protein. In some embodiments, the biomarker for HD is the level of mutant huntingtin (HTT) mRNA and / or protein. In some embodiments, the biomarker for HD is the level of neuronal filament light chain (NFL) protein.

[0255] The compositions of this disclosure (e.g., RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects who are asymptomatic to HD. Subjects may be asymptomatic but may undergo predictive genetic testing or biomarker evaluation to determine whether they are at risk for HD, and / or subjects may have family members (e.g., mother, father, brother, sister, aunt, uncle, grandparent) who have been diagnosed with HD. In some embodiments, subjects who are asymptomatic to HD have a mutant HTT allele containing 27–35 CAG repeats (e.g., 27, 28, 29, 30, 31, 32, 33, 34, and 35 CAG repeats).

[0256] The compositions of this disclosure (e.g., RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects in the early stages of HD. In the early stages, subjects may have subtle changes in coordination, some chorea, mood changes such as irritability and depression, difficulty in problem-solving, and / or a decline in the subject's ability to perform normal daily activities.

[0257] The compositions of this disclosure (e.g., RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects in the intermediate stage of HD. In the intermediate stage, subjects experience increased motor impairment, decreased speech, dysphagia, and greater difficulty in normal activities. At this stage, subjects may have occupational and physical therapists to help maintain control of spontaneous movement, and subjects may have speech and speech pathologists.

[0258] The compositions of this disclosure (e.g., RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects in the later stages of hemodialysis (HD). In the later stages, subjects with HD are no longer able to walk or speak, and are almost completely or completely dependent on others for care. Subjects generally still understand language and recognize family and friends, but choking is a major concern.

[0259] The compositions of the present disclosure (including, for example, RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects having juvenile HD, which develops before the age of 21, or subjects susceptible to developing juvenile HD, for example, subjects having at least 56 CAG repeats in exon 1 of the HTT gene (e.g., 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, or more than 75 CAG repeats). In some of these embodiments, subjects are administered the compositions of the present disclosure before the age of 21, or in some embodiments, before the age of 18.

[0260] The compositions of the present disclosure (including, for example, RGN or nucleic acid molecules encoding RGN, as well as guide RNA or nucleic acid molecules encoding guide RNA, RGN systems, RNP complexes, or vectors) may be administered to subjects having HD with full penetrating properties in which the mutant HTT allele has more than 40 CAG repeats in exon 1 (e.g., 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, or more than 90 CAG repeats). In some embodiments, a subject requiring treatment with the compositions of the present disclosure has a mutant HTT allele having at least 36 CAG repeats in exon 1 (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more than 50 CAG repeats).

[0261] The compo...

Claims

1. It is an RNA-inducible nuclease (RGN) system, a) A guide RNA comprising a spacer and a backbone, wherein the spacer has a nucleotide sequence that is one or two nucleotides different from the nucleotide sequence of SEQ ID NO: 80 or 81, or a guide RNA, or a nucleic acid molecule encoding the guide RNA, b) An RGN system comprising an RGN polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 7, or a nucleic acid molecule encoding the RGN polypeptide.

2. The RGN system according to claim 1, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 80 or 81.

3. The RGN system according to claim 1, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

4. The RGN system according to any one of claims 1 to 3, wherein the guide RNA binds to a target sequence within the mutant huntingtin (mutHTT) allele.

5. The RGN system according to claim 4, wherein the RGN system can bind to and cleave a target sequence within the mutHTT allele, and the guide RNA forms a complex with the RGN polypeptide, and the complex can be directed to the target sequence for binding and cleavage.

6. The RGN system according to claim 4 or 5, wherein the target sequence has the nucleotide sequence of sequence number 75 or 76.

7. The RGN system according to any one of claims 1 to 6, wherein the RGN system can recognize a protospacer adjacent motif (PAM) having the sequence NNRYA, which is produced by a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

8. The RGN system according to any one of claims 1 to 7, wherein the RGN polypeptide comprises a PAM interaction domain that binds to a protospacer adjacent motif (PAM) having the sequence of NNRYA produced by a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

9. The RGN system according to claim 8, wherein the PAM interaction domain comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

134.

10. The RGN system according to claim 8 or 9, wherein the PAM interaction domain comprises the amino acid sequence shown as Sequence ID No.

134.

11. The RGN system according to any one of claims 1 to 10, wherein the RGN polypeptide comprises at least one nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 147, 148, 149, and 150.

12. The RGN system according to any one of claims 1 to 11, wherein the RGN system is unable to cleave the wild-type HTT allele.

13. The RGN system according to any one of claims 1 to 12, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

7.

14. The RGN system according to any one of claims 1 to 13, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

7.

15. The RGN system according to any one of claims 1 to 14, wherein the RGN polypeptide further comprises at least one nuclear localization signal.

16. The RGN system according to claim 15, wherein the at least one nuclear localization signal includes an SV40 nuclear localization signal.

17. The RGN system according to claim 16, wherein the SV40 nuclear localization signal has the sequence shown as sequence number 86.

18. The RGN system according to claim 15, wherein the at least one nuclear localization signal includes a c-Myc nuclear localization signal.

19. The RGN system according to claim 18, wherein the c-Myc nuclear localization signal has the sequence shown as sequence number 125.

20. The RGN system according to any one of claims 15 to 19, wherein the NLS linker protein connects the RGN polypeptide to the at least one nuclear localization signal.

21. The RGN system according to claim 20, wherein the NLS linker protein has the sequence shown as Sequence ID No.

127.

22. The RGN system according to any one of claims 1 to 21, wherein the backbone of the guide RNA has a nucleotide length of 66 to 90.

23. The RGN system according to any one of claims 1 to 22, wherein the backbone of the guide RNA comprises a nucleotide sequence having at least 80% sequence identity with SEQ ID NO: 140 or 141.

24. The RGN system according to any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106, or a nucleotide sequence that differs from SEQ ID NO: 8 or 106 by one or two nucleotides, and a tracrRNA having a nucleotide sequence that has at least 90% sequence identity with SEQ ID NO: 9 or 107.

25. The RGN system according to any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat having a nucleotide sequence that differs by one nucleotide from SEQ ID NO: 8 or 106, and a tracrRNA having a nucleotide sequence that has at least 95% sequence identity with SEQ ID NO: 9 or 107.

26. The RGN system according to any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

27. The RGN system according to any one of claims 1 to 26, wherein the guide RNA is a single guide RNA.

28. The RGN system according to claim 27, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

29. A ribonucleoprotein (RNP) complex comprising the RGN polypeptide and the guide RNA according to any one of claims 1 to 28.

30. A nucleic acid molecule containing or encoding the nucleotide sequence of sequence number 80 or 81, or a guide RNA containing a spacer having a nucleotide sequence that differs by one or two nucleotides from sequence number 80 or 81.

31. The nucleic acid molecule according to claim 30, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 80 or 81.

32. The nucleic acid molecule according to claim 30, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

33. The nucleic acid molecule according to any one of claims 30 to 32, wherein the guide RNA binds to a target sequence within the mutant huntingtin (mutHTT) allele.

34. The nucleic acid molecule according to claim 33, wherein the target sequence has the nucleotide sequence of sequence number 75 or 76.

35. The nucleic acid molecule according to any one of claims 30 to 34, wherein the guide RNA binds to an RNA-induced nuclease (RGN) polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

7.

36. A nucleic acid molecule comprising or encoding a guide RNA that binds to a target sequence in a mutant huntingtin (mutHTT) allele, wherein the target sequence binds to an RNA-induced nuclease (RGN) polypeptide having the nucleotide sequence of SEQ ID NO: 75 or 76 and an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

7.

37. The nucleic acid molecule according to claim 36, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81, or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by one or two nucleotides.

38. The nucleic acid molecule according to claim 37, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 80 or 81.

39. The nucleic acid molecule according to claim 37, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

40. The nucleic acid molecule according to any one of claims 35 to 39, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

7.

41. The nucleic acid molecule according to any one of claims 35 to 40, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

7.

42. The nucleic acid molecule according to any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 8 or 106, and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO: 9 or 107.

43. The nucleic acid molecule according to any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat having a nucleotide sequence that differs by one nucleotide from sequence number 8 or 106, and a tracrRNA having a nucleotide sequence that has at least 95% sequence identity with sequence number 9 or 107.

44. The nucleic acid molecule according to any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

45. The nucleic acid molecule according to any one of claims 30 to 44, wherein the guide RNA is a single guide RNA.

46. The nucleic acid molecule according to claim 45, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

47. The nucleic acid molecule according to any one of claims 30 to 46, wherein the nucleic acid molecule encoding the guide RNA is operably linked to an RNA polymerase III promoter.

48. The nucleic acid molecule according to claim 47, wherein the RNA polymerase III promoter is a U6 promoter.

49. The nucleic acid molecule according to claim 48, wherein the U6 promoter is a cleaved U6 promoter.

50. The nucleic acid molecule according to claim 49, wherein the cleaved U6 promoter has a nucleotide sequence shown as sequence number 89 or 128.

51. A vector comprising a nucleic acid molecule according to any one of claims 30 to 35, wherein the nucleic acid molecule encodes the guide RNA.

52. A vector comprising a nucleic acid molecule according to any one of claims 36 to 50, wherein the nucleic acid molecule encodes the guide RNA.

53. The vector according to claim 51 or 52, wherein the vector is a viral vector.

54. The vector according to claim 53, wherein the viral vector is a lentiviral vector, a baculovirus vector, or an adeno-associated virus (AAV) vector.

55. The vector according to claim 54, wherein the viral vector is an AAV vector and includes an AAV inverted terminal repeat.

56. The vector according to claim 55, wherein the AAV inverted end repeat is an AAV2, AAV5, or AAV6 inverted end repeat.

57. The vector according to any one of claims 52 to 56, wherein the vector further comprises a nucleic acid molecule encoding the RGN polypeptide.

58. The vector according to claim 57, further comprising an RNA polymerase II promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide.

59. The vector according to claim 58, wherein the RNA polymerase II promoter is a constitutive promoter.

60. The vector according to claim 59, wherein the constitutive promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, a cleaved CMV promoter, an elongation factor 1α short (EFS) promoter, and a JeT promoter.

61. The vector according to claim 60, wherein the constitutive promoter is a Jet promoter.

62. The vector according to claim 61, wherein the JeT promoter has a nucleotide sequence shown as sequence number 92.

63. The vector according to claim 58, wherein the RNA polymerase II promoter is a tissue-specific promoter.

64. The vector according to claim 63, wherein the tissue-specific promoter is a brain or neuron-specific promoter.

65. The vector according to claim 64, wherein the brain or neuron-specific promoter is selected from the group consisting of human synapsin I (Syn) promoter, 67 kDa glutamate decarboxylase (GAD67) promoter, 65 kDa glutamate decarboxylase (GAD65) promoter, homeobox Dlx5 / 6 promoter, preprotachykinin 1 (Tac1) promoter, neuron-specific enolase (NSE) promoter, dopaminergic receptor 1 (Drd1a) promoter, dopaminergic receptor 2 (DRD2) promoter, and glial fibrillary acidic protein (GFAP) promoter.

66. The vector according to claim 65, wherein the neuron-specific promoter is a Syn promoter.

67. The vector according to claim 66, wherein the Syn promoter has a nucleotide sequence shown as Sequence ID No.

93.

68. The vector according to any one of claims 57 to 67, wherein the nucleic acid molecule encoding the RGN polypeptide includes a polyadenylated (poly-A) tail.

69. The vector according to claim 68, wherein the poly-A tail is an SV40 poly-A tail or a bovine growth hormone (bGH) poly-A tail.

70. The vector according to claim 69, wherein the SV40 polyA tail has a nucleotide sequence indicated as sequence number 94.

71. The vector according to claim 69, wherein the bGH polyA tail has the sequence indicated as sequence number 95.

72. The vector according to claim 57 or 58, comprising: a cleaved U6 promoter operably linked to the nucleic acid molecule encoding the guide RNA; a CMVeb promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide; c-Myc NLS at the N-terminus and C-terminus of the RGN polypeptide; an NLS linker protein connecting the c-Myc NLS and the RGN polypeptide; and an SV40 polyA tail.

73. The vector according to claim 72, wherein the cleaved U6 promoter has the sequence shown as SEQ ID NO: 128, the CMVeb promoter has the sequence shown as SEQ ID NO: 90, the c-Myc NLS has the sequence shown as SEQ ID NO: 125, the NLS linker protein has the sequence shown as SEQ ID NO: 127, and the SV40 polyA tail has the sequence shown as SEQ ID NO:

94.

74. The vector according to claim 72 or 73, wherein the sgRNA has the sequence shown as SEQ ID NO: 26, and the nucleic acid molecule encoding the RGN polypeptide has the sequence shown as SEQ ID NO:

88.

75. The vector according to any one of claims 72 to 74, wherein the vector includes the sequence indicated as sequence number 123.

76. The vector according to any one of claims 57 to 75, wherein the RGN polypeptide is operably linked to at least one nuclear localization signal.

77. The vector according to claim 76, wherein the at least one nuclear localization signal includes an SV40 nuclear localization signal.

78. The vector according to claim 77, wherein the SV40 nuclear localization signal has the sequence shown as sequence number 86.

79. The vector according to claim 76, wherein the at least one nuclear localization signal includes a c-Myc nuclear localization signal.

80. The vector according to claim 79, wherein the c-Myc nuclear localization signal has the sequence shown as sequence number 125.

81. The vector according to any one of claims 76 to 80, wherein the NLS linker protein links the RGN polypeptide to the at least one nuclear localization signal.

82. The vector according to claim 81, wherein the NLS linker protein has the sequence shown as Sequence ID No.

127.

83. The vector according to any one of claims 57 to 82, wherein the vector has an sequence represented as one of sequence numbers 32 to 39 or 121 to 123.

84. It is an RNA-inducible nuclease (RGN) system, a) A guide RNA having the nucleotide sequence of SEQ ID NO: 82 or 83, or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by one or two nucleotides, or a nucleic acid molecule encoding the guide RNA, b) An RGN system comprising an RGN polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 3, or a nucleic acid molecule encoding the RGN polypeptide.

85. The RGN system according to claim 84, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 82 or 83.

86. The RGN system according to claim 84, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

87. The RGN system according to any one of claims 84 to 86, wherein the guide RNA binds to a target sequence within the mutant huntingtin (mutHTT) allele.

88. The RGN system according to claim 87, wherein the system can bind to and cleave a target sequence within the mutHTT allele, and the guide RNA can form a complex with the RGN polypeptide, and the complex can be directed to the target sequence for binding and cleavage.

89. The RGN system according to claim 87 or 88, wherein the target sequence has the nucleotide sequence of sequence number 77 or 78.

90. The RGN system according to any one of claims 84 to 89, wherein the RGN system can recognize a protospacer adjacent motif (PAM) having an NNNNCC sequence produced by a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

91. The RGN system according to any one of claims 84 to 90, wherein the RGN includes a PAM interaction domain that binds to a protospacer adjacent motif (PAM) having an NNNNCC sequence produced by a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

92. The RGN system according to claim 91, wherein the PAM interaction domain comprises an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

133.

93. The RGN system according to claim 91 or 92, wherein the PAM interaction domain includes the amino acid sequence shown as Sequence ID No.

133.

94. The RGN system according to any one of claims 84 to 93, wherein the RGN polypeptide comprises at least one nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NOs: 143, 144, 145, and 146.

95. The RGN system according to any one of claims 84 to 94, wherein the RGN system is unable to cleave the wild-type HTT allele.

96. The RGN system according to any one of claims 84 to 95, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

3.

97. The RGN system according to any one of claims 84 to 96, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

3.

98. The RGN system according to any one of claims 84 to 97, wherein the RGN polypeptide comprises at least one nuclear localization signal.

99. The RGN system according to claim 98, wherein the at least one nuclear localization signal includes an SV40 nuclear localization signal.

100. The RGN system according to claim 99, wherein the SV40 nuclear localization signal has the sequence shown as sequence number 86.

101. The RGN system according to claim 98, wherein the at least one nuclear localization signal includes a c-Myc nuclear localization signal.

102. The RGN system according to claim 101, wherein the c-Myc nuclear localization signal has the sequence shown as sequence number 125.

103. The RGN system according to any one of claims 98 to 102, wherein the NLS linker protein connects the RGN polypeptide to the at least one nuclear localization signal.

104. The RGN system according to claim 103, wherein the NLS linker protein has the sequence shown as Sequence ID No.

127.

105. The RGN system according to any one of claims 84 to 104, wherein the backbone of the guide RNA has a nucleotide length of 94 to 110.

106. The RGN system according to any one of claims 84 to 105, wherein the backbone of the guide RNA comprises a nucleotide sequence having at least 80% sequence identity with SEQ ID NO:

142.

107. A ribonucleoprotein (RNP) complex comprising the RGN polypeptide and the guide RNA according to any one of claims 84 to 106.

108. A nucleic acid molecule containing or encoding a guide RNA that includes the nucleotide sequence of sequence number 82 or 83, or a spacer having a nucleotide sequence that differs by one or two nucleotides from sequence number 82 or 83.

109. The nucleic acid molecule according to claim 108, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 82 or 83.

110. The nucleic acid molecule according to claim 108, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

111. The nucleic acid molecule according to any one of claims 108 to 110, wherein the guide RNA binds to a target sequence within the mutant huntingtin (mutHTT) allele.

112. The nucleic acid molecule according to claim 111, wherein the target sequence has the nucleotide sequence of sequence number 77 or 78.

113. The nucleic acid molecule according to any one of claims 108 to 112, wherein the guide RNA binds to an RNA-induced nuclease (RGN) polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

3.

114. A nucleic acid molecule comprising or encoding a guide RNA that binds to a target sequence in a mutant huntingtin (mutHTT) allele, wherein the target sequence binds to an RNA-induced nuclease (RGN) polypeptide having the nucleotide sequence of SEQ ID NO: 77 or 78 and an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

3.

115. The nucleic acid molecule according to claim 114, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83, or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by one or two nucleotides.

116. The nucleic acid molecule according to claim 115, wherein the guide RNA includes a spacer having a nucleotide sequence that differs by one nucleotide from sequence number 82 or 83.

117. The nucleic acid molecule according to claim 115, wherein the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

118. The nucleic acid molecule according to any one of claims 113 to 117, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity with SEQ ID NO:

3.

119. The nucleic acid molecule according to any one of claims 113 to 118, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

3.

120. The nucleic acid molecule according to any one of claims 108 to 119, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 4, and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO:

5.

121. The nucleic acid molecule according to any one of claims 108 to 119, wherein the guide RNA comprises a crRNA repeat having a nucleotide sequence that differs by one nucleotide from SEQ ID NO: 4, and a tracrRNA having a nucleotide sequence that has at least 95% sequence identity with SEQ ID NO:

5.

122. The nucleic acid molecule according to any one of claims 108 to 119, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4 and a tracrRNA having the nucleotide sequence of SEQ ID NO:

5.

123. The nucleic acid molecule according to any one of claims 108 to 122, wherein the guide RNA is a single guide RNA.

124. The nucleic acid molecule according to claim 123, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

125. The nucleic acid molecule according to any one of claims 108 to 124, wherein the nucleic acid molecule encoding the guide RNA is operably linked to an RNA polymerase III promoter.

126. The nucleic acid molecule according to claim 125, wherein the RNA polymerase III promoter is a U6 promoter.

127. The nucleic acid molecule according to claim 126, wherein the U6 promoter is a cleaved U6 promoter.

128. The nucleic acid molecule according to claim 127, wherein the cleaved U6 promoter has a nucleotide sequence shown as sequence number 89 or 128.

129. A vector comprising a nucleic acid molecule according to any one of claims 108 to 112, wherein the nucleic acid molecule encodes the guide RNA.

130. A vector comprising a nucleic acid molecule according to any one of claims 113 to 128, wherein the nucleic acid molecule encodes the guide RNA.

131. The vector according to claim 129 or 130, wherein the vector is a viral vector.

132. The vector according to claim 131, wherein the viral vector is a lentiviral vector, a baculovirus vector, or an adeno-associated virus (AAV) vector.

133. The vector according to claim 132, wherein the viral vector is an AAV vector and includes an AAV inverted terminal repeat.

134. The vector according to claim 133, wherein the AAV inverted end repeat is an AAV2, AAV5, or AAV6 inverted end repeat.

135. The vector according to any one of claims 130 to 134, wherein the vector further comprises a nucleic acid molecule encoding the RGN polypeptide.

136. The vector according to claim 135, further comprising an RNA polymerase II promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide.

137. The vector according to claim 136, wherein the RNA polymerase II promoter is a constitutive promoter.

138. The vector according to claim 137, wherein the constitutive promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, a cleaved CMV promoter, an elongation factor 1α short (EFS) promoter, and a JeT promoter.

139. The vector according to claim 138, wherein the constitutive promoter is a Jet promoter.

140. The vector according to claim 139, wherein the Jet promoter has a nucleotide sequence shown as Sequence ID No.

92.

141. The vector according to claim 136, wherein the RNA polymerase II promoter is a tissue-specific promoter.

142. The vector according to claim 141, wherein the tissue-specific promoter is a brain or neuron-specific promoter.

143. The vector according to claim 142, wherein the brain or neuron-specific promoter is selected from the group consisting of human synapsin I (Syn) promoter, 67 kDa glutamate decarboxylase (GAD67) promoter, 65 kDa glutamate decarboxylase (GAD65) promoter, homeobox Dlx5 / 6 promoter, preprotachykinin 1 (Tac1) promoter, neuron-specific enolase (NSE) promoter, dopaminergic receptor 1 (Drd1a) promoter, dopaminergic receptor 2 (DRD2) promoter, and glial fibrillary acidic protein (GFAP) promoter.

144. The vector according to claim 143, wherein the neuron-specific promoter is a Syn promoter.

145. The vector according to claim 144, wherein the Syn promoter has a nucleotide sequence shown as Sequence ID No.

93.

146. The vector according to any one of claims 135 to 145, wherein the nucleic acid molecule encoding the RGN polypeptide includes a polyadenylated (poly-A) tail.

147. The vector according to claim 146, wherein the poly-A tail is an SV40 poly-A tail or a bovine growth hormone (bGH) poly-A tail.

148. The vector according to claim 147, wherein the SV40 poly-A tail has a nucleotide sequence indicated as sequence number 94.

149. The vector according to claim 147, wherein the bGH polyA tail has the sequence shown as sequence number 95.

150. The vector according to any one of claims 135 to 149, wherein the RGN polypeptide is operably linked to at least one nuclear localization signal.

151. The vector according to claim 150, wherein the at least one nuclear localization signal includes an SV40 nuclear localization signal.

152. The vector according to claim 151, wherein the SV40 nuclear localization signal has the sequence shown as sequence number 86.

153. The vector according to claim 150, wherein the at least one nuclear localization signal includes a c-Myc nuclear localization signal.

154. The vector according to claim 153, wherein the c-Myc nuclear localization signal has the sequence shown as sequence number 125.

155. The vector according to any one of claims 150 to 154, wherein the NLS linker protein links the RGN polypeptide to the at least one nuclear localization signal.

156. The vector according to claim 155, wherein the NLS linker protein has the sequence shown as Sequence ID No.

127.

157. A cell comprising a nucleic acid molecule according to any one of claims 30 to 50 and 108 to 128, or a vector according to any one of claims 51 to 83 and 129 to 156.

158. A pharmaceutical composition comprising a nucleic acid molecule according to any one of claims 30 to 50 and 108 to 128, a vector according to any one of claims 51 to 83 and 129 to 156, an RGN system according to any one of claims 1 to 28 and 84 to 106, or an RNP complex according to claim 29 or 107.

159. The pharmaceutical composition according to claim 158, having a purity of at least 95%.

160. The pharmaceutical composition according to claim 158 or 159, having an undetectable level of endotoxin or other impurities.

161. A pharmaceutical composition according to any one of claims 158 to 160, further comprising poloxamer 188.

162. A pharmaceutical composition according to any one of claims 158 to 161, which is a solution.

163. A pharmaceutically acceptable composition according to any one of claims 158 to 161, which is freeze-dried or lyophilized.

164. A vector comprising the RGN system described in any one of claims 1 to 28.

165. A vector comprising the RGN system described in any one of claims 84 to 106.

166. The vector according to claim 164 or 165, wherein the vector includes an adeno-associated vector (AAV) inverted terminal repeat.

167. The vector according to claim 166, wherein the AAV inverted end repeat is an AAV2, AAV5, or AAV6 inverted end repeat.

168. The vector according to claim 167, wherein the AAV inverted end repeat is an AAV5 inverted end repeat.

169. Use of nucleic acid molecules according to any one of claims 30-50 and 108-128, vectors according to any one of claims 51-83, 129-156 and 164-168, RGN systems according to any one of claims 1-28 and 84-106, or RNP complexes according to claim 29 or 107 for reducing levels of mutHTT mRNA and / or mutHTT protein in cells.

170. Use of a nucleic acid molecule according to any one of claims 30-50 and 108-128, a vector according to any one of claims 51-83, 129-156 and 164-168, an RGN system according to any one of claims 1-28 and 84-106, or an RNP complex according to claim 29 or 107 for the treatment of Huntington's disease.

171. A method for cleaving a mutant huntingtin (mutHTT) allele in a cell, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, and a protospacer adjacent motif (PAM) having the nucleotide sequence NNRYA comprises the first SNP allele, and the method comprises the cell, a) RNA-induced nuclease (RGN) polypeptide having at least 90% sequence identity with SEQ ID NO: 7, or a nucleic acid molecule encoding the RGN polypeptide, and i) A nucleic acid molecule comprising or encoding the guide RNA described in any one of claims 30 to 50, or ii) The vector according to any one of claims 51 to 56, b) The vector according to any one of claims 57-83, 164, and 166-168, c) The RGN system according to any one of claims 1 to 28, or d) A method comprising introducing the RNP complex described in claim 29.

172. A method for cleaving a mutant huntingtin (mutHTT) allele in a cell, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, and a protospacer adjacent motif (PAM) having the NNNNCC nucleotide sequence comprises the first SNP allele, and the method comprises the cell, a) RNA-induced nuclease (RGN) polypeptide having at least 90% sequence identity with SEQ ID NO: 3, or a nucleic acid molecule encoding the RGN polypeptide, and i) A nucleic acid molecule comprising or encoding the guide RNA described in any one of claims 108 to 128, or ii) The vector according to any one of claims 129 to 134, b) The vector according to any one of claims 135 to 156 and 165 to 168, c) The RGN system according to any one of claims 84 to 106, or d) A method comprising introducing the RNP complex described in claim 107.

173. The method according to claim 171 or 172, wherein the RGN polypeptide can recognize the PAM and cleave the mutHTT allele.

174. The method according to any one of claims 171 to 173, wherein the mutHTT allele has at least 36 CAG repeats in exon 1.

175. The method according to any one of claims 171 to 173, wherein the mutHTT allele has at least 40 CAG repeats in exon 1.

176. The method according to any one of claims 171 to 175, wherein the cells are assayed to determine whether the mutHTT allele contains the first SNP allele before introducing the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

177. The method according to any one of claims 171 to 176, wherein the cells contain a wild-type HTT (wtHTT) allele that includes a second SNP allele in which the PAM is absent, and thereby the cells are heterozygous for the SNP.

178. The method according to claim 177, wherein the cells are assayed to determine whether they are heterozygous for SNPs before introducing the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

179. The method according to any one of claims 171 to 178, wherein the mutHTT allele is edited, thereby producing a gene-modified cell containing the edited mutHTT allele.

180. The method according to claim 179, wherein the editing includes introducing an insertion and / or deletion (indel) of the SNP or its vicinity.

181. The method according to claim 179, wherein the editing includes introducing an early stop codon near the SNP.

182. The method according to any one of claims 179 to 181, wherein the gene-modified cells are gene-modified stem cells.

183. The method according to claim 182, wherein the gene-modified stem cells are gene-modified induced pluripotent stem cells (iPSCs) or gene-modified mesenchymal stem cells (MSCs).

184. The method according to claim 183, further comprising differentiating the gene-modified iPSC or MSC within a nerve cell.

185. The method according to any one of claims 179 to 184, wherein the level of mutHTT mRNA is reduced by at least 40% compared to the level of HTT mRNA in unmodified cells or the level of wild-type HTT mRNA.

186. The method according to any one of claims 179 to 185, wherein the level of mutHTT protein is reduced by at least 40% compared to the level of HTT protein in unmodified cells or the level of wild-type HTT protein.

187. The method according to any one of claims 179 to 186, further comprising selecting the gene-modified cells.

188. Genetically modified cells produced by the method of claim 187.

189. The method according to any one of claims 171 to 187, wherein the introduction comprises administering to a subject including the cells a composition comprising the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

190. The method according to claim 189, wherein the cell is a eukaryotic cell.

191. The method according to claim 190, wherein the eukaryotic cell is a mammalian cell.

192. The method according to claim 191, wherein the mammalian cell is a human cell.

193. The method according to claim 191 or 192, wherein the mammalian cell or human cell is a stem cell.

194. The method according to claim 191 or 192, wherein the mammalian or human cell is a forebrain neuron, striatal neuron, medium spiny neuron, cortical neuron, or glial cell.

195. The method according to claim 191 or 192, wherein the mammalian or human cells are located in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or a combination thereof.

196. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the subject is a) At least 36 CAG repeats within exon 1, b) A first single nucleotide polymorphism (SNP) allele in exon 50, comprising a first SNP allele and a mutant huntingtin (mutHTT) allele, wherein a protospacer adjacent motif (PAM) having the nucleotide sequence NNRYA contains the first SNP allele, The above method applies to the above target, a) RNA-induced nuclease (RGN) polypeptide having at least 90% sequence identity with SEQ ID NO: 7, or a nucleic acid molecule encoding the RGN polypeptide, and i) A nucleic acid molecule comprising or encoding the guide RNA described in any one of claims 30 to 50, or ii) The vector according to any one of claims 51 to 56, b) The vector according to any one of claims 57-83, 164, and 166-168, c) The RGN system according to any one of claims 1 to 28, or d) comprising administering the RNP complex described in claim 29, A method for reducing the level of the mutHTT protein encoded by the mutHTT allele compared to the level of a control HTT protein or a wild-type HTT protein.

197. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the subject is a) At least 36 CAG repeats within exon 1, b) A first single nucleotide polymorphism (SNP) allele in exon 50, comprising a first SNP allele and a mutant huntingtin (mutHTT) allele, wherein a protospacer adjacent motif (PAM) having the nucleotide sequence NNNNCC contains the first SNP allele, The above method applies to the above target, a) RNA-induced nuclease (RGN) polypeptide having at least 90% sequence identity with SEQ ID NO: 3, or a nucleic acid molecule encoding the RGN polypeptide, and i) A nucleic acid molecule comprising or encoding the guide RNA described in any one of claims 108 to 128, or ii) The vector according to any one of claims 129 to 134, b) The vector according to any one of claims 135 to 156 and 165 to 168, c) The RGN system according to any one of claims 84 to 106, or d) comprising administering the RNP complex according to claim 107, A method for reducing the level of the mutHTT protein encoded by the mutHTT allele compared to the level of a control HTT protein or a wild-type HTT protein.

198. The method according to claim 196 or 197, wherein the RGN polypeptide recognizes the PAM and cleaves and edits the mutHTT allele.

199. The method according to any one of claims 196 to 198, wherein the administration includes intrastriatal, intraparenchymal, intrathecal, intracerebral, intraventricular, intrathalamic, or intracisional injection.

200. The method according to any one of claims 196 to 199, wherein the subject is assayed to determine whether the mutHTT allele contains the first SNP allele containing the PAM before administration of the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

201. The method according to any one of claims 196 to 200, wherein the subject comprises a wild-type HTT (wtHTT) allele containing a second SNP allele in which the PAM is absent, and thereby the subject is heterozygous to the SNP.

202. The method according to claim 201, wherein the subject is assayed to determine whether the subject is heterozygous to the SNP before administration of the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

203. The method according to any one of claims 196 to 202, wherein the PAM is present only in the mutHTT allele and not in the wild-type HTT allele.

204. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the subject is a) At least 36 CAG repeats within exon 1, b) A first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele contains thymine at the position corresponding to position 151 of SEQ ID NO: 1, and a mutant huntingtin (mutHTT) allele comprising: The above method involves administering the above method to the target by intrastriatal injection. a) A first nucleic acid molecule encoding an RNA-induced nuclease (RGN) polypeptide having the amino acid sequence of Sequence ID No. 7, b) Administering an AAV5 vector comprising a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 25 or 26, A method wherein, four weeks after administration of the AAV5 vector, the level of the mutHTT protein encoded by the mutHTT allele in the subject is reduced compared to the level of the HTT protein in the control subject or the level of the wild-type HTT protein.

205. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the subject is a) At least 36 CAG repeats within exon 1, b) A first single nucleotide polymorphism (SNP) allele in exon 50, wherein the first SNP allele contains cytosine at the position corresponding to position 151 of SEQ ID NO: 2, and a mutant huntingtin (mutHTT) allele comprising the first SNP allele, The above method involves administering the above method to the target by intrastriatal injection. a) A first nucleic acid molecule encoding an RNA-induced nuclease (RGN) polypeptide having the amino acid sequence of Sequence ID No. 3, b) Administering an AAV5 vector comprising a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 27 or 28, A method wherein, four weeks after administration of the AAV5 vector, the level of the mutHTT protein encoded by the mutHTT allele in the subject is reduced compared to the level of the HTT protein in the control subject or the level of the wild-type HTT protein.

206. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject who requires such alleviation or delay, wherein the method is a) i) At least 36 CAG repeats within exon 1, ii) Select a subject that includes a mutant huntingtin (mutHTT) allele comprising a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele contains thymine at the position corresponding to position 151 of SEQ ID NO: 1, b) By intrastriatal injection, the subject, i) A first nucleic acid molecule encoding an RNA-induced nuclease (RGN) polypeptide having the amino acid sequence of Sequence ID No. 7, ii) Administering an AAV5 vector comprising a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 25 or 26, A method wherein, four weeks after administration of the AAV5 vector, the level of the mutHTT protein encoded by the mutHTT allele in the subject is reduced compared to the level of the HTT protein in the control subject or the level of the wild-type HTT protein.

207. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject who requires such alleviation or delay, wherein the method is a) i) At least 36 CAG repeats within exon 1, ii) Select a subject that includes a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele contains cytosine at the position corresponding to position 151 of SEQ ID NO: 2, and a mutant huntingtin (mutHTT) allele, b) By intrastriatal injection, the subject, i) A first nucleic acid molecule encoding an RNA-induced nuclease (RGN) polypeptide having the amino acid sequence of Sequence ID No. 3, ii) Administering an AAV5 vector comprising a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 27 or 28, A method wherein, four weeks after administration of the AAV5 vector, the level of the mutHTT protein encoded by the mutHTT allele in the subject is reduced compared to the level of the HTT protein in the control subject or the level of the wild-type HTT protein.

208. The method according to any one of claims 204 to 207, wherein the RGN polypeptide recognizes, cleaves, and edits the mutHTT allele.

209. The method according to any one of claims 204 to 208, wherein the editing includes introducing an indel into or near the SNP.

210. The method according to any one of claims 204 to 208, wherein the editing includes introducing an early stop codon near the SNP.

211. The method according to any one of claims 204 to 210, wherein the mutHTT allele has at least 40 CAG repeats in exon 1.

212. The method according to any one of claims 204 to 210, wherein the mutHTT allele has at least 56 CAG repeats in exon 1, and the subject is younger than 18 years of age.

213. The method according to any one of claims 204 to 212, wherein the administration is performed before the onset of symptoms of Huntington's disease.

214. The method according to any one of claims 204 to 213, wherein the method comprises preventing the onset of one or more symptoms of Huntington's disease.

215. The method according to any one of claims 204 to 214, wherein the subject has at least one symptom of Huntington's disease.

216. The method according to any one of claims 204 to 215, wherein a reduction of at least 40% in the level of mutant HTT mRNA is observed compared to the level of control HTT mRNA or the level of wild-type HTT mRNA.

217. The method according to any one of claims 204 to 216, wherein a decrease in the level of mutHTT protein is observed within 12 weeks after administration of the vector.

218. The method according to claim 217, wherein a reduction of at least 40% in the level of mutHTT protein is observed compared to the level of a control HTT protein or the level of wild-type HTT protein.

219. The method according to any one of claims 204 to 218, wherein a decrease in the level of mutHTT mRNA is observed in at least 50% of the striatal cells of the subject.

220. The method according to any one of claims 204 to 219, wherein a decrease in the level of mutHTT protein is observed in at least 50% of the striatal cells of the subject.

221. The method according to any one of claims 204 to 220, wherein only the mutated HTT allele is edited.

222. A method for cleaving an intracellular mutant huntingtin (mutHTT) allele, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, and the protospacer adjacent motif (PAM) comprises the first SNP allele, the method comprising introducing into the cell (i) an RNA-induced nuclease (RGN) polypeptide, or a nucleic acid molecule encoding the RGN polypeptide, and (ii) a guide RNA, or a nucleic acid molecule encoding the guide RNA.

223. The method according to claim 222, wherein the RGN polypeptide can recognize the PAM and cleave the mutHTT allele.

224. The method according to claim 223, wherein the mutHTT allele has at least 36 CAG repeats in exon 1.

225. The method according to claim 223 or 224, wherein the cells are assayed to determine whether the mutHTT allele contains the first SNP allele before introducing (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

226. The method according to any one of claims 223 to 225, wherein the cells contain a wild-type HTT (wtHTT) allele that includes a second SNP allele in which the PAM is absent, and thereby the cells are heterozygous for the SNP.

227. The method according to claim 226, wherein the cells are assayed to determine whether they are heterozygous for SNPs before introducing (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

228. The method according to any one of claims 222 to 227, wherein the mutHTT allele is edited, thereby producing a gene-modified cell containing the edited mutHTT allele.

229. The method according to claim 228, wherein the editing includes introducing an insertion and / or deletion (indel) of the SNP or its vicinity.

230. The method according to claim 228, wherein the editing includes introducing an early codon near the SNP.

231. The method according to any one of claims 228 to 230, wherein the gene-modified cells are gene-modified stem cells.

232. The method according to claim 231, wherein the gene-modified stem cells are gene-modified induced pluripotent stem cells (iPSCs) or gene-modified mesenchymal stem cells (MSCs).

233. The method according to claim 232, further comprising differentiating the gene-modified iPSC or MSC within a nerve cell.

234. The method according to any one of claims 228 to 233, wherein the level of mutHTT mRNA is reduced in the genetically modified cells compared to the level of HTT mRNA in unmodified cells or the level of wild-type HTT mRNA.

235. The method according to any one of claims 228 to 234, wherein the level of the mutHTT protein encoded by the mutHTT allele is reduced in the genetically modified cell compared to the level of the HTT protein in the unmodified cell or the level of the wild-type HTT protein.

236. The method according to any one of claims 228 to 235, further comprising selecting the gene-modified cells.

237. Genetically modified cells produced by the method described in claim 236.

238. The method according to any one of claims 222 to 227, wherein the introduction comprises (i) administering a composition comprising the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA, to a subject including the cells.

239. The method according to claim 238, wherein the cell is a eukaryotic cell.

240. The method according to claim 239, wherein the eukaryotic cell is a mammalian cell.

241. The method according to claim 240, wherein the mammalian cell is a human cell.

242. The method according to claim 240 or 241, wherein the mammalian cell or human cell is a stem cell.

243. The method according to claim 240 or 241, wherein the mammalian or human cell is a forebrain neuron, striatal neuron, medium spiny neuron, cortical neuron, or glial cell.

244. The method according to claim 240 or 241, wherein the mammalian or human cells are located in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or a combination thereof.

245. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the subject is a) At least 36 CAG repeats within exon 1, b) A first single nucleotide polymorphism (SNP) allele in exon 50, comprising a first SNP allele and a mutant huntingtin (mutHTT) allele, wherein the protospacer adjacent motif (PAM) includes the first SNP allele, The method comprises administering to the subject (i) an RNA-induced nuclease (RGN) polypeptide or a nucleic acid molecule encoding the RGN polypeptide, and (ii) a guide RNA or a nucleic acid molecule encoding the guide RNA, wherein the level of the mutHTT protein encoded by the mutHTT allele is reduced compared to the level of the HTT protein of a control or the level of the wild-type HTT protein.

246. The method according to claim 245, wherein the RGN polypeptide recognizes the PAM and cleaves and edits the mutHTT allele.

247. The method according to claim 245, wherein the mutHTT allele has at least 40 CAG repeats in exon 1.

248. The method according to claim 245, wherein the mutHTT allele has at least 56 CAG repeats in exon 1, and the subject is younger than 18 years of age.

249. The method according to any one of claims 245 to 248, wherein the administration is performed before the onset of symptoms of Huntington's disease.

250. The method according to any one of claims 245 to 249, wherein the method includes preventing the onset of one or more symptoms of Huntington's disease.

251. The method according to any one of claims 245 to 250, wherein the subject has at least one symptom of Huntington's disease.

252. The method according to any one of claims 245 to 251, wherein the administration includes intrastriatal, intraparenchymal, intrathecal, intracerebral, intraventricular, intrathalamic, or intracisional injection.

253. The method according to any one of claims 245 to 252, comprising assaying the subject to determine whether the mutHTT allele containing the PAM contains the first SNP allele before administration of (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

254. The method according to any one of claims 245 to 253, wherein the subject comprises a wild-type HTT (wtHTT) allele containing a second SNP allele in which the PAM is absent, and thereby the subject is heterozygous to the SNP.

255. The method according to claim 254, wherein the subject is assayed to determine whether the subject is heterozygous to the SNP before administration of (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

256. The method according to any one of claims 245 to 255, wherein the level of mutHTT mRNA is reduced by at least 40% compared to the level of control HTT mRNA or wild-type HTT mRNA.

257. The method according to any one of claims 245 to 256, wherein the level of mutHTT protein is reduced by at least 40% compared to the level of a control HTT protein or the level of wild-type HTT protein.

258. The method according to any one of claims 245 to 257, wherein a decrease in the level of mutHTT protein is observed within 12 weeks after administration.

259. The method according to any one of claims 245 to 258, wherein a decrease in the level of mutHTT protein is observed in at least 50% of the striatal cells of the subject.

260. The method according to any one of claims 222 to 259, wherein the PAM has a nucleotide sequence selected from the group consisting of NNNNCC, NNRYA, NNGRR, and NNGG.

261. The method according to claim 260, wherein the PAM sequence NNRYA comprises the first SNP allele, and the first SNP allele is a thymine at the position corresponding to position 151 of SEQ ID NO:

1.

262. The method according to claim 261, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with SEQ ID NO:

7.

263. The method according to claim 261 or 262, wherein the RGN polypeptide comprises the amino acid sequence of SEQ ID NO:

7.

264. The method according to claim 261 or 262, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 8 or 106, and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO: 9 or 107.

265. The method according to claim 263, wherein the guide RNA comprises a crRNA repeat having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

266. The method according to any one of claims 262 to 265, wherein the guide RNA includes a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 75 or 76.

267. The method according to claim 266, wherein the spacer has a nucleotide sequence of sequence number 80 or 81, or a nucleotide sequence that differs from sequence number 80 or 81 by one or two nucleotides.

268. The method according to claim 267, wherein the spacer has the nucleotide sequence of sequence number 80 or 81.

269. The method according to any one of claims 261 to 268, wherein the guide RNA is a single guide RNA.

270. The method according to claim 269, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

271. The method according to any one of claims 222 to 236 and 238 to 270, wherein the method comprises introducing a vector comprising the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA, the vector comprising a cleaved U6 promoter that modulates the expression of sgRNA, a CMVeb promoter that modulates the expression of the RGN polypeptide, c-Myc NLS located at the N-terminus and C-terminus of the RGN polypeptide, an NLS linker protein that connects the c-Myc NLS to the RGN polypeptide, and an SV40 polyA tail.

272. The method according to claim 271, wherein the cleaved U6 promoter has the sequence shown as SEQ ID NO: 128, the CMVeb promoter has the sequence shown as SEQ ID NO: 90, the c-Myc NLS has the sequence shown as SEQ ID NO: 125, the NLS linker protein has the sequence shown as SEQ ID NO: 127, and the SV40 polyA tail has the sequence shown as SEQ ID NO:

94.

273. The method according to claim 271 or 272, wherein the sgRNA has the sequence shown as SEQ ID NO: 26, and the nucleic acid molecule encoding the RGN polypeptide has the sequence shown as SEQ ID NO:

88.

274. The method according to any one of claims 271 to 273, wherein the vector includes the sequence shown as sequence number 123.

275. The method according to any one of claims 238 to 270, wherein the method comprises administering to a subject a vector comprising the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA, the vector comprising a cleaved U6 promoter that modulates the expression of sgRNA, a CMVeb promoter that modulates the expression of the RGN polypeptide, c-Myc NLS located at the N-terminus and C-terminus of the RGN polypeptide, an NLS linker protein that connects the c-Myc NLS to the RGN polypeptide, and an SV40 polyA tail.

276. The method according to claim 275, wherein the cleaved U6 promoter has the sequence shown as SEQ ID NO: 128, the CMVeb promoter has the sequence shown as SEQ ID NO: 90, the c-Myc NLS has the sequence shown as SEQ ID NO: 125, the NLS linker protein has the sequence shown as SEQ ID NO: 127, and the SV40 polyA tail has the sequence shown as SEQ ID NO:

94.

277. The method according to claim 275 or 276, wherein the sgRNA has the sequence shown as SEQ ID NO: 26, and the nucleic acid molecule encoding the RGN has the sequence shown as SEQ ID NO:

88.

278. The method according to any one of claims 275 to 277, wherein the vector includes the sequence shown as sequence number 123.

279. The method according to any one of claims 222 to 260, wherein the PAM sequence selected from the group consisting of NNNNCC, NNGRR, and NNGG comprises the first SNP allele, and the first SNP allele is cytosine at the position corresponding to position 151 of SEQ ID NO:

1.

280. The method according to claim 279, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NOs: 3, 11, and 15.

281. The method according to claim 279 or 280, wherein the RGN polypeptide comprises one amino acid sequence from SEQ ID NOs: 3, 11, and 15.

282. The RGN polypeptide and the guide RNA are a) A guide RNA comprising an RGN polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 3, a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 4, and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO: 5, b) A guide RNA comprising an RGN polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 11, a crRNA repeat having the nucleotide sequence of SEQ ID NO: 12, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 12, and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO: 13 or 120, and c) The method according to claim 279, selected from the group comprising: an RGN polypeptide having an amino acid sequence having at least 90% sequence identity with SEQ ID NO: 15; a crRNA repeat having the nucleotide sequence of SEQ ID NO: 16, or a nucleotide sequence having one or two nucleotides different from SEQ ID NO: 16; and a tracrRNA having a nucleotide sequence having at least 90% sequence identity with SEQ ID NO:

17.

283. The RGN polypeptide and the guide RNA are a) A guide RNA comprising an RGN polypeptide containing the amino acid sequence of SEQ ID NO: 3, a crRNA repeat having the nucleotide sequence of SEQ ID NO: 4, and a tracrRNA having the nucleotide sequence of SEQ ID NO:

5. b) A guide RNA comprising an RGN polypeptide containing the amino acid sequence of SEQ ID NO: 11, a crRNA repeat having the nucleotide sequence of SEQ ID NO: 12, and a tracrRNA having the nucleotide sequence of SEQ ID NO: 13 or 120, and c) The method according to claim 282, selected from the group comprising an RGN polypeptide having the amino acid sequence of SEQ ID NO: 15, and a guide RNA having a crRNA repeat having the nucleotide sequence of SEQ ID NO: 16 and a tracrRNA having the nucleotide sequence of SEQ ID NO:

17.

284. The method according to claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and guide RNA according to claim 282(a) or 283(a), and the guide RNA includes a spacer having a nucleotide sequence complementary to the target sequence of SEQ ID NO: 77 or 78.

285. The method according to claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and guide RNA according to claim 282(a) or 283(a), and the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83, or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by one or two nucleotides.

286. The method according to claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and guide RNA according to claim 282(a) or 283(a), and the guide RNA includes a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

287. The method according to any one of claims 279 to 286, wherein the guide RNA is a single guide RNA.

288. The method according to claim 287, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and guide RNA according to claim 282(a) or 283(a), and the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

289. The method according to any one of claims 222 to 288, wherein the PAM is present only in the mutHTT allele and not in the wild-type HTT allele.

290. The method according to any one of claims 222 to 289, wherein the nucleic acid molecule encoding the RGN polypeptide is mRNA.

291. The method according to any one of claims 222 to 289, wherein the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA are located within the viral vector.

292. The vector according to claim 291, wherein the viral vector is a lentiviral vector, a baculovirus vector, or an adeno-associated virus (AAV) vector.

293. The method according to claim 292, wherein the AAV vector is AAV5.

294. A method for detecting mutant huntingtin (mutHTT) protein and wild-type HTT (wtHTT) protein in a sample, wherein the method is a) Applying the modified sample to a capillary tube containing a sieving medium, b) Applying a voltage difference to the capillary tube to separate the proteins in the sample by molecular weight via electrophoresis, c) Immobilizing the separated protein within the capillary tube, d) Applying a first antibody or fragment thereof that can bind to both mutHTT and wtHTT to the capillary tube, e) Applying a second antibody or a fragment thereof that can bind to the first antibody to the capillary tube, wherein the second antibody includes a detectable label. f) A method comprising detecting the detectable sign.

295. The method according to claim 294, wherein the sieving medium is a hydrophilic polymer matrix.

296. The method according to claim 294 or 295, wherein the sample is a biological sample.

297. The method according to any one of claims 294 to 296, wherein the detectable label is a chemiluminescent label or a fluorescent label.

298. The method according to any one of claims 294 to 297, wherein the mutHTT and wtHTT proteins in the sample are quantified by comparison with a standard curve.

299. The method according to any one of claims 294 to 298, wherein the method can degrade the mutHTT protein from the wtHTT.

300. A method for alleviating or delaying the onset of one or more symptoms of Huntington's disease (HD) in a subject, wherein the method applies to the subject, a) A guide RNA having crRNA of sequence number 8 or 106 and tracrRNA of sequence number 9 or 107, or b) A method comprising delivering an adeno-associated virus (AAV) 5 vector comprising a single guide RNA having the nucleotide sequence of SEQ ID NO: 25 or 26.

301. The aforementioned subject is, a) At least 36 CAG repeats within exon 1, b) The method according to claim 300, comprising a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele contains thymine at a position corresponding to position 151 of sequence number 1, and a mutant huntingtin (mutHTT) allele.

302. The method according to claim 300 or 301, wherein the method comprises administering the AAV5 vector by intrastriatal injection.

303. The method according to any one of claims 300 to 302, wherein the subject has a reduced level of the mutHTT protein encoded by the mutHTT allele compared to the level of the control HTT protein or the level of the wild-type HTT protein.

304. The method according to claim 303, wherein a reduction of at least 40% in the level of mutHTT protein is observed compared to the level of a control HTT protein or the level of wild-type HTT protein.

305. The method according to claim 303 or 304, wherein a decrease in the level of mutHTT protein is observed within 4, 6, 8, 10, or 12 weeks after administration of the vector.

306. The method according to any one of claims 303 to 305, wherein a decrease in the level of mutHTT protein is observed in at least 50% of the striatal cells of the subject.

307. The method according to any one of claims 300 to 306, wherein the level of mutHTT mRNA is reduced by at least 40% compared to the level of control HTT mRNA or wild-type HTT mRNA.

308. The method according to any one of claims 300 to 307, wherein the mutHTT allele has at least 40 CAG repeats in exon 1.

309. The method according to any one of claims 300 to 308, wherein the mutHTT allele has at least 56 CAG repeats in exon 1, and the subject is younger than 18 years of age.

310. The method according to any one of claims 300 to 309, wherein the administration is performed before the onset of symptoms of Huntington's disease.

311. The method according to any one of claims 300 to 310, wherein the method comprises preventing the onset of one or more symptoms of Huntington's disease.

312. The method according to any one of claims 300 to 311, wherein the subject has at least one symptom of Huntington's disease.

313. The method according to any one of claims 300 to 312, wherein only the mutated HTT allele is edited.