Compositions and methods for treating huntington's disease by editing mutant huntington genes

By using the CRISPR-Cas system to guide RNA-guided nucleases to recognize and cleave the mutant HTT allele, the problem of the inability to completely eliminate mutant huntingtin protein in existing technologies has been solved, resulting in a significant reduction in mutant HTT mRNA and protein and improving HD symptoms.

CN121241141APending Publication Date: 2025-12-30LIFEEDIT THERAPEUTICS INC
View PDF 39 Cites 0 Cited by

Patent Information

Application Number
CN202480035202.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2024-04-12
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing RNA interference methods cannot completely eliminate the expression of mutant huntingtin protein. Huntington's disease (HD) patients need more effective genome editing methods to reduce the content of mutant HTT protein in order to improve symptoms.

Method used

An RNA-guided nuclease (RGN) using the CRISPR-Cas system recognizes and cleaves the pre-spacer adjacent motif (PAM) in the mutant HTT allele, introducing insertions or deletions (INDELs) to reduce the expression of mutant HTT mRNA and protein in an allele-specific manner.

Benefits of technology

Significantly reduced the levels of mutant HTT mRNA and protein in vivo and in vitro, achieving a reduction of at least 40% for at least 12 weeks, without affecting the expression of wild-type HTT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121241141A_ABST
    Figure CN121241141A_ABST
Patent Text Reader

Abstract

The present invention provides compositions and methods for cleaving a mutant Huntington (mutHTT) allele. The composition comprises CRISPR RNA, guide RNA and nucleic acid molecules for coding the CRISPR RNA and the guide RNA. Vectors and host cells comprising the nucleic acid molecules are also provided. Further provided is an RNA-guided nuclease (RGN) system for cleaving mutHTT alleles, wherein the RGN system comprises an RNA-guided nuclease and a guide RNA. The compositions may be used to cleave or modify the mutHTT allele, and / or to modify the expression of the mutHTT allele. The compositions are further useful for the treatment of Huntington's Disease (HD), especially in an allele-specific manner.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to the following U.S. provisional patent applications: U.S. Provisional Application No. 63 / 495,725, filed April 12, 2023; U.S. Provisional Application No. 63 / 497,904, filed April 24, 2023; U.S. Provisional Application No. 63 / 518,231, filed August 8, 2023; U.S. Provisional Application No. 63 / 593,881, filed October 27, 2023; and U.S. Provisional Application No. 63 / 555,290, filed February 19, 2024, each of which is incorporated herein by reference in its entirety.

[0003] References to sequence lists submitted electronically as XML files

[0004] This application contains a sequence list, which was filed in XML format with the USPTO Patent Centre and is incorporated herein by reference in its entirety. The XML copy was created on April 10, 2024, and is named L103438_1300WO_0257_6_SL, with a size of 251,299 bytes. Technical Field

[0005] This invention relates to the fields of molecular biology and gene editing. Background Technology

[0006] Huntington's disease (HD) is a hereditary neurodegenerative disorder caused by an amplification of the cytosine-adenine-guanine (CAG) trinucleotide in the huntingtin (HTT) gene (Huntington's Disease Collaborative Research Group, 1993, Cell 72:971-983). The resulting polyglutamine (polyQ)-containing mutant HTT disrupts the function of the wild-type HTT (wtHTT), leading to neurological stress and functional impairment (Kaemmerer et al., 2019, Degenerative Neurological and Neuromuscular Disease 9:3-17). HD patients develop striatal atrophy and cognitive impairment, followed by progressive psychomotor deficits (Ross et al., 2014, Nature Review Neurol. 10:204-216). Mouse models expressing full-length or exon 1 of human HTT containing amplified CAG repeat sequences reproduced HD pathophysiology (Southwell et al., 2016, Human Molecular Genetics 25(17):3654-3675; Slow et al., 2003, Human Molecular Genetics 12(13):1555-1567; Raamsdonk et al., 2007, Neurobiology of Disease 26:189-200; Southwell et al., 2017, Human Molecular Genetics 26(6):1115-1132). Reducing the content of mutant HTT in HD animal models led to improvements in motor and neuropathological abnormalities, which supports HTT reduction as a treatment method (Miniarikova et al., 2016, Molecular Therapy - Nucleic Acids 5(3):e297; Caron et al., 2020, Nucleic Acids Research 48(1):36-54; Spronck et al., 2019, Molecular Therapy Methods Clin Dev 13:334-343; Stanek et al., 2014, Human Gene Therapy 25(5):461-474).

[0007] While RNA interference is being developed to reduce the levels of mutant huntingtin protein, RNAi therapy does not completely eliminate the expression of mutant huntingtin protein. Targeted genome editing offers the opportunity to introduce changes at the genomic level. Targeted genome editing or modification is rapidly becoming an important tool in basic and applied research because it allows for modifications to the genome, such as cutting, deleting, inserting, or substituting nucleotides in nucleic acids, regulating gene expression at specific locations in the genome, and many other possible modifications. Initial efforts in genome editing involve designing nucleases capable of editing nucleic acids to recognize and specifically bind to the target nucleic acid sequence to be edited. However, the engineering of nucleases requires significant time and experimentation to obtain nucleases that can effectively edit specific sequences. Genome editing systems using RNA-guided nucleases, such as the CRISPR-Cas bacterial system's clustered regularly spaced short palindromic repeats (CRISPR)-associated (Cas) proteins, function by complexing the nuclease with guide RNA. Hybridization of the guide RNA with a specific target sequence allows for editing at specific locations in the genome. Therefore, genome editing systems using RNA-guided nucleases may be less costly and more efficient for editing genome sequences because nucleic acids are generally easier to design and redesign compared to nucleases.

[0008] Therefore, patients suffering from diseases associated with specific defects in the genome, such as Huntington's disease, will benefit from the development of RNA-guided nuclease systems that can edit genome defects for therapeutic purposes. Summary of the Invention

[0009] Compositions and methods for cleaving the mutant Huntington's disease (mutHTT) allele are provided. The composition comprises CRISPR RNA, guide RNA, and a nucleic acid molecule encoding it. A vector containing the nucleic acid molecule and a host cell are also provided. Further, an RNA-guided nuclease (RGN) system for cleaving the mutHTT allele is provided, wherein the RGN system comprises an RNA-guided nuclease and guide RNA. This composition can be used to cleave or modify the mutHTT allele, and / or modify the expression of the mutHTT allele. This composition is also suitable for treating Huntington's disease (HD), particularly in an allele-specific manner.

[0010] A method for lysing the mutHTT allele in cells includes introducing an RGN or a nucleic acid molecule encoding an RGN and a guide RNA or a nucleic acid molecule encoding a guide RNA, wherein the mutHTT allele contains a single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele generates a protospacer sequence adjacent motif (PAM), and wherein the RGN is able to recognize the PAM and lyse the mutHTT allele.

[0011] A method for improving one or more symptoms of HD or delaying onset in an individual in need comprises administering an RGN system to the individual, wherein the individual's mutHTT allele contains a SNP allele in exon 50 that generates a protospacer sequence adjacent motif (PAM) recognized by the RGN. The RGN then cleaves and edits the mutHTT allele, resulting in a reduced level of the mutHTT protein encoded by the mutHTT allele compared to control individuals or wild-type HTT protein. Brief description of the attached figures

[0012] Figure 1 This shows the percentage of insertions and / or deletions (INDELs) in patient fibroblasts nuclearly transfected with APG07433.1 nuclease and SGN002908 or SGN002911 guide RNA or APG05586 nuclease and SGN004282 guide RNA.

[0013] Figure 2A and 2B The percentage of INDELs and edited reads in patient fibroblasts transfected with RNA guided by APG05586 nuclease and SGN004282, SGN008949, or SGN007707 are shown, respectively. Figure 2B No C-allelic edited reads were detected in fibroblasts of patients who underwent any tests using APG05586 nuclease and SGN004282 guide RNA.

[0014] Figures 3A-3B Immunofluorescence analysis of neuronal marker genes in forebrain neurons derived from induced pluripotent stem cells (iPSCs) is provided. Figure 3A (Left image) shows forebrain neurons derived from iPSCs expressing the neuronal markers Tuj1 (β-tubulin III, green) and γ-aminobutyric acid (GABA) (red), magnification 10X; (Middle image) shows cells also expressing the neuronal markers Tuj1 (green) and microtubule-associated protein 2 (MAP2) (red), MAP2 being a marker of mature neurons, magnification 10X; (Right image) shows cells derived from iPSCs expressing Tuj1 (green) and the forebrain neuron-specific marker FoxG1 (red), magnification 20X. Cell nuclei are labeled with DAPI (blue). Figure 3B This shows the FACS analysis of forebrain neurons derived from iPSCs. Negative controls are shown in the leftmost image. The images from left middle to rightmost present the results stained with the following antibody combinations: Ki67 and neurofilament heavy chain (NEFH) (left middle); GABA and MAP2 (right middle); and GFAP and Tuj1 (rightmost).

[0015] Figure 4Provides immunofluorescence analysis of neuronal marker genes in medium-sized multispinous neurons derived from iPSCs: cAMP-regulated phosphoproteins (DARPP-32), GABA, MAP2, or Ctip2 with an apparent molecular weight of 32 kDa.

[0016] Figure 5 Provide display as Figure 1 The figure describes the percentage of INDEL induced in pluripotent stem cells (iPSCs), neural progenitor cells (NPCs), forebrain neuronal progenitor cells (FBPs), or forebrain neurons (FBNs) using the indicated guide RNA (and appropriate nucleases).

[0017] Figure 6 This study depicts the AAV5-mediated striatal delivery of APG07433.1, which elicits substantial amounts of AAV5 vector DNA in clinically relevant brain regions, the striatum, and the cortex. Dose response was observed at 4 weeks and 3 months post-administration. Each point represents an individual mouse and is shown as mean ± SE. Naïve and vector-treated animals were not depicted because they were below the lower limit of quantitation (LLOQ). vg = viral genome.

[0018] Figure 7 This demonstrates AAV5-mediated intrastriatal delivery of APG07433.1, which elicited strong APG07433.1 transgene expression 4 weeks post-administration. dPCR analysis from the right striatum. Each point represents an individual mouse and is shown as mean ± SE. [n=3–6 mice / group].

[0019] Figure 8 A dose-dependent reduction in mutant HTT protein in the striatum was confirmed at 4 weeks and 3 months after administration of AAV5-JeT-APG07433.1-SGN002908. Each point represents an individual mouse and is shown as mean ± SE. [4 weeks, n=2-6 mice / group; 3 months, n=2-10 mice / group]. **P<0.01, ****P<0.0001.

[0020] Figure 9 The results showed a dose-dependent decrease in mutant HTT mRNA 3 months after administration of AAV5-JeT-APG07433.1-SGN002908 to the striatum. Points represent individual mice and are shown as mean ± SE. [n=2–10 mice / group]. Compared with untreated animals, * P < 0.05; ** P < 0.01.

[0021] Figure 10INDEL next-generation sequencing (NGS) analysis results are provided, confirming editing in the striatum of animals treated with 6.4E10 vg and 3.6E11 vg AAV5-JeT-APG07433.1-SGN002908 at 4 weeks and 3 months post-treatment. Points represent individual mice and are shown as mean ± SE. [4 weeks, n=2–4 mice / group; 3 months, n=2–10 mice / group].

[0022] Figure 11 Intrastral injection of 1.72E11 vg AAV5-hU6-SGN004282-JeT-APG05586 into the striatum resulted in stable AAV5 vector content, inducing APG05586 nuclease editing and a 30% reduction in mutHTT protein. Points represent individual animals and are shown as mean ± SE. The untreated and vector groups showed below the quantitative limit of vector biodistribution. *P<0.05 compared to the untreated control group; unpaired t-test performed.

[0023] Figure 12 The vector genome biodistribution in the striatum and cortex after delivery into the striatum, and subsequent APG05586 nuclease expression, are shown. Each point represents an individual animal and is shown as mean ± SE. The untreated group was below the quantitative limit of vector biodistribution.

[0024] Figure 13 Intrastral administration of AAV5-hU6-SGN004282-JeT-APG05586 and AAV5-hU6-SGN004282-hSyn-APG05586 induced a reduction in mutHTT protein, confirming editing. Points represent individual animals and are shown as mean ± SE. An unpaired Student's t-test was used; **P < 0.01 compared to untreated animals.

[0025] Figure 14 The striatal delivery of 2.84E11 vg AAV5-hU6-SGN004282-hSyn-APG05586 showed widespread vector genome disposition within the striatum, leading to APG05586 nuclease expression. Each point represents an individual animal and is shown as mean ± SE. The untreated group was below the quantitative limit of vector biodistribution.

[0026] Figure 15Intrastral administration of AAV5-hU6-SGN004282-hSyn-APG05586 induced a decrease in mutHTT mRNA and protein in BACHD mice, confirming genome editing. Points represent individual animals and are shown as mean ± SE. Using an unpaired Student's t-test, *P<0.05 and ****P<0.0001 were observed compared to untreated animals.

[0027] Figure 16A and 16B This shows the biodistribution of AAV and nuclease expression in BACHD mice. Figure 16A This study demonstrates AAV5-mediated intrastriatal delivery of APG05586, a codon-optimized construct, in BACHD mice, resulting in clinically relevant parenchymal content in brain regions, the striatum, and the cortex. Deployment was evaluated 6 weeks post-administration. Points represent individual mice and are shown as mean ± SE. Untreated and mediator-treated animals are not depicted because they are below LLOQ. Figure 16B This shows nuclease expression induced by AAV5-mediated striatal delivery of the codon-optimized construct APG05586 in BACHD mice. Deployment was evaluated 6 weeks post-administration. Each point represents an individual mouse and is shown as mean + SE. Untreated and vector-treated animals are not depicted because they are below LLOQ.

[0028] Figure 17 Following intrastriatal administration in BACHD mice, the expression of AAV5 box-reduced mutant huntingtin proteins was shown to be SGN004282 driven by the hU6 (249–318 bp) promoter and APG05586 directed by multiple promoters (Jet, hSyn, CMVeb, EFS). Points represent individual mice and are shown as mean ± SE. *P<0.05, **P<0.01, ****P<0.0001, using one-way ANOVA with post-hoc Dunnett's evaluation.

[0029] Figure 18 The editing was confirmed by NGS INDEL analysis. The percentage of INDEL events in the striatum and cortex of BACHD mice after administration of an AAV5 box expressing SGN004282 driven by the hU6 (249–318 bp) promoter and APG05586, a mammalian codon-optimized promoter directed by multiple promoters (Jet, hSyn, CMVeb, EFS), confirmed genome editing. Points represent individual mice and are depicted as mean ± SE.

[0030] Figures 19A-19CWe provide results after introducing incremental AAV5-hU6-SGN004282-hSyn-APG05586mco-SV40pA with mammalian codon-optimized APG05586 into BACHD mice for 6 weeks. Figure 19A Provides information on the vector biodistribution in the striatum and cortex at increasing viral doses. Figure 19B The study showed a decrease in mutHTT protein in the striatum at increasing viral doses. Figure 19C Showing when the mutHTT percentage decreases Figure 19B Dose escalation studies. Additionally, Figure 19C The graph on the left side of the chart compares the 2.7e11 viral genome (vg) at the same dose with optimization study #2 at the same dose, demonstrating the effect of codon optimization with APG05586. Finally, the codon-optimized construct administered with 7.32e10 vg showed reproducibility in its effect on mutHTT protein content.

[0031] Figures 20A-20E This study demonstrates the specificity of capillary electrophoresis (CE) immunoassay for mutHTT and wild-type HTT. Figure 20A and 20B Provide electrophoresis images of SDS sample blanks, where Figure 20A Display full size and Figure 20B Show a magnified view. Figure 20C and 20D Electrophoresis images of Q73HTT (mutHTT) or Q7HTT (wild-type HTT) at multiple concentrations (0.12 ng / ml to 30 ng / ml) are provided. Figure 20E Showing respectively from Figure 20C and 20D Electrophoresis images of Q73HTT and Q7HTT were used as protein blots.

[0032] Figure 21A and 21B To demonstrate the specificity of CE immunoassay for mutHTT (Q73HTT) and wild-type HTT (Q7HTT) proteins. Figure 21A Provides a chart showing the linearity of Q7HTT and Figure 21B Provides a chart showing the linearity of the Q73HTT.

[0033] Figure 22A and 22B The results of CE immunoassays are shown for brain samples from BACHD mice treated with the AAV5 construct containing SGN004282 and codon-optimized APG05586, compared with untreated animals. Figure 22A Provides protein blotting and Figure 22B Provide electrophoresis images.

[0034] Figure 23A and 23B The results of CE immunoassay are shown for brain samples from BACHD mice treated with CMV (treatment 1 and group A) or EFS (treatment 2 and group B) or untreated mice (untreated or group G). Figure 23A It was confirmed that mutHTT was reduced in treated mice compared to untreated mice. Figure 23B Electrophoresis images from these studies are provided.

[0035] Figures 24A-24F Results are provided in the striatum of BACHD mice 12 weeks after the introduction of incremental doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA (each with mammalian codon-optimized APG05586). The low, intermediate, and high doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA were 2.05E10 vg, 7.28E10 vg, and 2.94E11 vg, respectively. The low, medium, and high doses of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are 2.05E10 vg, 7.28E10 vg, and 2.05E11 vg, respectively. Figure 24A , 24B 24C and 24D provide the biodistribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the striatum at increasing viral doses, respectively. Figure 24E Shows the percentage of INDEL formation in the striatum at increasing viral doses. Figure 24F The mutHTT protein in the striatum decreased with increasing viral doses. Figures 24A-24F In the various graphs, the results of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are displayed on the left side of the graph, and the results of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are displayed on the right side.

[0036] Figure 25A-25FResults are provided in the striatum of BACHD mice 12 weeks after the introduction of incremental doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA or AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA (each with mammalian codon-optimized APG05586) into the cortex. The low, intermediate, and high doses of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA were 2.05E10 vg, 7.28E10 vg, and 2.94E11 vg, respectively. The low, medium, and high doses of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are 2.05E10 vg, 7.28E10 vg, and 2.05E11 vg, respectively. Figure 25A , 25B 25C and 25D provide the biodistribution of the vector, guide RNA, APG05586 mRNA, and APG05586 protein in the cortex at increasing viral doses, respectively. Figure 25E The study showed a decrease in mutHTT protein in the cortex with increasing viral doses. Figure 25F Shows the percentage of INDEL formation in the cortex at increasing viral doses. Figure 25A-25F In the various graphs, the results of AAV5-hU6-SGN004282-CMVeb-APG05586mco-SV40pA are displayed on the left side of the graph, and the results of AAV5-hU6-SGN004282-EFS-APG05586mco-bGHpA are displayed on the right side.

[0037] Figures 26A-26E Results are provided in the striatum of animals treated with SEQ ID NO: 36, SEQ ID NO: 121, SEQ ID NO: 122 or SEQ ID NO: 123 6 weeks after administration of the test article. Figures 26A-26C The biodistribution of the vector, APG05586mco mRNA, and guide RNA are provided respectively. Figure 26D Show the percentage of INDEL formation and Figure 26E The results showed a decrease in mutHTT protein.

[0038] Figures 27A-27E The AAV construct with c-MYC NLS and NLS adaptor proteins showed improved INDEL production activity in HEK293t cells and iPSC-derived astrocytes using AAV5 or AAV6 serotypes. Figure 27A Provides immunofluorescence images of astrocytes derived from iPSCs. Figure 27BFlow cytometry analysis of iPSC-derived astrocytes stained with anti-glial fibrillary acidic protein (GFAP)-488 antibody. Figure 27C-27E Displayed in AAV6 ( Figure 27C and 27D ) or AAV5 ( Figure 27E Transducing astrocytes derived from iPSCs using SEQ ID NO: 36, 121, 122, or 123 ( Figure 27C ) or HEK293t cells ( Figure 27D and 27E The INDEL rate after that.

[0039] Figure 28 Biodistribution of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123) in adult cynomolgus monkeys following bilateral striatal administration. Animals were administered 225 μL / animal (75 μL / tail nucleus + 150 μL / putamen). Low dose (N=2); high dose (N=3). Vector genome was determined by qPCR using a primer-probe set targeting the APG05586mco sequence.

[0040] Figure 29 This study demonstrated the expression of APG05586mco in the brain of adult cynomolgus monkeys following bilateral striatal administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). Animals were administered 225 μL / animal (75 μL / tail nucleus + 150 μL / putamen). Low dose (N=2); high dose (N=3). mRNA transcripts were quantified by qPCR using a primer-probe set targeting the APG05586mco sequence.

[0041] Figure 30 Demonstrated expression of SGN004282 guide RNA in the brain of adult cynomolgus monkeys following bilateral striatal administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123). Animals were administered 225 μL / animal (75 μL / tail nucleus + 150 μL / putamen). Low dose (N=2); high dose (N=3). mRNA transcripts were quantified by qPCR using a primer-probe set targeting the SGN004282 sequence.

[0042] Figure 31 Biodistribution of pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123) in peripheral tissues after bilateral striatal administration in adult cynomolgus monkeys. Animals were administered 225 μL / animal (75 μL / tail nucleus + 150 μL / putamen). Low dose (N=2); High dose (N=3). Sample groups were used as recommended under ICH S12 guidelines: Nonclinical Biodistribution Considerations for Gene Therapy Products. Vector DNA was present in 3 of 12 tissues after low-dose administration; vector DNA was present in 6 of 12 tissues after high-dose administration. No evidence of vector DNA was found in the testes / ovaries. *: No evidence of vector DNA was found.

[0043] Figures 32A-32D This diagram shows the procedure for analyzing the immunogenicity of samples obtained from cynomolgus monkeys before and after administration of pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123) packaged with AAV5. CSF = Cerebrospinal fluid; PBMC = Peripheral blood mononuclear cells; DC = Dendritic cells; MHCII = Major histocompatibility complex II. IAV = Peptide library derived from influenza A virus. R10 = Negative control, culture medium alone. PHA / SEB = Positive control of phytohemagglutinin for T-cell stimulation-nonspecific TCR-independent stimulation.

[0044] Figure 33 This figure shows the pre-existing and post-treatment (day 29) anti-APG05586mco nuclease antibody measurements in cynomolgus monkey serum. No increase in serum antibody reactivity to APG05586mco was observed in cynomolgus monkeys (N=6) following administration of AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123).

[0045] Figure 34The total amount of anti-APG05586mco antibody in the cerebrospinal fluid of cynomolgus monkeys after administration of pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123) packaged with AAV5 was shown.

[0046] Figures 35A-35D This study showed dose-dependent distributions, nuclease transgene expression, and decreased muHTT protein levels in clinically relevant HD mouse models. Four weeks after administration of the vector or AAV5-packaged pAAV-hU6(249bp)-SGN004282-CMVeb-c-MYC-NLS-APG05586mco-c-MYC-NLS-SV40pA(179bp) (SEQ ID NO: 123) to the striatum of BACHD mice, striatal tissue was harvested, and reductions in AAV vector, nuclease transgene expression (mRNA and protein), and muHTT protein were assessed in bulk lysed tissue samples. Points represent mean ± SE, with each dose assessed in 4 to 6 animals. Animals treated with the vector showed lower LLOQs than those analyzed for both vector and transgene, with a zero percentage reduction in muHTT protein. Detailed Implementation

[0047] Benefiting from the teachings presented in the foregoing description and related drawings, those skilled in the art will conceive of many modifications and other embodiments of the invention set forth herein. Therefore, it should be understood that the invention is not limited to the specific embodiments disclosed, and modifications and other embodiments are intended to be included within the scope of the appended embodiments. Although specific terminology is used herein, it is used in a general and descriptive sense only and not for limiting purposes.

[0048] I. Overview

[0049] HD is an autosomal dominant disease that causes progressive degeneration of neural tissue in the brain. Huntington's disease results from the following: an amplified trinucleotide repeat sequence in the HTT gene, which leads to a significant increase in the repeat sequence of the trinucleotide (cytosine-adenine-guanine; CAG) motif in the mutant HTT protein, producing polyglutamine (polyQ) bundles and disrupting the function of wild-type huntingtin protein.

[0050] The compositions and methods disclosed in this invention utilize single nucleotide polymorphisms (SNPs) in the mutant huntingtin (mutHTT) allele, which generate adjacent motifs of the original spacer sequence of the mutHTT allele that allow cleavage via RNA-guided nucleases (RGNs). Unbound by theory, editing the mutHTT allele can cause frameshift INDELs, introducing premature stop codons, leading to degradation through nonsense-mediated decay, thereby knocking out the full-length mutHTT gene. Wild-type huntingtin protein has been shown to support key cellular and neural functions; therefore, a selective strategy targeting only disease-associated mutant HTT is slightly advantageous and is being explored in preclinical and clinical settings (O'Regan et al., 2020, *Scientific Reports* 10:17269; Tabrizi et al., 2019, *The New England Journal of Medicine* 380:2307-2316). Therefore, in some embodiments, the compositions and methods disclosed in this invention provide an allele-specific method that targets only the mutHTT allele and not the wtHTT allele. The allele-specific method disclosed in this invention targets cells and patients that are heterozygous for the SNP, wherein the SNP allele linked to the CAG amplification of the mutHTT allele is an RGN generating a PAM. Introducing the RGN that recognizes the PAM and a guide RNA targeting a sequence adjacent to the PAM into the cells or patients causes the mutHTT allele to cleave at or near the SNP, introducing an INDEL (insertion or deletion), resulting in a decrease in mutHTT mRNA and protein levels. Due to the heterozygous nature of the cells or patients at the SNP, only the mutHTT allele will cleave and only the mutHTT protein level will decrease, while the wtHTT allele and protein levels remain unchanged.

[0051] While other methods have demonstrated mutHTT editing in HD models, none have employed a SNP-derived PAM-dependent approach within exon 50, which allows for allele-specific reductions in mutHTT levels. This disclosure, for the first time, provides a single AAV delivery construct containing both an RNA-guided nuclease and a gRNA targeting the mutant allele in exon 50 of Huntington's disease using a SNP-derived PAM-dependent approach. This construct is delivered in vivo to both the striatum and cortex, regions known to be important for the pathogenesis of Huntington's disease. Importantly, this construct demonstrates allele-specific reductions in mutant HTT mRNA and protein both in vitro and in vivo. Notably, the method disclosed in this invention allows for a reduction of at least 40% in mutHTT mRNA and protein levels that persists for at least 12 weeks following treatment.

[0052] Therefore, the compositions and methods disclosed in this invention can be used to treat HD in individuals of need by reducing HTT levels, and in some embodiments, this reduction is allele-specific, wherein only mutHTT levels are reduced while wild-type HTT (wtHTT) expression remains unchanged. In some embodiments, the compositions and methods disclosed in this invention can reduce mutHTT mRNA and protein levels by at least 40% in at least 50% of striatal neurons.

[0053] II. Huntington (HTT) gene

[0054] Huntington's disease (HD) is a hereditary autosomal dominant disorder characterized by progressive degeneration of brain neurons caused by the amplification of the CAG repeat sequence in the first exon of the huntingtin gene on chromosome 4 (Huntington Disease Collaborative Research Group, 1993, Cell 72:971-983). The disruption of wild-type HTT protein by the polyglutamine-containing mutant HTT protein leads to neurological stress and dysfunction, ultimately resulting in striatal atrophy, cognitive impairment, and progressive mental and motor deficits (Kaemmerer et al., 2019, Degenerative Neuromuscular Diseases 9:3-17; Ross et al., 2014, Nature Review Neurology 10:204-216).

[0055] The huntingtin gene is large, spanning 180 kb and consisting of 67 exons. A non-limiting example of the HTT gene is the human HTT gene shown as NCBI gene ID No. 3064, and a non-limiting example of the HTT protein is the human huntingtin protein shown as NCBI reference sequence ID No. NP_001375421.1 and shown as SEQ ID NO: 85 in this document (both are incorporated herein by reference), which contains 21 glutamines in a poly-Q bundle and matches the GRCh38 reference genome.

[0056] The CAG triplet repeat region is located in exon 1 of the HTT gene, and individuals with more than 26 CAG repeat sequences are more likely to pass on the amplified CAG repeat sequences to their offspring. HTT genes with 27-35 CAG repeat sequences are considered intermediate alleles with approximately 0% probability of exhibiting the disease phenotype, but individuals with these intermediate alleles can pass on the amplified repeat sequences to their offspring. Individuals with 36-39 CAG repeat sequences have a higher probability of exhibiting disease symptoms, but this allele is considered incompletely penetrating. Huntington's disease patients have 40 or more CAG repeat sequences and approximately 100% probability of exhibiting disease symptoms. Those with 56 or more CAG repeat sequences in HD typically experience early onset of the disease in childhood or adolescence, classified as juvenile Huntington's disease or juvenile-onset Huntington's disease (JHD) (Tabrizi et al., 2022, Lancet Neurol 21:632-644). Therefore, in some embodiments, cells modified by the compositions and methods disclosed in this invention have an HTT gene with at least 27 CAG repeat sequences in a CAG repeat region in exon 1, referred herein as a mutant HTT gene or allele or mutHTT gene or allele. A CAG repeat region with at least 27 CAG repeat sequences is also referred herein as CAG repeat sequence amplification. Wild-type HTT genes or alleles or wtHTT genes or alleles have fewer than 27 CAG repeat sequences and typically have 15-20 CAG repeat sequences in exon 1. Individuals treated with the compositions and methods disclosed in this invention have an HTT gene containing at least 36 CAG repeat sequences, and in some embodiments at least 40 CAG repeat sequences. Most HD patients are heterozygous for amplified CAG repeat sequences and therefore carry a mutHTT allele with at least 36 or at least 40 CAG repeat sequences and a wtHTT allele with fewer than 27 CAG repeat sequences.

[0057] According to the present invention, the cell or individual further contains a single nucleotide polymorphism (SNP) allele within the mutHTT allele, which may be a major allele (present in most human populations) or a minor allele (present in a small number of populations). It should be noted that a particular genomic location may have multiple minor SNP alleles. As used herein, "single nucleotide polymorphism allele" or "SNP allele" refers to a single nucleotide difference at a specific site in the genome between population members, wherein the difference is the substitution of one nucleotide for another. The SNP present on the mutHTT allele can generate a PAM recognized by the RGN or be contained within a PAM. Therefore, the SNP allele generating the PAM is linked to the CAG repeat sequence amplification of mutHTT, or in other words, the SNP allele is present on the HTT gene (mutHTT allele) with the same copy of the CAG repeat sequence amplification.

[0058] Non-limiting examples of SNPs that can be targeted for allele-specific methods of treating HD are listed in Tables 1 and 2 herein and include NCBI dbSNP No. rs362331. In some embodiments, the SNP allele generates a PAM having the nucleotide sequences NNNNCC, NNRYA, NNGRR, and / or NNGG. In some embodiments, the method is used to lyse a mutant huntington (mutHTT) allele in cells, wherein the mutHTT allele contains a first single nucleotide polymorphism (SNP) allele in exon 1, and wherein the first SNP allele generates a protospacer adjacent motif (PAM) selected from NNNNCC, NNRYA, NNGRR, and / or NNGG, the method comprising introducing an RNA-guided nuclease (RGN) or a nucleic acid molecule encoding an RGN and a guide RNA or a nucleic acid molecule encoding a guide RNA, wherein the RGN is capable of recognizing the PAM and cleaving the mutHTT allele. In some of these embodiments, the RGN has at least 80%, 85%, 90%, 95%, or higher sequence identity with any one of SEQ ID NO: 7, 11, 13, and 15. The SNP allele may be located on either side of the CAG repeat sequence within exon 1. In these embodiments, the RGN that identifies the SNP allele on one side (5' or 3') of the CAG repeat sequence may be used in combination with another nuclease that cleaves the relative ends of the CAG repeat sequence to remove the CAG repeat region in-frame from the mutHTT allele.

[0059] In some implementations, the SNP alleles amplified or “in-phase” with the CAG repeat sequence of mutHTT are located in exon 50. The rs362331 SNP at position 151 in exon 50 of the HTT gene causes a nucleotide difference, shown herein as SEQ ID NO: 1 and 2. The major allele of the rs362331 SNP at position 151 in exon 50 of the HTT gene contains T or thymine (SEQ ID NO: 1), and the minor allele contains C or cytosine at that position (SEQ ID NO: 2). In those genomes containing the rs362331 SNP “T” allele (major allele) in the HTT gene, the presence of T at position 151 in exon 50 is RGN APG05586 (shown as SEQ ID NO: 7) or its active variant or fragment that generates the PAM motif (NNRYA). The RGN APG05586 PAM motif resides on the reverse complementary sequence of a DNA strand containing the T allele, where the "A" in NNRYA PAM is the complementary base of the T allele. Therefore, in individuals containing at least one rs362331 SNP "T" allele, APG05586, along with its corresponding guide RNA complementary to the target sequence upstream (5') of the PAM motif, can bind to and cleave the rs362331 SNP "T" allele.

[0060] Alternatively, in those genomes containing the rs362331 SNP "C" allele (minor allele) in the HTT gene, the presence of C at position 151 in exon 50 (SEQ ID NO: 2) generates a PAM motif for each of the following: RGNAPG07433.1 (shown as SEQ ID NO: 3; PAM of NNNNCC), APG01604 (shown as SEQ ID NO: 11; PAM of NNGRR), and LPG10145 (shown as SEQ ID NO: 15; PAM of NNGG), or an active variant or fragment thereof. Therefore, in individuals containing at least one rs362331 SNP "C" allele, APG07433.1, APG01604, or LPG10145, along with a corresponding guide RNA complementary to the target sequence upstream (5') of the PAM motif, can bind to and cleave the rs362331 SNP "C" allele.

[0061] Cleavage of the mutHTT allele near the SNP site can generate frameshift mutations within the HTT gene, leading to premature termination. Compared to cells or individuals lacking RGN and its homologous guide RNA, the levels of mutHTT mRNA and protein are reduced.

[0062] In most cases, individuals containing the mutHTT allele are heterozygous for that allele and contain both the mutHTT allele and the wild-type HTT (wtHTT) allele. In some embodiments, PAM generated by the SNP allele in mutHTT is not found in the wild-type HTT (wtHTT) allele; therefore, this method is allele-specific because the introduced or applied RGN, along with its homologous guide RNA, only cleaves and edits the mutHTT allele, not the wtHTT allele, and only reduces the levels of mutHTT mRNA and protein. An RGN peptide or RGN system that cannot cleave the wild-type HTT allele means that the RGN peptide or RGN system cannot cleave the wild-type HTT allele at all or cleaves it to a negligible degree, such that the levels of wtHTT mRNA and / or wtHTT protein are not significantly reduced. There is a non-significant reduction, wherein, for example, wtHTT can maintain support for key cellular and neural functions and / or, in the in vivo context (i.e., in individuals who are heterozygous for the mutHTT allele and who have been administered the RGN peptide or RGN system), there are no Huntington's disease symptoms. In some embodiments, the RGN peptide or RGN system is cleaved to a negligible degree, such that the levels of wtHTT mRNA and / or wtHTT protein are reduced by 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.9% or less, 0.8% or less, 0.7% or less, 0.6% or less, 0.5% or less, 0.4% or less, 0.3% or less, 0.2% or less, or 0.1% or less.

[0063] III. Guide RNA

[0064] This disclosure provides a guide RNA and a polynucleotide encoding the same for targeting a target nucleotide sequence in a mutant HTT allele with an associated RNA-guided nuclease (RGN). The term "guide RNA" comprises a nucleotide sequence (i.e., a spacer) that is sufficiently complementary to the target nucleotide sequence in the mutant HTT allele to hybridize with the target sequence and to guide the associated RGN sequence to specifically bind to the target nucleotide sequence. In some embodiments, when the target nucleotide sequence is double-stranded as in the case of DNA, the target nucleotide sequence comprises a non-target strand (which contains a PAM sequence) and a target strand that hybridizes with the spacer of the guide RNA. In these embodiments, the guide RNA is sufficiently complementary to the target strand of the double-stranded target sequence (e.g., the target DNA sequence in the mutant HTT allele) such that the guide RNA hybridizes with the target strand and guides the associated RGN sequence to specifically bind to the target sequence (e.g., the target DNA sequence in the mutant HTT allele). Thus, in some embodiments, the guide RNA comprises a spacer identical to the sequence of the non-target strand, except that uracil (U) replaces thymine (T) in the guide RNA.

[0065] Each guide RNA of an RGN is one or more RNA molecules (usually one or two) that can bind to the RGN and guide it to bind to a specific target sequence, and in embodiments where the RGN has nicking or nuclease activity, it also cleaves the target strand and / or non-target strand. Generally, guide RNAs comprise CRISPR RNA (crRNA) and trans-activating CRISPR RNA (tracrRNA), although some RGNs do not require tracrRNA. Native guide RNAs containing both crRNA and tracrRNA typically consist of two separate RNA molecules that hybridize with each other through a repeating sequence of crRNA and an inverted repeating sequence of tracrRNA. In some embodiments, crRNA and tracrRNA are linked together by a polynucleotide linker (e.g., a tetranucleotide linker) to form a single guide RNA molecule, wherein crRNA and tracrRNA hybridize with each other through a repeating sequence of crRNA and an inverted repeating sequence of tracrRNA. Therefore, guide RNA encompasses single guide RNA (sgRNA), where the crRNA segment and the tracrRNA segment are located in the same RNA molecule or strand. Guide RNA may include non-naturally occurring, chemically modified guide RNA not found in nature, crRNA and / or tracrRNA molecules not found in nature, and / or sequences not found in naturally occurring corresponding molecules.

[0066] This invention provides a CRISPR RNA (crRNA) or a polynucleotide encoding CRISPR RNA that enables an associated RGN to target a target sequence in a mutant HTT allele. As used herein, the term "crRNA" refers to an RNA molecule or a portion thereof comprising: a spacer, which is a nucleotide sequence that hybridizes directly with the target strand of the target sequence; and a CRISPR repeat sequence containing a nucleotide sequence that, alone or together with the hybridized tracrRNA, forms a structure recognized by the RGN molecule. As used herein, the term "tracrRNA" or "trans-activated crRNA" refers to an RNA molecule comprising an inverted repeat sequence having complementarity sufficient to hybridize with at least a portion of the CRISPR repeat sequence of the crRNA to form a structure recognized by the RGN molecule. In some embodiments, additional secondary structures (e.g., stem-loop) within the tracrRNA molecule are required to bind to the RGN.

[0067] crRNA comprises a spacer and a CRISPR repeat sequence. The "spacer" has a nucleotide sequence that directly hybridizes to the target strand of the target sequence of interest (e.g., the target DNA sequence in a mutant HTT allele). The spacer is engineered to have complete or partial complementarity to the target strand of the target sequence of interest (e.g., the target DNA sequence in a mutant HTT allele). In some embodiments, the spacer may contain about 8 nucleotides to about 30 nucleotides, or more. For example, the spacer length may be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, or more nucleotides. In some embodiments, the spacer length is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, the spacer length is about 10 to about 26 nucleotides, or about 12 to about 30 nucleotides. In some embodiments, the spacer length is about 30 nucleotides. In some embodiments, the spacer length is 30 nucleotides. In some implementations, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the spacer and the target strand of the target sequence (e.g., the target DNA sequence in the mutant HTT allele) is between 50% and 99% or greater, including but not limited to about 50%, about 60%, about 70%, about 75%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99% or greater. In some implementations, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the spacer and the target strand of the target sequence (e.g., the target DNA sequence in a mutant HTT allele) is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater. In some implementations, the spacer sequence may be identical to the non-target strand of the target sequence. In some of those implementations where the target sequence is the target DNA sequence, the spacer sequence may be identical to the non-target strand of the target DNA sequence, unless thymine (T) in the target strand is replaced by uracil (U) in the spacer.In the implementation scheme, the spacer does not contain secondary structures, which can be predicted using any suitable multinucleotide folding algorithm known in the art, including but not limited to mFold (see, for example, Zuker and Stiegler (1981) Nucleic Acids Res. 9:133-148) and RNAfold (see, for example, Gruber et al. (2008) Cell 106(1):23-24).

[0068] In some embodiments, the spacer of this disclosure has a nucleotide sequence shown as SEQ ID NO: 80 or differing from SEQ ID NO: 80 by 1 or 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 80 by 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 80 by 1 nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 80. In some embodiments, the spacer of this disclosure has a nucleotide sequence shown as SEQ ID NO: 81 or differing from SEQ ID NO: 81 by 1 or 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 81 by 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 81 by 1 nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 81. In some embodiments, the spacer of this disclosure has a nucleotide sequence shown as SEQ ID NO: 82 or differing from SEQ ID NO: 82 by 1 or 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 82 by 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 82 by 1 nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 82. In some embodiments, the spacer of this disclosure has a nucleotide sequence shown as SEQ ID NO: 83 or differing from SEQ ID NO: 83 by 1 or 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 83 by 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO: 83 by 1 nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO: 83. In some embodiments, the spacer of this disclosure has a nucleotide sequence shown as SEQ ID NO:84 or differing from SEQ ID NO:84 by 1 or 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO:84 by 2 nucleotides. In some embodiments, the spacer has a nucleotide sequence differing from SEQ ID NO:84 by 1 nucleotide. In some embodiments, the spacer has a nucleotide sequence shown as SEQ ID NO:84.

[0069] In addition to the spacer, the crRNA further comprises a CRISPR RNA (crRNA) repeat sequence. The CRISPR RNA repeat sequence comprises a nucleotide sequence that, alone or together with hybrid tracrRNA, forms a structure recognized by the RGN molecule. In some embodiments, the CRISPR RNA repeat sequence may contain about 8 nucleotides to about 30 nucleotides or more. For example, the length of the CRISPR repeat sequence may be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some implementations, the CRISPR repeat sequence is 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides long. In some implementations, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA inverted repeat sequence is approximately or greater than approximately 50%, approximately 60%, approximately 70%, approximately 75%, approximately 80%, approximately 81%, approximately 82%, approximately 83%, approximately 84%, approximately 85%, approximately 86%, approximately 87%, approximately 88%, approximately 89%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, or greater. In a specific implementation, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA inverted repeat sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater.

[0070] In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence of any one of SEQ ID NO: 4, 8, 12, or 16, or an active variant or fragment thereof, which, when included within a guide RNA, directs the associated RNA-guided nuclease sequence provided herein to specifically bind to a target DNA sequence within the mutant HTT allele disclosed herein. In some embodiments, the active CRISPR repeat sequence variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity with any of the following: SEQ ID NO: 4, 8, 12, 16, and 106. In some embodiments, the active CRISPR repeat sequence fragment comprises at least 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, or 22 consecutive nucleotides of any of the following: SEQ ID NO: 4, 8, 12, 16, and 106. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 4 by 1 or 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 4 by 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 4 by 1 nucleotide. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence of SEQ ID NO: 4. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 8 by 1 or 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 8 by 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 8 by 1 nucleotide. In some embodiments, the CRISPR repeat sequence comprises the nucleotide sequence shown as SEQ ID NO: 8. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 12 by 1 or 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 12 by 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 12 by 1 nucleotide. In some embodiments, the CRISPR repeat sequence comprises the nucleotide sequence shown as SEQ ID NO: 12.In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 16 by 1 or 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 16 by 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 16 by 1 nucleotide. In some embodiments, the CRISPR repeat sequence comprises the nucleotide sequence shown as SEQ ID NO: 16. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 106 by 1 or 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 106 by 2 nucleotides. In some embodiments, the CRISPR repeat sequence comprises a nucleotide sequence differing from SEQ ID NO: 106 by 1 nucleotide. In some embodiments, the CRISPR repeat sequence comprises the nucleotide sequence shown as SEQ ID NO: 106.

[0071] In some implementations, the crRNA is an engineered sequence that is not naturally occurring. In some implementations, the specific CRISPR repeat sequence is not inherently linked to the engineered spacer and the CRISPR repeat sequence is considered heterologous to the spacer. In some implementations, the spacer is an engineered sequence that is not naturally occurring.

[0072] The guide RNA disclosed in this invention comprises crRNA and trans-activated CRISPR RNA (tracrRNA), while some compositions and methods disclosed in this invention utilize RGN polypeptides that do not require tracrRNA. The tracrRNA molecule contains a nucleotide sequence comprising a region referred herein as an inverted repeat sequence, which has complementarity sufficient to hybridize with the CRISPR repeat sequence of the crRNA. In some embodiments, the tracrRNA molecule further comprises a region having secondary structures (e.g., stem-loop) or forming secondary structures after hybridization with its corresponding crRNA. In embodiments, the region of the tracrRNA that is fully or partially complementary to the CRISPR repeat sequence contains secondary structures at the 5' end of the molecule and at the 3' end of the tracrRNA. This secondary structure region generally contains several hairpin structures, including nexus hairpins, which are found adjacent to the inverted repeat sequence. The nexus forms the core of the interaction between the guide RNA and the RGN and is located at the intersection of the guide RNA, RGN, and target sequence. The nexus hairpin typically has a conserved nucleotide sequence in the bases of the hairpin stem, and the motif UNANNC is found in many nexus hairpins in the tracrRNA. In some embodiments, the guide RNA or RGN system of this disclosure uses a plurality of tracrRNAs, including UNANNG and CNANNC, which contain atypical sequences in the bases of the hairpin stem connecting the hairpin. In some embodiments, the guide RNA or RGN system of this disclosure uses a single tracrRNA that includes the atypical sequence UNANNG or CNANNC in the bases of the hairpin stem connecting the hairpin. A terminal hairpin is typically present at the 3' end of the tracrRNA; its structure and number may vary, but it typically contains a GC-rich p-independent transcription termination hairpin, followed by a string of Us at the 3' end. See, for example, Briner et al. (2014) Molecular Cell 56:333-339; Briner and Barrangou (2016) Cold Spring Harbor Laboratory Manual; doi:10.1101 / pdb.top090902; and U.S. Patent No. 2017 / 0275648, each incorporated herein by reference in its entirety.

[0073] In some implementations, the inverted repeat sequence of the tracrRNA, which is fully or partially complementary to the CRISPR repeat sequence, contains about 8 nucleotides to about 30 nucleotides or more. For example, the length of the base-pairing region between the tracrRNA inverted repeat sequence and the CRISPR repeat sequence can be about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30 or more nucleotides. In some embodiments, the length of the base-pairing region between the tracrRNA inverted repeat sequence and the CRISPR repeat sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 or more nucleotides. In some embodiments, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA inverted repeat sequence is approximately or greater than approximately 50%, approximately 60%, approximately 70%, approximately 75%, approximately 80%, approximately 81%, approximately 82%, approximately 83%, approximately 84%, approximately 85%, approximately 86%, approximately 87%, approximately 88%, approximately 89%, approximately 90%, approximately 91%, approximately 92%, approximately 93%, approximately 94%, approximately 95%, approximately 96%, approximately 97%, approximately 98%, approximately 99%, or greater. In some implementations, when optimal alignment is performed using a suitable alignment algorithm, the complementarity between the CRISPR repeat sequence and its corresponding tracrRNA inverted repeat sequence is 50%, 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater.

[0074] In some implementations, the entire tracrRNA may contain from about 60 nucleotides to more than about 210 nucleotides. For example, the length of tracrRNA may be about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, about 105, about 110, about 115, about 120, about 125, about 130, about 135, about 140, about 150, about 160, about 170, about 180, about 190, about 200, about 210 or more nucleotides. In some implementations, the tracrRNA is 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 150, 160, 170, 180, 190, 200, 210 or more nucleotides long. In some implementations, the tracrRNA is about 70 to about 105 nucleotides in length, including lengths of about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92, about 93, about 94, about 95, about 96, about 97, about 98, about 99, about 100, about 101, about 102, about 103, about 104, and about 105 nucleotides. In the implementation scheme, the tracrRNA is 70 to 105 nucleotides in length, including lengths of 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, and 105 nucleotides.

[0075] In some embodiments, the tracrRNA comprises a nucleotide sequence of any one of SEQ ID NO: 5, 9, 13, 17, 107, and 120, or an active variant or fragment thereof, which, when included within a guide RNA, directs the associated RNA-guided nuclease sequence provided herein to specifically bind to a target sequence within a mutant HTT allele. In some embodiments, the active tracrRNA sequence variant comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity with any one of the following nucleotide sequences shown: SEQ ID NO: 5, 9, 13, 17, 107, and 120. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 5, 9, 13, 17, 107, and 120. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 5. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 9. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 13. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 17. In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 107.In some embodiments, the active tracrRNA sequence fragment comprises at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides as shown in SEQ ID NO: 120. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 5. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 9. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 13. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 17. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 107. In some embodiments, the active tracrRNA sequence fragment comprises the nucleotide sequence shown in SEQ ID NO: 120.

[0076] When two polynucleotide sequences hybridize under stringent conditions, they are considered substantially complementary. Similarly, if a guide RNA bound to an RGN binds to a target sequence under stringent conditions, the RGN is considered to bind to the specific target sequence in a sequence-specific manner. "Stringent conditions" or "stringent hybridization conditions" are intended to be conditions under which two polynucleotide sequences will hybridize to a detectably greater extent than other sequences (e.g., at least 2-fold compared to background). Stringent conditions are sequence-dependent and will vary depending on the circumstances. Typically, stringent conditions will be conditions with a salt concentration of less than about 1.5 M Na+ ions at pH 7.0 to 8.3, typically about 0.01 to 1.0 M Na+ ion concentration (or other salts), and a temperature of at least about 30°C for short sequences (e.g., 10 to 50 nucleotides) and at least about 60°C for long sequences (e.g., greater than 50 nucleotides). Stringent conditions can also be achieved by adding a destabilizing agent (such as formamide). Exemplary low-toughness conditions include hybridization at 37°C in a buffer of 30% to 35% formamide, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), followed by washing at 50°C to 55°C in 1X to 2X SSC (20X SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Exemplary medium-toughness conditions include hybridization at 37°C in a buffer of 40% to 45% formamide, 1.0 M NaCl, and 1% SDS, followed by washing at 55°C to 60°C in 0.5X to 1X SSC. Exemplary high-toughness conditions include hybridization at 37°C in a buffer of 50% formamide, 1 M NaCl, and 1% SDS, followed by washing at 60°C to 65°C in 0.1X SSC. Depending on the application, the wash buffer may contain approximately 0.1% to approximately 1% SDS. Hybridization duration is generally less than approximately 24 hours, typically approximately 4 to approximately 12 hours. Washing time will be sufficient to allow for at least a period of time to reach equilibrium.

[0077] Tm is the temperature at which 50% of the complementary target sequence hybridizes with the perfectly matched sequence (at defined ionic strength and pH). For DNA-DNA hybrids, Tm can be estimated by the equation from Meinkoth and Wahl (1984), *Analytical Biochemistry*, 138:267-284: Tm = 81.5 °C + 16.6 (log M) + 0.41 (GC%) - 0.61 (formamide%) - 500 / L; where M is the molar concentration of the monovalent cation, GC% is the percentage of guanine and cytosine nucleotides in the DNA, formamide% is the percentage of formamide in the hybridization solution, and L is the base pair length of the hybrid. Generally, stringent conditions are selected to be approximately 5 °C lower than the melting temperature (Tm) of the specific sequence and its complement at defined ionic strength and pH. However, severely stringent conditions can be achieved by hybridization and / or washing at temperatures 1, 2, 3, or 4 °C lower than the melting temperature (Tm); moderately stringent conditions can be achieved by hybridization and / or washing at temperatures 6, 7, 8, 9, or 10 °C lower than the melting temperature (Tm); and lightly stringent conditions can be achieved by hybridization and / or washing at temperatures 11, 12, 13, 14, 15, or 20 °C lower than the melting temperature (Tm). By using equations, hybridization and washing compositions, and the desired Tm, those skilled in the art will understand that variations in hybridization stringency and / or washing solutions are inherently described. Detailed guidelines for nucleic acid hybridization can be found in Tijssen (1993), *Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes*, Part I, Chapter 2 (Elsevier, New York); and Ausubel et al., eds., *Current Protocols in Molecular Biology* (1995), Chapter 2 (Greene Publishing and Wiley-Interscience, New York). See also Sambrook et al. (1989), *Molecular Cloning: A Laboratory Manual* (2nd ed., Cold Spring Harbor Laboratory Press, Plainview, New York).

[0078] The term "sequence specificity" can also refer to the fact that the RGN peptide has a greater affinity for the target sequence than for the randomized background sequence.

[0079] Guide RNA can be a single guide RNA (sgRNA) or a dual guide RNA (dgRNA). A single guide RNA comprises crRNA and tracrRNA on a single RNA molecule, while a dual guide RNA system comprises crRNA and tracrRNA on two different RNA molecules. The crRNA and tracrRNA hybridize to each other through at least a portion of the CRISPR repeat sequence of the crRNA and at least a portion of the tracrRNA (i.e., an inverted repeat sequence), which may be fully or partially complementary to the CRISPR repeat sequence of the crRNA. In embodiments where the guide RNA is a single guide RNA, the crRNA and tracrRNA are separated by a linker nucleotide sequence. The crRNA repeat sequence and tracrRNA linked by the nucleotide linker can be referred to as the backbone of the sgRNA. The backbone can also refer to the crRNA repeat sequence and tracrRNA of the dgRNA.

[0080] The backbone of the guide RNA may contain a nucleotide sequence of any one of SEQ ID NO: 140, 141, and 142, or an active variant or fragment thereof, which, when included within the guide RNA, directs the associated RNA-guided nuclease sequence provided herein to specifically bind to a target sequence within the mutant HTT allele. The backbone of the sgRNA or dgRNA may be engineered to be shorter or longer than its native length while retaining its function. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 2 to approximately 30 nucleotides shorter than the original backbone before engineering. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, or more nucleotides shorter than the original backbone. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 2 to 18 nucleotides shorter than the original backbone. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 2, 4, 6, 8, 10, 12, 14, 16, or 18 nucleotides shorter than the original backbone. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 14 nucleotides shorter than the original backbone. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 16 nucleotides shorter than the original backbone. In some embodiments, the engineered sgRNA or dgRNA backbone is approximately 20 nucleotides shorter than the original backbone.

[0081] In some embodiments, the active backbone variant of the guide RNA disclosed herein may comprise from about 60 nucleotides to more than about 120 nucleotides. For example, the backbone length may be about 60, about 61, about 62, about 63, about 64, about 65, about 66, about 67, about 68, about 69, about 70, about 71, about 72, about 73, about 74, about 75, about 76, about 77, about 78, about 79, about 80, about 81, about 82, about 83, about 84, about 85, about 86, about 87, about 88, about 89, about 90, about 91, about 92. Approximately 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, and approximately 120 or more nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain from about 66 nucleotides to more than about 110 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 66 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 70 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 76 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 90 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 94 nucleotides. In some embodiments, the active backbone variant of the guide RNA of this disclosure may contain 110 nucleotides.

[0082] The active backbone fragment of the guide RNA disclosed herein may comprise at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 7 3, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110 or more consecutive nucleotides: SEQ ID NO: 140, 141 or 142. In some embodiments, the active backbone fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85 or more consecutive nucleotides as shown in SEQ ID NO: 140. In some embodiments, the active backbone fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more consecutive nucleotides as shown in SEQ ID NO: 141. In some embodiments, the active backbone fragment of the guide RNA of this disclosure comprises at least 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105 or more consecutive nucleotides as shown in SEQ ID NO: 142.

[0083] The active backbone variant of the sgRNA disclosed herein may have at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with any of the following: SEQ ID NO: 140-142. In some embodiments, the active backbone variant of the sgRNA disclosed herein has at least 80% sequence identity with any of the following: SEQ ID NO: 140-142. In some embodiments, the active backbone variant of the sgRNA disclosed herein has at least 85% sequence identity with any of the following: SEQ ID NO: 140-142. In some embodiments, the active backbone variant of the sgRNA disclosed herein has at least 90% sequence identity with any of the following: SEQ ID NO: 140-142. In some embodiments, the active backbone variant of the sgRNA of this disclosure has at least 95% sequence identity with any of the following: SEQ ID NO: 140-142. In some embodiments, the active backbone variant of the sgRNA of this disclosure has a nucleotide sequence shown as any of the following: SEQ ID NO: 140-142.

[0084] Generally, the linker nucleotide sequence connecting crRNA and tracrRNA is a nucleotide sequence that does not include complementary bases to avoid the formation of secondary structures within the linker nucleotide sequence or the formation of secondary structures containing the linker nucleotide sequence. In some embodiments, the length of the linker nucleotide sequence between crRNA and tracrRNA is at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12 or more nucleotides. In some embodiments, the linker nucleotide sequence of a single guide RNA is at least 4 nucleotides long. In some embodiments, the linker nucleotide sequence of a single guide RNA is 4 nucleotides long. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence indicated as AAAA, GAAA, ACUU, and CAAAGG. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence indicated as AAAA. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence indicated as GAAA. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence indicated as ACUU. In some embodiments, the linker nucleotide sequence includes a nucleotide sequence indicated as CAAAGG.

[0085] In some implementations, the sgRNA has a nucleotide sequence shown as any of the following: SEQ ID NO: 6, 10, 14, 18 and 25-29.

[0086] Single or dual guide RNAs can be chemically synthesized or transcribed in vitro. Assays for determining sequence-specific binding between RGN and guide RNA are known in the art and include, but are not limited to, in vitro binding assays of expressed RGN to guide RNA, which can be labeled with a detectable marker (e.g., biotin) and used in pull-down assays, wherein the guide RNA:RGN complex is captured via a detectable marker (e.g., streptavidin beads). Control guide RNAs with sequences or structures unrelated to the guide RNA can be used as negative controls for nonspecific binding of RGN to RNA.

[0087] In some embodiments, the guide RNA may be introduced into the target cell in the form of an RNA molecule. The guide RNA may be transcribed in vitro or synthesized chemically. In some embodiments, a nucleic acid molecule encoding the guide RNA is introduced into the target cell. In some embodiments, the nucleic acid molecule encoding the guide RNA is operatively linked to a promoter (e.g., the RNA polymerase III promoter). The promoter may be a native promoter or a promoter heterologous to the nucleic acid molecule encoding the guide RNA.

[0088] In some implementations, as described herein, the guide RNA may be introduced into the target cell as part of a ribonucleoprotein complex, wherein the guide RNA binds to the RGN polypeptide.

[0089] Guide RNA directs the associated RGN to a specific target nucleotide sequence of interest through hybridization with the target sequence of interest. The target sequence can be bound to a nuclease via RNA guidance (and cleaved in some embodiments) in vitro or in cells. The target sequence can comprise DNA, RNA, or a combination of both and can be single-stranded or double-stranded. In some embodiments, the target sequence can be genomic DNA (i.e., chromosomal DNA), plasmid DNA, free DNA, or RNA molecules (e.g., messenger RNA, ribosomal RNA, transfer RNA, microRNA, small interfering RNA). In those embodiments where the target sequence is a chromosomal sequence, the chromosomal sequence can be a nuclear or mitochondrial chromosomal sequence. In the compositions and methods disclosed in this invention, the target sequence is within a double-stranded target nucleic acid molecule (e.g., a target DNA sequence). More specifically, the target sequence is within a mutant HTT allele. In some embodiments, the target sequence is unique within the target genome. In some embodiments, the target sequence comprises a target strand and a non-target strand, and the target sequence has a nucleotide sequence shown as any one of SEQ ID NOs: 75-79 and 130.

[0090] The target sequence is adjacent to the original spacer sequence adjacent motif (PAM), and the non-target strand of the target sequence is the strand containing the PAM. The PAM is immediately adjacent to the target sequence and typically contains N, where N represents any nucleotide. In some embodiments, the PAM contains about 1 to about 10 N, including about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 N. In some embodiments, the PAM contains 1 to 10 N, including 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 N. The PAM may be at the 5' or 3' of the target sequence on the non-target strand. In some embodiments, for the guide RNA and RGN system disclosed in this invention, the PAM is at the 3' of the target sequence on the non-target strand. Generally, the PAM is a common sequence of about 3 to 4 nucleotides, but in some embodiments, its length may be 2, 3, 4, 5, 6, 7, 8, 9, or more nucleotides.

[0091] In some embodiments, the PAM sequence adjacent to the target sequence disclosed in this invention on a non-target chain includes a common sequence shown as NNNCC, NNRYA, NNGRR, and NNGG. In some embodiments, the PAM sequence adjacent to the target sequence on a non-target chain includes a common sequence shown as NNNCC. In some embodiments, the PAM sequence adjacent to the target sequence on a non-target chain includes a common sequence shown as NNRYA. In some embodiments, the PAM sequence adjacent to the target sequence on a non-target chain includes a common sequence shown as NNGRR. In some embodiments, the PAM sequence adjacent to the target sequence on a non-target chain includes a common sequence shown as NNGG. In some embodiments, the PAM sequence is located at the 3' of the target sequence on the non-target chain.

[0092] As is well known in the art, the PAM sequence specificity of a given nuclease is affected by the enzyme concentration (see, for example, Karvelis et al. (2015) Genome Biology 16:253), which can be modified by altering the amount of the promoter used to express RGN or the ribonucleoprotein complex delivered to the cell.

[0093] After identifying its corresponding PAM sequence, the RGN can cleave one or both strands of the target sequence at a specific cleavage site. As used herein, the cleavage site consists of two specific nucleotides within the target sequence, between which the target strand and / or non-target strand of the target sequence is cleaved by the RGN. The cleavage site may contain nucleotides 1 and 2, 2 and 3, 3 and 4, 4 and 5, 5 and 6, 7 and 8, or 8 and 9 from the PAM in the 5' or 3' direction. In some embodiments, the cleavage site may be more than 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the PAM in the 5' or 3' direction. Because the RGN can cleave the target sequence to produce staggered ends, in some embodiments, the cleavage site is defined based on the distance between two nucleotides from the PAM on the non-target strand of the target sequence and, for the target strand, the distance between two nucleotides from the complementary sequence of the PAM and the distance of the PAM in the non-target strand of the target sequence.

[0094] IV. RNA-guided nucleases and other nucleases

[0095] In some embodiments of a method for cleaving the mutHTT allele and treating Huntington's disease, RGN is a type II CRISPR-Cas polypeptide. In some embodiments, RGN is a type V CRISPR-Cas polypeptide. In some embodiments, RGN is a Cas9, CasX, CasY, Cpfl, C2cl, C2c2, C2c3, GeoCas9, CjCas9, Casl2a, Casl2b, Casl2g, Casl2h, Casl2i, Casl3b, Casl3c, Casl3d, Casl4, Csn2, xCas9, SpCas9-NG, LbCasl2a, AsCasl2a, Cas9-KKH, circularly arranged Cas9, Argonaute (Ago), SmacCas9, or Spy-macCas9 domain.

[0096] This document provides an RNA-guided nuclease system comprising the guide RNA disclosed herein. The term RNA-guided nuclease (RGN) refers to a polypeptide that, by binding to a guide RNA molecule that hybridizes to the target strand of a target sequence (e.g., the target DNA sequence in a mutant HTT allele), guides the target to that specific target sequence in a sequence-specific manner. The active fragment of a naturally occurring RGN or a variant thereof maintains RNA-guided, sequence-specific binding to the target nucleotide sequence. RGN cleaves the target strand of the target sequence, causing either single-strand or double-strand breaks, but generally double-strand breaks.

[0097] The RGN system disclosed in this invention comprises an RGN incorporating the target sequence disclosed herein. In some embodiments, the RGN recognizes a PAM (where N is A, C, T, or G; R is G or A; Y is C or T) at the 3' end of the target sequence on a non-target strand, having a common nucleotide sequence including NNNCC, NNRYA, NNGRR, and NNGG, and its active fragment or variant thereof. In some embodiments, the RGN recognizes a PAM (where N is A, C, T, or G) at the 3' end of the target sequence on a non-target strand, having a common nucleotide sequence including NNNCC, and its active fragment or variant thereof. In some embodiments, the RGN recognizes a PAM (where N is A, C, T, or G; R is G or A; Y is C or T) at the 3' end of the target sequence on a non-target strand, having a common nucleotide sequence including NNRYA, and its active fragment or variant thereof. In some embodiments, the RGN recognizes a PAM (where N is A, C, T, or G; R is G or A) with a shared nucleotide sequence including NNGRR at the 3' end of the target sequence on the non-target strand, and its active fragment or variant. In some embodiments, the RGN recognizes a PAM (where N is A, C, T, or G) with a shared nucleotide sequence including NNGG at the 3' end of the target sequence on the non-target strand, and its active fragment or variant. In some embodiments, the active fragment or variant of the RGN that recognizes such PAM sequences is capable of binding to and, in some embodiments, cleaving or nicking the target sequence.

[0098] The disclosed RGN polypeptide may include a linker domain 1 (L1), a linker domain 2 (L2), a wedge-shaped (WED) domain, a RuvC nuclease domain, an HNH nuclease domain, a double helix bridge (BH) domain, a Rec domain, or a PAM interaction (PI) domain. In some embodiments, the RuvC domain is a RuvCIII domain. The Rec or recognition leaflet mediates nucleic acid binding by sensing nucleic acid through multiple Rec domains (e.g., Rec1-3), regulates HNH conformational transitions, and locks the catalytic HNH domain at the cleavage site. The wedge-shaped domain is responsible for recognizing and guiding the RNA backbone. The arginine-rich double helix bridge (BH) domain connects the nuclease leaflet and the recognition leaflet.

[0099] Non-limiting examples of domains within the APG07433.1 RGN polypeptide shown as SEQ ID NO: 3 include: RuvC-I from amino acid residues 1-54; BH from amino acid residues 55-83; REC1 from amino acid residues 84-244; REC2 from amino acid residues 245-462; RuvC-II from amino acid residues 463-521; L1 from amino acid residues 522-552; HNH from amino acid residues 553-672; L2 from amino acid residues 673-685; RuvC-III from amino acid residues 686-833; WED from amino acid residues 834-938; and PI from amino acid residues 939-1071, all with reference to SEQ ID NO: 3.

[0100] Non-limiting examples of the domains within the APG05586 RGN polypeptide shown as SEQ ID NO: 7 have the following domains: RuvC-I from amino acid residues 1-33; BH from amino acid residues 34-71; REC1 from amino acid residues 72-232; REC2 from amino acid residues 233-468; RuvC-II from amino acid residues 469-517; L1 from amino acid residues 518-552; HNH from amino acid residues 553-672; L2 from amino acid residues 673-687; RuvC-III from amino acid residues 688-837; WED from amino acid residues 838-998; and PI from amino acid residues 999-1150, all with reference to SEQ ID NO: 7.

[0101] Non-limiting examples of domains within the APG01604 RGN polypeptide shown as SEQ ID NO: 11 have the following domains: RuvC-I from amino acid residues 1-40; BH from amino acid residues 41-74; REC1 from amino acid residues 75-223; REC2 from amino acid residues 224-430; RuvC-II from amino acid residues 431-483; L1 from amino acid residues 484-516; HNH from amino acid residues 517-631; L2 from amino acid residues 632-651; RuvC-III from amino acid residues 652-775; WED from amino acid residues 776-909; and PI from amino acid residues 910-1052, all with reference to SEQ ID NO: 11.

[0102] Non-limiting examples of domains within the LPG10145 RGN polypeptide shown as SEQ ID NO: 15 have the following domains: RuvC-I from amino acid residues 1-42; BH from amino acid residues 43-79; REC1 from amino acid residues 80-236; REC2 from amino acid residues 237-476; RuvC-II from amino acid residues 477-524; L1 from amino acid residues 525-560; HNH from amino acid residues 561-676; L2 from amino acid residues 677-690; RuvC-III from amino acid residues 691-828; WED from amino acid residues 829-976; and PI from amino acid residues 977-1130, all with reference to SEQ ID NO: 15.

[0103] The RGN system disclosed in this invention may include an RGN comprising a PAM-interacting domain that facilitates recognition and binding to a PAM site. In a particular embodiment, the PAM-interacting domain of the RGN has the sequence shown in SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNCC. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNCC.

[0104] In some embodiments, the PAM-interacting domain of the RGN has the sequence shown in SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 134 and recognizing the PAM sequence NNRYA.

[0105] In some embodiments, the PAM-interacting domain of the RGN has the sequence shown in SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0106] In some embodiments, the PAM-interacting domain of the RGN has the sequence shown in SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 136 and recognizing the PAM sequence NNGG. The PAM interaction domains of APG07433.1 nuclease (shown as SEQ ID NO: 3), APG05586 nuclease (shown as SEQ ID NO: 7), APG01604 nuclease (shown as SEQ ID NO: 11), and LPG10145 nuclease (shown as SEQ ID NO: 15) were determined by comparing nuclease sequences with known RNA-guided nucleases with resolved structures, including Staphylococcus aureus (PDB: 5CZZ-chain-A), Neisseria meningitidis 1 (PDB: 6JDV_1|chain), and Streptococcus thermophilus (6M0W_4|chain) and identifying regions with similar positions in protein alignment.

[0107] The RGN system disclosed in this invention may include an RGN polypeptide comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. The nuclease domain may comprise a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 143, 144, 145, and 146.

[0108] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 147, 148, 149, and 150.

[0109] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 151, 152, 153, and 154.

[0110] The RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 155 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 155, 156, 157, and 158.

[0111] The RGN system disclosed in this invention may include an RGN polypeptide comprising a PAM interaction domain that facilitates recognition and binding to a PAM site and further comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence shown as any of SEQ ID NO: 143, 144, 145, and 146.

[0112] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 147, 148, 149, and 150.

[0113] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 151, 152, 153, and 154.

[0114] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 155, 156, 157, and 158.

[0115] In some embodiments, the compositions and methods disclosed herein use an RGN or an active variant or fragment thereof capable of binding to a target sequence adjacent to a PAM common sequence shown as NNNNCC, NNRYA, NNGRR, and NNGG (i.e., capable of recognizing the PAM common sequence). In some embodiments, the PAM sequence is located at the 3' of the target sequence on a non-target strand. In some embodiments, the RGN binds to a guide RNA having a sequence shown as any of the following: SEQ ID NO: 6, 10, 14, 18, and 25-29. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 3 is capable of recognizing the PAM sequence of NNNNCC and binds to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof; and a tracrRNA shown as SEQ ID NO: 5, or an active variant or fragment thereof. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 7 is capable of recognizing the PAM sequence of NNRYA and binding to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof; and a tracrRNA shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof. In some embodiments, an RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 11 is capable of recognizing the PAM sequence of NNGRR and binding to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof; and a tracrRNA shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof. In some embodiments, the RGN having at least 90% sequence identity with the amino acid sequence shown as SEQ ID NO: 15 is able to recognize the PAM sequence of NNGG and bind to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof; and a tracrRNA shown as SEQ ID NO: 17, or an active variant or fragment thereof.

[0116] Non-limiting examples of RGNs applicable to the methods and compositions disclosed herein include APG07433.1, APG05586, APG01604, and LPG10145 RNA-guided nucleases, whose amino acid sequences are shown as SEQ ID NO: 3, 7, 11, and 15, respectively, and their active fragments or variants that retain the ability to bind to the target sequence in an RNA-guided sequence-specific manner. In some embodiments, the active variants of the RGNs disclosed herein comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater sequence identity with the amino acid sequence shown as SEQ ID NO: 3. In some embodiments, the active variants of the RGN disclosed herein comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the amino acid sequence shown below: SEQ ID NO: 7. In some embodiments, the active variants of the RGN disclosed herein comprise an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the amino acid sequence shown below: SEQ ID NO: 11. In some embodiments, the active variant of the RGN disclosed herein comprises an amino acid sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the amino acid sequence shown below: SEQ ID NO: 15. In some embodiments, the active fragment of the APG07433.1 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues shown below: SEQ ID NO: 3.In some embodiments, the active fragment of APG05586 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues as shown in SEQ ID NO: 7. In some embodiments, the active fragment of APG01604 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues as shown in SEQ ID NO: 11. In some embodiments, the active fragment of LPG10145 RGN comprises at least 50, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050 or more consecutive amino acid residues as shown in SEQ ID NO: 15.

[0117] The compositions and methods disclosed in this invention may comprise an RGN capable of binding to the target sequence of this disclosure or an RGN having the amino acid sequence shown as SEQ ID NO: 3, or an active variant or fragment thereof, wherein the RGN is capable of binding to the target sequence adjacent to the PAM concordance sequence shown as NNNNCC. In some embodiments, the PAM sequence is at the 3' of the target sequence on a non-target strand. In some embodiments, the RGN binds to a guide RNA having the sequence shown as SEQ ID NO: 27 or 28. In some embodiments, the RGN binds to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 5, or an active variant or fragment thereof.

[0118] The compositions and methods disclosed in this invention may comprise an RGN capable of binding to the target sequence of this disclosure or an RGN having the amino acid sequence shown as SEQ ID NO: 7, or an active variant or fragment thereof, wherein the RGN is capable of binding to a target sequence adjacent to a PAM concordant sequence shown as NNRYA. In some embodiments, the PAM sequence is at the 3' of the target sequence on a non-target strand. In some embodiments, the RGN binds to a guide RNA having the sequence shown as SEQ ID NO: 25 or 26. In some embodiments, the RGN binds to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof.

[0119] The compositions and methods disclosed in this invention may comprise an RGN capable of binding to the target sequence of this disclosure or an RGN having the amino acid sequence shown as SEQ ID NO: 11, or an active variant or fragment thereof, wherein the RGN is capable of binding to the target sequence adjacent to the PAM concordance sequence shown as NNGRR. In some embodiments, the PAM sequence is at the 3' of the target sequence on the non-target strand. In some embodiments, the RGN binds to a guide RNA having the sequence shown as SEQ ID NO: 29. In some embodiments, the RGN binds to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof.

[0120] The compositions and methods disclosed in this invention may comprise an RGN capable of binding to the target sequence of this disclosure or an RGN having the amino acid sequence shown as SEQ ID NO: 15, or an active variant or fragment thereof, wherein the RGN is capable of binding to the target sequence adjacent to the PAM concordance sequence shown as NNGG. In some embodiments, the PAM sequence is at the 3' of the target sequence on the non-target strand. In some embodiments, the RGN binds to a guide RNA having the sequence shown as SEQ ID NO: 18. In some embodiments, the RGN binds to a guide RNA comprising: a CRISPR repeat sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof, and a tracrRNA shown as SEQ ID NO: 17, or an active variant or fragment thereof.

[0121] In some embodiments, the compositions and methods disclosed in this invention use nucleases other than RGN. These nucleases bind to or within the relative ends of an amplified trinucleotide repeat sequence of the HTT gene from the target sequence disclosed in this invention. As used herein, the term "nuclease" refers to an enzyme that catalyzes the cleavage of phosphodiester bonds between nucleotides in a nucleic acid molecule. Generally, nucleases are endonucleases that enable the cleavage of phosphodiester bonds between nucleotides within a nucleic acid molecule. In some embodiments, sequence-specific nucleases are selected from the group consisting of: broad-spectrum nucleases, zinc finger nucleases, TAL effector DNA-binding domain-nuclease fusion proteins (TALENs), and RNA-guided nucleases (RGNs), or variants thereof with reduced or inhibited nuclease activity.

[0122] As used herein, the term "macro-nuclease" or "homing endonuclease" refers to an endonuclease that binds to a recognition site of 12 to 40 bp in length within double-stranded DNA. A non-limiting example of a macro-nuclease is a macro-nuclease belonging to the LAGLIDADG family containing the conserved amino acid motif LAGLIDADG (SEQ ID NO: 139). The term "macro-nuclease" can also refer to dimerizing or single-stranded macro-nucleases.

[0123] As used in this article, the term "zinc finger nuclease" or "ZFN" refers to a chimeric protein that contains a zinc finger DNA-binding domain and a nuclease domain.

[0124] As used herein, the terms “TAL effector DNA-binding domain-nuclease fusion protein” or “TALEN” refer to a chimeric protein containing both a TAL effector DNA-binding domain and a nuclease domain.

[0125] According to the present invention, the target sequence within the mutant HTT allele disclosed herein is bound by an RGN. The target strand of the target sequence hybridizes with the guide RNA associated with the RGN. If the polypeptide has nuclease activity, the target strand and / or non-target strand of the target sequence (e.g., the target DNA sequence) can subsequently be cleaved by the RGN. The term "cleave" or "cleavage" refers to the hydrolysis of at least one phosphodiester bond within the backbone of one or both strands of a double-stranded target sequence (e.g., the target DNA sequence), which can cause single-strand or double-strand breaks within the target DNA sequence. Cleavage of the target sequence disclosed herein can result in staggered breaks or blunt ends.

[0126] The compositions and methods disclosed in this invention can utilize RGNs or other nucleases containing at least one nuclear localization signal (NLS) to enhance the transport of RGNs to the cell nucleus. Nuclear localization signals are known in the art and generally comprise a stretch of a basic amino acid (see, for example, Lange et al., *Journal of Biol. Chem.* (2007) 282:5101-5105). In some embodiments, the RGN contains 2, 3, 4, 5, 6, or more nuclear localization signals. The nuclear localization signals can be heterologous NLSs. Non-limiting examples of nuclear localization signals suitable for the RGNs disclosed in this invention are nuclear localization signals for the SV40 large T-antigen, nucleoplasmic protein, and c-Myc (see, for example, Ray et al. (2015), *Bioconjug Chemistry* 26(6):1004-7). In embodiments, the RGN contains an NLS sequence shown as SEQ ID NO: 86, 87, or 125. RGN or other nucleases may include one or more NLS sequences at their N-terminus, C-terminus, or both. For example, an RGN may include two NLS sequences at its N-terminus and four NLS sequences at its C-terminus. In some embodiments, the RGN or other nuclease includes an SV40 NLS (such as the sequence shown in SEQ ID NO: 86) at its N-terminus and a nucleoplasmic protein NLS (such as the sequence shown in SEQ ID NO: 87) at its C-terminus. In some embodiments, the RGN or other nuclease includes a c-Myc NLS (such as the sequence shown in SEQ ID NO: 125) at both its N-terminus and C-terminus. When an NLS is attached to the N-terminus, C-terminus, or both of the RGN or other nuclease, an NLS adaptor protein may be present to separate the RGN or other nuclease from the NLS. In some embodiments, the NLS adaptor protein links the RGN polypeptide or other nuclease to the NLS. These NLS adaptor proteins can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more amino acids long. In some embodiments, the NLS adaptor protein is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7 or at least 8 amino acids long. In some embodiments, the NLS adaptor protein connecting NLS to RGN or other nucleases, or linking NLS and RGN or other nucleases, has the sequence shown as SEQ ID NO: 127.In some embodiments, the RGN or other nuclease includes c-Myc NLS (such as the sequence shown in SEQ ID NO: 125) at its N-terminus, separated from the nuclease protein by an NLS adaptor protein having the sequence shown in SEQ ID NO: 127, and includes c-Myc NLS (such as the sequence shown in SEQ ID NO: 125) at its C-terminus, separated from the nuclease protein by an NLS adaptor protein having the sequence shown in SEQ ID NO: 127. The RGN polypeptide or other nuclease may be linked to c-Myc NLS (such as the sequence shown in SEQ ID NO: 125) at its N-terminus and to c-Myc NLS (such as the sequence shown in SEQ ID NO: 125) at its C-terminus, wherein the RGN polypeptide or other nuclease is linked to each of the N-terminal c-Myc NLS and the C-terminal c-Myc NLS by an NLS adaptor protein having the sequence shown in SEQ ID NO: 127.

[0127] In some embodiments, the compositions and methods disclosed in this invention utilize RGN or other nucleases comprising at least one cell-penetrating domain that facilitates cellular uptake of RGN. A cell-penetrating domain is known in the art and generally comprises an extension of positively charged amino acid residues (i.e., a polycationic cell-penetrating domain), alternating polar and nonpolar amino acid residues (i.e., an amphipathic cell-penetrating domain), or hydrophobic amino acid residues (i.e., a hydrophobic cell-penetrating domain) (see, for example, Milletti F. (2012) Drug Discov Today 17:850-860). A non-limiting example of a cell-penetrating domain is trans-activated transcription activator (TAT) from human immunodeficiency virus 1.

[0128] Nuclear localization signals and / or cell penetration domains may be located at the N-terminus, C-terminus, and / or internal location of the RGN.

[0129] V. Nucleic acid molecules encoding RNA-guided nucleases, single guide RNA, CRISPR RNA, and / or tracrRNA

[0130] This disclosure provides nucleic acid molecules that contain or encode the RGN, crRNA, tracrRNA, and / or sgRNA disclosed herein.

[0131] The use of the terms "polynucleotide" or "nucleic acid molecule" is not intended to limit this disclosure to polynucleotides containing DNA. Those skilled in the art will recognize that polynucleotides can comprise ribonucleotides (RNA) and combinations of ribonucleotides and deoxyribonucleotides. Such deoxyribonucleotides and ribonucleotides include both naturally occurring molecules and synthetic analogs. These include peptide nucleic acids (PNAs), PNA-DNA chimeras, locked nucleic acids (LNAs), and phosphate-thioester linker sequences. The polynucleotides disclosed herein also encompass all sequence forms, including but not limited to single-stranded, double-stranded, DNA-RNA hybrids, triple-stranded structures, stem-loop structures, and similar forms.

[0132] In some embodiments of the compositions and methods disclosed in this invention that include a nucleic acid molecule encoding an RGN, the nucleic acid molecule is an mRNA (messenger RNA) molecule. mRNA refers to any polynucleotide that encodes a polypeptide of interest and is capable of translation to produce the encoded polypeptide of interest in vitro, in vivo, in situ, or ex vivo. In some embodiments, the basic components of the mRNA molecule include at least a coding region, a 5' UTR, a 3' UTR, a 5' cap, and a poly-A tail. In some embodiments, the mRNA encoding an RGN suitable for use in the methods and compositions disclosed in this invention may include one or more structural and / or chemical modifications or alterations that impart suitable properties to the polynucleotide. For example, suitable properties of the mRNA include a lack of substantial induction of an innate immune response to the cell in which the mRNA is introduced. A “structural” feature or modification is a feature or modification in which two or more linking nucleotides are inserted, deleted, duplicated, inverted, or randomized in the mRNA without significant chemical modification of the nucleotides themselves. Because chemical bonds will necessarily break and reform to achieve structural modification, structural modification is chemical in nature and therefore a chemical modification. However, structural modification will produce different nucleotide sequences. Chemical modifications to mRNA may involve the inclusion of 5-methylcytosine, N1-methyl-pseuuridine, pseudouridine, 2-thiouridine, 4-thiouridine, 5-methoxyuridine, 2'-fluoroguanidine, 2'-fluorouridine, 5-bromouridine, 5-(2-methoxycarbonylvinyl)uridine, 5-[3(1-E-propenylamino)]uridine, α-thiocytidine, N6-methyladenosine, 5-methylcytidine, N4-acetylcytidine, 5-formylcytidine, or combinations thereof in the mRNA.

[0133] Nucleic acid molecules encoding RGNs can be codon-optimized for expression in organisms of interest, such as mammals. “Codon-optimized” coding sequences are polynucleotide sequences whose codon usage frequencies are designed to mimic the optimal codon usage frequencies or transcriptional conditions of a particular host cell. Expression in a particular host cell or organism is enhanced by one or more codon changes at the nucleic acid level without altering the translated amino acid sequence. Nucleic acid molecules can be codon-optimized, either fully or partially. Codon tables and other references providing information on preferences for various organisms are available in the art (see, for example, Gaspar et al. (2012) Bioinformatics 28(20):2683–2684; Komar et al. (1998) Biol. Chem. 379(10):1295–1300; and Inouye et al. (2015) Protein Expr. Purif. 109:47–54). A non-limiting example of a codon-optimized coding sequence for an RGN suitable for the compositions and methods disclosed in this invention is shown in SEQ ID NO: 88.

[0134] The polynucleotides encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein may be provided in an expression cassette for in vitro expression or expression in cells, embryos, or organisms of interest. The cassette will include 5' and 3' regulatory sequences operatively linked to the polynucleotides encoding the RGN, crRNA, tracrRNA, and / or sgRNA provided herein, allowing expression of the polynucleotides. The cassette may additionally contain at least one additional gene or genetic element to be co-transformed into an organism. Where an additional gene or element is included, this component is operatively linked. The term “operatively linked” is intended to mean a functional link between two or more elements. For example, an operative link between a promoter and a coding region of interest (e.g., a region encoding RGN, crRNA, tracrRNA, and / or sgRNA) is a functional link that allows expression of the coding region of interest. The operatively linked elements may be contiguous or non-contiguous. When used to refer to the conjugation of two protein-coding regions, “operatively linked” or “operatively fused” means that the coding regions are in the same reading frame. For example, "operably fused" peptides may mean that the structure and / or biological activity of each individual peptide are also present in the fusion. Alternatively, additional genes or elements may be provided on multiple expression cassettes. For example, the nucleotide sequence encoding the RGN disclosed in this invention may be present on one expression cassette, while the nucleotide sequence encoding crRNA, tracrRNA, or complete guide RNA may be on separate expression cassettes. Such expression cassettes provide multiple restriction sites and / or recombination sites for the insertion of polynucleotides into the transcriptional regulation of the regulatory region. The expression cassette may additionally contain optional marker genes.

[0135] The expression cassette will include, in the 5'-3' direction of transcription, a transcription (and in some embodiments, translation) initiation region (i.e., promoter), a polynucleotide encoding RGN, crRNA, tracrRNA, and / or sgRNA as disclosed herein, and a transcription (and in some embodiments, translation) termination region (i.e., termination region) functional in the organism of interest. The promoter of this disclosure is capable of directing or driving the expression of the coding sequence in the host cell. Regulatory regions (e.g., promoter, transcription regulatory region, and translation termination region) may be endogenous or heterologous to the host cell or to each other. As used herein, “heterogeneous” when referring to a sequence means a sequence derived from a foreign species, or, if derived from the same species, modified substantially from its native form by deliberate human intervention of the composition and / or genomic loci. As used herein, a chimeric gene comprises a coding sequence operatively linked to a transcription initiation region heterologous to the coding sequence.

[0136] Appropriate termination regions include those from simian virus (SV40), human growth hormone (hGH), bovine growth hormone (BGH), and rabbit β-globin (rbGlob). See also Proudfoot (1991) Cell 64:671-674; Munroe et al. (1990) Gene 91:151-158; Schek et al. (1992) Molecular and Cellular Biology 12(12):5386-5393; Gil and Proudfoot (1987) Cell 49(3):399-406; Goodwin and Rottman (1992) The Journal of Biological Chemistry 267(23):16330-16334; and Lanoix and Acheson (1988) EMBO J. 7(8):2515-2522.

[0137] Additional regulatory signals include, but are not limited to, transcription initiation sites, operons, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, and analogues. See, for example, Sambrook et al. (1992), *Molecular Cloning: A Laboratory Manual*, edited by Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY), hereinafter referred to as "Sambrook 11"; Davis et al. (1980), *Advanced Bacterial Genetics* (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY), and the references cited therein.

[0138] During expression cassette preparation, various DNA fragments can be manipulated to provide the DNA sequence in the appropriate reading frame at the appropriate orientation and when appropriate. For this purpose, adaptors or linkers can be used to attach the DNA fragments, or other manipulations may be involved to provide facilitated restriction sites, remove redundant DNA, remove restriction sites, or similar techniques. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, and substitution, such as transitions and transversions, may be employed.

[0139] Many promoters are available for use in the practice of this invention. Promoters can be selected based on desired results. Generally, RGN expression is controlled by the RNA polymerase II promoter, and the RGN coding sequence can therefore be operatively linked to the RNA polymerase II promoter. The expression of crRNA, tracrRNA, or sgRNA is generally controlled by the RNA polymerase III promoter, and the coding sequences of these elements can therefore be operatively linked to the RNA polymerase III promoter. Non-limiting examples of RNA polymerase III promoters suitable for expressing crRNA, tracrRNA, and sgRNA include mammalian U6, U3, H1, and 7SL RNA promoters and rice U6 and U3 promoters, such as the human U6 small nucleus promoter or a truncated version thereof, such as sequences shown as SEQ ID NO: 89 or 128, and promoters disclosed in U.S. Provisional Application No. 63 / 209,660, filed June 11, 2021, and International Application No. PCT / US2022 / 032940, filed June 10, 2022 (each incorporated herein by reference in its entirety), including promoters shown herein as SEQ ID NO: 96-105.

[0140] Nucleic acids can be combined with constitutive, inducible, growth stage-specific, cell type-specific, tissue-biased, tissue-specific promoters or other promoters for expression in the organism of interest.

[0141] Exemplary constitutive promoters for expression in cells according to this disclosure include: the SV40 early promoter; the mouse mammary tumor virus long terminal repeat (LTR) promoter; and the adenovirus major late promoter (Ad... MLP); herpes simplex virus (HSV) promoters; cytomegalovirus (CMV) promoters, such as the early promoter region of CMV (CMVIE); Rous sarcoma virus (RSV) promoters; human ubiquitin C promoter (UBC); human U6 small nucleus promoter (U6); truncated U6 promoters; enhanced U6 promoters; human H1 promoter from RNA polymerase III (H1); human elongation factor 1α promoter (EF1A); human β-actin promoter (ACTB); human or mouse phosphoglycerate kinase 1 promoter (PGK); chicken β-actin promoter coupled to the early enhancer of CMV (CAGG); yeast transcription elongation factor promoter (TEF1); elongation factor 1α short (EFS) promoter; JeT promoter (see, for example, U.S. Publication No. 2002 / 0098547, which is incorporated herein by reference in its entirety); and analogues thereof. See, for example, Miyagishi et al. (2002), *Nature Biotechnology* 20:497-500; Xia et al. (2003), *Nucleic Acid Research* 31(17):e100-e100; Pasleau et al. (1985), *Gene* 38:227–232; Martin-Gallardo et al. (1988), *Gene* 70:51–56; Oellig and Seliger (1990), *Journal of Neuroscience Research* 26:390–396; Manthorpe et al. (1993), *Human Gene Therapy*. Ther) 4:419–431; Yew et al. (1997) Human Gene Therapy 8:575–584; Xu et al. (2001) Gene 272:149–156; Nguyen et al. (2008) Journal of Surgical Research 148:60–66; Costa et al. (2005) Natural Methods 2:259–260; Lam and Truong (2020) ACS Synthetic Biology 9(10):2625–2631.In some implementations, the RGN coding sequence is operatively linked to a constitutive promoter, which may be a cytomegalovirus (CMV) promoter, a truncated CMV promoter, such as the CMVeb promoter shown in SEQ ID NO: 90, the extension factor 1α short (EFS) promoter shown in SEQ ID NO: 91, or the JeT promoter shown in SEQ ID NO: 92.

[0142] Examples of inducible promoters include: stress-regulated promoters, such as the Hsp70 and Hsp90 promoters (Wurm et al. (1986), Proc. Natl. Acad. Sci. USA. 83:5414-5418; Norr. L. Heat Shock Response. CRC Press; Boca Raton, FL, USA:1991); metal-regulated promoters (Mayo et al. (1982), Cell 29:99-108; Searle et al. (1985), Molecular Cell Biology 5:1480–1489); and hormone-responsive promoters, including glucocorticoid-responsive promoters (Hynes et al. (1981), Proc. Natl. Acad. Sci. USA. 83:5414-5418; Norr. L. Heat Shock Response. CRC Press; Boca Raton, FL, USA:1991); metal-regulated promoters (Mayo et al. (1982), Cell 29:99-108; Searle et al. (1985), Molecular Cell Biology 5:1480–1489); and hormone-responsive promoters, including glucocorticoid-responsive promoters (Hynes et al. (1981), Proc. Natl. Acad. Sci. USA. 83:5414-5418; Norr. L. Heat Shock Response. CRC Press; Boca Raton, FL, USA:1991). USA. 78:2038–2042; Klock et al. (1987) Nature. 329:734–736. Chemically regulated promoters from prokaryotes that have been used include the isopropyl-β-D-thiogalactoside (IPTG) regulator, the lactose regulator, and the tetracycline regulator (see, for example, Gossen et al. (1993) Trends Biochem Sci. 18:471–475; Gossen and Bujard (1992) Proceedings of the National Academy of Sciences of the United States of America 89:5547–5551; Zhou et al. (2006) Gene Ther. 13:1382–1390).Inducible expression can be achieved using operon systems including AlcR / acetaldehyde, ArgR / L-arginine, BirA / biotinyl-AMP, CymR / cumate, EthR / 2-phenylethylbutyrate, HdnoR / 6-hydroxynicotinic acid, HucR / uric acid, MphR(A) / macrolide, PIP / streptocin, Rex / NADH, RheA / heat, ScbR / SCB1, TraR / 3-oxo-C8-HSL, and TtgR / phlorizin; see, for example, U.S. Patent No. 8,728,759B2; U.S. Patent No. 7,745,592B2; Weber and Fussenegger (2004), Methods in Molecular Biology. Biol.) 267:451-466; Hartenbach et al. (2007) Nucleic Acid Research 35:e136; Weber et al. (2009) Metabolic Engineering 11:117-124; Weber et al. (2008) Proceedings of the National Academy of Sciences of the United States of America 105:9994-9998; Malphettes et al. (2005) Nucleic Acid Research 33:e107; Kemmer et al. (2010) Nature Biotechnology 28:355-360; Weber et al. (2002) Nature Biotechnology 20:901-907; Fussenegger et al. (2000) Nature Biotechnology Biotechnol. 18:1203-1208; Weber et al. (2006) Metabolic Engineering 8:273-280; Weber et al. (2003) Nucleic Acid Research 31:e69; Weber et al. (2003) Nucleic Acid Research 31:e71; Neddermann et al. (2003) EMBO Rep. 4:159-165; and Gitzinger et al. (2009) Proceedings of the National Academy of Sciences of the United States of America 106:10638-10643.Inducible expression can be achieved using protein-protein interaction systems, including: rapamycin-induced interaction between FKBP12 (FK506-binding protein 12) and mTOR (Rivera et al. (1996) *Nature Medicine* 2:1028-1032; Belshaw et al. (1996) *Proceedings of the National Academy of Sciences of the United States of America* 93:4604-46077); ABA-regulated interaction between PYL1 (abscisic acid receptor) and ABI1 (protein phosphatase 2C56) (Liang et al. (2011) *Science Signal* 4(164):rs2-rs2); and light-induced protein-protein interaction systems (Wang et al. (2012) *Nature Methods*). Methods. 9:266-269; Yamada et al. (2018) Cell Reports 25:487-500.

[0143] Tissue-specific or tissue-preferred promoters can be used to target the expression of constructs within specific tissues. In embodiments, tissue-specific or tissue-preferred promoters are active in mammalian tissues. Examples of tissue-specific or tissue-preferred promoters include promoters that preferentially initiate transcription in certain tissues, such as the brain. A “tissue-specific” promoter is a promoter that initiates transcription only in certain tissues. Unlike constitutive gene expression, tissue-specific expression is the result of several levels of interaction in gene regulation. Therefore, promoters from homologous or closely related species are preferred for achieving efficient and reliable expression of transgenes in specific tissues. In some embodiments, expression includes tissue-preferred promoters. A “tissue-preferred” promoter is a promoter that preferentially, but not necessarily completely or only initiating transcription in certain tissues, such as the brain.

[0144] In the implementation scheme, the nucleic acid molecule encoding RGN, crRNA, tracrRNA, and / or sgRNA contains a cell-type-specific promoter. A “cell-type-specific” promoter is a promoter that primarily drives the expression of certain cell types in one or more organs. Examples of cell-type-specific promoters that primarily have active cells include, for example, neurons. The nucleic acid molecule may also include a cell-type-preferred promoter. A “cell-type-preferred” promoter is a promoter that primarily, but not necessarily entirely or only in certain cell types in one or more organs, drives expression. Examples of cell-type-preferred promoters that preferentially have active cells include, for example, neurons. Neurons may include neural progenitor cells, forebrain neuronal progenitor cells, striatal neurons, medium-sized polyspinous neurons, and cortical neurons. Cell-type-preferred promoters may preferentially have activity in non-neuronal cells (such as glial cells) in the brain. Glial cells may include microglia, astrocytes, and oligodendrocytes. In some implementations, cell type-preferred promoters may be preferentially active in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or combinations thereof.

[0145] The RGN coding sequence can be operatively linked to brain- or neuron-specific promoters, such as the human synaptic protein I (Syn) promoter, the 65 kDa or 67 kDa glutamate decarboxylase (GAD65 or GAD67, respectively) promoter, the homeobox Dlx5 / 6 promoter, the protakininogen 1 (Tac1) promoter, the neuron-specific enolase (NSE), the dopamine receptor 1 (Drd1a) promoter or the dopamine receptor 2 (DRD2) promoter, the glial fibrillary acidic protein (GFAP) promoter or the 32 kDa dopamine and cyclic AMP-regulated phosphoprotein (DARP32) promoter (see, for example, Delzor et al., 2012, HumGene Ther Methods, 23(4):242-254, which is incorporated herein by reference in its entirety). A non-limiting example of a neuron-specific promoter that can be used to drive RGN expression in the compositions and methods disclosed herein is the human synaptic protein I (Syn) promoter. The Syn promoter may have the nucleotide sequence shown as SEQ ID NO: 93.

[0146] Nucleic acid sequences encoding RGN, crRNA, tracrRNA, and / or sgRNA can be operatively linked to a promoter sequence recognized by, for example, a phage RNA polymerase used for in vitro mRNA synthesis. In embodiments, the in vitro transcribed RNA can be purified for use in the methods described herein. For example, the promoter sequence can be a T7, T3, or SP6 promoter sequence or a variant of the T7, T3, or SP6 promoter sequence. In embodiments, the expressed protein and / or RNA can be purified for use in the genome modification methods described herein.

[0147] In some embodiments, the polynucleotide encoding RGN, crRNA, tracrRNA, and / or sgRNA may also be linked to a polyadenylation (polyA) signal and / or at least one transcription termination sequence. In some embodiments, the coding sequence (e.g., a nucleic acid molecule encoding RGN, crRNA, tracrRNA, and / or sgRNA) is linked to a polyA tail such as that of simian virus (SV40) shown in SEQ ID NO: 94, or a polyA tail such as that of bovine growth hormone shown in SEQ ID NO: 95. See, for example, Proudfoot (1991) Cell 64:671-674; Munroe et al. (1990) Gene 91:151-158; Schek et al. (1992) Molecular and Cellular Biology 12(12):5386-5393; Gil and Proudfoot (1987) Cell 49(3):399-406; Goodwin and Rottman (1992) The Journal of Biological Chemistry 267(23):16330-16334; and Lanoix and Acheson (1988) European Organization for Molecular Biology 7(8):2515-2522.

[0148] Additionally, the sequence encoding RGN may also be linked to a sequence encoding at least one nuclear localization signal, at least one cell penetration domain, and / or at least one signal peptide capable of transporting the protein to a specific subcellular location, as described elsewhere in this document.

[0149] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA may be present in one or more vectors. A “vector” is a polynucleotide composition used to transfer, deliver, or introduce nucleic acids into a host cell. Suitable vectors include plasmid vectors, phage particles, granules, artificial / miniature chromosomes, transposons, and viral vectors (e.g., lentiviral vectors, adeno-associated virus vectors, baculovirus vectors). Vectors may contain additional expression control sequences (e.g., enhancer sequences, Kozak sequences, polyadenylation sequences, transcription termination sequences), optional marker sequences (e.g., antibiotic resistance genes), origin of replication, and similar sequences. Additional information can be found in *Current Protocols in Molecular Biology*, Ausubel et al., John Wiley & Sons, New York, 2003, or *Molecular Cloning: A Laboratory Manual*, Sambrook and Russell, Cold Spring Harbor Press, Cold Spring Harbor, NY, 3rd edition, 2001.

[0150] Vectors may also contain optional marker genes for selecting transformed cells. Selective marker genes are used to select transformed cells or tissues. Marker genes include genes encoding antibiotic resistance, such as those encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT). Marker genes may include genes that allow selection for growth on specific nutrients or substances, such as dihydrofolate reductase (DHFR; Simonsen and Levinson (1983), Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA) 80:2495-2499), histidine dehydrogenase (hisD; Hartman and Mulligan (1988), Proceedings of the National Academy of Sciences (Proc. Natl. Acad. Sci. USA) 85:8047-8051), and puromycin-N-acetyltransferase (PAC or puromycin; de la... Luna et al. (1988) Gene 62:121-126, thymidine kinase (TK; Littlefield (1964) Science 145:709-710) and xanthine-guanine phosphoribosyltransferase (XGPRT or gpt; Mulligan and Berg (1981) Proceedings of the National Academy of Sciences of the United States of America (Proc. Natl. Acad. Sci. USA) 78:2072-2076).

[0151] As indicated, expression constructs comprising nucleotide sequences encoding RGN, crRNA, tracrRNA, and / or sgRNA can be used to transform an organism of interest. Transformation methods involve introducing the nucleotide construct into the organism of interest. "Introduction" is intended to introduce the construct into a host cell in a manner that allows the nucleotide construct to enter the host cell. The methods disclosed herein do not require a specific method for introducing the nucleotide construct into the host organism; they only require that the nucleotide construct can enter the interior of at least one cell of the host organism. The host cell can be a eukaryotic or prokaryotic cell. In some embodiments, the eukaryotic host cell is a mammalian cell, avian cell, or insect cell. In some embodiments, the eukaryotic cell comprising or expressing the crRNA, tracrRNA, sgRNA, and / or RGN disclosed in this invention, or modified by the RGN system disclosed in this invention, is a human cell. In some embodiments, the eukaryotic cell comprising or expressing the crRNA, tracrRNA, sgRNA, and / or RGN disclosed in this invention, or modified by the RGN system disclosed in this invention, is a stem cell, including induced pluripotent stem cells. In some embodiments, mammalian or human cells containing or expressing the crRNA, tracrRNA, sgRNA and / or RGN disclosed in this invention, or modified by the RGN system disclosed in this invention, are cardiomyocytes, neurons, glial cells or retinal ganglion cells.

[0152] Methods for introducing nucleotide constructs into host cells are known in the art, including but not limited to stable transformation methods, transient transformation methods and virus-mediated methods.

[0153] The method disclosed in this invention can produce transformed organisms or cell lines derived from these transformed cells.

[0154] "Transgenic organism," "transformed organism," or "stable transformed" refers to an organism, cell, or tissue that has incorporated or integrated polynucleotides encoding the RGN, crRNA, tracrRNA, and / or sgRNA disclosed herein. It should be recognized that other exogenous or endogenous nucleic acid sequences or DNA fragments may also be incorporated into host cells. Transformation of host cells can be achieved through infection, conjugation, transfection, microinjection, electroporation, microprojection, gene gun or particle bombardment, electroporation, silica / carbon fiber, ultrasound-mediated, PEG-mediated, calcium phosphate coprecipitation, polycationic DMSO technology, DEAE polydextrose program, and viral-mediated, liposome-mediated, and similar methods. Viral-mediated introduction of polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA includes introduction and expression mediated by retroviruses, lentiviruses, adenoviruses, and adeno-associated viruses.

[0155] Transformation can result in either stable or transient incorporation of nucleic acids into the cell. "Stable transformation" refers to the integration of the introduced nucleotide construct into the host cell's genome and its inheritance by its offspring. "Transient transformation" refers to the introduction of a polynucleotide into the host cell without integration into its genome.

[0156] In some implementations, the transformed cells can be introduced into an organism. These cells can be derived from an organism in which the cells are transformed in vitro. These cells can be autologous (derived from and returned to the same individual) or allogeneic (the donor and recipient individuals belong to the same species). Generally, the donor and recipient of allogeneic cells are complete or partial HLA matches.

[0157] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA, or containing crRNA, tracrRNA, and / or sgRNA, can also be used to transform any prokaryotic species, including but not limited to archaea and bacteria (e.g., Bacillus sp., Klebsiella sp., Streptomyces sp., Rhizobium sp., Escherichia sp., Pseudomonas sp., Salmonella sp., Shigella sp., Vibrio sp., Yersinia sp., Mycoplasma sp., Agrobacterium, Lactobacillus sp.)).

[0158] Polynucleotides encoding RGN, crRNA, tracrRNA, and / or sgRNA, or containing crRNA, tracrRNA, and / or sgRNA, can be used to transform any eukaryotic species, including but not limited to animals (e.g., mammals, humans, mice, rats, non-human primates, insects, fish, birds, and reptiles), fungi, amoebas, algae, and yeast.

[0159] Conventional viral and non-viral gene transfer methods can be used to introduce nucleic acids into mammalian, insect, or avian cells or target tissues. These methods can be used to administer nucleic acids encoding components of the RGN system to cells in culture or in a host organism. Non-viral vector delivery systems include DNA plasmids, RNA (such as transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with delivery media such as liposomes. Viral vector delivery systems include DNA and RNA viruses that possess a free or integrated genome after delivery to cells. For reviews of gene therapy procedures, see Anderson, Science 256:808-813 (1992); Nabel and Feigner, Trends in Biotechnology 11:211-217 (1993); Mitani and Caskey, Trends in Biotechnology 11:162-166 (1993); Dillon, Trends in Biotechnology 11:167-175 (1993); Miller, Nature 357:455-460 (1992); Van Brunt, Biotechnology 6(10):1149-1154 (1988); Vigne, Restorative Neurology and Neuroscience 8:35-36 (1995); Kremer and Perricaudet, British Medical Bulletin. Bulletin, 51(1):31-44 (1995); Haddad et al., Current Topics in Microbiology and Immunology, Doerfler and Bohm (eds.) (1995); and Yu et al., Gene Therapy, 1:13-26 (1994).

[0160] Non-viral methods for delivering nucleic acids include liposome transfection, nuclear transfection, microinjection, gene gun, virions, liposomes, immunoliposomes, polycationic or lipid:nucleic acid conjugates, naked DNA, artificial viral particles, and drug-enhanced DNA uptake. Liposome transfection is described, for example, in U.S. Patents 5,049,386, 4,946,787; and 4,897,355, and liposome transfection reagents are commercially available (e.g., Transfectam™ and Lipofectin™). Cationic and neutral lipids suitable for liposome transfection of polynucleotides with effective receptor recognition include those listed in Feigner, WO 91 / 17424; WO 91 / 16024. These can be delivered to cells (e.g., in vitro or ex vivo) or to target tissues (e.g., in vivo). The preparation of lipid:nucleic acid complexes (including targeted liposomes, such as immunolipid complexes) is well known to those skilled in the art (see, for example, Crystal, Science 270:404-410 (1995); Blaese et al., Cancer Gene Therapy 2:291-297 (1995); Behr et al., Bioconjugate Chem. 5:382-389 (1994); Remy et al., Bioconjugate Chem. 5:647-654 (1994); Gao et al., Gene Therapy 2:710-722 (1995); Ahmad et al., Cancer Research 52:4817-4820 (1992); U.S. Patent Nos. 4,186,183, 4,217,344, 4,235,871, 4,261,975, 4,485,054, 4,501,728, 4,774,085, 4,837,028 and 4,946,787.

[0161] The delivery of nucleic acids using RNA or DNA virus-based systems leverages highly evolved processes to target viruses to specific cells within the body and transport viral loads to the cell nucleus. Viral vectors can be administered directly to patients (in vivo) or used to process cells in vitro, and modified cells can be administered to patients as appropriate (ex vivo). Common virus-based systems include retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, and herpes simplex virus vectors for gene transfer. Using retroviruses, lentiviruses, and adeno-associated viruses for gene transfer, integration into the host genome is possible, typically resulting in long-term expression of the inserted transgene. Furthermore, high transduction efficiency has been observed in many different cell types and target tissues.

[0162] Retroviral tropism can be altered by incorporating foreign envelope proteins and amplifying potential target populations of cells. Lentiviral vectors are retroviral vectors capable of transducing or infecting non-dividing cells and typically produce high viral titers. Therefore, the choice of retroviral gene transfer system depends on the target tissue. Retroviral vectors contain cis-acting long terminal repeats (LTRs) capable of packaging foreign sequences up to 6-10 kb. Minimal cis-acting LTRs are sufficient for vector replication and packaging, followed by integration of therapeutic genes into target cells to provide permanent transgenic expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibberish leukemia virus (GaLV), simmon immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof (see, for example, Buchscher et al., Journal of Virology 66:2731-2739 (1992); Johann et al., Journal of Virology 66:1635-1640 (1992); Sommnerfelt et al., Virology 176:58-59 (1990); Wilson et al., Journal of Virology 63:2374-2378 (1989); Miller et al., Journal of Virology 65:2220-2224 (1991); PCT / US94 / 05700).

[0163] For applications requiring good transient expression, adenovirus-based systems can be used. Adenovirus-based vectors exhibit extremely high transduction efficiency in many cell types and do not require cell division. High titers and high expression levels have been achieved with these vectors. These vectors can be mass-produced in relatively simple systems. Adeno-associated virus (“AAV”) vectors can also be used, for example, in the in vitro production of nucleic acids and peptides and in in vivo and in vitro gene therapy procedures, to transduce cells with target nucleic acids (see, for example, West et al., Virology 160:38-47 (1987); U.S. Patent No. 4,797,368; WO 93 / 24641; Katin, Human Gene Therapy 5:793-801 (1994); Muzyczka, Journal of Clinical Research 94:1351 (1994). The construction of recombinant AAV vectors has been described in numerous publications, including U.S. Patent No. 5,173,414; Tratschin et al., Molecular Cell Biology 5:3251-3260 (1985); Tratschin et al., Molecular Cell Biology 5:3251-3260 (1985). Biol. 4:2072-2081 (1984); Hermonat and Muzyczka, Proceedings of the National Academy of Sciences (PNAS) 81:6466-6470 (1984); and Samulski et al., Journal of Virology 63:03822-3828 (1989). Packaging cells are commonly used to form viral particles capable of infecting host cells. These cells include 293 cells for packaging adenoviruses and ψJ2 or PA317 cells for packaging retroviruses.

[0164] Viral vectors used in gene therapy are typically produced by generating cell lines that package nucleic acid vectors into viral particles. The vectors usually contain the minimum viral sequence required for packaging and subsequent integration into the host, with other viral sequences replaced by expression cassettes for the polynucleotides to be expressed. Missing viral functions are typically supplied trans-form by the packaging cell lines. For example, AAV vectors used in gene therapy typically contain only the ITR sequence from the AAV genome, which is required for packaging and integration into the host genome. The viral DNA is packaged in cell lines containing helper plasmids encoding other AAV genes (i.e., rep and cap) but lacking the ITR sequence.

[0165] Adenoviruses can also be used as helper viruses to infect cell lines. Helper viruses promote AAV vector replication and AAV gene expression from helper plasmids. Due to the lack of an ITR sequence, helper plasmids are not packaged in large quantities. Adenovirus contamination can be reduced, for example, by heat treatment; adenoviruses are more sensitive to heat treatment than AAVs. Additional methods for delivering nucleic acids into cells are known to those skilled in the art. See, for example, US20030087817, which is incorporated herein by reference.

[0166] Non-limiting examples of AAV vectors suitable for the compositions and methods disclosed in this invention are AAV2, AAV3, AAV5, AAV6, and AAV9 vectors (see, for example, Pupo et al., 2022, Molecular Therapy 30(12): pp. 3515-3541, which is incorporated herein by reference in its entirety). In some embodiments, the vector is AAV5 or AAV6. In some embodiments, the AAV vector has a sequence shown as any one of SEQ ID NO: 30-39 or 121-123 and the AAV is packaged with the vector sequence.

[0167] In some embodiments, host cells are transiently or non-transiently transfected with one or more nucleic acid molecules or vectors described herein. In some embodiments, transfection is performed while the cells are naturally present in the individual. In some embodiments, the transfected cells are obtained from an individual, such as a Huntington's disease patient. In some embodiments, the cells are derived from cells obtained from an individual, such as cell lines. In some embodiments, the cell lines may be mammalian, insect, or avian cells. Various cell lines for tissue culture are known in the art. Examples of cell lines include, but are not limited to, C8161, CCRF-CEM, MOLT, mIMCD-3, NHDF, HeLaS3, Huhl, Huh4, Huh7, HUVEC, HASMC, HEKn, HEKa, MiaPaCell, Panel, PC-3, TFl, CTLL-2, CIR, Rat6, CVI, RPTE, AlO, T24, 182, A375, ARH-77, Calul, SW480, SW620, SKOV3, SK-UT, CaCo2, P388Dl, SEM-K2, WEHI-231, HB56, TIB55, lurkat, 145.01, LRMB, Bcl-1, BC-3, IC21, DLD2, Raw264.7, NRK, NRK-52E, MRC5, MEF, Hep G2, HeLa B, and HeLa T4. COS, COS-1, COS-6, COS-M6A, BS-C-1 monkey kidney epithelial cells, BALB / 3T3 mouse embryonic fibroblasts, 3T3 Swiss, 3T3-Ll, 132-d5 human fetal fibroblasts; 10.1 mouse fibroblasts, 293-T, 3T3, 721, 9L, A2780, A2780ADR, A2780cis, A172, A20, A253, A431, A-549, ALC, B16, B35, BCP-I cells, BEAS-2B, bEnd.3, BHK-21, BR 293, BxPC3, C3H-10Tl / 2, C6 / 36, Cal-27, CHO, CHO-7, CHO-IR, CHO-Kl, CHO-K2, CHO-T, CHO Dhfr- / -, COR-L23, COR-L23 / CPR, COR-L235010, CORL23 / R23, COS-7, COV-434, CML Tl, CMT, CT26, D17, DH82, DU145, DuCaP, EL4, EM2, EM3, EMT6 / AR1, EMT6 / AR10.0, FM3, H1299, H69, HB54, HB55, HCA2, HEK-293, HeLa, Hepalclc7, HL-60, HMEC, HT-29, lurkat, lY cells, K562 cells, Ku812, KCL22, KGl, KYOl, LNCap, Ma-Mel 1-48, MC-38, MCF-7, MCF-10A, MDA-MB-231, MDA-MB-468, MDA-MB-435, MDCKII, MDCKII, MOR / 0.2R, MONO-MAC6, MTD-IA, MyEnd, NCI-H69 / CPR, NCI-H69 / LX10, NCI-H69 / LX20, NCI-H69 / LX4, NIH-3T3, NALM-1, NW-145, OPCN / OPCT cell lines, Peer, PNT-IA / PNT 2, RenCa, RIN-5F, RMA / RMAS, Saos-2 cells, Sf-9, SkBr3, T2, T-47D, T84, THP1 cell lines, U373, U87, U937, VCaP, Vero cells, WM39, WT-49, X63, YAC-1, YAR, and their transgenic variants. Cell lines can be obtained from a variety of sources known to those skilled in the art (see, for example, the United States Type Culture Collection (ATCC) (Manassas, Va.)).

[0168] In some embodiments, novel cell lines comprising one or more vector-derived sequences are established using cells transfected with one or more nucleic acid molecules or vectors described herein. In some embodiments, novel cell lines comprising cells containing modified sequences but not any other exogenous sequences are established using cells transiently transfected with components of an RGN system as described herein (such as transient transfection with one or more vectors or with RNA) and modified by the activity of the RGN system.

[0169] In some embodiments, non-human transgenic animals are produced using one or more nucleic acid molecules or vectors described herein. In some embodiments, the transgenic animals are mammals, such as mice, rats, hamsters, rabbits, cattle, or pigs.

[0170] VI. Variants and fragments of polypeptides and polynucleotides

[0171] This disclosure provides active variants and fragments of the crRNA repeat sequences, crRNA, tracrRNA, sgRNA, and RGN disclosed herein. Active variants or fragments of naturally occurring (i.e., wild-type) RGN bind to the target sequence described herein within the mutant HTT allele in an RNA-guided sequence-specific manner. In some embodiments, the target sequence described herein comprises a nucleotide sequence shown as any one of SEQ ID NOs: 75-79 and 130. In some embodiments, this disclosure provides active variants and fragments of RGN having amino acid sequences shown as any one of SEQ ID NO: 3, 7, 11, and 15, and active variants and fragments of naturally occurring CRISPR repeat sequences, including sequences shown as any one of SEQ ID NO: 4, 8, 12, 16, and 106; active variants and fragments of naturally occurring tracrRNA, such as any one of sequences shown as any one of SEQ ID NO: 5, 9, 13, 17, 107, and 120; and active variants and fragments of sgRNA, such as sequences shown as any one of SEQ ID NO: 25-29, and polynucleotides encoding them. In some embodiments, the sgRNA of this disclosure includes sgRNA shown as SEQ ID NO: 6, 10, 14, or 18, wherein the sgRNA comprises any spacer suitable for targeting a target sequence in the mHTT allele and a backbone that can be bound by the RGN polypeptide or its active variant or fragment, respectively, of SEQ ID NO: 3, 7, 11, or 15.

[0172] Although the activity of variants or fragments may change relative to the polynucleotide or peptide of interest, variants and fragments should retain the functionality of the polynucleotide or peptide of interest. For example, variants or fragments may have increased activity, decreased activity, a different range of activity, or any other change in activity compared to the polynucleotide or peptide of interest.

[0173] Fragments and variants of naturally occurring RGN polypeptides (such as those disclosed herein) retain sequence-specific RNA-guided DNA binding activity. In embodiments, fragments and variants of naturally occurring RGN polypeptides (such as those disclosed herein) retain nuclease activity (single-stranded or double-stranded).

[0174] Fragments and variants of naturally occurring CRISPR repeat sequences (such as those disclosed herein) retain the ability to bind in a sequence-specific manner and guide nucleases (complexed with the guide RNA) to the target sequence when used as part of a guide RNA (including tracrRNA).

[0175] Fragments and variants of naturally occurring tracrRNAs (such as those disclosed herein) retain the ability to guide nucleases (complexed with the guide RNA) to the target sequence in a sequence-specific manner when used as part of guide RNA (including CRISPR RNA).

[0176] Fragments and variants of sgRNA (such as those disclosed herein) will retain the ability to guide RNA to nucleases (complexed with sgRNA) to the target sequence in a sequence-specific manner.

[0177] The term "fragment" refers to a portion of the polynucleotide or polypeptide sequence disclosed herein. A "fragment" or "bioactive portion" includes a polynucleotide containing a sufficient number of consecutive nucleotides to retain biological activity (i.e., to bind in a sequence-specific manner and guide the RGN to the target sequence when included within guide RNA). A "fragment" or "bioactive portion" includes a polypeptide containing a sufficient number of consecutive amino acid residues to retain biological activity (i.e., to bind in a sequence-specific manner to the target sequence when complexed with guide RNA). Fragments of the RGN protein include fragments shorter than the full-length sequence due to the use of an alternative downstream initiation site. The bioactive portion of an RGN protein may be a polypeptide comprising, for example, an RGN that binds to the target nucleotide sequence disclosed herein, or an RGN having an amino acid sequence shown as any of the following: SEQ ID NO: 3, 7, 11, and 15. Such bioactive portions may be prepared using recombinant techniques and evaluated for sequence-specific RNA-guided DNA binding activity. The biologically active fragment of a CRISPR repeat sequence may comprise at least eight consecutive nucleotides of any one of the following: SEQ ID NO: 4, 8, 12, 16, and 106. The biologically active portion of a CRISPR repeat sequence may be a polynucleotide comprising, for example, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 consecutive nucleotides of any one of the following: SEQ ID NO: 4, 8, 12, 16, and 106. The biologically active portion of tracrRNA may be a polynucleotide comprising, for example, any one of the following: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or more consecutive nucleotides: SEQ ID NO: 5, 9, 13, 17, 107 and 120.The biologically active portion of sgRNA may be a polynucleotide comprising, for example, any one of the following: 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more consecutive nucleotides: SEQ ID NO: 6, 10, 14, 18 and 25-29.

[0178] Generally, "variant" is intended to refer to substantially similar sequences. For polynucleotides, a variant comprises the deletion and / or addition of one or more nucleotides at one or more internal sites within the native polynucleotide and / or the substitution of one or more nucleotides at one or more sites within the native polynucleotide. As used herein, "native" or "wild-type" polynucleotides or polypeptides comprise naturally occurring nucleotide or amino acid sequences, respectively. For polynucleotides, conserved variants include those sequences that encode the native amino acid sequence of the gene of interest due to genetic code degeneracy. Naturally occurring allelic variants (such as these allelic variants) can be identified using well-known molecular biology techniques, such as polymerase chain reaction (PCR) and hybridization techniques as outlined below. Variant polynucleotides also include synthetically derived polynucleotides, such as those produced by, for example, site-directed mutagenesis, but which still encode the polypeptide or polynucleotide of interest. Generally, as determined by the sequence alignment procedures and parameters described elsewhere in this document, the variants of the specific polynucleotide disclosed herein will have at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with that specific polynucleotide.

[0179] Variants of the specific polynucleotides disclosed herein (i.e., reference polynucleotides) can also be assessed by comparing the percentage of sequence identity between the polypeptide encoded by the variant polynucleotide and the polypeptide encoded by the reference polynucleotide. The percentage of sequence identity between any two polypeptides can be calculated using the sequence alignment procedures and parameters described elsewhere herein. When any given polynucleotide pair disclosed herein is assessed by comparing the percentage of sequence identity shared by two polypeptides encoded by the polynucleotide, the percentage of sequence identity between the two encoding polypeptides is at least about 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater.

[0180] In some embodiments, the polynucleotides disclosed in this invention encode an RNA-guided nuclease polypeptide comprising an amino acid sequence that encodes an RGN that binds to the target sequence disclosed herein, or an amino acid sequence shown as any one of the following, having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity. SEQ ID NO: 75-79 and 130.

[0181] Bioactive variants of the RGN polypeptide disclosed herein may differ by as few as about 1-15 amino acid residues, as few as about 1-10, such as about 6-10, as few as 5, as few as 4, as few as 3, as few as 2 or as few as 1 amino acid residue. In some embodiments, the polypeptide may include an N-terminal or C-terminal truncation, which may include a deletion of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700 or more amino acids at the N-terminus or C-terminus of the polypeptide.

[0182] In some embodiments, the polynucleotides disclosed in this invention comprise or encode crRNA repeat sequences that contain nucleotide sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity with nucleotide sequences shown as any one of the following: SEQ ID NO: 4, 8, 12, 16, and 106.

[0183] The polynucleotides disclosed in this invention may comprise or encode tracrRNA, which comprises a nucleotide sequence having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity with any of the following nucleotide sequences: SEQ ID NO: 5, 9, 13, 17, 107 and 120.

[0184] The polynucleotides disclosed in this invention may comprise or encode sgRNAs that contain nucleotide sequences having at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater identity with any of the following nucleotide sequences: SEQ ID NO: 6, 10, 14, 18 and 25-29.

[0185] The bioactive variants of the CRISPR repeat sequences, crRNA, tracrRNA, or sgRNA disclosed herein may differ by as few as about 1-15 amino acid residues, as few as about 1-10, such as about 6-10, as few as 5, as few as 4, as few as 3, as few as 2, or as few as 1 nucleotide. In some embodiments, the polynucleotide may contain a 5' or 3' truncation, which may contain at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 90, 95, 100, 105, 110, or more nucleotides deleted from the 5' or 3' end of the polynucleotide.

[0186] As used herein, a sequence that differs from the parental sequence by a certain number of amino acids or nucleotides can be attributed to amino acid or nucleotide substitutions, additions, and / or deletions. For example, a nucleotide sequence that differs from the parental nucleotide sequence by one nucleotide may have a single nucleotide substitution, may be one nucleotide longer than the parental sequence, or may be one nucleotide shorter than the parental sequence.

[0187] It is recognized that the RGN peptides, CRISPR repeat sequences, crRNA, tracrRNA, and sgRNA provided herein can be modified to produce variant proteins and polynucleotides. Artificially designed changes can be introduced using site-directed mutagenesis. Alternatively, native, unknown, or unidentified polynucleotides and / or peptides belonging to the aspects of this disclosure that are structurally and / or functionally related to the sequences disclosed herein can be identified. Conserved amino acid substitutions can be performed in non-conserved regions without altering the function of the RGN protein. Alternatively, modifications can be made to improve RGN activity.

[0188] Variant polynucleotides and proteins also encompass sequences and proteins derived from mutagenesis and recombination procedures, such as DNA shuffling. Using such procedures, one or more different RGN proteins disclosed herein (e.g., SEQ ID NO: 3, 7, 11, or 15) are manipulated to produce novel RGN proteins with desired properties. In this way, recombinant polynucleotide libraries are generated from a population of related sequence polynucleotides containing sequence regions with substantial sequence identity and capable of homologous recombination in vitro or in vivo. For example, using this method, sequence motifs encoding the domain of interest can be shuffled between the RGN sequences provided herein and other known RGN genes to obtain sequences encoding modified properties of interest (such as increased K in the case of enzymes). mThis involves creating new genes for proteins. Such DNA shuffling strategies are known in the art. See, for example, Stemmer (1994), Proceedings of the National Academy of Sciences of the United States of America (PNAS) 91:10747-10751; Stemmer (1994), Nature 370:389-391; Crameri et al. (1997), Nature Biotechnology 15:436-438; Moore et al. (1997), Journal of Molecular Biology 272:336-347; Zhang et al. (1997), Proceedings of the National Academy of Sciences of the United States of America (PNAS) 94:4504-4509; Crameri et al. (1998), Nature 391:288-291; and U.S. Patent Nos. 5,605,793 and 5,837,458. "Recombined" nucleic acids are nucleic acids generated through a recombining procedure (such as any recombining procedure described herein). Recombined nucleic acids are generated by, for example, manually and, where appropriate, recursively recombinating (physical or virtual) two or more nucleic acids (or strings). Generally, one or more screening steps are used during the recombining process to identify the nucleic acids of interest; these screening steps may be performed before or after any recombination steps. In some (but not all) recombining implementations, multiple rounds of recombination are required before selection to increase the diversity of the library to be screened. The overall process of recombination and selection is repeated recursively, where appropriate. Depending on the context, recombining may refer to the entire process of recombination and selection, or alternatively, may refer only to the recombination portion of the overall process.

[0189] As used herein, “sequence identity” or “identity” in the context of two polynucleotide or polypeptide sequences refers to identical residues in two sequences when aligned according to maximum correspondence within a specified comparison window. It is recognized that dissimilar residue positions are often distinguished by conserved amino acid substitutions, where the amino acid residue substitutes for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore does not alter the functional properties of the molecule. Protein sequences that differ due to such conserved substitutions are said to have “sequence similarity” or “similarity.” Methods for measuring sequence similarity are well known to those skilled in the art. Typically, this involves scoring conserved substitutions as partial mismatches rather than complete mismatches. Thus, for example, where identical amino acids are assigned a score of 1 and non-conservative substitutions are assigned a score of 0, conserved substitutions are assigned a score between 0 and 1. The scoring of conserved substitutions is calculated, for example, in the program PC / GENE (Intelligenetics, Mountain View, California).

[0190] As used herein, "sequence identity percentage" refers to a value determined by comparing two best-aligned sequences within a comparison window. The polynucleotide sequence portion of the comparison window may contain additions or deletions (i.e., gaps) compared to a reference sequence (which does not contain additions or deletions) to best align the two sequences. The percentage is calculated as follows: determine the number of positions in the two sequences where there are identical nucleic acid bases or amino acid residues, obtain the number of matching positions, divide this number of matching positions by the total number of positions in the comparison window, and multiply the result by 100 to obtain the sequence identity percentage.

[0191] Unless otherwise stated, the sequence identity / similarity values ​​provided herein refer to values ​​obtained using GAP version 10 with the following parameters: nucleotide sequence identity % and similarity % using a 50-fold gap weight and a 3-fold length weight, and an nwsgapdna.cmp scoring matrix; amino acid sequence identity % and similarity % using an 8-fold gap weight and a 2-fold length weight, and a BLOSUM62 scoring matrix; or any equivalent procedure thereof. “Equivalent procedure” means any sequence comparison procedure that, for any two sequences in discussion, generates an alignment that has the same nucleotide or amino acid residue match and the same percentage of sequence identity compared to the corresponding alignment generated by GAP version 10.

[0192] When two sequences are aligned using a defined amino acid substitution matrix (e.g., BLOSUM62), a gap presence penalty, and a gap expansion penalty to achieve the highest possible score for that pair of sequences, the sequence is considered "optimally aligned." Amino acid substitution matrices and their use in quantifying the similarity between two sequences are well-known in the art and described, for example, in Dayhoff et al. (1978), "A model of evolutionary change in proteins," *Atlas of Protein Sequence and Structure*, Vol. 5, Supplement 3 (edited by MO Dayhoff), pp. 345-352. Also cited in *Proceedings of the National Academy of Sciences of the United States of America* (Proc. Natl. Acad. Sci. USA) 89:10915-10919, by Henikoff et al. (1992). The BLOSUM62 matrix is ​​commonly used as a preset score substitution matrix in sequence alignment schemes. For a single amino acid gap introduced into one of the aligned sequences, a gap penalty is applied, and a gap expansion penalty is applied for each additional empty amino acid position inserted into an already opened gap. Alignment is defined by the start and end amino acid positions of each sequence and, as appropriate, by inserting one or more gaps in one or two sequences to achieve the highest possible score. Although optimal alignment and scoring can be performed manually, this process is facilitated by the use of computer-implemented alignment algorithms, such as gapped BLAST 2.0 described in Altschul et al. (1997) Nucleic Acid Research 25:3389-3402, which is publicly available on the website of the National Center for Biotechnology Information (www.ncbi.nlm.nih.gov). Optimal alignments, including multiple alignments, can be made using, for example, pSI-BLAST, which is available from www.ncbi.nlm.nih.gov and described in Altschul et al. (1997) Nucleic Acid Research 25:3389-3402.

[0193] Regarding the optimal alignment of a nucleotide or amino acid sequence with a reference sequence, the nucleotide or amino acid residue "corresponds" to the position in the reference sequence that pairs with that nucleotide or residue in the alignment. This "position" is indicated by the numbering of the nucleotides in the reference nucleotide sequence based on their position relative to the 5' end, or by the numbering of the amino acids in the reference amino acid sequence based on their position relative to the N-terminus. Because deletions, insertions, truncations, fusions, etc., must be considered when determining the optimal alignment, the nucleotide position or amino acid residue number in the test sequence determined by counting only from the 5' end or N-terminus will generally not be the same as its corresponding position number in the reference sequence. For example, in the case of a deletion in the aligned test sequence, there will be no nucleotide or amino acid at the position corresponding to the deletion site in the reference sequence. In the case of an insertion in the aligned reference sequence, the insertion will not correspond to any nucleotide or amino acid position in the reference sequence. In the case of truncation or fusion, there may be nucleotide or amino acid extensions in the reference or aligned sequence that do not correspond to any nucleotide or amino acid in the corresponding sequence.

[0194] VII. RGN system for binding to target sequences of interest and ribonucleoprotein complex and their preparation methods

[0195] This disclosure provides an RGN system for binding to a target sequence in a mutant HTT allele. As used herein, the RGN system comprises at least one RGN polypeptide or a polynucleotide comprising a nucleotide sequence encoding an RGN polypeptide and one or more guide RNAs capable of forming a complex (ribonucleoprotein (RNP) complex) with the RGN polypeptide. The RGN system comprises: a) one or more guide RNAs or one or more polynucleotides comprising one or more nucleotide sequences encoding one or more guide RNAs; and b) an RGN polypeptide or a polynucleotide comprising a nucleotide sequence encoding an RGN polypeptide, wherein the one or more guide RNAs are capable of forming a complex with the RGN polypeptide to direct the binding of the RGN polypeptide to a target sequence in the mutant HTT allele. The guide RNA hybridizes to the target strand of the target sequence in the mutant HTT allele and also forms a complex with the RGN polypeptide, thereby directing the binding of the RGN polypeptide to the target sequence. In some embodiments, the target sequence comprises a nucleotide sequence shown as any one of SEQ ID NO: 75-79 and 130. In some embodiments, the RGN is capable of recognizing a common PAM sequence shown as NNNNCC, NNRYA, NNGRR, or NNGG. In some embodiments, the RGN comprises an amino acid sequence shown as any one of SEQ ID NO: 3, 7, 11, and 15, or an active variant or fragment thereof. In some embodiments, the RGN comprises an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95%, or greater sequence identity with any one of SEQ ID NO: 3, 7, 11, and 15. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising a nucleotide sequence shown as any one of SEQ ID NO: 4, 8, 12, 16, and 106, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising any nucleotide sequence shown as any one of SEQ ID NO: 5, 9, 13, 17, 107, and 120, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA containing any nucleotide sequence, or an active variant or fragment thereof, of any one of SEQ ID NOs: 6, 10, 14, 18, and 25-29. The guide RNA of the system may be a single guide RNA or dual guide RNA. In some embodiments, the system comprises an RNA-guided nuclease heterologous to the guide RNA, wherein the RGN and the guide RNA are found to be essentially non-complexed (i.e., do not bind to each other).

[0196] In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown in SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC.

[0197] In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown in SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 134 and recognizing the PAM sequence NNRYA.

[0198] In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown in SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0199] In some embodiments, the PAM interaction domain of the RGN polypeptide has the sequence shown in SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 136 and recognizing the PAM sequence NNGG.

[0200] The RGN system disclosed in this invention may include an RGN polypeptide comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. The nuclease domain may comprise a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 143, 144, 145, and 146.

[0201] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 147, 148, 149, and 150.

[0202] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 151, 152, 153, and 154.

[0203] The RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 155 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 155, 156, 157, and 158.

[0204] The RGN system disclosed in this invention may include an RGN polypeptide comprising a PAM interaction domain that facilitates recognition and binding to a PAM site and further comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence shown as any of SEQ ID NO: 143, 144, 145, and 146.

[0205] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 147, 148, 149, and 150.

[0206] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 151, 152, 153, and 154.

[0207] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 155, 156, 157, and 158.

[0208] The system provided herein for binding a target sequence of interest may be a ribonucleoprotein complex, which is at least one RNA molecule that binds to at least one protein. The ribonucleoprotein complex provided herein comprises at least one guide RNA as an RNA component and an RNA-guided nuclease as a protein component. Such ribonucleoprotein complexes can be purified from cells or organisms that naturally express RGN peptides and have been engineered to express specific guide RNAs that are specific to the target sequence of interest (e.g., the target sequence in a mutant HTT allele). Alternatively, the ribonucleoprotein complex can be purified from cells or organisms that have been transformed with a polynucleotide (e.g., mRNA) encoding both the RGN peptide and the guide RNA and cultured under conditions allowing the expression of the RGN peptide and the guide RNA. In some embodiments, the ribonucleoprotein complex is purified from cells or organisms that have been transformed with a polynucleotide (e.g., mRNA) encoding the RGN peptide and in which a synthetically derived gRNA has been introduced. Therefore, methods for manufacturing RGN peptides or RGN ribonucleoprotein complexes are provided. Such methods involve culturing cells containing a nucleotide sequence encoding the RGN peptide and, in some embodiments, a nucleotide sequence encoding the guide RNA, under conditions expressing the RGN peptide (and, in some embodiments, the guide RNA). The RGN polypeptide or RGN ribonucleoprotein can then be purified from the lysate of cultured cells. In some embodiments, the nucleotide sequence encoding the RGN polypeptide includes mRNA (messenger RNA). In some embodiments, the method for assembling the RNP complex includes combining one or more of the guide RNAs disclosed herein with one or more of the RGN polypeptides disclosed herein under conditions suitable for forming the RNP complex.

[0209] Methods for purifying RGN peptides or RGN ribonucleoprotein complexes from lysates of biological samples are known in the art (e.g., size exclusion and / or affinity chromatography, 2D-PAGE, HPLC, reversed-phase chromatography, immunoprecipitation). In certain methods, the RGN peptide is generated recombinantly and includes purification tags to aid its purification. These purification tags include, but are not limited to, glutathione S-transferase (GST), chitin-binding protein (CBP), maltose-binding protein, thioredoxin (TRX), poly(NANP), tandem affinity purification (TAP) tags, myc, AcV5, AU1, AU5, E, ECS, E2, FLAG (e.g., 3X FLAG tags), HA, nus, Softag 1, Softag 3, Strep, SBP, Glu-Glu, HSV, KT3, S, S1, T7, V5, VSV-G, 6xHis, 10xHis, biotinylate carboxyl carrier protein (BCCP), and calmodulin. Generally, immobilized metal affinity chromatography is used to purify labeled RGN peptides or RGN ribonucleoprotein complexes. It should be understood that other similar methods known in the art, including other forms of chromatography or, for example, immunoprecipitation, can be used alone or in combination.

[0210] "Isolated" or "purified" polypeptides, or their biologically active portions, are substantially or substantially free of components typically found in their natural environment that accompany or interact with the polypeptide. Therefore, when produced by recombinant technology, isolated or purified polypeptides are substantially free of other cellular material or culture media, or when chemically synthesized, substantially free of chemical precursors or other chemicals. Proteins substantially free of cellular material include protein formulations having less than about 30%, 20%, 10%, 5%, or 1% (dry weight) of contaminating proteins. When the proteins of this disclosure, or their biologically active portions, are produced recombinantly, preferably, the culture medium represents less than about 30%, 20%, 10%, 5%, or 1% (dry weight) of chemical precursors or chemicals not of interest. Similarly, "isolated" polynucleotides or nucleic acid molecules are removed from their natural environment. When chemically synthesized or removed from genomic loci by breaking phosphodiester bonds, isolated polynucleotides are substantially free of chemical precursors or other chemicals. Isolated polynucleotides may be part of a carrier, a composition of substances, or may be contained within cells, provided the cells are not the original environment of the polynucleotides.

[0211] The specific methods provided herein for binding and / or cleaving target sequences of interest involve the use of an in vitro assembled RGN ribonucleoprotein complex. The in vitro assembly of the RGN ribonucleoprotein complex can be performed using any method known in the art, wherein the RGN peptide is contacted with the guide RNA under conditions that allow binding to the guide RNA. As used herein, “contact,” “contacting,” and “contacted” mean placing the components of the desired reaction together under conditions suitable for carrying out the desired reaction. The RGN peptide may be purified from a biological sample, cell lysate, or culture medium, generated via in vitro translation, or chemically synthesized. The guide RNA may be purified from a biological sample, cell lysate, or culture medium, transcribed in vitro, or chemically synthesized. The RGN peptide and guide RNA may be contacted in a solution (e.g., a buffered saline solution) to allow for the in vitro assembly of the RGN ribonucleoprotein complex.

[0212] Some aspects of this disclosure provide kits comprising one or more elements of the RGN system described herein, including: a guide RNA (i.e., crRNA, tracrRNA, and / or sgRNA), an RGN, and / or a polynucleotide encoding it; cells; and the complete RGN system, and, in some embodiments, another type of nuclease. In some embodiments, the kit includes suitable reagents, buffers, and / or instructions for using one or more elements of the RGN system, such as instructions for in vitro or in vivo nucleic acid editing. The reagents are available in any suitable container, such as vials, bottles, or tubes. The reagents can be used in methods utilizing one or more elements of the RGN system. For example, a restriction enzyme may be included for cloning a polynucleotide encoding an RGN or guide RNA into a vector. In some embodiments, the kit includes instructions for designing and using a suitable guide RNA (i.e., crRNA, tracrRNA, and / or sgRNA) to target and edit nucleic acid sequences. The reagents may be provided in a form suitable for a particular assay or in a form requiring the addition of one or more other components prior to use (e.g., in a concentrate or lyophilized form). The buffer can be any buffer, including but not limited to sodium carbonate buffer, sodium bicarbonate buffer, borate buffer, Tris buffer, MOPS buffer, HEPES buffer, and combinations thereof. In some embodiments, the buffer is alkaline. In some embodiments, the buffer has a pH of about 7 to about 10.

[0213] Kits incorporating one or more elements of the RGN system of this disclosure are useful in a variety of applications involving target polynucleotides in a wide range of cell types, including modifications such as deletions, insertions, translocations, inactivations, and activations. Therefore, kits incorporating one or more elements of the RGN system of this disclosure are suitable for applications such as gene therapy, drug screening, disease diagnosis, and prognosis.

[0214] In some embodiments, the kits disclosed herein include pharmaceutical kits comprising the pharmaceutical compositions described herein. In some embodiments, the pharmaceutical kit may include: (a) a container containing a lyophilized form of a composition of the present disclosure and (b) a second container containing a pharmaceutically acceptable diluent for injection (e.g., sterile water). The pharmaceutically acceptable diluent may be used to reconstitute or dilute the lyophilized compound of the present disclosure. As appropriate, the precautions associated with such containers may be in the form prescribed by a government agency regulating the manufacture, use, or sale of pharmaceutical or biological products, reflecting approval for human administration by such agency.

[0215] VIII. Methods for binding, cleaving, and / or modifying target sequences

[0216] This disclosure provides a method for binding, cleaving, and / or modifying (i.e., editing) a target sequence in a mutant HTT allele. The method includes introducing an RGN system comprising at least one guide RNA or a polynucleotide encoding thereof and at least one RGN polypeptide or a polynucleotide encoding thereof into a cell containing the target sequence. In some embodiments, the delivery is ex vivo, and the cell containing the target sequence may be a stem cell, zygote, embryonic cell, or gamete. In some embodiments, the stem cell is an induced pluripotent stem cell (iPSC) or a mesenchymal stem cell (MSC). In some embodiments, the method includes in vivo delivery of the RGN system, and the cell containing the target sequence is in vivo. In some embodiments, the target sequence in the mutant HTT allele has a nucleotide sequence shown as any one of SEQ ID NOs: 75-79 and 130. In some embodiments, the RGN is capable of recognizing a common PAM sequence shown as any one of NNNNCC, NNRYA, NNGRR, and NNGG. In some embodiments, the RGN comprises an amino acid sequence shown as any one of SEQ ID NOs: 3, 7, 11, and 15, or an active variant or fragment thereof. The system's guide RNA can be a single guide RNA or two guide RNAs.

[0217] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 3. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 5, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 27 or 28, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC.

[0218] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 7. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 25 or 26, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM-interacting domain having an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 134 and recognizing the PAM sequence NNRYA.

[0219] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 11. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 29, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0220] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 15. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 17, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 136 and recognizing the PAM sequence NNGG.

[0221] The RGN system disclosed in this invention may include an RGN polypeptide comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. The nuclease domain may comprise a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 143, 144, 145, and 146.

[0222] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 147, 148, 149, and 150.

[0223] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 151, 152, 153, and 154.

[0224] The RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 155 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 155, 156, 157, and 158.

[0225] The RGN system disclosed in this invention may include an RGN polypeptide comprising a PAM interaction domain that facilitates recognition and binding to a PAM site and further comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence shown as any of SEQ ID NO: 143, 144, 145, and 146.

[0226] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 147, 148, 149, and 150.

[0227] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 151, 152, 153, and 154.

[0228] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 155, 156, 157, and 158.

[0229] In some implementations, the RGN and / or guide RNA are heterologous to the cell in which the RGN and / or guide RNA (or a polynucleotide encoding at least one of the RGN and guide RNA) is introduced.

[0230] In embodiments where the method includes delivering a polynucleotide encoding a guide RNA and / or an RGN polypeptide, the cells may subsequently be cultured under conditions expressing the guide RNA and / or the RGN polypeptide. In some embodiments, the method includes contacting a target nucleic acid molecule with an RGN ribonucleoprotein complex. In some embodiments, the method includes introducing an RGN ribonucleoprotein complex into a cell containing the target nucleic acid molecule. The RGN ribonucleoprotein complex may be a complex purified from a biological sample, recombinantly generated and subsequently purified or assembled in vitro as described herein. In embodiments where the RGN ribonucleoprotein complex in contact with the target nucleic acid molecule or a cell containing the target nucleic acid molecule has been assembled in vitro, the method may further include assembling the complex in vitro prior to contact with the target nucleic acid molecule or a cell containing the target nucleic acid molecule.

[0231] The purified or in vitro assembled RGN ribonucleoprotein complex can be introduced into cells using any method known in the art, including but not limited to electroporation. Alternatively, RGN polypeptides and / or polynucleotides encoding or containing guide RNA can be introduced into cells using any method known in the art (e.g., electroporation).

[0232] After delivery to or contact with the target nucleic acid molecule or a cell containing the target nucleic acid molecule, the guide RNA directs the RGN to bind to the target sequence within the target nucleic acid molecule in a sequence-specific manner. In embodiments where the RGN has nuclease activity, the RGN peptide cleaves the target sequence after binding. The target sequence can then be modified (i.e., edited) via endogenous repair mechanisms such as non-homologous end joining (NHEJ).

[0233] Methods for measuring the binding of RGN peptides to target sequences are known in the art and include chromatin immunoprecipitation analysis, gel migration variation analysis, DNA pull-down analysis, reporter gene analysis, and microplate capture and detection analysis. Similarly, methods for measuring the cleavage or modification of target nucleic acid molecules containing target sequences are known in the art and include in vitro or in vivo cleavage analysis, where cleavage is confirmed using PCR, sequencing, or gel electrophoresis, with or without appropriate labeling (e.g., radioisotopes, fluorescent substances) attached to the target sequence to facilitate the detection of degradation products. Alternatively, nicking-triggered exponential amplification reaction (NTEXPAR) analysis can be used (see, for example, Zhang et al. (2016), *Chem. Sci.* 7:4951-4957). In vivo cleavage can be assessed using Surveyor analysis (Guschin et al. (2010), *Methods in Molecular Biology* 649:247-256).

[0234] This method may involve using only one RGN and only one guide RNA. Single double-strand cleavage within or near amplified trinucleotide repeat sequences has been shown to result in loss or reduction of the repeat region's length, which may be attributed to the instability of the repeat bundle (Richard et al., PLoS ONE (2014), 9(4):e95611; Mittelman et al., Proceedings of the National Academy of Sciences (2009), 106(24):9607-12; van Agtmaal et al., Molecular Therapy, January 4, 2017; 25(1):24-43). In some embodiments, the RGN and its associated guide RNA recognize PAMs generated by SNPs in the mutant HTT allele, causing the mutant HTT allele to be cleaved by the RGN system, potentially resulting in reduced levels of mutant HTT protein and / or mutant HTT mRNA.

[0235] This method may involve using a single type of RGN compounded with more than one guide RNA. In some embodiments, the method involves using two types of RGN, each compounded with a guide RNA. More than one guide RNA can target different regions of the mutated HTT allele. For example, a first guide RNA can target near the 5' end of an amplified trinucleotide repeat sequence in the HTT gene; and a second guide RNA can target near the 3' end of the amplified trinucleotide repeat sequence to allow for excision of the amplified trinucleotide repeat sequence.

[0236] Double-strand breaks introduced by RGN peptides can be repaired via non-homologous end joining (NHEJ) repair. Due to the error-prone nature of NHEJ, double-strand break repair can induce mutations in the target sequence. In some embodiments, a “mutation” in a nucleic acid molecule refers to a change in the nucleotide sequence of the nucleic acid molecule, which may be the deletion, insertion, or substitution of one or more nucleotides, or a combination thereof. In some embodiments, cleavage of the mutHTT allele causes the introduction of INDELs (insertions and / or deletions) and premature termination, resulting in a decrease in mutHTT mRNA and / or protein levels. In some embodiments, cleavage of the mutHTT allele causes the introduction of premature stop codons, resulting in a decrease in mutHTT protein levels.

[0237] In some embodiments, the cells incorporating RGN and / or guide RNA (or polynucleotides encoding at least one of RGN and guide RNA) have a mutant HTT allele containing at least 27 CAG repeat sequences (e.g., 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, or more than 36 CAG repeat sequences) in exon 1. The cells may contain a mutant HTT allele containing at least 36 CAG repeat sequences (e.g., 36, 37, 38, 39, 40, or more than 40 CAG repeat sequences) in exon 1. In other embodiments, the cells contain a mutant HTT allele containing at least 40 CAG repeat sequences (e.g., 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, or more than 56 CAG repeat sequences) in exon 1. Cells may also contain mutant HTT alleles containing at least 56 CAG repeat sequences (e.g., 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70 or more) in exon 1.

[0238] The method may include the use of an RGN peptide or RGN system that cannot cleave the wild-type HTT allele. An RGN peptide or RGN system that cannot cleave the wild-type HTT allele means that the RGN peptide or RGN system cannot cleave the wild-type HTT allele, or cleaves it to a negligible degree such that the levels of wtHTT mRNA and / or wtHTT protein are not significantly reduced, where, for example, wtHTT can maintain support for key cellular and neural functions and / or the absence of Huntington's disease symptoms in the in vivo context (i.e., in individuals heterozygous for the mutHTT allele and administered the RGN peptide or RGN system). In some embodiments, the RGN peptide or RGN system is cleaved to a negligible degree, such that the content of wtHTT mRNA and / or wtHTT protein is reduced by 5% or less, 4% or less, 3% or less, 2% or less, 1% or less, 0.9% or less, 0.8% or less, 0.7% or less, 0.6% or less, 0.5% or less, 0.4% or less, 0.3% or less, 0.2% or less, or 0.1% or less.

[0239] IX. Cells containing polynucleotide gene modifications

[0240] This document provides cells and organisms containing a target sequence in a mutant HTT allele modified using methods mediated by RGN, crRNA, tracrRNA, and / or sgRNA as described herein. Cells containing the modified target sequence in the mutant HTT allele may be Huntington's disease patient cells, stem cells, zygotes, embryonic cells, or gametes. In some embodiments, the stem cells are induced pluripotent stem cells (iPSCs) or mesenchymal stem cells (MSCs). In some embodiments, the cells are derived from induced pluripotent stem cells (iPSCs) or mesenchymal stem cells (MSCs). In some embodiments, the cells containing the modified target sequence in the mutant HTT allele are in vitro or ex vivo. In some embodiments, the cells containing the modified target sequence in the mutant HTT allele are in vivo. The modified cells (e.g., embryonic cells, zygotes, gametes) can develop into an organism under appropriate conditions. The RGN introduced into the cells to modify the target sequence in the mutant HTT allele recognizes a common PAM sequence including any one of NNNNCC, NNRYA, NNGRR, and NNGG. In some implementations, the target sequence in the mutated HTT allele has a nucleotide sequence shown as any one of SEQ ID NO: 75-79 and 130. The system's guide RNA can be a single guide RNA or two guide RNAs.

[0241] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 3. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 4, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 5, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 27 or 28, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizes the PAM sequence NNNNCC. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC.

[0242] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 7. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 8 or 106, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 9 or 107, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 25 or 26, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizes the PAM sequence NNRYA. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 134 and recognizing the PAM sequence NNRYA.

[0243] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 11. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 12, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 13 or 120, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises sgRNA comprising the nucleotide sequence shown as SEQ ID NO: 29, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizes the PAM sequence NNGRR. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 135 and recognizes the PAM sequence NNGRR.

[0244] RGN may comprise an amino acid sequence having at least 80%, at least 85%, at least 90%, at least 95% or greater sequence identity with the following sequence: SEQ ID NO: 15. In some embodiments, the guide RNA comprises a CRISPR repeat sequence comprising the nucleotide sequence shown as SEQ ID NO: 16, or an active variant or fragment thereof. In some embodiments, the guide RNA comprises a tracrRNA comprising the nucleotide sequence shown as SEQ ID NO: 17, or an active variant or fragment thereof. In some embodiments, the PAM interacting domain of the RGN polypeptide has the sequence shown as SEQ ID NO: 136 and recognizes the PAM sequence NNGG. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM-interacting domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizes the PAM sequence NNGG. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 includes a PAM-interacting domain having the sequence shown as SEQ ID NO: 136 and recognizing the PAM sequence NNGG.

[0245] The RGN system disclosed in this invention may include an RGN polypeptide comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. The nuclease domain may comprise a RuvC or HNH domain. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any one of SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 143, 144, 145, and 146.

[0246] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 147, 148, 149, and 150.

[0247] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 151, 152, 153, and 154.

[0248] The RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 155 comprises a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a nuclease domain having a sequence shown as any one of SEQ ID NO: 155, 156, 157, and 158.

[0249] The RGN system disclosed in this invention may include an RGN polypeptide comprising a PAM interaction domain that facilitates recognition and binding to a PAM site and further comprising at least one nuclease domain, wherein each nuclease domain is responsible for cleaving single strands of nucleic acid molecules. For example, an RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 143, 144, 145, and 146. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 3 comprises a PAM interaction domain having the sequence shown as SEQ ID NO: 133 and recognizing the PAM sequence NNNNCC, and may further comprise a nuclease domain having an amino acid sequence shown as any of SEQ ID NO: 143, 144, 145, and 146.

[0250] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 147, 148, 149, and 150. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 7 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 134 and recognizing the PAM sequence NNRYA, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 147, 148, 149, and 150.

[0251] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 151, 152, 153, and 154. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 11 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 135 and recognizing the PAM sequence NNGRR, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 151, 152, 153, and 154.

[0252] The RGN polypeptide of this disclosure having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 may comprise a PAM-interacting domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having an amino acid sequence having at least 95% sequence identity with SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence having at least 95% sequence identity with any of the following: SEQ ID NO: 155, 156, 157, and 158. In some embodiments, the RGN polypeptide having at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO: 15 comprises a PAM interaction domain having the sequence shown in SEQ ID NO: 136 and recognizing the PAM sequence NNGG, and may further comprise a nuclease domain having an amino acid sequence shown in any of SEQ ID NO: 155, 156, 157, and 158.

[0253] The modified cells may be eukaryotic (e.g., mammalian). In some embodiments, the cells modified by the methods disclosed in this invention include stem cells (e.g., induced pluripotent stem cells, mesenchymal stem cells), neurons, and glial cells. Stem cells are pluripotent, totipotent, or multipotent cells capable of differentiating into one or more different cell types. The term "pluripotent" refers to the ability of a cell to differentiate into any type of cell in a morphological organism and into extraembryonic material (such as the placenta). The term "pluripotent" refers to a cell line capable of differentiating into any terminally differentiated cell type. The term "multipotent" refers to a cell line capable of differentiating into at least two terminally differentiated cell types. The term "induced pluripotent stem cells" or "iPSCs" refers to a type of pluripotent stem cell similar to embryonic stem cells, formed by introducing certain embryonic genes (such as transgenic OCT4, SOX2, and KLF4) into somatic (e.g., adult) cells (see, for example, Takahashi and Yamanaka, Cell, 126, 663-676 (2006), which are incorporated herein by reference). Examples of somatic cells include, but are not limited to, bone marrow cells, epithelial cells, fibroblasts, hematopoietic cells, hepatocytes, intestinal cells, mesenchymal cells, bone marrow progenitor cells, neurons, glial cells, and spleen cells. Alternatively, iPSCs can be generated by reprogramming somatic cells into an embryonic stem cell-like state through forced expression of factors important for maintaining the "stemness" of embryonic stem cells (ESCs). Reprogramming factors can be expressed from expression cassettes contained in one or more vectors, such as integration vectors, chromosomal non-integrating RNA virus vectors, or free vectors, such as EBV element-based systems (Yu et al. (2009), Science, 324(5928):797-801). In some embodiments, reprogramming proteins or RNAs (such as mRNA or miRNA) can be introduced directly into somatic cells via protein or RNA transfection (Yakubov et al. (2010), Biochemical and Biophysical Research Communications, 394(1):189-193).

[0254] Mesenchymal stem cells (MSCs) can generate connective tissue, bone, cartilage, and cells in the circulatory and lymphatic systems. MSCs are found in mesenchyme, a portion of the embryonic mesoderm containing loosely packed, spindle-shaped, or stellate unspecialized cells. In some embodiments, MSCs include CD34. -Stem cells. MSCs can be isolated from various sources, including bone marrow, umbilical cord blood, (mobilized) peripheral blood, and adipose tissue (Horwitz et al., "Clarification of the nomenclature for MSC: the International Society for Cellular Therapy position statement." Cytotherapy (2005) 7:393-395).

[0255] Stem cells (e.g., iPSCs or MSCs) containing target sequences in mutant HTT alleles modified via the described RGN system can differentiate into neurons or glial cells. Differentiated cells refer to the change of a predetermined cell type (genotype and / or phenotype) to a non-predetermined cell type (genotype and / or phenotype). For example, differentiated stem cells (e.g., iPSCs or MSCs) refer to the induction of stem cells to divide into daughter cells with characteristics different from those of stem cells (e.g., iPSCs or MSCs), such as genotype (i.e., changes in gene expression, as determined by gene analysis such as microarrays) and / or phenotype (i.e., changes in protein expression). One or more of small molecules, growth factor proteins, and other growth conditions can be used to promote an unspecialized state, such as the conversion of stem cells to a more specialized cell fate (e.g., neuronal cells). In some embodiments, stem cell differentiation leads to the entry of stem cells into cellular pathways that produce somatic cells. For example, factors that differentiate stem cells into neurons may include: Wnt activator; SMAD inhibitors (e.g., Noggin peptide, SB-431542); neuronal growth factors (e.g., brain-derived neurotrophic factor, nerve growth factor, glial-derived neurotrophic factor); and / or the introduction of polynucleotides to express neuronal genes (e.g., neuron-2, NeuroD1). Differentiation may involve culturing pluripotent stem cells and / or their progeny cells in attached or suspended cultures.

[0256] Neuronal cells that can be modified using the methods described herein with RGN peptides, crRNA, tracrRNA, and / or guide RNA may include neural progenitor cells, forebrain neuronal progenitor cells, striatal neurons, medium-sized polyspinous neurons, and cortical neurons. Non-neuronal brain cells that can be modified using the methods described herein with RGN peptides, crRNA, tracrRNA, and / or guide RNA include glial cells. Glial cells may include microglia, astrocytes, and oligodendrocytes. In some embodiments, cells suitable for use in this disclosure include mammalian or human cells present in the putamen, caudate nucleus, striatum, cerebral cortex, globus pallidus, hippocampus, amygdala, thalamus, hypothalamus, subthalamic nucleus, substantia nigra, cerebellum, brainstem, or combinations thereof.

[0257] It also provides embryonic cells, zygotes, or gametes containing a mutant HTT allele modified by methods described herein using RGN, crRNA, tracrRNA, and / or sgRNA. The modified cells and organisms may be heterozygous or homozygous with respect to the modified mutant HTT allele. In some embodiments, the modified cells and organisms are heterozygous with respect to the modified mutant HTT allele.

[0258] Chromosomal modification of cells containing the target sequence in the mutant HTT allele using the RGN system of this disclosure can induce the expression of mutant HTT protein and / or downregulation of mutant HTT mRNA. In some embodiments, chromosomal modification causes a reduction or elimination of mutant HTT mRNA compared to the mutant HTT mRNA content in cells that have not been chromosomally modified using the RGN system. In some embodiments, chromosomal modification causes a reduction or elimination of mutant HTT protein compared to the mutant HTT protein content in cells that have not been chromosomally modified using the RGN system. The mutant HTT protein content can be measured by an analysis including immunoassays using antibodies that can distinguish between wild-type and mutant HTT proteins (e.g., Western blotting, ELISA, single-molecule counting immunoassay, immunoprecipitation assay in combination with flow cytometry, time-resolved fluorescence energy transfer (TR-FRET)), and including JESS capillary blot analysis as described herein in the examples. The mutant HTT mRNA content can be measured by, for example, RT-qPCR or array-based methods.

[0259] X. Pharmaceutical Composition

[0260] A pharmaceutical composition is provided comprising: the crRNA disclosed in this invention and its active variants and fragments, or the polynucleotide encoding thereof; the tracrRNA disclosed in this invention and its active variants and fragments, or the polynucleotide encoding thereof; the sgRNA disclosed in this invention and its active variants and fragments, or the polynucleotide encoding thereof; the RGN polypeptide disclosed in this invention and its active variants and fragments, or the polynucleotide encoding thereof; the RGN system disclosed in this invention; or the RNP complex disclosed in this invention comprising the RGN polypeptide and gRNA; or the vector disclosed in this invention (e.g., a viral vector); and a pharmaceutically acceptable carrier.

[0261] The pharmaceutical composition is a composition for the prevention, reduction, cure or otherwise treatment of a target symptom or disease, comprising an active ingredient (i.e., an RGN polypeptide, a polynucleotide encoding RGN, gRNA, a polynucleotide encoding gRNA, an RGN system, an RNP complex or a carrier) and a pharmaceutically acceptable carrier.

[0262] As used herein, a "pharmaceutically acceptable carrier" is a material that does not significantly irritate an organism and does not eliminate the activity and properties of the active ingredient (i.e., RGN polypeptide, RGN-encoding polynucleotide, gRNA, gRNA-encoding polynucleotide, RGN system, RNP complex, or carrier). The carrier must be of sufficiently high purity and sufficiently low toxicity to be suitable for administration to the individual being treated. The carrier may be inert or may have pharmaceutical benefits. In some embodiments, a pharmaceutically acceptable carrier comprises one or more compatible solid or liquid fillers, diluents, or encapsulating substances suitable for administration to humans or other vertebrates. In some embodiments, the pharmaceutically acceptable carrier is not naturally occurring. In some embodiments, the pharmaceutically acceptable carrier and active ingredient are not found together in nature.

[0263] The pharmaceutical compositions used in the methods disclosed in this invention can be formulated with suitable carriers, excipients, and other agents that provide suitable transfer, delivery, tolerability, and similar properties. A variety of suitable formulations are known to those skilled in the art. See, for example, Remington, The Science and Practice of Pharmacy (21st edition, 2005). Suitable formulations include, for example, powders, pastes, ointments, gels, waxes, oils, lipids, lipid-containing (cationic or anionic) vesicles (such as LIPOFECTIN vesicles), lipid nanoparticles, DNA conjugates, anhydrous absorbent pastes, oil-in-water and water-in-oil emulsions, carbowax emulsions (polyethylene glycol of various molecular weights), semi-solid gels, and semi-solid mixtures containing carbowax. Pharmaceutical compositions for oral or parenteral use can be prepared in unit dose dosage forms appropriate to the dosage of the active ingredient. Such unit dose dosage forms include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc.

[0264] Pharmaceutical compositions containing an active ingredient (i.e., an RGN polypeptide, a polynucleotide encoding an RGN, a gRNA, a polynucleotide encoding a gRNA, an RGN system or an RNP complex, or a carrier) may be emulsified or presented as a liposome composition, subject to the limitation that the emulsification process does not adversely affect the active ingredient or the patient.

[0265] Additional pharmaceutical agents included in a pharmaceutical composition may include pharmaceutically acceptable salts. Pharmaceutically acceptable salts include acid addition salts (formed from the free amino group of a polypeptide) formed from inorganic acids (such as hydrochloric acid or phosphoric acid) or organic acids (such as acetic acid, tartaric acid, mandelic acid, and similar acids). Salts formed with a free carboxyl group may also be derived from inorganic bases such as sodium hydroxide, potassium hydroxide, ammonium hydroxide, calcium hydroxide, or ferric hydroxide; and organic bases such as isopropylamine, trimethylamine, 2-ethylaminoethanol, histidine, procaine, and similar bases.

[0266] Physiologically tolerable and pharmaceutically acceptable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions containing no materials other than the active ingredient and water, or containing buffers such as sodium phosphate at physiological pH, physiological saline, or both, such as phosphate-buffered saline. Furthermore, aqueous carriers may contain more than one buffer salt, as well as salts such as sodium chloride and potassium chloride, dextrose, sucrose, mannose, polyethylene glycol, and other solutes. Liquid compositions may also contain not only water but also a water-excluded liquid phase. Such additional exemplary liquid phases are glycerin, vegetable oils (such as cottonseed oil), and water-oil emulsions. The amount of an active compound used in the cellular composition that is effective in treating a specific disease or symptom may be determined by standard clinical techniques, depending on the nature of the disease or symptom.

[0267] In some embodiments, the pharmaceutical composition comprises one or more molecules with surfactant properties to allow them to interact with biological membranes, such as pluronic or poloxamer, such as PLURONIC F68 (poloxamer 188, P188). In some embodiments, the pharmaceutical composition comprises between 0.001% and 0.1% poloxamer. In some embodiments, the pharmaceutical composition comprises about 0.001% poloxamer.

[0268] The RGN peptides, guide RNAs, RGN systems, polynucleotides encoding them, RNP complexes, or carriers disclosed in this invention can be formulated with pharmaceutically acceptable excipients, such as carriers, solvents, stabilizers, adjuvants, diluents, etc., depending on the specific administration route and dosage form. In some embodiments, these pharmaceutical compositions are formulated to obtain a physiologically compatible pH, and depending on the formulation and route of administration, within the range of pH from about 3 to about 11, and from about pH 3 to about pH 7. In some embodiments, the pH can be adjusted to the range of about pH 5.0 to about pH 8. In some embodiments, the composition may contain a therapeutically effective amount of at least one active ingredient as described herein (i.e., an RGN peptide, a polynucleotide encoding an RGN, gRNA, a polynucleotide encoding gRNA, an RGN system, an RNP complex, or a carrier) and one or more pharmaceutically acceptable excipients. In some embodiments, the composition comprises a combination of the active ingredients described herein, or includes a second active ingredient suitable for treating or preventing bacterial growth (e.g., but not limited to antibacterial or antimicrobial agents), or includes a combination of reagents disclosed herein.

[0269] Suitable excipients include, for example, carrier molecules, including large, slowly metabolizing macromolecules such as proteins, polysaccharides, polylactic acid, polyglycolic acid, polymeric amino acids, amino acid copolymers, and inactivated viral particles. Other exemplary excipients may include antioxidants (e.g., but not limited to ascorbic acid), chelating agents (e.g., but not limited to EDTA), carbohydrates (e.g., but not limited to dextrin, hydroxyalkyl cellulose, and hydroxyalkyl methyl cellulose), stearic acid, liquids (e.g., but not limited to oils, water, saline, glycerol, and ethanol), wetting agents or emulsifiers, pH buffers, and the like.

[0270] In some embodiments, the formulation is provided in single-dose or multi-dose containers (e.g., sealed ampoules and vials) and can be stored under lyophilized (freeze-dried) conditions, requiring the addition of a sterile liquid carrier, such as saline, water for injection, semi-liquid foam, or gel, just before use. Ready-to-use injectable solutions and suspensions can be prepared from the types of sterile powders, granules, and tablets previously described. In some embodiments, the active ingredient is dissolved in a buffer solution, frozen in single-dose or multi-dose containers, and later thawed for injection or kept / stabilized under refrigeration until use.

[0271] Therapeutic agents can be contained in controlled-release systems. To prolong the action of a drug, it is often necessary to slow the absorption of drugs injected subcutaneously, intrathecally, or intramuscularly. This can be achieved by using liquid suspensions of poorly water-soluble crystalline or amorphous materials. The absorption rate of the drug depends on its dissolution rate, which in turn depends on the crystal size and crystal form. Alternatively, parenteral administration of the drug can achieve delayed absorption by dissolving or suspending the drug in an oily medium. In some embodiments, the use of long-term sustained-release implants is particularly suitable for treating chronic conditions. Long-term sustained-release implants are well known to those skilled in the art.

[0272] Therapeutic agents (such as the crRNA and its active variants and fragments, or the polynucleotides encoding them, disclosed in this invention; the tracrRNA and its active variants and fragments, or the polynucleotides encoding them, disclosed in this invention; the sgRNA and its active variants and fragments, or the polynucleotides encoding them, disclosed in this invention; the RGN polypeptide and its active variants and fragments, or the polynucleotides encoding them, disclosed in this invention; the RGN system disclosed in this invention; the RNP complex containing the RGN polypeptide and gRNA disclosed in this invention; or the vectors disclosed in this invention (e.g., viral vectors)) may be separable from components that are normally present in their natural environment and usually accompany or interact with the therapeutic agent, or from chemical precursors or other chemicals used to synthesize the therapeutic agent, or from the culture medium when produced by recombinant technology, or from endotoxins and / or related pyrogenic substances, or substantially or substantially free of them. Endotoxins include toxins confined within the microorganism and released only upon microbial decomposition or death. Pyrogenic substances also include heat-stable substances (glycoproteins) from the outer membranes of bacteria and other microorganisms that induce fever. Both substances can cause fever, hypotension, and shock if administered to humans. Due to potential adverse effects, even low levels of endotoxins must be removed from intravenously administered pharmaceutical solutions. The U.S. Food and Drug Administration (“FDA”) has set an upper limit of 5 endotoxin units (EUs) / dose / kg body weight for intravenous drug administration over a one-hour period (The United States Pharmaceutical Convention, pharmaceuticalform) 26(1):223 (2000)). In certain specific aspects, the endotoxin and pyrogen content in the composition is less than about 1 EU / mg, or less than about 0.1 EU / mg, or less than about 0.01 EU / mg, or less than about 0.001 EU / mg. In some embodiments, the endotoxin and pyrogen content in the composition is 0.0138 EU / mg or less.

[0273] The therapeutic agent or pharmaceutical composition may have a purity of at least 80%, 85%, 90%, 95%, or greater. The therapeutic agent or pharmaceutical composition may contain low or undetectable levels of endotoxins or other impurities.

[0274] The pharmaceutical composition may be frozen, refrigerated, or stored at room temperature. Storage conditions may be below freezing, for example, below about -10°C, or below about -20°C, or below about -40°C, or below about -70°C. Storage conditions are generally below room temperature, for example, below about 32°C, or below about 30°C, or below about 27°C, or below about 25°C, or below about 20°C, or below about 15°C. In some embodiments, the formulation is stored at 2°C–8°C. For example, the formulation may be isotonic with blood or have an ionic strength that mimics physiological conditions.

[0275] In some embodiments, the pharmaceutical composition is stable under storage conditions. Stability can be measured using any suitable method in the art. Generally, a stable formulation is one exhibiting an increase of less than 5% in degradation products or impurities. In some embodiments, the formulation is stable under storage conditions for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, at least about 6 months, at least about 1 year, or at least about 2 years or longer. In some embodiments, the formulation is stable at 25°C for at least about 1 month, at least about 2 months, at least about 3 months, at least about 4 months, at least about 5 months, at least about 6 months, or at least about 1 year or longer.

[0276] When intended for internal administration, the pharmaceutical compositions of this disclosure shall be sterile. Formulations of this disclosure may be sterilized by a variety of sterilization methods, including sterile filtration, irradiation, and similar methods. In one aspect, the formulation is sterilized by filtration through a pre-sterilized 0.22-micron filter. Sterile compositions for injection may be formulated according to routine pharmaceutical practice, as described in Remington: The Science & Practice of Pharmacy, 21st edition, Lippincott Williams & Wilkins, (2005).

[0277] XI. Treatment methods for Huntington's disease

[0278] This article provides a method for treating Huntington's disease (HD) in an individual in need. The method comprises administering to the individual in need a crRNA or polynucleotide encoding the present invention, a tracrRNA or polynucleotide encoding the present invention, an sgRNA or polynucleotide encoding the present invention, an RGN or polynucleotide encoding the present invention, an RGN system disclosed in the present invention, an RNP complex disclosed in the present invention, or a vector disclosed in the present invention, or a pharmaceutical composition comprising any of these. In some embodiments, the therapeutic composition may reduce or inhibit the expression of the mutant HTT gene, reduce or inhibit the production of the mutant HTT protein, and / or reduce or prevent one or more symptoms of HD in the individual in need, thereby providing therapeutic treatment for HD.

[0279] In some embodiments, the treatment comprises in vivo gene editing by administering the RGN system disclosed herein, the polynucleotide encoding it, the RNP complex disclosed herein, or the vector disclosed herein. In some embodiments, the treatment comprises in vitro gene editing of zygotes, embryonic cells, or gametes to correct genetic errors early in life. In some embodiments, a therapeutic composition comprising the RGN system, the polynucleotide encoding it, the RNP complex, or the vector disclosed herein is targeted in vivo to cells of an individual. In some embodiments, the cells targeted for gene editing of the mutated HTT allele include stem cells, neurons (e.g., intermediate-sized polyspinous neurons, cortical neurons), and glial cells (e.g., astrocytes, oligodendrocytes, and microglia).

[0280] Huntington's disease (HD) is a single-gene, fatal neurodegenerative disease characterized by progressive chorea (involuntary movements), neuropsychiatric dysfunction, and cognitive impairment. Symptoms typically appear between the ages of 35 and 44, and the life expectancy after onset is 10 to 25 years.

[0281] HD is known to be caused by the amplification of the triplet cytosine-adenine-guanine (CAG) repeat sequence at the end of exon 1 of the huntingtin (HTT) gene. The CAG repeat sequence encodes polyglutamine at the N-terminus of the huntingtin (HTT) protein. Normal HTT alleles contain 15-20 CAG repeat sequences, while alleles containing 27-35 CAG repeat sequences are considered intermediate alleles with a very low probability of exhibiting the disease phenotype. HTT alleles containing 35 or more CAG repeat sequences are considered alleles that may cause HD and carry a risk of developing the disease. Alleles containing 36-39 CAG repeat sequences are considered incomplete penetrance, and individuals with those alleles may or may not develop the disease (or may develop symptoms later in life), while alleles containing 40 or more CAG repeat sequences are considered complete penetrance. An integrated classification system for Huntington's disease (HD-ISS) has been developed, which defines disease stages from birth to death and considers clinical criteria, clinical biomarkers and functional assessments to classify individuals with HD (Tabrizi et al., Lancet Neurology, 2022; 21:632-644).

[0282] Juvenile episodic HD (JHD) is a form of HD that affects children and adolescents. Individuals with juvenile episodic HD (<21 years of age) are typically found to have 60 or more CAG repeat sequences. JHD symptoms include changes in personality, coordination, behavior, speech, or cognitive abilities. Physical changes also occur and include rigidity, leg stiffness, clumsiness, bradykinesia, tremors, or myoclonus. Seizures and rigidity are more common than in adult HD, while chorea is less common. JHD has a faster rate of progression than adult HD and can result in death within 10 years of onset.

[0283] The mutated HTT allele is usually inherited as a dominant trait from one parent. If the other parent does not have the disease, any child born to a parent with HD has a 50% chance of developing the disease. In some cases, parents may have an intermediate HD allele and be asymptomatic, while the child may develop the disease due to amplification of the repetitive sequence. Additionally, the HD allele can exhibit a phenomenon known as prediction, where increased severity or a decreased age of onset is observed over several generations due to instability of the repetitive region during spermatogenesis.

[0284] Amplification of this repetitive sequence produces a mutant HTT protein, which can form aggregates in cells, interfering with normal cellular function and / or interacting pathologically with other molecules. Ultimately, the presence of the mutant HTT protein causes striatal neurodegeneration, which progresses to widespread brain atrophy.

[0285] Trinucleotide amplification in the HTT gene leads to neuronal loss of GABA-projecting neurons in the striatum, with neuronal loss also occurring in the neocortex. In some implementations, medium-sized polyspinous neurons (MSNs) containing enkephalins and projecting to the outer globus pallidus and / or MSNs containing substance P and projecting to the inner globus pallidus are affected. Other brain regions significantly affected in individuals with HD include the substantia nigra, cortices 3, 5, and 6, CA1 area of ​​the hippocampus, angular gyri in the parietal lobe, Purkinje cells in the cerebellum, lateral tuberculous nuclei in the hypothalamus, and the centromedial parafascicular complex in the thalamus (Walker (2007) Lancet 369:218-228). Currently, there is no curative treatment for HD, but experimental approaches based on drugs, cell therapy, and gene therapy are under investigation.

[0286] The RGN system, a polynucleotide encoding a component of the RGN system, an RNP complex, a vector, or a composition comprising any of these can be used to modify the mutant HTT allele in vivo in cells of patients with HD. In some embodiments, modifying the mutant HTT allele includes modifying a target sequence in exon 50 of the HTT gene. For example, APG07433.1 RGN is used with a suitable guide RNA selected from SEQ ID NO: 27 and 28 to modify the target sequence in exon 50 of the mutant HTT allele. As another example, APG05586 RGN is used with a suitable guide RNA selected from SEQ ID NO: 25 and 26 to modify the target sequence in exon 50 of the mutant HTT allele. As yet another example, APG01604 RGN is used with a suitable guide RNA having the nucleotide sequence shown as SEQ ID NO: 29 to modify the target sequence in exon 50 of the mutant HTT allele.

[0287] Modifying the target sequence in the mutant HTT allele involves cleaving the mutant HTT allele in vivo in cells of patients with HD. In some embodiments, cleavage occurs in exon 50 of the mutant HTT allele. In some embodiments, after cleavage of the target sequence, non-homologous end joining (NHEJ) occurs, causing nucleotide insertions and / or deletions (indels) at the cleavage site and disrupting the mutant HTT coding sequence, resulting in reduced levels of mutant HTT mRNA and / or protein. In some embodiments, cleavage of the mutHTT allele causes the introduction of a premature stop codon, resulting in reduced levels of mutHTT protein.

[0288] As used herein, the term "individual" refers to any individual requiring diagnosis, treatment, or therapy. In some embodiments, the individual is an animal. In some embodiments, the individual is a mammal. In some embodiments, the individual is a human.

[0289] The methods for treating HD in this disclosure utilize SNPs present in the human genome. SNP alleles that generate PAM sites recognized by the RGN system and RNP complex of this disclosure have been identified. When a mutant HTT allele contains an SNP allele that generates a PAM recognized by the RGN system of this disclosure, these RGN systems can be used to cleave the mutant HTT allele, resulting in a reduction in the levels of mutant HTT mRNA and / or protein. In some embodiments, the mutant HTT allele comprises a PAM containing an SNP allele. Therefore, individuals in need suitable for treatment with the described compositions (e.g., RGN or a nucleic acid molecule encoding RGN, and guide RNA or a nucleic acid molecule encoding guide RNA; the RGN system; or the RNP complex) can be analyzed prior to treatment to determine whether the mutant HTT allele contains an SNP allele that generates a PAM recognized by the RGN system described herein. In some embodiments, individuals in need suitable for treatment with the described RGN system further possess a wild-type HTT allele that does not contain a PAM-generating SNP allele and is therefore heterozygous for the SNP. It is desirable to have a wild-type HTT allele that does not contain a SNP allele that generates PAM, such that the applied composition cleaves only the mutant HTT allele containing the SNP allele that generates PAM and does not cleave the wild-type HTT allele. Therefore, in some embodiments, PAM is present only on the mutant HTT allele and not on the wild-type HTT allele, and the treatment is considered allele-specific. The method may include using an RGN system that cannot cleave the wild-type HTT allele, a polynucleotide encoding a component of the RGN system, an RNP complex, a vector, or a composition comprising any of these. In some implementations, the RGN system that cannot cleave the wild-type HTT allele, the polynucleotide encoding a component of the RGN system, the RNP complex, the vector, or a composition containing any of these cannot cleave the wild-type HTT allele, or cleaves it to a negligible degree such that the levels of wtHTT mRNA and / or wtHTT protein are not significantly reduced, wherein, for example, wtHTT can maintain support for key cellular and neural functions and / or the absence of Huntington's disease symptoms in the in vivo context (i.e., in individuals who are heterozygous for the mutHTT allele and are administered the RGN system).

[0290] Determining whether an individual possesses a mutant HTT allele containing an SNP allele that generates a PAM recognized by the RGN system described herein, and / or whether the SNP is heterozygous, can be achieved using sequencing, array-based hybridization, and / or PCR-based methods (e.g., long-range PCR) on a biological sample obtained from the individual. In some embodiments, the biological sample may include blood, cells, and / or cerebrospinal fluid. In some embodiments, the mutant HTT allele comprises a PAM containing an SNP allele.

[0291] The composition administered to an individual in need comprises an RGN that recognizes a PAM generated by an SNP allele (i.e., containing, for example, exon 50 of the mutHTT allele), which may include the NNNCC, NNRYA, NNGRR, and NNGG PAM sequences. The SNP generating the PAM may be located within exon 50 of the mutated HTT allele. In some embodiments, the SNP is a thymine corresponding to position 151 of SEQ ID NO: 1, generating the PAM sequence NNRYA. In some embodiments, the NNRYA PAM sequence contains an SNP that is a thymine corresponding to position 151 of SEQ ID NO: 1. In some embodiments, the NNRYA PAM sequence is recognized by an RGN comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the following sequence: SEQ ID NO: 7. In some embodiments, the NNRYA PAM sequence is recognized by an RGN comprising the amino acid sequence of SEQ ID NO: 7. In some embodiments, the NNRYA PAM sequence is recognized by an RGN binding to a guide RNA comprising a crRNA repeat sequence having a nucleotide sequence having SEQ ID NO: 8 or 106 or a nucleotide sequence differing from SEQ ID NO: 8 or 106 by 1 or 2 nucleotides. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA comprising a tracrRNA having a nucleotide sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity: SEQ ID NO: 9 or 107. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107. In some embodiments, the NNRYA PAM sequence is recognized by an RGN that binds to a guide RNA comprising a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 75 or 76. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 80 or 81 or a nucleotide sequence differing from SEQ ID NO: 80 or 81 by 1 or 2 nucleotides. In some implementations, the guide RNA is a single guide RNA.In some implementations, the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

[0292] The SNP that generates PAM can be located within exon 50 of the mutant HTT allele. In some embodiments, the SNP is a cytosine at position 151 of SEQ ID NO: 2, generating the PAM sequences NNNNCC, NNGRR, and NNGG. In some embodiments, the PAM sequences NNNNCC, NNGRR, and NNGG contain an SNP that is a cytosine at position 151 of SEQ ID NO: 2. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN containing an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the following: SEQ ID NO: 3. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN containing the amino acid sequence of SEQ ID NO: 3. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 or a nucleotide sequence differing from SEQ ID NO: 4 by 1 or 2 nucleotides. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA comprising a tracrRNA having a nucleotide sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity: SEQ ID NO: 5. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 5. In some embodiments, the NNNNCC PAM sequence is recognized by an RGN that binds to a guide RNA comprising a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 77 or 78. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 82 or 83, or a nucleotide sequence differing from SEQ ID NO: 82 or 83 by 1 or 2 nucleotides. In some embodiments, the guide RNA is a single guide RNA. In some embodiments, the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

[0293] In some embodiments, the NNGRR PAM sequence is recognized by an RGN comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the following sequence: SEQ ID NO: 11. In some embodiments, the NNGRR PAM sequence is recognized by an RGN comprising the amino acid sequence of SEQ ID NO: 11. In some embodiments, the NNGRR PAM sequence is recognized by an RGN binding to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 12 or a nucleotide sequence differing from SEQ ID NO: 12 by 1 or 2 nucleotides. In some embodiments, the NNGRR PAM sequence is recognized by an RGN bound to a guide RNA comprising a tracrRNA having a nucleotide sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the following: SEQ ID NO: 13 or 120. In some embodiments, the NNGRR PAM sequence is recognized by an RGN bound to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 12 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 13 or 120. In some embodiments, the NNGRR PAM sequence is recognized by an RGN bound to a guide RNA comprising a spacer having a nucleotide sequence complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 79. In some embodiments, the spacer has the nucleotide sequence of SEQ ID NO: 84 or a nucleotide sequence differing from SEQ ID NO: 84 by 1 or 2 nucleotides. In some embodiments, the guide RNA is a single guide RNA. In some implementations, the single guide RNA has the nucleotide sequence of SEQ ID NO: 29.

[0294] In some embodiments, the NNGG PAM sequence is recognized by an RGN comprising an amino acid sequence having at least 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity with the following sequence: SEQ ID NO: 15. In some embodiments, the NNGG PAM sequence is recognized by an RGN comprising the amino acid sequence of SEQ ID NO: 15. In some embodiments, the NNGG PAM sequence is recognized by an RGN that binds to a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 16 or a nucleotide sequence differing from SEQ ID NO: 16 by 1 or 2 nucleotides. In some embodiments, the NNGG PAM sequence is recognized by an RGN that binds to a guide RNA, the guide RNA comprising a tracrRNA having a nucleotide sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or greater sequence identity: SEQ ID NO: 17. In some embodiments, the NNGGPAM sequence is recognized by an RGN that binds to a guide RNA, the guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 16 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 17. In some embodiments, the guide RNA is a single guide RNA.

[0295] As used herein, the terms "treatment" or "treating" refer to a method for obtaining a beneficial or desired outcome, including but not limited to therapeutic and / or preventative benefits. A therapeutic benefit is defined as any treatment-related improvement or effect against one or more diseases, symptoms, or conditions being treated. For preventative benefits, the composition may be administered to an individual at risk of developing a specific disease, symptom, or condition, or to an individual reporting one or more of the physiological symptoms or disease biomarkers (e.g., altered neuronal cell phenotypes) of a disease, even if the disease, symptom, or condition may not yet be present. In some embodiments, an individual at risk of developing HD is an individual containing a mutant HTT allele as defined herein. In some embodiments, an individual at risk of developing HD is defined as an individual having a mutant HTT allele containing more than 35 CAG repeat sequences. In some embodiments, an individual at risk of developing HD is defined as an individual having a mutant HTT allele containing at least 40 CAG repeat sequences. In some embodiments, an individual at risk of developing HD is defined as an individual having a mutant HTT allele containing more than 56 CAG repeat sequences. In some embodiments, treatment may be administered after one or more symptoms have appeared and / or after the disease has been diagnosed. In some embodiments, treatment may be administered in the absence of symptoms, for example, to prevent or delay the onset of symptoms or to suppress the onset or progression of the disease. For example, treatment may be administered to susceptible individuals before the onset of symptoms (e.g., based on a history of symptoms and / or based on genetic or other susceptibility factors). Treatment may also continue after symptoms have subsided, for example, to prevent or postpone their prevention or recurrence. In some embodiments, the number of CAG repeat sequences influences the decision of when to treat an individual. In some embodiments, individuals with a mutated HTT allele containing more than 56 CAG repeat sequences are treated as adolescents or young adults before any HD symptoms onset or the presentation of physiological biomarkers.

[0296] In some embodiments, the compositions and methods disclosed in this invention are used to improve (i.e., reduce) or delay the onset of one or more symptoms of Huntington's disease in an individual in need. Symptoms may be reduced by approximately 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, and 100%, or by at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, 30-100%, 4 0-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100% or 95-100%. Compared to control values ​​or control individuals (e.g., individuals who have not yet received treatment), the onset of one or more symptoms of Huntington's disease can be delayed for approximately 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 12 years, 15 years, 20 years, 25 years, 30 years, or longer. In some embodiments, the compositions and methods disclosed in this invention are capable of preventing the occurrence of one or more symptoms of Huntington's disease in an individual.

[0297] In some embodiments, the compositions and methods disclosed in this invention are used to improve (i.e., reduce) or delay the onset of one or more biomarkers of Huntington's disease in individuals in need. Compared to control values ​​or control individuals (e.g., individuals not yet receiving treatment), biomarkers may be reduced by about 20%, 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, and 100%, or by at least 20-30%, 20-40%, 20-50%, 20-60%, 20-70%, 20-80%, 20-90%, 20-95%, 20-100%, 30-40%, 30-50%, 30-60%, 30-70%, 30-80%, 30-90%, 30-95%, and 30-100%. 40-50%, 40-60%, 40-70%, 40-80%, 40-90%, 40-95%, 40-100%, 50-60%, 50-70%, 50-80%, 50-90%, 50-95%, 50-100%, 60-70%, 60-80%, 60-90%, 60-95%, 60-100%, 70-80%, 70-90%, 70-95%, 70-100%, 80-90%, 80-95%, 80-100%, 90-95%, 90-100% or 95-100%. Compared to control values ​​or control individuals (e.g., individuals who have not yet received treatment), the onset of one or more biomarkers of Huntington's disease can be delayed for approximately 6 months, 1 year, 2 years, 3 years, 4 years, 5 years, 6 years, 7 years, 8 years, 9 years, 10 years, 12 years, 15 years, 20 years, 25 years, 30 years, or longer. In some embodiments, the compositions and methods disclosed in this invention are capable of preventing the development of one or more biomarkers of Huntington's disease in an individual.

[0298] Individuals at risk of developing HD or who have HD can be identified in various ways, including cognitive assessments and / or neurological or neuropsychiatric examinations, motor tests, sensory tests, psychiatric assessments, brain imaging, family history, and / or genetic testing. Individuals may have symptoms of HD and / or be diagnosed with HD without symptoms.

[0299] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP complex; or a vector) can be applied to individuals classified using the Huntington's Disease Integrated Classification System (HD-ISS) (Tabrizi et al., *Lancet Neurol*, 2022; 21:632–644, the contents of which are incorporated herein by reference in their entirety). In this system, stage 0 HD individuals have ≥40 CAG repeat sequences. Stage 1 HD individuals have ≥40 CAG repeat sequences and pathogenesis biomarkers (e.g., putamen volume and / or caudate nucleus volume). Stage 2 HD individuals have ≥40 CAG repeat sequences, pathogenesis biomarkers, and signs or symptoms (e.g., as measured by total motor score (TMS) and / or symbolic digit modality test (SDMT)). Stage 3 HD individuals have ≥40 CAG repeat sequences, pathogenesis biomarkers, signs or symptoms, and functional changes (e.g., as measured by independent scales and / or total functional capacity (TFC)).

[0300] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP; or a vector) can be applied to individuals diagnosed using the Huntington's Disease Prognostic Index or a derivative thereof (Long JD et al., Movement Disorders, 2017, 32(2), 256-263, the contents of which are incorporated herein by reference in their entirety). The prognostic index uses four components to predict the probability of a motor diagnosis: (1) the total motor score (TMS) from the Unified Huntington's Disease Rating Scale (UHDRS); (2) the Symbolic Digit Modality Test (SDMT); (3) baseline age; and (4) CAG amplification. In some embodiments, the prognostic index for HD is calculated using the following formula: PI HD =51×TMS+(−34)×SDMT+7×age×(CAG-34), where PI HD Larger values ​​indicate a greater risk of diagnosis or symptom onset. In some implementations, the prognostic index for HD is calculated using the following normalized formula, which gives the standard deviation in units to be interpreted in the context of a 50% 10-year survival rate: PIN HD =(PI HD -883) / 1044, where PIN HD <0 indicates a 10-year survival rate greater than 50%, and the PIN... HDA value >0 indicates a 10-year survival rate of less than 50%. In some embodiments, the prognostic index can be used to identify individuals who will develop HD symptoms within several years but do not yet have clinically diagnosable symptoms. Additionally, these asymptomatic patients can be selected and receive treatment using the compositions of this disclosure during the asymptomatic period.

[0301] The compositions disclosed herein (e.g., comprising RGN or a nucleic acid molecule encoding RGN, and guide RNA or a nucleic acid molecule encoding guide RNA; an RGN system; an RNP complex; or a vector) can be applied to individuals who have undergone biomarker evaluation. Potential blood biomarkers for HD include, but are not limited to, 8-hydroxy-2-deoxyguanosine (8-OhdG) oxidative stress markers, metabolic markers (e.g., creatine kinase, branched-chain amino acids), cholesterol metabolites (e.g., 24-OH cholesterol), immune and inflammatory proteins (e.g., clusteringins, complement components, interleukins 6 and 8), gene expression changes (e.g., transcriptome markers), endocrine markers (e.g., cortisol, gastric growth hormone-releasing hormone, and leptin), brain-derived neurotrophic factor (BDNF), and adenosine 2A receptor. Potential brain imaging biomarkers for HD include, but are not limited to, striatal volume, putamen volume, caudate nucleus volume, subcortical white matter volume, cortical thickness, whole brain volume, and ventricular volume. Brain imaging can be performed using functional imaging (e.g., functional MRI), positron emission tomography (PET) (e.g., with fluorodeoxyglucose), and magnetic resonance spectroscopy (e.g., lactate). Potential biomarkers for quantitative clinical tools for HD include, but are not limited to, quantitative motor assessments, motor physiology assessments (e.g., transcranial magnetic stimulation), and quantitative eye movement measurements. Non-limiting examples of quantitative clinical biomarker assessments include tongue force variability, metronome-guided tapping, gripping force, eye movement assessments, and cognitive tests. Non-limiting examples of multicenter observational studies include PREDICT-HD and TRACK-HD. In some embodiments, the biomarker for HD is the level of wild-type huntingtin (HTT) mRNA and / or protein. In some embodiments, the biomarker for HD is the level of mutant huntingtin (HTT) mRNA and / or protein. In some embodiments, the biomarker for HD is the level of neurofilament light chain (NFL) protein.

[0302] The compositions disclosed herein (e.g., comprising RGN or a nucleic acid molecule encoding RGN, and guide RNA or a nucleic acid molecule encoding guide RNA; an RGN system; an RNP complex; or a vector) can be administered to individuals without HD symptoms. The individual may be asymptomatic but may have undergone predictive genetic testing or biomarker assessment to determine their risk of HD and / or may have family members diagnosed with HD (e.g., mother, father, brother, sister, aunt, uncle, grandparents). In some embodiments, the asymptomatic individual has a mutated HTT allele comprising 27-35 CAG repeat sequences (e.g., CAG repeat sequences 27, 28, 29, 30, 31, 32, 33, 34, and 35).

[0303] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP complex; or a vector) can be applied to individuals in the early stages of HD. In the early stages, individuals may exhibit subtle changes in coordination, some chorea, mood changes such as irritability and depression, problem-solving difficulties, and / or a reduced ability to function in their normal daily lives.

[0304] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP complex; or a vector) can be applied to individuals in the intermediate stage of HD. In the intermediate stage, individuals experience increased motor impairment, diminished speech, dysphagia, and greater difficulty performing normal activities. At this stage, individuals may have occupational and physical therapists to help maintain control of voluntary movement, and individuals may have a speech-language pathologist.

[0305] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP complex; or a vector) can be administered to individuals in the late stages of HD. In the late stages, individuals with HD are almost entirely or completely dependent on caregivers because they are no longer able to walk or speak. Individuals can generally still understand language and recognize family and friends, but choking is a major problem.

[0306] The compositions disclosed herein (e.g., comprising an RGN or a nucleic acid molecule encoding an RGN, and a guide RNA or a nucleic acid molecule encoding a guide RNA; an RGN system; an RNP complex; or a vector) can be administered to individuals with juvenile HD (HD onset before age 21) or who are susceptible to juvenile HD, such as those individuals having at least 56 CAG repeat sequences (e.g., 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75 or more CAG repeat sequences) in exon 1 of the HTT gene. In some of these embodiments, the compositions disclosed herein are administered to the individual before age 21, or in some embodiments before age 18.

[0307] The compositions disclosed herein (e.g., comprising RGN or a nucleic acid molecule encoding RGN, and guide RNA or a nucleic acid molecule encoding guide RNA; RGN system; RNP complex; or vector) can be administered to individuals with full penetrance HD, wherein the mutant HTT allele has more than 40 CAG repeat sequences in exon 1 (e.g., 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90 or more CAG repeat sequences). In some embodiments, individuals in need who can be treated with the compositions of this disclosure have a mutated HTT allele containing at least 36 CAG repeat sequences (e.g., 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more than 50 CAG repeat sequences) in exon 1.

[0308] The compositions disclosed herein (e.g., comprising RGN or a nucleic acid molecule encoding RGN, and guide RNA or a nucleic acid molecule encoding guide RNA; RGN system; RNP complex; or vector) can be administered to individuals with HD having incomplete penetrance, wherein the mutant HTT allele has between 36 and 39 CAG repeat sequences (e.g., 36, 37, 38, and 39 CAG repeat sequences).

[0309] In some implementations, the control reference or healthy individual has no more than 15-20 HTT alleles with CAG trinucleotide repeat sequences. In some implementations, the control reference or healthy individual has no more than 26 HTT alleles with CAG trinucleotide repeat sequences.

[0310] The term "effective amount" or "therapeutic effective amount" is the amount of a pharmaceutical agent sufficient to achieve a beneficial or desired clinical outcome. Therapeutic effective amounts can vary depending on one or more of the following: the individual being treated and the condition of the disease, the individual's weight and age, the severity of the disease, the method of administration, and similar factors, which can be readily determined by a person skilled in the art. Specific dosages can vary depending on one or more of the following: the particular pharmaceutical agent selected, the dosing regimen followed, whether it is administered in combination with other compounds, the timing of administration, and the delivery system used to deliver it.

[0311] Treatment effic...

Claims

1. An RNA-guided nuclease (RGN) system comprising: a) a guide RNA comprising a spacer and a backbone, wherein the spacer has the nucleotide sequence of SEQ ID NO: 80 or 81 or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 or 2 nucleotides, or a nucleic acid molecule encoding the guide RNA; and b) an RGN polypeptide having an amino acid sequence that has at least 90% sequence identity to SEQ ID NO: 7, or a nucleic acid molecule encoding the RGN polypeptide.

2. The RGN system of claim 1, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 nucleotide.

3. The RGN system of claim 1, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

4. The RGN system of any one of claims 1 to 3, wherein the guide RNA binds to a target sequence in a mutant huntingtin (mutHTT) allele.

5. The RGN system of claim 4, wherein the RGN system is capable of binding and cleaving a target sequence in the mutHTT allele, and wherein the guide RNA is capable of forming a complex with the RGN polypeptide and directing the complex to the target sequence for binding and cleavage.

6. The RGN system of claim 4 or 5, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 75 or 76.

7. The RGN system of any one of claims 1 to 6, wherein the RGN system is capable of recognizing a protospacer adjacent motif (PAM) having the sequence of NNRYA resulting from a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

8. The RGN system of any one of claims 1 to 7, wherein the RGN polypeptide comprises a PAM interaction domain that binds a protospacer adjacent motif (PAM) having the sequence of NNRYA resulting from a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

9. The RGN system of claim 8, wherein the PAM interaction domain comprises an amino acid sequence that has at least 95% sequence identity to SEQ ID NO:

134.

10. The RGN system of claim 8 or 9, wherein the PAM interaction domain comprises the amino acid sequence set forth as SEQ ID NO:

134.

11. The RGN system of any one of claims 1 to 10, wherein the RGN polypeptide comprises at least one nuclease domain comprising an amino acid sequence that has at least 95% sequence identity to any one of SEQ ID NOs: 147, 148, 149, and 150.

12. The RGN system of any one of claims 1 to 11, wherein the RGN system is incapable of cleaving a wild-type HTT allele.

13. The RGN system of any one of claims 1 to 12, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity to SEQ ID NO:

7.

14. The RGN system of any one of claims 1 to 13, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

7.

15. The RGN system of any one of claims 1 to 14, wherein the RGN polypeptide further comprises at least one nuclear localization signal.

16. The RGN system of claim 15, wherein the at least one nuclear localization signal comprises an SV40 nuclear localization signal.

17. The RGN system of claim 16, wherein the SV40 nuclear localization signal has the sequence set forth as SEQ ID NO:

86.

18. The RGN system of claim 15, wherein the at least one nuclear localization signal comprises a c-Myc nuclear localization signal.

19. The RGN system of claim 18, wherein the c-Myc nuclear localization signal has the sequence set forth as SEQ ID NO:

125.

20. The RGN system of any one of claims 15 to 19, wherein a NLS linker protein links the RGN polypeptide and the at least one nuclear localization signal.

21. The RGN system of claim 20, wherein the NLS linker protein has the sequence set forth as SEQ ID NO:

127.

22. The RGN system of any one of claims 1 to 21, wherein the backbone of the guide RNA is 66 to 90 nucleotides in length.

23. The RGN system of any one of claims 1 to 22, wherein the backbone of the guide RNA comprises a nucleotide sequence having at least 80% sequence identity to SEQ ID NO: 140 or 141.

24. The RGN system of any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106 or a nucleotide sequence differing by 1 or 2 nucleotides from SEQ ID NO: 8 or 106 and a tracrRNA having a nucleotide sequence having at least 90% sequence identity to SEQ ID NO: 9 or 107.

25. The RGN system of any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat sequence having a nucleotide sequence differing by 1 nucleotide from SEQ ID NO: 8 or 106 and a tracrRNA having a nucleotide sequence having at least 95% sequence identity to SEQ ID NO: 9 or 107.

26. The RGN system of any one of claims 1 to 23, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

27. The RGN system of any one of claims 1 to 26, wherein the guide RNA is a single guide RNA.

28. The RGN system of claim 27, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

29. A ribonucleoprotein (RNP) complex comprising the RGN polypeptide and the guide RNA of the RGN system of any one of claims 1 to 28.

30. A nucleic acid molecule comprising or encoding a guide RNA comprising a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81 or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 or 2 nucleotides.

31. The nucleic acid molecule of claim 30, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 nucleotide.

32. The nucleic acid molecule of claim 30, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

33. The nucleic acid molecule of any one of claims 30 to 32, wherein the guide RNA binds to a target sequence in a mutant huntingtin (mutHTT) allele.

34. The nucleic acid molecule of claim 33, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 75 or 76.

35. The nucleic acid molecule of any one of claims 30 to 34, wherein the guide RNA binds to a RNA-guided nuclease (RGN) polypeptide having an amino acid sequence that has at least 90% sequence identity to SEQ ID NO:

7.

36. A nucleic acid molecule comprising or encoding a guide RNA that binds to a target sequence in a mutant huntingtin (mutHTT) allele, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 75 or 76 and binds to a RNA-guided nuclease (RGN) polypeptide having an amino acid sequence that has at least 90% sequence identity to SEQ ID NO:

7.

37. The nucleic acid molecule of claim 36, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81 or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 or 2 nucleotides.

38. The nucleic acid molecule of claim 37, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 nucleotide.

39. The nucleic acid molecule of claim 37, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 80 or 81.

40. The nucleic acid molecule of any one of claims 35 to 39, wherein the RGN polypeptide has an amino acid sequence that has at least 95% sequence identity to SEQ ID NO:

7.

41. The nucleic acid molecule of any one of claims 35 to 40, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

7.

42. The nucleic acid molecule of any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106, or a nucleotide sequence that differs from SEQ ID NO: 8 or 106 by 1 or 2 nucleotides, and a tracrRNA having a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 9 or 107.

43. The nucleic acid molecule of any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat sequence having a nucleotide sequence that differs from SEQ ID NO: 8 or 106 by 1 nucleotide, and a tracrRNA having a nucleotide sequence with at least 95% sequence identity to SEQ ID NO: 9 or 107.

44. The nucleic acid molecule of any one of claims 30 to 41, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106, and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

45. The nucleic acid molecule of any one of claims 30 to 44, wherein the guide RNA is a single guide RNA.

46. The nucleic acid molecule of claim 45, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

47. The nucleic acid molecule of any one of claims 30 to 46, wherein the nucleic acid molecule encoding the guide RNA is operably linked to an RNA polymerase III promoter.

48. The nucleic acid molecule of claim 47, wherein the RNA polymerase III promoter is a U6 promoter.

49. The nucleic acid molecule of claim 48, wherein the U6 promoter is a truncated U6 promoter.

50. The nucleic acid molecule of claim 49, wherein the truncated U6 promoter has the nucleotide sequence set forth as SEQ ID NO: 89 or 128.

51. A vector comprising the nucleic acid molecule of any one of claims 30 to 35, wherein the nucleic acid molecule encodes the guide RNA.

52. A vector comprising the nucleic acid molecule of any one of claims 36 to 50, wherein the nucleic acid molecule encodes the guide RNA.

53. The vector of claim 51 or 52, wherein the vector is a viral vector.

54. The vector of claim 53, wherein the viral vector is a lentiviral vector, a baculoviral vector, or an adeno-associated viral (AAV) vector.

55. The vector of claim 54, wherein the viral vector is an AAV vector and comprises an AAV inverted terminal repeat sequence.

56. The vector of claim 55, wherein the AAV inverted terminal repeat sequence is an AAV2, AAV5, or AAV6 inverted terminal repeat sequence.

57. The vector of any one of claims 52 to 56, wherein the vector further comprises a nucleic acid molecule encoding the RGN polypeptide.

58. The vector of claim 57, wherein the vector further comprises a RNA polymerase II promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide.

59. The vector of claim 58, wherein the RNA polymerase II promoter is a constitutive promoter.

60. The vector of claim 59, wherein the constitutive promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, a truncated CMV promoter, an elongation factor 1 alpha short (EFS) promoter, and a JeT promoter.

61. The vector of claim 60, wherein the constitutive promoter is a JeT promoter.

62. The vector of claim 61, wherein the JeT promoter has the nucleotide sequence set forth as SEQ ID NO:

92.

63. The vector of claim 58, wherein the RNA polymerase II promoter is a tissue-specific promoter.

64. The vector of claim 63, wherein the tissue-specific promoter is a brain or neuron-specific promoter.

65. The vector of claim 64, wherein the brain or neuron-specific promoter is selected from the group consisting of a human synapsin I (Syn) promoter, a 67 kDa glutamate decarboxylase (GAD67) promoter, a 65 kDa glutamate decarboxylase (GAD65) promoter, a homeobox Dlx5 / 6 promoter, a pre-pro-tachykinin 1 (Tac1) promoter, a neuron-specific enolase (NSE) promoter, a dopamine receptor 1 (Drd1a) promoter, a dopamine receptor 2 (DRD2) promoter, and a glial fibrillary acidic protein (GFAP) promoter.

66. The vector of claim 65, wherein the neuron-specific promoter is a Syn promoter.

67. The vector of claim 66, wherein the Syn promoter has the nucleotide sequence set forth as SEQ ID NO:

93.

68. The vector of any one of claims 57-67, wherein the nucleic acid molecule encoding the RGN polypeptide comprises a polyadenylation (polyA) tail.

69. The vector of claim 68, wherein the polyA tail is an SV40 polyA tail or a bovine growth hormone (bGH) polyA tail.

70. The vector of claim 69, wherein the SV40 polyA tail has the nucleotide sequence set forth as SEQ ID NO:

94.

71. The vector of claim 69, wherein the bGH polyA tail has the sequence set forth as SEQ ID NO:

95.

72. The vector of claim 57 or 58, wherein the vector comprises: a truncated U6 promoter operably linked to the nucleic acid molecule encoding the guide RNA; a CMVeb promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide; a c-Myc NLS at the N-terminus or C-terminus of the RGN polypeptide; an NLS linker protein linking the c-Myc NLS and the RGN polypeptide; and an SV40 polyA tail.

73. The vector of claim 72, wherein the truncated U6 promoter has the sequence set forth as SEQ ID NO: 128, the CMVeb promoter has the sequence set forth as SEQ ID NO: 90, the c-Myc NLS has the sequence set forth as SEQ ID NO: 125, the NLS linker protein has the sequence set forth as SEQ ID NO: 127, and the SV40 polyA tail has the sequence set forth as SEQ ID NO:

94.

74. The vector of claim 72 or 73, wherein the sgRNA has the sequence set forth as SEQ ID NO: 26 and the nucleic acid molecule encoding the RGN polypeptide has the sequence set forth as SEQ ID NO:

88.

75. The vector of any one of claims 72 to 74, wherein the vector comprises the sequence set forth as SEQ ID NO:

123.

76. The vector of any one of claims 57 to 75, wherein the RGN polypeptide is operably linked to at least one nuclear localization signal.

77. The vector of claim 76, wherein the at least one nuclear localization signal comprises an SV40 nuclear localization signal.

78. The vector of claim 77, wherein the SV40 nuclear localization signal has the sequence set forth as SEQ ID NO:

86.

79. The vector of claim 76, wherein the at least one nuclear localization signal comprises a c-Myc nuclear localization signal.

80. The vector of claim 79, wherein the c-Myc nuclear localization signal has the sequence set forth as SEQ ID NO:

125.

81. The vector of any one of claims 76 to 80, wherein a NLS linker protein links the RGN polypeptide and the at least one nuclear localization signal.

82. The vector of claim 81, wherein the NLS linker protein has the sequence set forth as SEQ ID NO:

127.

83. The vector of any one of claims 57 to 82, wherein the vector has the sequence set forth as any one of SEQ ID NO: 32-39 or 121-123.

84. An RNA-guided nuclease (RGN) system, comprising: a) a guide RNA comprising a nucleotide sequence having SEQ ID NO: 82 or 83 or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 or 2 nucleotides, or a nucleic acid molecule encoding the guide RNA; and b) an RGN polypeptide having an amino acid sequence that has at least 90% sequence identity to SEQ ID NO: 3, or a nucleic acid molecule encoding the RGN polypeptide.

85. The RGN system of claim 84, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 nucleotide.

86. The RGN system of claim 84, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

87. The RGN system of any one of claims 84-86, wherein the guide RNA binds to a target sequence in a mutant huntingtin (mutHTT) allele.

88. The RGN system of claim 87, wherein the system is capable of binding and cleaving a target sequence in the mutHTT allele, and wherein the guide RNA is capable of forming a complex with the RGN polypeptide and directing the complex to the target sequence for binding and cleavage.

89. The RGN system of claim 87 or 88, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 77 or 78.

90. The RGN system of any one of claims 84-89, wherein the RGN system is capable of recognizing a protospacer adjacent motif (PAM) of the sequence NNNNCC resulting from a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

91. The RGN system of any one of claims 84-90, wherein the RGN comprises a PAM interaction domain that binds a protospacer adjacent motif (PAM) of the sequence NNNNCC resulting from a single nucleotide polymorphism (SNP) in exon 50 of the HTT gene.

92. The RGN system of claim 91, wherein the PAM interaction domain comprises an amino acid sequence having at least 95% sequence identity to SEQ ID NO:

133.

93. The RGN system of claim 91 or 92, wherein the PAM interaction domain comprises the amino acid sequence set forth as SEQ ID NO:

133.

94. The RGN system of any one of claims 84-93, wherein the RGN polypeptide comprises at least one nuclease domain comprising an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 143, 144, 145, and 146.

95. The RGN system of any one of claims 84-94, wherein the RGN system is incapable of cleaving a wild-type HTT allele.

96. The RGN system of any one of claims 84-95, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity to SEQ ID NO:

3.

97. The RGN system of any one of claims 84-96, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

3.

98. The RGN system of any one of claims 84-97, wherein the RGN polypeptide comprises at least one nuclear localization signal.

99. The RGN system of claim 98, wherein the at least one nuclear localization signal comprises an SV40 nuclear localization signal.

100. The RGN system of claim 99, wherein the SV40 nuclear localization signal has the sequence set forth as SEQ ID NO:

86.

101. The RGN system of claim 98, wherein the at least one nuclear localization signal comprises a c-Myc nuclear localization signal.

102. The RGN system of claim 101, wherein the c-Myc nuclear localization signal has the sequence set forth as SEQ ID NO:

125.

103. The RGN system of any one of claims 98-102, wherein a NLS linker protein links the RGN polypeptide and the at least one nuclear localization signal.

104. The RGN system of claim 103, wherein the NLS linker protein has the sequence set forth as SEQ ID NO:

127.

105. The RGN system of any one of claims 84-104, wherein the backbone of the guide RNA is 94-110 nucleotides in length.

106. The RGN system of any one of claims 84-105, wherein the backbone of the guide RNA comprises a nucleotide sequence having at least 80% sequence identity to SEQ ID NO:

142.

107. A ribonucleoprotein (RNP) complex comprising the RGN polypeptide and the guide RNA of the RGN system of any one of claims 84-106.

108. A nucleic acid molecule comprising or encoding a guide RNA comprising a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83 or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 or 2 nucleotides.

109. The nucleic acid molecule of claim 108, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 nucleotide.

110. The nucleic acid molecule of claim 108, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

111. The nucleic acid molecule of any one of claims 108-110, wherein the guide RNA binds to a target sequence in a mutant huntingtin (mutHTT) allele.

112. The nucleic acid molecule of claim 111, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 77 or 78.

113. The nucleic acid molecule of any one of claims 108-112, wherein the guide RNA binds to a RNA-guided nuclease (RGN) polypeptide having an amino acid sequence with at least 90% sequence identity to SEQ ID NO:

3.

114. A nucleic acid molecule comprising or encoding a guide RNA that binds to a target sequence in a mutant huntingtin (mutHTT) allele, wherein the target sequence has the nucleotide sequence of SEQ ID NO: 77 or 78 and binds to a RNA-guided nuclease (RGN) polypeptide having an amino acid sequence with at least 90% sequence identity to SEQ ID NO:

3.

115. The nucleic acid molecule of claim 114, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83 or a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 or 2 nucleotides.

116. The nucleic acid molecule of claim 115, wherein the guide RNA comprises a spacer having a nucleotide sequence that differs from SEQ ID NO: 82 or 83 by 1 nucleotide.

117. The nucleic acid molecule of claim 115, wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

118. The nucleic acid molecule of any one of claims 113-117, wherein the RGN polypeptide has an amino acid sequence having at least 95% sequence identity to SEQ ID NO:

3.

119. The nucleic acid molecule of any one of claims 113-118, wherein the RGN polypeptide has the amino acid sequence of SEQ ID NO:

3.

120. The nucleic acid molecule of any one of claims 108-119, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 or a nucleotide sequence that differs from SEQ ID NO: 4 by 1 or 2 nucleotides and a tracrRNA having a nucleotide sequence having at least 90% sequence identity to SEQ ID NO:

5.

121. The nucleic acid molecule of any one of claims 108-119, wherein the guide RNA comprises a crRNA repeat sequence having a nucleotide sequence that differs from SEQ ID NO: 4 by 1 nucleotide and a tracrRNA having a nucleotide sequence having at least 95% sequence identity to SEQ ID NO:

5.

122. The nucleic acid molecule of any one of claims 108-119, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 and a tracrRNA having the nucleotide sequence of SEQ ID NO:

5.

123. The nucleic acid molecule of any one of claims 108-122, wherein the guide RNA is a single guide RNA.

124. The nucleic acid molecule of claim 123, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

125. The nucleic acid molecule of any one of claims 108-124, wherein the nucleic acid molecule encoding the guide RNA is operably linked to an RNA polymerase III promoter.

126. The nucleic acid molecule of claim 125, wherein the RNA polymerase III promoter is a U6 promoter.

127. The nucleic acid molecule of claim 126, wherein the U6 promoter is a truncated U6 promoter.

128. The nucleic acid molecule of claim 127, wherein the truncated U6 promoter has the nucleotide sequence set forth as SEQ ID NO: 89 or 128.

129. A vector comprising the nucleic acid molecule of any one of claims 108-112, wherein the nucleic acid molecule encodes the guide RNA.

130. A vector comprising the nucleic acid molecule of any one of claims 113-128, wherein the nucleic acid molecule encodes the guide RNA.

131. The vector of claim 129 or 130, wherein the vector is a viral vector.

132. The vector of claim 131, wherein the viral vector is a lentiviral vector, a baculoviral vector, or an adeno-associated viral (AAV) vector.

133. The vector of claim 132, wherein the viral vector is an AAV vector and comprises an AAV inverted terminal repeat sequence.

134. The vector of claim 133, wherein the AAV inverted terminal repeat sequence is an AAV2, AAV5, or AAV6 inverted terminal repeat sequence.

135. The vector of any one of claims 130 to 134, wherein the vector further comprises a nucleic acid molecule encoding the RGN polypeptide.

136. The vector of claim 135, wherein the vector further comprises a RNA polymerase II promoter operably linked to the nucleic acid molecule encoding the RGN polypeptide.

137. The vector of claim 136, wherein the RNA polymerase II promoter is a constitutive promoter.

138. The vector of claim 137, wherein the constitutive promoter is selected from the group consisting of a cytomegalovirus (CMV) promoter, a truncated CMV promoter, an elongation factor 1 alpha short (EFS) promoter, and a JeT promoter.

139. The vector of claim 138, wherein the constitutive promoter is a JeT promoter.

140. The vector of claim 139, wherein the JeT promoter has the nucleotide sequence set forth as SEQ ID NO:

92.

141. The vector of claim 136, wherein the RNA polymerase II promoter is a tissue-specific promoter.

142. The vector of claim 141, wherein the tissue-specific promoter is a brain or neuron-specific promoter.

143. The vector of claim 142, wherein the brain or neuron-specific promoter is selected from the group consisting of a human synapsin I (Syn) promoter, a 67 kDa glutamate decarboxylase (GAD67) promoter, a 65 kDa glutamate decarboxylase (GAD65) promoter, a homeobox Dlx5 / 6 promoter, a pre-pro-tachykinin 1 (Tac1) promoter, a neuron-specific enolase (NSE) promoter, a dopamine receptor 1 (Drd1a) promoter, a dopamine receptor 2 (DRD2) promoter, and a glial fibrillary acidic protein (GFAP) promoter.

144. The vector of claim 143, wherein the neuron-specific promoter is a Syn promoter.

145. The vector of claim 144, wherein the Syn promoter has the nucleotide sequence set forth as SEQ ID NO:

93.

146. The vector of any one of claims 135 to 145, wherein the nucleic acid molecule encoding the RGN polypeptide comprises a polyadenylation (polyA) tail.

147. The vector of claim 146, wherein the polyA tail is an SV40 polyA tail or a bovine growth hormone (bGH) polyA tail.

148. The vector of claim 147, wherein the SV40 polyA tail has the nucleotide sequence set forth as SEQ ID NO:

94.

149. The vector of claim 147, wherein the bGH polyA tail has the sequence set forth as SEQ ID NO:

95.

150. The vector of any one of claims 135-149, wherein the RGN polypeptide is operably linked to at least one nuclear localization signal.

151. The vector of claim 150, wherein the at least one nuclear localization signal comprises an SV40 nuclear localization signal.

152. The vector of claim 151, wherein the SV40 nuclear localization signal has the sequence set forth as SEQ ID NO:

86.

153. The vector of claim 150, wherein the at least one nuclear localization signal comprises a c-Myc nuclear localization signal.

154. The vector of claim 153, wherein the c-Myc nuclear localization signal has the sequence set forth as SEQ ID NO:

125.

155. The vector of any one of claims 150-154, wherein a NLS linker protein links the RGN polypeptide and the at least one nuclear localization signal.

156. The vector of claim 155, wherein the NLS linker protein has the sequence set forth as SEQ ID NO:

127.

157. A cell comprising the nucleic acid molecule of any one of claims 30-50 and 108-128 or the vector of any one of claims 51-83 and 129-156.

158. A pharmaceutical composition comprising the nucleic acid molecule of any one of claims 30-50 and 108-128, the vector of any one of claims 51-83 and 129-156, the RGN system of any one of claims 1-28 and 84-106, or the RNP complex of claim 29 or 107.

159. The pharmaceutical composition of claim 158, having a purity of at least 95%.

160. The pharmaceutical composition of claim 158 or 159, having an undetectable amount of endotoxin or other impurities.

161. The pharmaceutical composition of any one of claims 158-160, further comprising poloxamer 188.

162. The pharmaceutical composition of any one of claims 158-161, which is a solution.

163. The pharmaceutical composition of any one of claims 158-161, which is lyophilized or freeze-dried.

164. A vector comprising the RGN system of any one of claims 1-28.

165. A vector comprising the RGN system of any one of claims 84-106.

166. The vector of claim 164 or 165, wherein the vector comprises an adeno-associated vector (AAV) inverted terminal repeat sequence.

167. The vector of claim 166, wherein the AAV inverted terminal repeat sequence is an AAV2, AAV5, or AAV6 inverted terminal repeat sequence.

168. The vector of claim 167, wherein the AAV inverted terminal repeat sequence is an AAV5 inverted terminal repeat sequence.

169. Use of a nucleic acid molecule of any one of claims 30 to 50 and 108 to 128, a vector of any one of claims 51 to 83, 129 to 156, and 164 to 168, a RGN system of any one of claims 1 to 28 and 84 to 106, or an RNP complex of claim 29 or 107 for reducing the amount of mutHTT mRNA and / or mutHTT protein in a cell.

170. Use of a nucleic acid molecule of any one of claims 30 to 50 and 108 to 128, a vector of any one of claims 51 to 83, 129 to 156, and 164 to 168, a RGN system of any one of claims 1 to 28 and 84 to 106, or an RNP complex of claim 29 or 107 for treating Huntington’s disease.

171. A method for cleaving a mutant Huntington (mutHTT) allele in a cell, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer-adjacent motif (PAM) comprising a nucleotide sequence of NNRYA comprises the first SNP allele, wherein the method comprises introducing into the cell: a) an RNA-guided nuclease (RGN) polypeptide having at least 90% sequence identity to SEQ ID NO: 7 or a nucleic acid molecule encoding the RGN polypeptide, and i) a nucleic acid molecule comprising or encoding a guide RNA of any one of claims 30 to 50; or ii) a vector of any one of claims 51 to 56; b) a vector of any one of claims 57 to 83, 164, and 166 to 168; c) a RGN system of any one of claims 1 to 28; or d) an RNP complex of claim 29.

172. A method for cleaving a mutant Huntington (mutHTT) allele in a cell, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer-adjacent motif (PAM) comprising a nucleotide sequence of NNNNCC comprises the first SNP allele, wherein the method comprises introducing into the cell: a) an RNA-guided nuclease (RGN) polypeptide having at least 90% sequence identity to SEQ ID NO: 3 or a nucleic acid molecule encoding the RGN polypeptide, and i) a nucleic acid molecule comprising or encoding a guide RNA of any one of claims 108 to 128; or ii) a vector of any one of claims 129 to 134; b) a vector of any one of claims 135 to 156, and 165 to 168; c) a RGN system of any one of claims 84 to 106; or d) an RNP complex of claim 107.

173. The method of claim 171 or 172, wherein the RGN polypeptide is capable of recognizing the PAM and cleaving the mutHTT allele.

174. The method of any one of claims 171 to 173, wherein the mutHTT allele has at least 36 CAG repeat sequences in exon 1.

175. The method of any one of claims 171 to 173, wherein the mutHTT allele has at least 40 CAG repeat sequences in exon 1.

176. The method of any one of claims 171 to 175, wherein the cell has been analyzed to determine whether the mutHTT allele comprises the first SNP allele prior to introducing the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

177. The method of any one of claims 171 to 176, wherein the cell comprises a wild-type HTT (wtHTT) allele that comprises a second SNP allele that is not present in the PAM, and the cell is therefore heterozygous for the SNP.

178. The method of claim 177, wherein the cell has been analyzed to determine whether the cell is heterozygous for the SNP prior to introducing the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

179. The method of any one of claims 171 to 178, wherein the mutHTT allele is edited, thereby generating a genetically modified cell comprising the edited mutHTT allele.

180. The method of claim 179, wherein the edit comprises introducing an insertion and / or deletion (INDEL) at or near the SNP.

181. The method of claim 179, wherein the edit comprises introducing a premature stop codon at or near the SNP.

182. The method of any one of claims 179 to 181, wherein the genetically modified cell is a genetically modified stem cell.

183. The method of claim 182, wherein the genetically modified stem cell is a genetically modified induced pluripotent stem cell (iPSC) or a genetically modified mesenchymal stem cell (MSC).

184. The method of claim 183, wherein the method further comprises differentiating the genetically modified iPSC or MSC into a neuronal cell.

185. The method of any one of claims 179 to 184, wherein the mutHTT mRNA content is reduced by at least 40% compared to the HTT mRNA content in a non-genetically modified cell or compared to a wild-type HTT mRNA content.

186. The method of any one of claims 179-185, wherein the amount of mutHTT protein is reduced by at least 40% as compared to the amount of HTT protein in a cell that is not genetically modified or as compared to the amount of wild-type HTT protein.

187. The method of any one of claims 179-186, further comprising selecting the genetically modified cell.

188. A genetically modified cell produced by the method of claim 187.

189. The method of any one of claims 171-187, wherein the introducing comprises administering to an individual comprising the cell a composition comprising the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

190. The method of claim 189, wherein the cell is a eukaryotic cell.

191. The method of claim 190, wherein the eukaryotic cell is a mammalian cell.

192. The method of claim 191, wherein the mammalian cell is a human cell.

193. The method of claim 191 or 192, wherein the mammalian cell or human cell is a stem cell.

194. The method of claim 191 or 192, wherein the mammalian cell or human cell is a forebrain neuron, a striatal neuron, a medium spiny neuron, a cortical neuron, or a glial cell.

195. The method of claim 191 or 192, wherein the mammalian cell or human cell is present in the putamen, the caudate, the striatum, the cerebral cortex, the globus pallidus, the hippocampus, the amygdala, the thalamus, the hypothalamus, the subthalamic nucleus, the substantia nigra, the cerebellum, the brainstem, or a combination thereof.

196. A method for ameliorating or delaying the onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the individual comprises a mutant huntingtin (mutHTT) allele, the mutHTT allele comprising: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer adjacent motif (PAM) comprising the nucleotide sequence of NNRYA comprises the first SNP allele; wherein the method comprises administering to the individual: a) an RNA-guided nuclease (RGN) polypeptide having at least 90% sequence identity to SEQ ID NO: 7 or a nucleic acid molecule encoding the RGN polypeptide, and i) a nucleic acid molecule comprising or encoding a guide RNA of any one of claims 30-50; or ii) a vector of any one of claims 51-56; b) a vector of any one of claims 57-83, 164, and 166-168; c) a RGN system of any one of claims 1-28; or d) an RNP complex of claim 29; and wherein the amount of mutHTT protein encoded by the mutHTT allele is reduced as compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 197. A method for ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the individual comprises a mutant Huntingtin (mutHTT) allele, which mutHTT allele comprises: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer adjacent motif (PAM) comprising a nucleotide sequence of NNNNCC comprises the first SNP allele; wherein the method comprises administering to the individual: a) an RNA-guided nuclease (RGN) polypeptide having at least 90% sequence identity to SEQ ID NO: 3 or a nucleic acid molecule encoding the RGN polypeptide, and i) a nucleic acid molecule comprising or encoding a guide RNA of any one of claims 108-128; or ii) a vector of any one of claims 129-134; b) a vector of any one of claims 135-156 and 165-168; c) an RGN system of any one of claims 84-106; or d) an RNP complex of claim 107; and wherein the amount of mutHTT protein encoded by the mutHTT allele is reduced compared to the amount of HTT protein or wild-type HTT protein in a control individual.

198. The method of any one of claims 196 or 197, wherein the RGN polypeptide recognizes the PAM and cleaves and edits the mutHTT allele.

199. The method of any one of claims 196-198, wherein the administering comprises intrastriatal, intraparenchymal, intrathecal, intracerebral, intracerebroventricular, intrathalamic, or intracerebellomedullary cisterna injection.

200. The method of any one of claims 196-199, wherein the individual has been analyzed to determine whether the mutHTT allele includes the first SNP allele comprising the PAM prior to administering the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

201. The method of any one of claims 196-200, wherein the individual comprises a wild-type HTT (wtHTT) allele, which wild-type HTT (wtHTT) allele comprises a second SNP allele that is absent the PAM, and the individual is therefore heterozygous for the SNP.

202. The method of claim 201, wherein the individual has been analyzed to determine whether the individual is heterozygous for the SNP prior to administering the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide and the guide RNA or the nucleic acid molecule encoding the guide RNA, the vector, the RGN system, or the RNP complex.

203. The method of any one of claims 196-202, wherein the PAM is present only on the mutHTT allele and is absent from the wild-type HTT allele. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 204. A method of ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the individual comprises a mutant Huntingtin (mutHTT) allele comprising: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele comprises a thymine at a position corresponding to position 151 of SEQ ID NO: 1; wherein the method comprises administering to the individual an AAV5 vector comprising: a) a first nucleic acid molecule encoding an RNA-guided nuclease (RGN) polypeptide having the amino acid sequence of SEQ ID NO: 7; and b) a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 25 or 26; and wherein four weeks after administration of the AAV5 vector, the amount of mutHTT protein encoded by the mutHTT allele in the individual is reduced as compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual.

205. A method of ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the individual comprises a mutant Huntingtin (mutHTT) allele comprising: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the first SNP allele comprises a cytosine at a position corresponding to position 151 of SEQ ID NO: 2; wherein the method comprises administering to the individual an AAV5 vector comprising: a) a first nucleic acid molecule encoding an RNA-guided nuclease (RGN) polypeptide having the amino acid sequence of SEQ ID NO: 3; and b) a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 27 or 28; and wherein four weeks after administration of the AAV5 vector, the amount of mutHTT protein encoded by the mutHTT allele in the individual is reduced as compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual.

206. A method of ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, the method comprising: a) selecting an individual comprising a mutant Huntingtin (mutHTT) allele comprising: i) at least 36 CAG repeat sequences in exon 1; ii) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele comprises a thymine at a position corresponding to position 151 of SEQ ID NO: 1; and b) administering to the individual an AAV5 vector comprising: i) a first nucleic acid molecule encoding an RNA-guided nuclease (RGN) polypeptide having the amino acid sequence of SEQ ID NO: 7; and ii) a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 25 or 26; and wherein four weeks after administration of the AAV5 vector, the amount of mutHTT protein encoded by the mutHTT allele in the individual is reduced compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual.

207. A method of ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the method comprises: a) selecting an individual comprising a mutant Huntingtin (mutHTT) allele, the mutHTT allele comprising: i) at least 36 CAG repeat sequences in exon 1; and ii) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele comprises a cytosine at a position corresponding to position 151 of SEQ ID NO: 2; and b) administering to the individual an AAV5 vector by intrastriatal injection, the AAV5 vector comprising: i) a first nucleic acid molecule encoding an RNA-guided nuclease (RGN) polypeptide having the amino acid sequence of SEQ ID NO: 3; and ii) a second nucleic acid molecule encoding a guide RNA having the nucleotide sequence of SEQ ID NO: 27 or 28; and wherein four weeks after administration of the AAV5 vector, the amount of mutHTT protein encoded by the mutHTT allele in the individual is reduced compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual.

208. The method of any one of claims 204-207, wherein the RGN polypeptide recognizes and cleaves and edits the mutHTT allele.

209. The method of any one of claims 204-208, wherein the edit comprises introducing an INDEL at or near the SNP.

210. The method of any one of claims 204-208, wherein the edit comprises introducing a premature stop codon at or near the SNP.

211. The method of any one of claims 204-210, wherein the mutHTT allele has at least 40 CAG repeat sequences in exon 1.

212. The method of any one of claims 204-210, wherein the mutHTT allele has at least 56 CAG repeat sequences in exon 1 and wherein the individual is less than 18 years old.

213. The method of any one of claims 204-212, wherein the administration is performed prior to onset of symptoms of Huntington’s Disease.

214. The method of any one of claims 204-213, wherein the method comprises preventing onset of one or more symptoms of Huntington’s Disease.

215. The method of any one of claims 204-214, wherein the individual has at least one symptom of Huntington’s Disease.

216. The method of any one of claims 204-215, wherein a reduction of at least 40% in the amount of mutant HTT mRNA as compared to the amount of HTT mRNA or the amount of wild-type HTT mRNA in a control individual is observed.

217. The method of any one of claims 204-216, wherein a reduction in the amount of the mutHTT protein is observed 12 weeks after administration of the vector.

218. The method of claim 217, wherein a reduction of at least 40% in the amount of the mutHTT protein as compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual is observed.

219. The method of any one of claims 204-218, wherein a reduction in the amount of mutHTT mRNA is observed in at least 50% of striatal cells in the individual.

220. The method of any one of claims 204-219, wherein a reduction in the amount of the mutHTT protein is observed in at least 50% of striatal cells in the individual.

221. The method of any one of claims 204-220, wherein only the mutant HTT allele is edited.

222. A method for cleaving a mutant huntingtin (mutHTT) allele in a cell, wherein the mutHTT allele comprises a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer adjacent motif (PAM) comprises the first SNP allele, wherein the method comprises introducing into the cell: (i) an RNA-guided nuclease (RGN) polypeptide or a nucleic acid molecule encoding the RGN polypeptide, and (ii) a guide RNA or a nucleic acid molecule encoding the guide RNA.

223. The method of claim 222, wherein the RGN polypeptide is capable of recognizing the PAM and cleaving the mutHTT allele.

224. The method of claim 223, wherein the mutHTT allele has at least 36 CAG repeat sequences in exon 1.

225. The method of claim 223 or 224, wherein the cell has been analyzed to determine whether the mutHTT allele comprises the first SNP allele prior to introducing (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

226. The method of any one of claims 223-225, wherein the cell comprises a wild-type HTT (wtHTT) allele, which wild-type HTT (wtHTT) allele comprises a second SNP allele that is not present in the PAM, and the cell is therefore heterozygous for the SNP.

227. The method of claim 226, wherein the cell has been analyzed to determine whether the cell is heterozygous for the SNP prior to introducing (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

228. The method of any one of claims 222-227, wherein the mutHTT allele is edited, thereby producing a genetically modified cell comprising the edited mutHTT allele.

229. The method of claim 228, wherein the edit comprises introducing an insertion and / or deletion (INDEL) at or near the SNP.

230. The method of claim 228, wherein the edit comprises introducing a premature codon at or near the SNP.

231. The method of any one of claims 228-230, wherein the genetically modified cell is a genetically modified stem cell.

232. The method of claim 231, wherein the genetically modified stem cell is a genetically modified induced pluripotent stem cell (iPSC) or a genetically modified mesenchymal stem cell (MSC).

233. The method of claim 232, wherein the method further comprises differentiating the genetically modified iPSC or MSC into a neuronal cell.

234. The method of any one of claims 228-233, wherein the mutHTT mRNA content in the genetically modified cell is reduced as compared to the HTT mRNA content in a non-genetically modified cell or as compared to the wild-type HTT mRNA content.

235. The method of any one of claims 228-234, wherein the content of mutHTT protein encoded by the mutHTT allele in the genetically modified cell is reduced as compared to the HTT protein content in a non-genetically modified cell or as compared to the wild-type HTT protein content.

236. The method of any one of claims 228-235, further comprising selecting the genetically modified cell.

237. A genetically modified cell produced by the method of claim 236.

238. The method of any one of claims 222-227, wherein the introducing comprises administering to an individual comprising the cell a composition comprising: (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA.

239. The method of claim 238, wherein the cell is a eukaryotic cell.

240. The method of claim 239, wherein the eukaryotic cell is a mammalian cell.

241. The method of claim 240, wherein the mammalian cell is a human cell.

242. The method of claim 240 or 241, wherein the mammalian cell or human cell is a stem cell.

243. The method of claim 240 or 241, wherein the mammalian cell or human cell is a forebrain neuron, a striatal neuron, a medium spiny neuron, a cortical neuron, or a glial cell.

244. The method of claim 240 or 241, wherein the mammalian cell or human cell is present in the putamen, the caudate, the striatum, the cerebral cortex, the globus pallidus, the hippocampus, the amygdala, the thalamus, the hypothalamus, the subthalamic nucleus, the substantia nigra, the cerebellum, the brainstem, or a combination thereof.

245. A method for ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the individual comprises a mutant Huntingtin (mutHTT) allele, which mutHTT allele comprises: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein a protospacer adjacent motif (PAM) comprises the first SNP allele; wherein the method comprises administering to the individual (i) an RNA-guided nuclease (RGN) polypeptide or a nucleic acid molecule encoding the RGN polypeptide, and (ii) a guide RNA or a nucleic acid molecule encoding the guide RNA, and wherein the amount of mutHTT protein encoded by the mutHTT allele is reduced compared to the amount of HTT protein or wild-type HTT protein in a control individual.

246. The method of claim 245, wherein the RGN polypeptide recognizes the PAM and cleaves and edits the mutHTT allele.

247. The method of claim 245, wherein the mutHTT allele has at least 40 CAG repeat sequences in exon 1.

248. The method of claim 245, wherein the mutHTT allele has at least 56 CAG repeat sequences in exon 1 and wherein the individual is less than 18 years old.

249. The method of any one of claims 245-248, wherein the administration is performed prior to onset of symptoms of Huntington’s Disease.

250. The method of any one of claims 245-249, wherein the method comprises preventing onset of one or more symptoms of Huntington’s Disease.

251. The method of any one of claims 245-250, wherein the individual has at least one symptom of Huntington’s Disease.

252. The method of any one of claims 245-251, wherein the administration comprises intrastriatal, intraparenchymal, intrathecal, intracerebral, intracerebroventricular, intrathalamic, or intracerebellomedullary cisternal injection.

253. The method of any one of claims 245-252, wherein the individual has been analyzed prior to administration of (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA to determine whether the mutHTT allele comprising the PAM comprises the first SNP allele.

254. The method of any one of claims 245-253, wherein the individual comprises a wild-type HTT (wtHTT) allele, which wtHTT allele comprises a second SNP allele that is not present in the PAM, and the individual is therefore heterozygous for the SNP.

255. The method of claim 254, wherein the individual has been analyzed prior to administration of (i) the RGN polypeptide or the nucleic acid molecule encoding the RGN polypeptide, and (ii) the guide RNA or the nucleic acid molecule encoding the guide RNA to determine whether the individual is heterozygous for the SNP. ​ ​ ​ 256. The method of any of claims 245-255, wherein the amount of mutHTT mRNA is reduced by at least 40% as compared to the amount of HTT mRNA or the amount of wild-type HTT mRNA in a control individual.

257. The method of any of claims 245-256, wherein the amount of mutHTT protein is reduced by at least 40% as compared to the amount of HTT protein or the amount of wild-type HTT protein in a control individual.

258. The method of any of claims 245-257, wherein a reduction in the amount of mutHTT protein is observed by 12 weeks post administration.

259. The method of any of claims 245-258, wherein a reduction in the amount of mutHTT protein is observed in at least 50% of striatal cells in the individual.

260. The method of any of claims 222-259, wherein the PAM has a nucleotide sequence selected from the group consisting of: NNNNCC, NNRYA, NNGRR, and NNGG.

261. The method of claim 260, wherein the PAM sequence NNRYA comprises the first SNP allele, and wherein the first SNP allele is a thymine at a position corresponding to position 151 of SEQ ID NO:

1.

262. The method of claim 261, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to SEQ ID NO:

7.

263. The method of claim 261 or 262, wherein the RGN polypeptide comprises the amino acid sequence of SEQ ID NO:

7.

264. The method of claim 261 or 262, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106, or a nucleotide sequence that differs from SEQ ID NO: 8 or 106 by 1 or 2 nucleotides, and a tracrRNA having a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 9 or 107.

265. The method of claim 263, wherein the guide RNA comprises a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 8 or 106, and a tracrRNA having the nucleotide sequence of SEQ ID NO: 9 or 107.

266. The method of any of claims 262-265, wherein the guide RNA comprises a spacer comprising a nucleotide sequence that is complementary to a target sequence having the nucleotide sequence of SEQ ID NO: 75 or 76.

267. The method of claim 266, wherein the spacer has the nucleotide sequence of SEQ ID NO: 80 or 81, or a nucleotide sequence that differs from SEQ ID NO: 80 or 81 by 1 or 2 nucleotides.

268. The method of claim 267, wherein the spacer has the nucleotide sequence of SEQ ID NO: 80 or 81.

269. The method of any one of claims 261-268, wherein the guide RNA is a single guide RNA.

270. The method of claim 269, wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 25 or 26.

271. The method of any one of claims 222-236 and 238-270, wherein the method comprises introducing a vector comprising the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA, and wherein the vector comprises: a truncated U6 promoter regulating expression of sgRNA; a CMVeb promoter regulating expression of RGN polypeptide; a c-Myc NLS at the N- or C-terminus of the RGN polypeptide; a NLS linker protein linking the c-Myc NLS to the RGN polypeptide; and a SV40 polyA tail.

272. The method of claim 271, wherein the truncated U6 promoter has the sequence set forth as SEQ ID NO: 128, the CMVeb promoter has the sequence set forth as SEQ ID NO: 90, the c-Myc NLS has the sequence set forth as SEQ ID NO: 125, the NLS linker protein has the sequence set forth as SEQ ID NO: 127, and the SV40 polyA tail has the sequence set forth as SEQ ID NO:

94.

273. The method of claim 271 or 272, wherein the sgRNA has the sequence set forth as SEQ ID NO: 26 and the nucleic acid molecule encoding the RGN polypeptide has the sequence set forth as SEQ ID NO:

88.

274. The method of any one of claims 271-273, wherein the vector comprises the sequence set forth as SEQ ID NO:

123.

275. The method of any one of claims 238-270, wherein the method comprises administering to the individual a vector comprising the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA, and wherein the vector comprises: a truncated U6 promoter regulating expression of sgRNA; a CMVeb promoter regulating expression of RGN polypeptide; a c-Myc NLS at the N- or C-terminus of the RGN polypeptide; a NLS linker protein linking the c-Myc NLS to the RGN polypeptide; and a SV40 polyA tail.

276. The method of claim 275, wherein the truncated U6 promoter has the sequence set forth as SEQ ID NO: 128, the CMVeb promoter has the sequence set forth as SEQ ID NO: 90, the c-Myc NLS has the sequence set forth as SEQ ID NO: 125, the NLS linker protein has the sequence set forth as SEQ ID NO: 127, and the SV40 polyA tail has the sequence set forth as SEQ ID NO:

94.

277. The method of claim 275 or 276, wherein the sgRNA has the sequence set forth as SEQ ID NO: 26 and the nucleic acid molecule encoding the RGN has the sequence set forth as SEQ ID NO:

88.

278. The method of any one of claims 275-277, wherein the vector comprises the sequence set forth as SEQ ID NO:

123.

279. The method of any one of claims 222-260, wherein the PAM sequence selected from the group consisting of NNNNCC, NNGRR, and NNGG comprises the first SNP allele, and wherein the first SNP allele is cytosine at a position corresponding to position 151 of SEQ ID NO:

1.

280. The method of claim 279, wherein the RGN polypeptide comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 3, 11, and 15.

281. The method of claim 279 or 280, wherein the RGN polypeptide comprises the amino acid sequence of any one of SEQ ID NOs: 3, 11, and 15.

282. The method of claim 279, wherein the RGN polypeptide and the guide RNA are selected from the group consisting of: a) a RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 3 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 or a nucleotide sequence that differs from SEQ ID NO: 4 by 1 or 2 nucleotides, and a tracrRNA having a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 5; b) a RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 11 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 12 or a nucleotide sequence that differs from SEQ ID NO: 12 by 1 or 2 nucleotides, and a tracrRNA having a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 13 or 120; and c) a RGN polypeptide comprising an amino acid sequence having at least 90% sequence identity to SEQ ID NO: 15 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 16 or a nucleotide sequence that differs from SEQ ID NO: 16 by 1 or 2 nucleotides, and a tracrRNA having a nucleotide sequence with at least 90% sequence identity to SEQ ID NO:

17.

283. The method of claim 282, wherein the RGN polypeptide and the guide RNA are selected from the group consisting of: a) an RGN polypeptide comprising the amino acid sequence of SEQ ID NO: 3 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 4 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 5; b) an RGN polypeptide comprising the amino acid sequence of SEQ ID NO: 11 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 12 and a tracrRNA having the nucleotide sequence of SEQ ID NO: 13 or 120; and c) an RGN polypeptide comprising the amino acid sequence of SEQ ID NO: 15 and a guide RNA comprising a crRNA repeat sequence having the nucleotide sequence of SEQ ID NO: 16 and a tracrRNA having the nucleotide sequence of SEQ ID NO:

17.

284. The method of claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and the guide RNA of claim 282(a) or 283(a), and wherein the guide RNA comprises a spacer having a nucleotide sequence complementary to a target sequence of SEQ ID NO: 77 or 78.

285. The method of claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and the guide RNA of claim 282(a) or 283(a), and wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83 or a nucleotide sequence differing by 1 or 2 nucleotides from SEQ ID NO: 82 or 83.

286. The method of claim 282 or 283, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and the guide RNA of claim 282(a) or 283(a), and wherein the guide RNA comprises a spacer having the nucleotide sequence of SEQ ID NO: 82 or 83.

287. The method of any one of claims 279-286, wherein the guide RNA is a single guide RNA.

288. The method of claim 287, wherein the RGN polypeptide and the guide RNA are the RGN polypeptide and the guide RNA of claim 282(a) or 283(a), and wherein the single guide RNA has the nucleotide sequence of SEQ ID NO: 27 or 28.

289. The method of any one of claims 222-288, wherein the PAM is present only on the mutHTT allele and is not present on the wild-type HTT allele.

290. The method of any one of claims 222-289, wherein the nucleic acid molecule encoding the RGN polypeptide is an mRNA.

291. The method of any one of claims 222-289, wherein the nucleic acid molecule encoding the RGN polypeptide and the nucleic acid molecule encoding the guide RNA are in a viral vector.

292. The method of claim 291, wherein the viral vector is a lentiviral vector, a baculoviral vector, or an adeno-associated viral (AAV) vector.

293. The method of claim 292, wherein the AAV vector is AAV5.

294. A method for detecting mutant huntingtin (mutHTT) protein and wild type HTT (wtHTT) protein in a sample, the method comprising: a) applying a denatured sample to a capillary comprising a sieving medium; b) applying a voltage difference to the capillary to separate proteins within the sample by molecular weight via electrophoresis; c) immobilizing the separated proteins within the capillary; d) applying to the capillary a first antibody or fragment thereof capable of binding to both mutHTT and wtHTT; e) applying to the capillary a second antibody or fragment thereof capable of binding to the first antibody, wherein the second antibody comprises a detectable label; and f) detecting the detectable label.

295. The method of claim 294, wherein the sieving medium is a hydrophilic polymer matrix.

296. The method of claim 294 or 295, wherein the sample is a biological sample.

297. The method of any one of claims 294 to 296, wherein the detectable label is a chemiluminescent label or a fluorescent label.

298. The method of any one of claims 294 to 297, wherein the mutHTT and wtHTT proteins in the sample are quantified by comparison to a standard curve.

299. The method of any one of claims 294 to 298, wherein the method is capable of resolving the mutHTT protein from the wtHTT.

300. A method of ameliorating or delaying onset of one or more symptoms of Huntington’s Disease (HD) in an individual in need thereof, wherein the method comprises delivering to the individual an adeno-associated viral (AAV) 5 vector comprising: a) a guide RNA having a crRNA of SEQ ID NO: 8 or 106 and a tracrRNA of SEQ ID NO: 9 or 107; or b) a single guide RNA having a nucleotide sequence of SEQ ID NO: 25 or 26.

301. The method of claim 300, wherein the individual comprises a mutant huntingtin (mutHTT) allele comprising: a) at least 36 CAG repeat sequences in exon 1; and b) a first single nucleotide polymorphism (SNP) allele in exon 50, wherein the SNP allele comprises a thymine at a position corresponding to position 151 of SEQ ID NO:

1.

302. The method of claim 300 or 301, wherein the method comprises administering the AAV5 vector by intrastriatal injection.

303. The method of any one of claims 300 to 302, wherein the amount of mutHTT protein encoded by the mutHTT allele in the individual is reduced compared to the amount of HTT protein or the amount of wild type HTT protein in a control individual.

304. The method of claim 303, wherein a reduction in the amount of mutHTT protein of at least 40% as compared to the amount of HTT protein or wild-type HTT protein in a control individual is observed.

305. The method of claim 303 or 304, wherein the reduction in the amount of mutHTT protein is observed 4 weeks, 6 weeks, 8 weeks, 10 weeks, or 12 weeks after administration of the vector.

306. The method of any one of claims 303-305, wherein the reduction in the amount of mutHTT protein is observed in at least 50% of striatal cells in the individual.

307. The method of any one of claims 300-306, wherein the amount of mutHTT mRNA is reduced by at least 40% as compared to the amount of HTT mRNA or wild-type HTT mRNA in a control individual.

308. The method of any one of claims 300-307, wherein the mutHTT allele has at least 40 CAG repeat sequences in exon 1.

309. The method of any one of claims 300-308, wherein the mutHTT allele has at least 56 CAG repeat sequences in exon 1 and wherein the individual is less than 18 years old.

310. The method of any one of claims 300-309, wherein the administration is performed prior to the onset of symptoms of Huntington’s disease.

311. The method of any one of claims 300-310, wherein the method comprises preventing the onset of one or more symptoms of Huntington’s disease.

312. The method of any one of claims 300-311, wherein the individual has at least one symptom of Huntington’s disease.

313. The method of any one of claims 300-312, wherein only the mutant HTT allele is edited.

Citation Information

Patent Citations

  • Method for measurement of total protein content and detection of protein via immunoassay in a microfluidic device

    US11143650B1

  • Method for measurement of total protein content and detection of protein via immunoassay in a microfluidic device

    US11237164B2

  • Methods and systems for analysis of samples containing particles used for gene delivery

    US11535900B1

  • Apparatus, systems, and methods for capillary electrophoresis

    US11933759B2

  • JeT promoter

    US20020098547A1