Compositions and methods for treatment of huntington's disease

A CRISPR system targeting specific genome coordinates selectively edits the mutant HTT gene to reduce CAG repeats, effectively treating Huntington's disease by minimizing impact on wild-type alleles.

WO2025172421A1PCT designated stage Publication Date: 2025-08-21ASTRAZENECA AB
View PDF 17 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/053831
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-14
Filing Date
2025-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

There is currently no cure for Huntington's disease, a monogenic, autosomal dominant neurodegenerative disorder caused by pathologic expansion of CAG trinucleotide repeats in the huntingtin (HTT) gene, leading to progressive neurodegeneration.

Method used

A CRISPR system comprising specific guide and scaffold sequences targets the human genome coordinates chr4:3078423-3078443, selectively reducing or removing CAG repeats in the mutant HTT gene while sparing the wild-type allele, using Cas9 from Staphylococcus aureus (SaCas9) to edit the mutant allele.

Benefits of technology

The CRISPR system effectively reduces or removes CAG repeats in the mutant allele, lowering RNA and protein expression by at least 30% without significantly affecting the wild-type allele, providing a potential therapeutic approach for Huntington's disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025053831_21082025_PF_FP_ABST
    Figure EP2025053831_21082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to polynucleotides, expression vectors, compositions, and methods for treatment of Huntington's Disease (HD). In some embodiments, the polynucleotides, expression vectors, and compositions provided herein are components of a CRISPR system.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] COMPOSITIONS AND METHODS FOR TREATMENT OF HUNTINGTON’S DISEASE FIELD

[0001] The present disclosure provides polynucleotides, expression vectors, compositions, and methods for treatment of Huntington’s Disease (HD). In some embodiments, the polynucleotides, expression vectors, and compositions provided herein are components of a CRISPR system. In some embodiments, the disclosure provides a polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises any one of SEQ ID NOs:21-95. BACKGROUND

[0002] Huntington’s disease (HD) is an autosomal dominant neurodegenerative disease that manifests in adults (adult-onset HD) or children (Juvenile HD). HD is a monogenic, neurodegenerative disease, generally understood to be caused by a pathologic expansion of CAG trinucleotide repeats (also called CAG repeats) within exon 1 of the huntingtin (HTT) gene. Patients with HD develop progressive neurodegeneration leading to death, generally within 20 years of onset. There is currently no cure for HD.

[0003] Previous studies using genetically modified mouse models showed that HD-like disease phenotypes can be resolved if mutant huntingtin expression is eliminated, even at advanced disease stages. See, e.g., Yamamoto et al.. Cell 101:57-66, 2000; and Diaz- Hernandez et al., J Neurosci 25:9773-9781, 2005. SUMMARY

[0004] In some embodiments, the present disclosure provides a polynucleotide including a guide sequence and a scaffold sequence, wherein the guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence includes any one of SEQ ID NOs:21-95.

[0005] In some embodiments, the guide sequence does not form a secondary structure. In some embodiments, the secondary structure is a stem loop. In some embodiments, the guide sequence includes any one of SEQ ID NOs:1-12.

[0006] In some embodiments, the present disclosure provides a polynucleotide including a guide sequence and a scaffold sequence, wherein the guide sequence includes any one of SEQ ID NOs:2-12, and wherein the guide sequence including SEQ ID NO:2 does not include a 5' guanine.

[0007] In some embodiments, the guide sequence includes SEQ ID NO:2 and does not include a 5' guanine. In some embodiments, the scaffold sequence is capable of binding to a Cas protein. In some embodiments, the Cas protein is Cas9 from Staphylococcus aureus (SaCas9). In some embodiments, the scaffold sequence does not include an early termination signal sequence. In some embodiments, the early termination signal sequence includes 4 to 6 consecutive thymine bases. In some embodiments, the scaffold sequence includes a stabilized secondary structure. In some embodiments, the stabilized secondary structure includes a locked loop.

[0008] In some embodiments, the scaffold sequence includes any one of SEQ ID NOs:20-95. In some embodiments, the scaffold sequence includes any one of SEQ ID NOs:20-22, 25, 28- 31, and 44-47.

[0009] In some embodiments, the present disclosure provides a polynucleotide including any one of SEQ ID NOs:97-1007. In some embodiments, the present disclosure provides a polynucleotide, including any one of SEQ ID NOs:97, 98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199. In some embodiments, the present disclosure provides a polynucleotide, including any one of SEQ ID NOs: 97, 98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

[0010] In some embodiments, the present disclosure provides a polynucleotide including a guide sequence and a scaffold sequence, wherein: (i) the guide sequence includes any one of SEQ ID NOs:2-12, and the scaffold sequence includes any one of SEQ ID NOs:20-95, wherein the guide sequence including SEQ ID NO:2 does not include a 5' guanine; or (ii) the guide sequence includes any one of SEQ ID NOs:1-12, and the scaffold sequence includes any one of SEQ ID NOs:21-95.

[0011] In some embodiments, the present disclosure provides an expression system including a nucleic acid sequence encoding the polynucleotide described herein.

[0012] In some embodiments, the nucleic acid sequence is a first nucleic acid sequence, the polynucleotide is a first polynucleotide, the guide sequence is a first guide sequence, and the scaffold sequence is a first scaffold sequence; and wherein the expression system further includes a second nucleic acid sequence encoding a second polynucleotide, wherein the second polynucleotide includes a second guide sequence and a second scaffold sequence, wherein the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862.

[0013] In some embodiments, the second guide sequence does not form a secondary structure. In some embodiments, the secondary structure is a stem loop. In some embodiments, the second guide sequence includes any one of SEQ ID NOs:13-19. In some embodiments, the second guide sequence includes SEQ ID NO:13.

[0014] In some embodiments, the second scaffold sequence is capable of binding to a Cas protein. In some embodiments, the Cas protein is SaCas9. In some embodiments, the second scaffold sequence does not include an early termination signal sequence. In some embodiments, the early termination signal sequence includes 4 to 6 consecutive thymine bases. In some embodiments, the second scaffold sequence includes a stabilized secondary structure. In some embodiments, the stabilized secondary structure includes a locked loop.

[0015] In some embodiments, the second scaffold sequence includes any one of SEQ ID NOs:20-95. In some embodiments, the second scaffold sequence includes any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

[0016] In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1539. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0017] In some embodiments, the present disclosure provides an expression system including: a first nucleic acid sequence encoding a first polynucleotide including any one of SEQ ID NOs:96-1007; and a second nucleic acid sequence encoding a second polynucleotide including any one of SEQ ID NOs:1008-1539.

[0018] In some embodiments, the first polynucleotide includes any one of SEQ ID NOs:96- 98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199. In some embodiments, the first polynucleotide includes any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172- 174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0019] In some embodiments, the expression system includes a vector. In some embodiments, wherein the vector is a viral vector. In some embodiments, the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector. In some embodiments, the first nucleic acid sequence and the second nucleic acid sequence are on a single vector. In some embodiments, each of the first nucleic acid sequence and the second nucleic acid sequence is on a separate vector.

[0020] In some embodiments, the expression system including the first nucleic acid sequence further includes a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide. In some embodiments, the third nucleic acid sequence is on a separate vector from the first nucleic acid sequence. In some embodiments, the first nucleic acid sequence and the third nucleic acid sequence are on a single vector. In some embodiments, the Cas protein is SaCas9.

[0021] In some embodiments, the expression system including the first nucleic acid sequence and the second nucleic acid sequence further includes a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide and / or the second polynucleotide. In some embodiments, the first, second, and third nucleic acid sequences are on a single vector. In some embodiments, the first and second nucleic acid sequences are on a first vector, and the third nucleic acid sequence is on a second vector; or wherein the first and third nucleic acid sequences are on a first vector, and the second nucleic acid sequence is on a second vector; or wherein the second and third nucleic acid sequences are on a first vector, and the first nucleic acid sequence is on a second vector. In some embodiments, each of the first, second, and third nucleic acid sequences is on a separate vector. In some embodiments, the Cas protein is SaCas9.

[0022] In some embodiments, the present disclosure provides a composition including a Cas protein, and one or both of: a) a first polynucleotide including a first guide sequence, wherein: (i) the first guide sequence targets human genome coordinates chr4:3078423- 3078443, and the first polynucleotide further includes a first scaffold sequence including any one of SEQ ID NOs:21-95; or (ii) the first guide sequence includes any one of SEQ ID NOs:2-12, wherein the guide sequence including SEQ ID NO:2 does not include a 5' guanine; and b) a second polynucleotide including a second guide sequence, wherein: (i) the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, and the second polynucleotide further includes a second scaffold sequence including any one of SEQ ID NOs:21-95; or (ii) the second guide sequence includes any one of SEQ ID NOs:13-19.

[0023] In some embodiments, the Cas protein is Cas9 from Staphylococcus aureus (SaCas9). In some embodiments, the first guide sequence and / or the second guide sequence does not form a secondary structure. In some embodiments, the secondary structure is a hairpin loop.

[0024] In some embodiments, the composition includes both the first and second polynucleotides. In some embodiments, the first polynucleotide includes the first guide sequence including any one of SEQ ID NOs:2-12, wherein the guide sequence including SEQ ID NO:2 does not include a 5' guanine, and wherein the first polynucleotide further includes a first scaffold sequence. In some embodiments, the second polynucleotide includes the second guide sequence including any one of SEQ ID NOs:13-19, and wherein the second polynucleotide further includes a second scaffold sequence.

[0025] In some embodiments, the first and second scaffold sequences are each capable of binding to the Cas protein. In some embodiments, the first scaffold sequence and the second scaffold sequence are identical. In some embodiments, the first scaffold sequence and the second scaffold sequence are different. In some embodiments, the first scaffold sequence and / or the second scaffold sequence does not include an early termination signal sequence. In some embodiments, the early termination signal sequence includes 4 to 6 consecutive thymine bases. In some embodiments, the first scaffold sequence and / or the second scaffold sequence includes a stabilized secondary structure. In some embodiments, the stabilized secondary structure includes a locked hairpin loop.

[0026] In some embodiments, the first scaffold sequence and / or the second scaffold sequence includes any one of SEQ ID NOs:20-96. In some embodiments, the first scaffold sequence and / or the second scaffold sequence includes any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

[0027] In some embodiments, the first polynucleotide includes any one of SEQ ID NOs:96- 1007. In some embodiments, the first polynucleotide includes any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 177, 180-183, 196-199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1539. In some embodiments, the second polynucleotide includes any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032-1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0028] In some embodiments, the first polynucleotide and the second polynucleotide are on a vector. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector.

[0029] In some embodiments, the present disclosure provides a delivery particle including the polynucleotide, the expression system, or the composition described herein, or combination thereof. In some embodiments, the delivery particle includes a lipid-based particle or a virus-like particle. In some embodiments, the delivery particle includes a liposome, a micelle, a vesicle, an exosome, or a lipid nanoparticle.

[0030] In some embodiments, the present disclosure provides a cell including the polynucleotide, the expression system, the composition, or the delivery particle described herein, or combination thereof.

[0031] In some embodiments, the present disclosure provides a method of reducing or removing CAG repeats in a mutant allele of a huntingtin (HTT) gene of a cell, comprising introducing to the cell: the polynucleotide, the expression system, the composition, or the delivery particle described herein, or combination thereof.

[0032] In some embodiments, the mutant allele includes greater than 35 CAG repeats prior to the introducing, and fewer than 35 CAG repeats following the introducing. In some embodiments, following the introducing, all CAG repeats of the mutant allele are removed. In some embodiments, following the introducing, exon 1 of the mutant allele is removed.

[0033] In some embodiments, the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%. In some embodiments, the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from the wild-type allele of the HTT gene by more than 50%.

[0034] In some embodiments, the present disclosure provides a method of treating Huntington’s Disease (HD) in a subject in need thereof, comprising administering to the subject: the polynucleotide, the expression system, the composition, or the delivery particle described herein, or combination thereof.

[0035] In some embodiments, the subject has greater than 35 CAG repeats in a mutant allele of a huntingtin (HTT) gene prior to the administering. In some embodiments, the subject has fewer than 35 CAG repeats in the mutant allele following the administering. In some embodiments, following the administering, all CAG repeats of the mutant allele are removed. In some embodiments, following the administering, exon 1 of the mutant allele is removed.

[0036] In some embodiments, the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%. In some embodiments, the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%. BRIEF DESCRIPTION OF THE FIGURES

[0037] FIG. 1 shows an exemplary selective editing strategy of the huntingtin (HTT) gene mutant allele as described in embodiments herein. In FIG. 1, a first guide RNA (gRNA) targets a sequence adjacent to SNP, which forms a PAM in the mutant allele but not in the wild-type allele; and a second gRNA targets a sequence adjacent to a PAM in both the mutant and wild-type alleles. The mutant allele is cleaved at both the first gRNA and second gRNA target sequences, thereby deleting exon 1 containing an expanded CAG repeats region. The wild-type allele is only cleaved at the second gRNA target sequence, leaving exon 1 intact upon repair of the cleaved second gRNA target sequence.

[0038] FIG. 2A shows an exemplary assay schematic for testing 7 different upstream gRNAs paired with a single downstream gRNA targeting the rs3856973 SNP in intron 1 of the HTT gene, as described in embodiments herein.

[0039] FIG. 2B shows representative results of the assay of FIG. 2A. The top panel shows a gel of the uncleaved and cleaved amplicon DNA. The bottom panel shows quantification of the editing band intensity for a relative comparison of each of the gRNA pairs tested.

[0040] FIG. 3A shows an exemplary assay schematic for testing additional upstream gRNAs paired with the same downstream gRNA as in FIG. 2A targeting the rs3856973 SNP in intron 1 of the HTT gene, as described in embodiments herein.

[0041] FIG. 3B shows representative results of the assay of FIG. 3A. The top panel shows a gel of the amplicon DNA spanning across gRNA cleavage sites, marked with arrows in FIG. 3A. The size of edited amplicon upon dual gRNA excision is shorter than the size of unedited amplicon. The bottom panel shows quantification of the editing band intensity for a relative comparison of each of the gRNA pairs tested. The experiment was performed in patient or healthy donor fibroblast cells. Synthetic gRNAs and SaCas9 mRNA were delivered by Neon electroporation system.

[0042] FIGS. 4A and 4B show representative quantitative assessment of the editing strategy with the lead gRNA pairs as shown in FIGS. 1, 2A, and 3A in patient-derived fibroblasts, with combinations of gRNA 1 targeting the SNP in intron 1 with gRNA 11, gRNA 25, gRNA 27, or gRNA 30, each of which targets a sequence upstream of HTT exon 1. FIG. 4A shows the percent mutant allele excision as measured by ddPCR assay that detects the corrected sequence upon exon 1 excision. FIG. 4B shows the percent knockdown of mRNA produced from the mutant vs. wild-type allele, normalized to a control group transfected with SaCas9 protein only.

[0043] FIG. 5A shows an exemplary schematic of a lentiviral vector design for expressing the first and second gRNAs and SaCas9, as described in embodiments herein.

[0044] FIGS. 5B and 5C show representative results of editing efficiency in HD patient- derived IPS neuron cells with the viral vector of FIG. 5A. FIG. 5B shows percent indels detected at each gRNA target site upon editing, measured by amplicon-sequencing.. FIG. 5C shows percent mutant allele excision as measured by ddPCR assay that detects the corrected sequence upon exon 1 excision.. MOI: multiplicity of infection.

[0045] FIG. 6A shows an exemplary predicted secondary structure of gRNA 1 as a synthetic gRNA, which does not contain the 5’ G. FIG. 6B shows an exemplary predicted secondary structure of gRNA 1 containing the 5’ G for expression from an adeno-associated viral (AAV) vector or lentiviral vector, which forms a stem loop in the spacer sequence.

[0046] FIG. 6C shows an exemplary predicted secondary structure of gRNA 1 with a mutation of the adenine directly following the 5’ G (A1) to uracil. FIG. 6D shows an exemplary predicted secondary structure of gRNA 1 with addition of a uracil directly following the 5’ G.

[0047] FIG. 6E shows representative results of an in vitro cleavage assay that tests gRNA 1 and gRNA 11 with or without a 5’ G. DMD 16-58 and EMX1-sg1 were positive control gRNAs.

[0048] FIG. 7A shows an exemplary sequence of gRNA 1 with part of the U6 promoter region, including three loop structures in the scaffold region, as described in embodiments herein. FIG. 7B shows a non-limiting list of modifications that were made to gRNA 1 to improve expression and efficiency, as described in embodiments herein.

[0049] FIG. 8A shows representative mutant allele excision as measured by ddPCR assay that detects the corrected sequence upon exon 1 excision, in patient IPS-derived neurons using lentiviruses expressing SaCas9, gRNA11 and gRNA1 unmodified or with various modifications described herein.

[0050] FIG. 8B shows representative percent indels detected at the gRNA1 target site upon editing, with gRNA11 and the same modified and unmodified gRNA1 variants shown in FIG. 8A, measured by amplicon-sequencing.

[0051] FIG. 8C shows representative percent indels detected at the gRNA11 target site upon editing, with gRNA11 and the same modified and unmodified gRNA1 variants shown in FIG. 8A, measured by amplicon-sequencing.

[0052] FIG. 9 shows representative percent mutant allele excision as measured by ddPCR assay that detects the corrected sequence upon exon 1 excision, in patient IPS-derived neurons using lentiviruses expressing SaCas9, gRNA11 and gRNA1 unmodified or with various modifications in combination described herein.

[0053] FIG. 10A shows representative percent indels detected at the gRNA1 target site upon editing, with gRNA11 and the same modified and unmodified gRNA1 variants shown in FIG. 8A, measured by amplicon-sequencing.

[0054] FIG. 10B shows representative percent indels detected at the gRNA11 target site upon editing, with gRNA11 and the same modified and unmodified gRNA1 variants shown in FIG. 8A, measured by amplicon-sequencing.

[0055] FIG. 11 shows an exemplary schematic with combinations of the different modifications described in embodiments herein for gRNA 1 and gRNA 11.

[0056] FIG. 12 A shows the fold change in HTT mRNA levels upon dual gRNA excision with the engineered guides and an optimized SaCas9 using lentiviral vector delivery in HD patient-derived IPS-Neurons. Measurements were obtained using allele specific SNP (rs362331) recognizing RT-qPCR assays that can distinguish wt and mutant HTT RNA.

[0057] FIG. 12 B shows the fold change of total and mutant HTT protein reduction achieved in HD patient-derived IPS-Neurons upon dual gRNA excision with the engineered guides and the optimized SaCas9 with lentivirus delivery. Gys1 refers to a guide RNA targeting the mouse glycogen synthase gene Gys1 as non-targeting control and HTT to HTT targeting optimized dual gRNAs. Measurements were obtained using SMCxPRO immunoassay system. Data is normalized to untransduced cells.

[0058] FIG. 13A shows the mutant HTT allele excision % in HD patient-derived iNeurons after treatment with escalating doses of AAVs (1E3, 1E4 and 1E5 AAV particles / cell) containing the optimized guides and the original or a codon-NLS optimized SaCas9..

[0059] FIG. 13B Shows the efficiency of editing (indels) at single sites (g1 top, g11 bottom) associated with the dual excision shown in FIG. 13A. Measurements were obtained by NGS using primers flanking each single cut site.

[0060] FIG 14A shows that using the optimized SaCas9 variant resulted in an approximately 50 % increased mutant HTT allele excision compared to the original SaCas9, reaching to 9% measured by ddPCR designed to detect the excision product.

[0061] FIG. 15B shows the corresponding indel % at single gRNA target sites, reaching to 15% measured by NGS using primers flanking each single cut site.

[0062] FIG. 16 shows exemplary sequences described herein. DETAILED DESCRIPTION

[0063] The present disclosure relates to treatment of Huntington’s Disease (HD) using the Clustered-Regularly Interspaced Short Palindromic Repeats (CRISPR) system. HD is a monogenic, autosomal dominant neurodegenerative disease, generally understood to be caused by a pathologic expansion of CAG trinucleotide repeats (also called CAG repeats) within exon 1 of the huntingtin (HTT) gene. The healthy number of CAG repeats is 26 or less. The term “expanded CAG repeats” means more than 26 CAG repeats in exon 1 of the HTT gene and is generally indicative of HD. CAG repeats between 27 and 35 will not develop symptoms, but the next generation is at a small risk to develop expansion, which may or may not be into the disease-causing range. CAG repeats between 36 and 39 are incompletely penetrant; individuals may develop symptoms but typically with a late age of onset. When CAG repeats are equal to or greater than 40, the disease is fully penetrant and symptoms of the disease will occur. If an individual has 60 or more CAG repeats, juvenile onset HD will occur. Those individuals with the earliest onset tend to have the largest expansion in the number of repeats, while a lower expansion of the repeat number correlates with onset late in life. Rate of disease progression is inversely related to repeat size.

[0064] Most patients are heterozygous for the disease-causing expanded CAG repeats, i.e., having a wild-type allele and a mutant allele. SNPs present only in one allele can be used to selectively reduce expression of the mutant HTT allele while leaving expression of the wild- type HTT allele substantially unaltered. In some embodiments, the present disclosure provides a SNP that can be targeted to treat HD.

[0065] The CRISPR system has revolutionized the field of genome engineering and gene editing. In naturally occurring CRISPR systems (e.g., the bacterial immunity system), foreign DNA (e.g., from an invading virus or plasmid) is incorporated into CRISPR arrays, which then produce CRISPR-RNAs (crRNAs). The crRNA includes protospacer sequences complementary to the foreign DNA and hybridizes with trans-activating CRISPR-RNA (tracrRNA). The tracrRNA forms secondary structures, e.g., stem loops, and is capable of binding to an RNA-guided nuclease (e.g., Cas9). The crRNA / tracrRNA / nuclease complex is capable of targeting foreign DNA bearing the protospacer sequences, thereby conferring immunity against the invading virus or plasmid.

[0066] Since its original discovery, extensive research has focused on CRISPR’s ability to perform site-specific cleavage of target polynucleotides. CRISPR systems used for gene editing typically include two components: (i) a single guide RNA (gRNA or sgRNA), which includes a “crRNA” portion, also referred to herein as “guide sequence” or “spacer,” that recognizes a target sequence; and a “tracrRNA” portion, also referred to herein as “scaffold sequence,” that binds to an RNA-guided nuclease, e.g., Cas9; and (ii) the RNA-guided nuclease, which associates with the sgRNA. In order for cleavage to occur, the target sequence generally requires a PAM that is adjacent or in proximity to the target sequence. The sequence and location of the PAM varies based on the type of RNA-guided nuclease. For example, the wild-type Cas9 protein from Streptococcus pyogenes (SpCas9) recognizes the PAM “NGG,” where N is any nucleotide; the wild-type Cas9 protein from Staphylococcus aureus (SaCas9) recognizes the PAM “NNGRRT,” where N is any nucleotide and R is a purine (e.g., A or G); and the wild-type Cas12a (formerly known as Cpf1) proteins from Acidaminococcus sp. BV3L6 (AsCas12a) and Lachnospiraceae bacterium ND2006 (LbCas12a) recognize the PAM “TTTV,” where V is G, C, or A. Further, RNA-guided nucleases such as Cas9 and Cas12a may be engineered to have altered PAM specificity. See, e.g., Kleinstiver et al., Nature 523:481-485, 2015, describing SpCas9 variants that recognize the PAMs “NGAN” and “NGNG.” See also, e.g., Gao et al., Nat Biotechnol 35(8):789-792, 2017, and Toth et al., Nucleic Acids Res 48(7):3722-3733, 2020, describing AsCas12a and LbCas12a variants with altered PAM specificities. CRISPR systems and their uses are further described in, e.g., Jinek et al., Science 337(6096):816-821, 2012; Cong et al., Science 339(6121):819-823, 2013; Mali et al., Science 339(6121):823-826, 2013; and Sander et al., Nat Biotechnol 32:347-355, 2014.

[0067] In some embodiments, the present disclosure provides CRISPR systems and components thereof, which are useful for the treatment of HD. Definitions

[0068] Unless otherwise defined herein, scientific and technical terms used in the present disclosure shall have the meanings that are commonly understood by one of ordinary skill in the art. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. As used herein, “a” or “an” may mean one or more. As used herein, when used in conjunction with the word “comprising,” the words “a” or “an” may mean one or more than one. As used herein, “another” or “a further” may mean at least a second or more.

[0069] Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the method / device being employed to determine the value, or the variation that exists among the study subjects. Typically, the term “about” is meant to encompass approximately or less than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19% or 20% variability, depending on the situation.

[0070] The use of the term “or” in the claims is used to mean “and / or”, unless explicitly indicated to refer only to alternatives or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.”

[0071] As used herein, the terms “comprising” (and any variant or form of comprising, such as “comprise” and “comprises”), “having” (and any variant or form of having, such as “have” and “has”), “including” (and any variant or form of including, such as “includes” and “include”) or “containing” (and any variant or form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited, elements or method steps. It is contemplated that any embodiment discussed in this specification can be implemented with respect to any protein, compositions, polynucleotides, vectors, cells, methods, and / or kits of the present disclosure. Furthermore, compositions, polynucleotides, vectors, cells, and / or kits of the present disclosure can be used to achieve methods and proteins of the present disclosure.

[0072] The use of the term “for example” and its corresponding abbreviation “e.g.” (whether italicized or not) means that the specific terms recited are representative examples and embodiments of the disclosure that are not intended to be limited to the specific examples referenced or cited unless explicitly stated otherwise.

[0073] As used herein, “between” is a range inclusive of the ends of the range. For example, a number between x and y explicitly includes the numbers x and y, and any numbers that fall within x and y.

[0074] A “nucleic acid,” “nucleic acid molecule,” “nucleotide,” “nucleotide sequence,” “oligonucleotide,” or “polynucleotide” means a polymeric compound including covalently linked nucleotides. The term “nucleic acid” includes ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) both of which may be single- or double-stranded. The polynucleotide may comprise naturally-occurring nucleobases (e.g., guanine, adenine, cytosine, thymine, and uracil), modified nucleobases (e.g., hypoxanthine, xanthine, 7- methylguanine, dihydrouracil, 5-methylcytosine, 5-hydroxymethylcytosine), and / or artificial nucleobases (e.g., isoguanine or isocytosine). Nucleic acids are transcribed from a 5’ end to a 3’ end. In some embodiments, the disclosure provides a polynucleotide comprising RNA and DNA nucleotides. Methods of producing a polynucleotide comprising both RNA and DNA nucleotides are known in the art and include, e.g., ligation or oligonucleotide synthesis methods. In some embodiments, the disclosure provides a polynucleotide capable of forming a complex with a Cas protein as described herein. In some embodiments, the disclosure provides a polynucleotide encoding any one of the proteins disclosed herein, e.g., a Cas protein.

[0075] A “gene” refers to an assembly of nucleotides that encode a polypeptide and includes cDNA and genomic DNA nucleic acid molecules. In some embodiments, “gene” also refers to a non-coding nucleic acid fragment that can act as a regulatory sequence preceding (i.e., 5’ or “upstream”) and following (i.e., 3’ or “downstream”) the coding sequence. Genes include exons, i.e., nucleotides that are included in mature mRNA following transcription, and introns, i.e., nucleotides that do not remain in the mature mRNA and do not code for amino acids in the protein encoded by the gene. Exons include coding and non-coding sequences. In some embodiments, the disclosure relates to compositions and methods for editing a gene involved in a disease described herein, e.g., Huntington’s Disease (HD).

[0076] As used herein, “locus” or its plural form “loci” refers to a specific, fixed position on a chromosome where a particular gene or genetic marker is located. Genes may possess multiple variants known as “alleles,” and an allele may also be referred to as residing at a particular locus. Genes that have the same allele at a given locus are “homozygous,” while genes that have different alleles at a given locus are “heterozygous.” Human genome locus positions are identified by a first number corresponding to the chromosome number; a letter corresponding to the p-arm or q-arm of the chromosome; and subsequent numbers indicating the chromosome position. For example, the locus of the huntingtin (HTT) gene is 4p16.3, i.e., chromosome 4, p-arm, position 16.3. Specific regions of genes or genetic markers can be identified by their genome coordinates, denoted as “chr” followed by the chromosome number and a base pair range, counting from the p-arm telomere. For example, the HTT gene is located at base pair 3,074,510 to base pair 3,243,960 of chromosome 4, wherein the base pair numbering is based on the reference sequence GRCh38.p14, denoted as chr4:3074510- 3243960 (GRCh38.p14). Unless specified otherwise, all genome coordinates used herein are based on the reference sequence GRCh38.p14 (GenBank ID 31457668).

[0077] A nucleic acid molecule is “hybridizable” or “hybridized” to another nucleic acid molecule, such as a cDNA, genomic DNA, or RNA, when a single stranded form of the nucleic acid molecule can anneal to the other nucleic acid molecule under the appropriate conditions of temperature and solution ionic strength. Hybridization and washing conditions are known and exemplified in Sambrook et al., Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 1989, particularly Chapter 11 and Table 11.1 therein. The conditions of temperature and ionic strength determine the stringency of the hybridization. The stringency of the hybridization conditions can be selected to provide selective formation or maintenance of a desired hybridization product of two complementary polynucleotides, in the presence of other potentially cross-reacting or interfering polynucleotides. Stringent conditions are sequence- dependent; typically, longer complementary sequences specifically hybridize at higher temperatures than shorter complementary sequences. Generally, stringent hybridization conditions are between about 5 °C to about 10 °C lower than the thermal melting point Tm (i.e., the temperature at which 50% of the sequences hybridize to a substantially complementary sequence) for a specific polynucleotide at a defined ionic strength, concentration of chemical denaturants, pH, and concentration of the hybridization partners. Generally, nucleotide sequences having a higher percentage of G and C bases hybridize under more stringent conditions than nucleotide sequences having a lower percentage of G and C bases. Generally, stringency can be increased by increasing temperature, increasing pH, decreasing ionic strength, and / or increasing the concentration of chemical nucleic acid denaturants (such as formamide, dimethylformamide, dimethylsulfoxide, ethylene glycol, propylene glycol and ethylene carbonate). Stringent hybridization conditions typically include salt concentrations or ionic strength of less than about 1 M, 500 mM, 200 mM, 100 mM or 50 mM; hybridization temperatures above about 20 °C, 30 °C, 40 °C, 60 °C or 80 °C; and chemical denaturant concentrations above about 10%, 20%, 30% 40% or 50%. Because many factors can affect the stringency of hybridization, the combination of parameters may be more significant than the absolute value of any parameter alone.

[0078] The term “complementary” is used to describe the relationship between nucleotide bases that are capable of hybridizing to one another. For example, with respect to DNA, adenosine is complementary to thymine and cytosine is complementary to guanine. When two nucleic acids are “complementary,” it is meant that a first nucleic acid or one or more regions thereof is capable of hydrogen bonding with a second nucleic acid or one or more regions thereof. Complementary nucleic acids may pair through canonical Watson-Crick base pairing, or through non-canonical base pairing, e.g., Hoogsteen base pairing. Complementary nucleic acids need not have complementarity at each nucleotide and may include one or more nucleotide mismatches, i.e., points at which hydrogen bonding does not occur. For example, complementary oligonucleotides can have at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of nucleotides hydrogen bond. By contrast, “fully complementary” or “100% complementary” in reference to oligonucleotides means that each nucleotide hydrogen bonds without any nucleotide mismatches.

[0079] A nucleic acid molecule that “targets” another nucleic acid molecule, in the context of guide sequences that “target” a target sequence described herein, means that the guide sequence is capable of hybridizing to a strand of the target sequence. In some embodiments, the target sequence is a target DNA sequence comprising a coding strand and a non-coding strand of a gene described herein, wherein the non-coding strand is complementary to the coding strand. In some embodiments, a guide sequence targeting the target DNA sequence is capable of hybridizing to the coding strand. In some embodiments, a guide sequence targeting the target DNA sequence is capable of hybridizing to the non-coding strand. In some embodiments, the guide sequence is capable of hybridizing to the target DNA sequence in its entirety. For example, a target DNA sequence may comprise 30 nucleotides, and the guide sequence hybridizes to all 30 nucleotides. In some embodiments, the guide sequence is capable of hybridizing to a region within the target DNA sequence. For example, a target DNA sequence may comprise 30 nucleotides, and the guide sequence hybridizes to a region within the target DNA sequence, e.g., hybridizes to greater than 10 nucleotides, greater than 15 nucleotides, greater than 20 nucleotides, greater than 25 nucleotides, etc. In some embodiments, the guide sequence hybridizes to a region within the target DNA sequence, e.g., hybridizes to 10 to 30 nucleotides, 15 nucleotides to 30 nucleotides, or 20 nucleotides to 30 nucleotides, etc. In some embodiments, the term “targets” includes a guide sequence that is not 100% complementary to the target sequence (either the coding or non-coding strand of the target sequence), e.g., it is not complementary in 1, 2, 3 or 4 bases. In some embodiments, the noncomplementary bases are at the 5’ or the 3’ end of the target sequence. In some embodiments, the guide sequence is longer than the target sequence. For example, in some embodiments, the guide sequence has bases that are not complementary to the target sequence but add stability to the guide sequence and / or prohibit formation of secondary structures.

[0080] As used herein, the term “operably linked” means that a polynucleotide of interest, e.g., the polynucleotide encoding a nuclease, is linked to the regulatory element in a manner that allows for expression of the polynucleotide. Regulatory elements can be cis-regulatory elements or trans-regulatory elements. Regulatory elements include, for example, promoters, enhancers, terminators, 5’ and 3’ UTRs, insulators, silencers, operators, and the like. In some embodiments, the regulatory element is a promoter. In some embodiments, a polynucleotide expressing a protein of interest is operably linked to a promoter on an expression vector.

[0081] As used herein, “promoter,” “promoter sequence,” or “promoter region” refers to a DNA regulatory region or polynucleotide capable of binding RNA polymerase and involved in initiating transcription of a downstream coding or non-coding sequence. In some embodiments, the promoter sequence includes the transcription initiation site and extends upstream to include the minimum number of bases or elements used to initiate transcription at levels detectable above background. In some embodiments, the promoter sequence includes a transcription initiation site, as well as protein binding domains responsible for the binding of RNA polymerase. Eukaryotic promoters typically contain “TATA” boxes and “CAT” boxes. Various promoters, including inducible promoters, may be used to drive expression of the various vectors of the present disclosure.

[0082] A “vector” is any means for the cloning of and / or transfer of a nucleic acid into a host cell. A vector may be a replicon to which another DNA segment may be attached so as to bring about the replication of the attached segment. A “replicon” is any genetic element (e.g., plasmid, phage, cosmid, chromosome, virus) that functions as an autonomous unit of DNA replication in vivo, i.e., capable of replication under its own control. In some embodiments, the vector is an episomal vector, which is removed / lost from a population of cells after a number of cellular generations, e.g., by asymmetric partitioning. The term “vector” includes both viral and non-viral means for introducing the nucleic acid into a cell in vitro, ex vivo, or in vivo. A large number of vectors known in the art may be used to manipulate nucleic acids, incorporate response elements and promoters into genes, etc. A vector may include one or more regulatory regions, and / or selectable markers useful in selecting, measuring, and monitoring nucleic acid transfer results.

[0083] Possible vectors include, for example, plasmids or modified viruses including, for example, bacteriophages such as lambda derivatives, or plasmids such as PBR322 or pUC plasmid derivatives, or the Bluescript vector. For example, the insertion of the DNA fragments corresponding to response elements and promoters into a suitable vector can be accomplished by ligating the appropriate DNA fragments into a chosen vector that has complementary cohesive termini. Alternatively, the ends of the DNA molecules may be enzymatically modified, or any site may be produced by ligating polynucleotides (linkers) into the DNA termini. Such vectors may be engineered to contain selectable marker genes that provide for the selection of cells that have incorporated the marker into the cellular genome. Such markers allow identification and / or selection of host cells that incorporate and express the proteins encoded by the marker.

[0084] Viral vectors, and particularly retroviral vectors, have been used in a wide variety of gene delivery applications in cells, as well as living animal subjects. Viral vectors that can be used include, but are not limited to, retrovirus, adenovirus, adeno-associated virus, pox, baculovirus, vaccinia, herpes simplex, Epstein-Barr, adenovirus, geminivirus, and caulimovirus vectors. In some embodiments, a viral vector is utilized to provide the polynucleotides described herein. In some embodiments, a viral vector is utilized to provide a polynucleotide coding for a protein described herein.

[0085] Vectors may be introduced into the desired host cells by known methods, including, but not limited to, transfection, transduction, cell fusion, and lipofection. Vectors can include various regulatory elements including promoters. In some embodiments, vector designs can be based on constructs designed by Mali et al., Nat Methods 10: 957-63, 2013.

[0086] Once a suitable host system and growth conditions are established, the polynucleotides and / or expression vectors described herein can be propagated and prepared in quantity. As described herein, the expression vectors which can be used include, but are not limited to, the following vectors or their derivatives: human or animal viruses such as vaccinia virus or adenovirus; insect viruses such as baculovirus; yeast vectors; bacteriophage vectors (e.g., lambda), and plasmid and cosmid DNA vectors.

[0087] The term “expression system” refers to components for expressing a protein and / or nucleic acid (e.g., RNA) of interest. Exemplary expression systems may include, without limitation, a vector (e.g., an expression vector described herein), a cell, a transfection reagent for introducing the vector into the cell, a reagent for inducing expression of the protein and / or nucleic acid of interest (e.g., an inducer for an inducible promoter), or combinations thereof. In some embodiments, the expression vector comprises a polynucleotide, wherein the polynucleotide is capable of being transcribed into an mRNA for the protein of interest, or wherein the polynucleotide is capable of being transcribed in the nucleic acid of interest. In some embodiments, an expression system of the present disclosure comprises one or more vectors. In some embodiments, the present disclosure provides an expression system for expressing polynucleotides (e.g., guide RNAs) described herein. In some embodiments, the present disclosure provides an expression system for expressing proteins (e.g., Cas proteins) described herein.

[0088] The term “plasmid” refers to an extra chromosomal element often carrying a gene that is not part of the central metabolism of the cell, and usually in the form of circular double- stranded DNA molecules. Such elements may be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear, circular, or supercoiled, of a single- or double-stranded DNA or RNA, derived from any source, in which a number of polynucleotides have been joined or recombined into a unique construction which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3’ untranslated sequence into a cell. In some embodiments, a plasmid is utilized to provide the polynucleotides described herein. In some embodiments, a plasmid is utilized to provide a polynucleotide coding for a protein described herein.

[0089] The term “transfection” as used herein means the introduction of an exogenous nucleic acid molecule, including a vector, into a cell. Transfection methods, e.g., for components of the CRISPR / Cas compositions described herein, are known to one of ordinary skill in the art. A “transfected” cell includes an exogenous nucleic acid molecule inside the cell and a “transformed” cell is one in which the exogenous nucleic acid molecule within the cell induces a phenotypic change in the cell. The transfected nucleic acid molecule can be integrated into the host cell’s genomic DNA and / or can be maintained by the cell, temporarily or for a prolonged period of time, extra-chromosomally. Host cells or organisms that express exogenous nucleic acid molecules or fragments are referred to herein as “recombinant,” “transformed,” or “transgenic” organisms. In some embodiments, the present disclosure provides a host cell comprising any of the expression vectors described herein, e.g., an expression vector comprising a polynucleotide that encodes a protein described herein.

[0090] The term “host cell” refers to a cell into which a recombinant expression vector has been introduced, or “host cell” may also refer to the progeny of such a cell. Because modifications may occur in succeeding generations, for example, due to mutation or environmental influences, the progeny may not be identical to the parent cell, but are still included within the scope of the term “host cell.”

[0091] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.

[0092] The start of the protein or polypeptide is known as the “N-terminus” (and also referred to as the amino-terminus, NH2-terminus, N-terminal end or amine-terminus), referring to the free amine (-NH2) group of the first amino acid residue of the protein or polypeptide. The end of the protein or polypeptide is known as the “C-terminus” (and also referred to as the carboxy-terminus, carboxyl-terminus, C-terminal end, or COOH-terminus), referring to the free carboxyl group (-COOH) of the last amino acid residue of the protein or polypeptide.

[0093] An “amino acid” as used herein refers to a compound including both a carboxyl (- COOH) and amino (-NH2) group. “Amino acid” refers to both natural and unnatural, i.e., synthetic, amino acids. Natural amino acids, with their three-letter and single-letter abbreviations, include: alanine (Ala; A); arginine (Arg, R); asparagine (Asn; N); aspartic acid (Asp; D); cysteine (Cys; C); glutamine (Gln; Q); glutamic acid (Glu; E ); glycine (Gly; G); histidine (His; H); isoleucine (Ile; I); leucine (Leu; L); lysine (Lys; K); methionine (Met; M); phenylalanine (Phe; F); proline (Pro; P); serine (Ser; S); threonine (Thr; T); tryptophan (Trp; W); tyrosine (Tyr; Y); and valine (Val; V). Unnatural or synthetic amino acids include a side chain that is distinct from the natural amino acids provided above and may include, e.g., fluorophores, post-translational modifications, metal ion chelators, photocaged and photocross-linking moieties, uniquely reactive functional groups, and NMR, IR, and x-ray crystallographic probes. Exemplary unnatural or synthetic amino acids are provided in, e.g., Mitra et al., Mater Methods 3:204, 2013 and Wals et al., Front Chem 2:15, 2014. Unnatural amino acids may also include naturally-occurring compounds that are not typically incorporated into a protein or polypeptide, such as, e.g., citrulline (Cit), selenocysteine (Sec), and pyrrolysine (Pyl).

[0094] An “amino acid substitution” refers to a polypeptide or protein including one or more substitutions of wild-type or naturally occurring amino acid with a different amino acid relative to the wild-type or naturally occurring amino acid at that amino acid residue. The substituted amino acid may be a synthetic or naturally occurring amino acid. In some embodiments, the substituted amino acid is a naturally occurring amino acid selected from the group consisting of: A, R, N, D, C, Q, E, G, H, I, L, K, M, F, P, S, T, W, Y, and V. In some embodiments, the substituted amino acid is an unnatural or synthetic amino acid. Substitution mutants may be described using an abbreviated system. For example, a substitution mutation in which the fifth (5th) amino acid residue is substituted may be abbreviated as “X5Y,” wherein “X” is the wild-type or naturally occurring amino acid to be replaced, “5” is the amino acid residue position within the amino acid sequence of the protein or polypeptide, and “Y” is the substituted, or non-wild-type or non-naturally occurring, amino acid.

[0095] An “isolated” polypeptide, protein, peptide, or nucleic acid is a molecule that has been removed from its natural environment. It is also understood that “isolated” polypeptides, proteins, peptides, or nucleic acids may be formulated with excipients such as diluents or adjuvants and still be considered isolated. As used herein, “isolated” does not necessarily imply any particular level purity of the polypeptide, protein, peptide, or nucleic acid.

[0096] The term “recombinant” when used in reference to a nucleic acid molecule, peptide, polypeptide, or protein means of, or resulting from, a new combination of genetic material that is not known to exist in nature. A recombinant molecule can be produced by any of the techniques available in the field of recombinant technology, including, but not limited to, polymerase chain reaction (PCR), gene splicing (e.g., using restriction endonucleases), and solid-phase synthesis of nucleic acid molecules, peptides, or proteins.

[0097] The term “exogenous” means that the referenced molecule or activity introduced into the host cell. The molecule can be introduced, for example, by introduction of an encoding nucleic acid into the host genetic material, such as by integration into a host chromosome or as non-chromosomal genetic material, e.g., a plasmid. An “exogenous” protein can be introduced into a host cell via an “exogenous” nucleic acid encoding the protein. The term “endogenous” refers to a referenced molecule or activity that is naturally present in the host cell. An “endogenous” protein is expressed by a nucleic acid contained within the host cell. The term “heterologous” refers to a molecule or activity derived from a source other than the referenced organism / species, whereas “homologous” refers to a molecule or activity derived from the host organism / species. Accordingly, exogenous expression of an encoding nucleic acid can utilize either or both of a heterologous or homologous encoding nucleic acid.

[0098] As used herein, the terms “sequence similarity” or “% similarity” refers to the degree of identity or correspondence between nucleic acid sequences or amino acid sequences. In the context of polynucleotides, “sequence similarity” may refer to nucleic acid sequences wherein changes in one or more nucleotide bases results in substitution of one or more amino acids, but do not affect the functional properties of the protein encoded by the polynucleotide. “Sequence similarity” may also refer to modifications of the polynucleotide, such as deletion or insertion of one or more nucleotide bases, that do not substantially affect the functional properties of the resulting transcript. It is therefore understood that the present disclosure encompasses more than the specific exemplary sequences. Methods of making nucleotide base substitutions are known, as are methods of determining the retention of biological activity of the encoded polypeptide.

[0099] Moreover, the skilled artisan recognizes that similar polynucleotides encompassed by the present disclosure are also defined by their ability to hybridize, under stringent conditions, with the sequences exemplified herein. Similar polynucleotides of the present disclosure are about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 99%, at least about 99%, or about 100% identical to the polynucleotides disclosed herein.

[0100] In the context of polypeptides, “sequence similarity” refers to two or more polypeptides wherein greater than about 40% of the amino acids are identical, or greater than about 60% of the amino acids are functionally identical. “Functionally identical” or “functionally similar” amino acids have chemically similar side chains. For example, amino acids can be grouped in the following manner according to functional similarity: (i) positively-charged side chains: Arg, His, Lys; (ii) negatively-charged side chains: Asp, Glu; (iii) polar, uncharged side chains: Ser, Thr, Asn, Gln; (iv) hydrophobic side chains: Ala, Val, Ile, Leu, Met, Phe, Tyr, Trp; and (v) others: Cys, Gly, Pro.

[0101] In some embodiments, similar polypeptides of the present disclosure have about 40%, at least about 40%, about 45%, at least about 45%, about 50%, at least about 50%, about 55%, at least about 55%, about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% identical amino acids. In some embodiments, similar polypeptides of the present disclosure have about 60%, at least about 60%, about 65%, at least about 65%, about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99%, or about 100% functionally identical amino acids.

[0102] Sequence similarity can be determined by sequence alignment using methods known in the field, such as, for example, BLAST, MUSCLE, Clustal (including ClustalW and ClustalX), and T-Coffee (including variants such as, for example, M-Coffee, R-Coffee, and Expresso).

[0103] Percent identity of polynucleotides or polypeptides can be determined when the polynucleotide or polypeptide sequences are aligned over a specified comparison window. In some embodiments, only specific portions of two or more sequences are aligned to determine sequence identity. In some embodiments, only specific domains of two or more sequences are aligned to determine sequence similarity. A comparison window can be a segment of at least 10 to over 1000 residues, at least 20 to about 1000 residues, or at least 50 to 500 residues in which the sequences can be aligned and compared. Methods of alignment for determination of sequence identity are well-known and can be performed using publicly available databases such as BLAST. For example, in some embodiments, “percent identity” of two amino acid sequences is determined using the algorithm of Karlin and Altschul, Proc Nat Acad Sci USA 87:2264-2268 (1990), modified as in Karlin and Altschul, Proc Nat Acad Sci USA 90:5873- 5877 (1993). Such algorithms are incorporated into BLAST programs, e.g., BLAST+ or the NBLAST and XBLAST programs described in Altschul et al., J Mol Biol, 215: 403-410 (1990). BLAST protein searches can be performed with programs such as, e.g., the XBLAST program, score=50, wordlength=3 to obtain amino acid sequences homologous to the protein molecules of the disclosure. Where gaps exist between two sequences, Gapped BLAST can be utilized as described in Altschul et al., Nucleic Acids Res 25(17): 3389-3402 (1997). When utilizing BLAST and Gapped BLAST programs, the default parameters of the respective programs (e.g., XBLAST and NBLAST) can be used.

[0104] In some embodiments, a polypeptide or polynucleotide has 70%, at least 70%, 75%, at least 75%, 80%, at least 80%, 85%, at least 85%, 90%, at least 90%, 95%, at least 95%, 97%, at least 97%, 98%, at least 98%, 99%, or at least 99% or 100% sequence identity with a reference polypeptide or polynucleotide (or a fragment of the reference polypeptide or polynucleotide) provided herein. In some embodiments, a polypeptide or polynucleotide have about 70%, at least about 70%, about 75%, at least about 75%, about 80%, at least about 80%, about 85%, at least about 85%, about 90%, at least about 90%, about 95%, at least about 95%, about 97%, at least about 97%, about 98%, at least about 98%, about 99%, at least about 99% or about 100% sequence identity with a reference polypeptide or polynucleotide (or a fragment of the reference polypeptide or nucleic acid molecule) provided herein.

[0105] As used herein, a “complex” refers to a group of two or more associated polynucleotides and / or polypeptides. In the context of complex formation, the terms “associate” or “association” refers to molecules bound to one another through electrostatic, hydrophobic / hydrophilic, and / or hydrogen bonding interaction, without being covalently attached. A molecule that comprises different moieties covalently attached to one another is known. In some embodiments, a complex is formed when all the components of the complex are present together, i.e., a self-assembling complex. In some embodiments, a complex is formed through chemical interactions between different components of the complex such as, for example, hydrogen-bonding. In some embodiments, the polynucleotides provided herein form a complex with the proteins provided herein through secondary structure recognition of the polynucleotide by the protein. In some embodiments, a scaffold sequence of the polynucleotides provided herein comprise a secondary structure recognized by a Cas protein provided herein, thereby forming a complex comprising the polynucleotide and the Cas protein. Polynucleotides

[0106] In some embodiments, the disclosure provides a polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence targets a locus comprising the huntingtin (HTT) gene. In some embodiments, the polynucleotide is a guide RNA (gRNA or sgRNA), e.g., a guide RNA for a CRISPR system. As discussed herein, the term “guide sequence” is used interchangeably with “spacer,” “spacer sequence,” “crRNA,” and “crRNA sequence”; and the term “scaffold sequence” is used interchangeably with “tracrRNA” and “tracrRNA sequence.” While the nucleic acid sequences of the polynucleotides described herein include the DNA nucleotide thymine (T), it will be understood by one of ordinary skill in the art that the polynucleotides also encompass the RNA nucleotide uracil (U) in place of T. In some embodiments, the locus targeted by the polynucleotide is 4p16.3 of the human genome. In some embodiments, the HTT gene is located at human genome coordinates chr4:3074510-3243960.

[0107] In some embodiments, the disclosure provides a polynucleotide comprising a guide sequence and a scaffold sequence. In some embodiments, the guide sequence targets a sequence that is within about 1 to about 50 nucleotides, or within about 2 to about 40 nucleotides, or within about 3 to about 30 nucleotides, or within about 4 to about 25 nucleotides, or within about 5 to about 20 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 30 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 25 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 20 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 15 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 10 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 5 nucleotides of a SNP in the HTT gene. In some embodiments, the guide sequence targets a sequence that is within 3 nucleotides of a SNP in the HTT gene. In some embodiments, guide sequence targets a sequence that comprises a SNP in the HTT gene. As described herein, a guide sequence that “targets” a target sequence, e.g., a sequence in the HTT gene, is capable of hybridizing to a strand of the target sequence. In some embodiments, the guide sequence is capable of hybridizing to a coding strand of the target sequence. In some embodiments, the guide sequence is capable of hybridizing to a non-coding strand of the target sequence. First guide sequence

[0108] In some embodiments, the polynucleotide of the present disclosure is a first polynucleotide, the guide sequence is a first guide sequence, and the scaffold sequence is a first scaffold sequence. The first polynucleotide is also referred to herein as the “first gRNA.” In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443. In some embodiments, chr4:3078423-3078443 is within locus 4p16.3. In some embodiments, chr4:3078423-3078443 is within the huntingtin (HTT) gene. In some embodiments, chr4:3078423-3078443 is downstream (i.e., 3’) of exon 1 in the HTT gene. As described herein, a mutant allele of the HTT gene comprises expanded CAG repeats, thereby causing Huntington’s Disease. In some embodiments, chr4:3078423-3078443 is in an intron of the HTT gene. In some embodiments, the first guide sequence is capable of hybridizing to a coding strand of chr4:3078423-3078443. In some embodiments, the first guide sequence is capable of hybridizing to a non-coding strand of chr4:3078423-3078443. In some embodiments, the first guide sequence is capable of hybridizing to the entire length of chr4:3078423-3078443.

[0109] In some embodiments, chr4:3078423-3078443 is adjacent to a protospacer adjacent motif (PAM) recognizable by an RNA-guided nuclease. In some embodiments, the RNA- guided nuclease is a Cas9 protein. In some embodiments, the RNA-guided nuclease is a Cas12a protein. In some embodiments, the Cas9 protein is Staphylococcus aureus Cas9 (SaCas9). In some embodiments, the PAM comprises NNGRRT, where N is any nucleotide and R is any purine (e.g., A or G). In some embodiments, the Cas9 protein is Streptococcus pyogenes Cas9 (SpCas9). In some embodiments, the PAM comprises NGG, wherein N is any nucleotide. In some embodiments, the PAM comprises NGAN or NGNG, e.g., NGAG or NGCG, where N is any nucleotide.

[0110] In some embodiments, a SNP of the HTT gene is located adjacent to chr4:3078423- 3078443, wherein the SNP causes the loss of the PAM in a wild-type HTT allele. In some embodiments, the SNP has the NCBI SNP Database (dbSNP) accession number rs3856973. In some embodiments, the PAM is present in a mutant allele of HTT and absent in a wild- type allele of HTT. In some embodiments, the mutant allele comprises the motif NGNG. In some embodiments, the mutant allele comprises the motif NNGRRT. In some embodiments, the mutant allele comprises the sequence TCGAGT, and the wild-type allele comprises the sequence TCAAGT. In some embodiments, the sequence TCGAGT is recognizable by SpCas9 or SaCas9, while the sequence TCAAGT is not recognizable by SpCas9 or SaCas9. In some embodiments, the sequence TCGAGT is recognizable efficiently by SaCas9, while the sequence TCAAGT is not recognizable efficiently by SaCas9. One of ordinary skill in the art would understand that a PAM not recognizable efficiently by a Cas protein, e.g., SaCas9, may be capable of some cleavage, but the degree of cleavage is much lower than with a PAM recognizable efficiently by the Cas protein. In some embodiments, the SpCas9 or SaCas9, when guided to the target sequence of the first guide sequence, is capable of selectively cleaving the mutant allele of HTT and not the wild-type allele of HTT.

[0111] In some embodiments, the first guide sequence provides improved targeting efficiency over conventional guide sequences, which may form secondary structures that reduce the guide sequence’s ability to hybridize to the target sequence. In some embodiments, the first guide sequence provides increased CRISPR editing efficiency. In some embodiments, the first guide sequence does not form a secondary structure. Non-limiting examples of nucleic acid secondary structures include stem loops (also known as hairpin loops), internal loops, bulge loops, pseudoknots, and the like. In some embodiments, the first guide sequence does not form a stem loop.

[0112] When designing gRNAs for CRISPR systems, a guanine is typically added to the 5’ end of a guide sequence (also referred to herein as a “5’ guanine”) to improve expression. It was discovered that a single 5’ guanine in certain guide sequences, e.g., the first guide sequence described herein, contributes to secondary structure formation and therefore decreased targeting efficiency. In some embodiments, the first guide sequence does not comprise a 5’ guanine.

[0113] In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:1- 12, provided that the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, provided that the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine. In some embodiments, the first guide sequence comprises SEQ ID NO:1. In some embodiments, the first guide sequence comprises SEQ ID NO:2 and does not comprise a 5’ guanine. SEQ ID NO:2 is identical to SEQ ID NO:1 except that SEQ ID NO:2 does not comprise a 5’ guanine. In some embodiments, a CRISPR system comprising a gRNA with SEQ ID NO:2 as the guide sequence has at least 1.5-fold, at least 2- fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8- fold, at least 9-fold, or at least 10-fold higher editing efficiency as compared to an otherwise identical CRISPR system except with SEQ ID NO:1 as the guide sequence.

[0114] In some embodiments, the first guide sequence comprises at its 5’ end: at least one guanine and at least one additional nucleotide that do not hybridize to the target sequence and that prevent formation of a secondary structure, e.g., a stem loop. In some embodiments, the first guide sequence comprises two or more guanines at its 5’ end. In some embodiments, the two or more guanines at the 5’ end of the first guide sequence prevent formation of secondary structures. In some embodiments, the first guide sequence comprises 2, 3, 4, 5, or more than 5 guanines at its 5’ end. In some embodiments, the first guide sequence comprises SEQ ID NO:3. In some embodiments, the first guide sequence comprises SEQ ID NO:4. In some embodiments, the first guide sequence comprises SEQ ID NO:5. In some embodiments, the first guide sequence comprises SEQ ID NO:6. In some embodiments, the first guide sequence comprises SEQ ID NO:7. In some embodiments, the first guide sequence comprises SEQ ID NO:8. In some embodiments, the first guide sequence comprises SEQ ID NO:9. In some embodiments, the first guide sequence comprises SEQ ID NO:10. In some embodiments, the first guide sequence comprises SEQ ID NO:11.

[0115] In some embodiments, the first guide sequence comprises a sequence at its 5’ end that forms a 5’ secondary structure. In some embodiments, the 5’ secondary structure prevents formation of further secondary structure in the guide sequence. In some embodiments, the sequence that forms the 5’ secondary structure comprises the sequence GGACTTCGGTCC (SEQ ID NO:1540). In some embodiments, the 5’ secondary structure is a stem loop. In some embodiments, the first guide sequence comprises SEQ ID NO:12.

[0116] In some embodiments, each of SEQ ID NOs:1-12 is capable of targeting chr4:3078423-3078443. In some embodiments, each of SEQ ID NOs:1-12 is capable of hybridizing to at least a portion of the non-coding strand of chr4:3078423-3078443. In some embodiments, each of SEQ ID NOs:1-12 targets a target sequence, wherein the 3’ end of the target sequence is within 3 nucleotides of a SNP with the dbSNP accession number rs3856973. In some embodiments, each of SEQ ID NOs:1-12 targets a sequence directly upstream of a PAM for SpCas9 or SaCas9. Second guide sequence

[0117] In some embodiments, the disclosure further provides a second polynucleotide comprising a second guide sequence and a second scaffold sequence. The second polynucleotide is also referred to herein as the “second gRNA.” In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346- 3074862. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3072397-3074774. In some embodiments, the second guide sequence targets a sequence of about 15 to 30 nucleotides, or about 17 to about 28 nucleotides, or about 19 to about 26 nucleotides, or about 20 to about 25 nucleotides, or about 21 to about 24 nucleotides in length within the region of chr4:3068346-3074862, e.g., within chr4:3072397-3074774. In some embodiments, chr4:3068346-3074862 is upstream (i.e., 5’) of exon 1 in the HTT gene. In some embodiments, chr4:3068346-3074862 is upstream of the region comprising the CAG repeats in exon 1 of the HTT gene. In some embodiments, the second guide sequence is capable of hybridizing to a coding strand of chr4:3068346- 3074862, e.g., chr4:3072397-3074774. In some embodiments, the second guide sequence is capable of hybridizing to a non-coding strand of chr4:3068346-3074862, e.g., chr4:3072397- 3074774.

[0118] In some embodiments, the second guide sequence hybridizes to a sequence adjacent to a PAM recognizable by an RNA-guided nuclease. In some embodiments, the RNA-guided nuclease is a Cas9 protein. In some embodiments, the RNA-guided nuclease is a Cas12a protein. In some embodiments, the Cas9 protein SaCas9. In some embodiments, the PAM comprises NNGRRT, where N is any nucleotide and R is any purine. In some embodiments, the Cas9 protein is SpCas9. In some embodiments, the PAM comprises NGG, wherein N is any nucleotide. In some embodiments, the PAM comprises NGAN or NGNG, e.g., NGAG or NGCG, where N is any nucleotide.

[0119] In some embodiments, the first guide sequence and the second guide sequence each targets a sequence adjacent to a PAM, wherein the PAM is recognizable by the same RNA- guided nuclease, e.g., Cas protein such as SaCas9 or SpCas9. In some embodiments, the PAM adjacent to the target sequence of the first guide sequence and the PAM adjacent to the target sequence of the second guide sequence are identical. In some embodiments, the PAM adjacent to the target sequence of the first guide sequence and the PAM adjacent to the target sequence of the second guide sequence are different. In some embodiments, the PAMs are different but are recognizable by the same RNA-guided nuclease, e.g., Cas protein such as SaCas9 or SpCas9. In some embodiments, the PAM adjacent to the target sequence of the first guide sequence and the PAM adjacent to the target sequence of the second guide sequence comprise the motif NNGRRT. In some embodiments, the PAM adjacent to the target sequence of the first guide sequence and the PAM adjacent to the target sequence of the second guide sequence comprise the motif NGAN or NGNG, e.g., NGAG or NGCG. In some embodiments, the PAM adjacent to the target sequence of the first guide sequence and the PAM adjacent to the target sequence of the second guide sequence comprise the motif NGG.

[0120] In some embodiments, the PAM adjacent to the target sequence of the first guide sequence is present in a mutant allele of the HTT gene and absent in a wild-type allele of the HTT gene. In some embodiments, the PAM adjacent to the target sequence of the second guide sequence is present in both wild-type and mutant alleles of the HTT gene. In some embodiments, the SpCas9 or SaCas9, when guided to the target sequence of the second guide sequence is capable of cleaving both the wild-type allele and the mutant allele of HTT. Thus, when a CRISPR system comprising both the first gRNA and the second gRNA is present in a cell comprising the mutant allele and the wild-type allele of the HTT gene, the mutant allele comprises two cleavage sites is cleaved at both the sequence targeted by the first guide sequence (i.e., chr4:3078423-3078443) and the sequence targeted by the second guide sequence (i.e., the sequence within chr4:3068346-3074862), while the wild-type allele is only cleaved at the sequence targeted by the second guide sequence. The two cleavage sites in the mutant allele flank the region comprising the disease-causing expanded CAG repeats, which allows the expanded CAG repeats to be excised upon cleavage of the mutant allele, while the normal, non-expanded CAG repeats region of the wild-type allele is not excised due to the presence of only one cleavage site, thereby allowing selective editing of the mutant HTT allele while the wild-type allele is substantially unaltered. The selective editing of the mutant HTT allele is illustrated in FIG. 1.

[0121] In some embodiments, the second guide sequence provides improved targeting efficiency over conventional guide sequences, which may form secondary structures that reduce the guide sequence’s ability to hybridize to the target sequence. In some embodiments, the second guide sequence provides increased CRISPR editing efficiency. In some embodiments, the second guide sequence does not form a secondary structure. Non-limiting examples of nucleic acid secondary structures include stem loops (also known as hairpin loops), internal loops, bulge loops, pseudoknots, and the like. In some embodiments, the second guide sequence does not form a stem loop.

[0122] In some embodiments, the second guide sequence comprises a sequence at its 5’ end that forms a 5’ secondary structure. In some embodiments, the 5’ secondary structure prevents formation of further secondary structure in the guide sequence. In some embodiments, the sequence that forms the 5’ secondary structure comprises the sequence GGACTTCGGTCC (SEQ ID NO:1540). In some embodiments, the 5’ secondary structure is a stem loop.

[0123] In some embodiments, the second guide sequence targets a target sequence within human genome coordinates chr4:3068346-3074862, as shown in Table 1. In some embodiments, the second guide sequence comprises a sequence as shown in Table 1. Table 1. Second Guide Sequence Chromosomal Locations and Sequence IDs Target Sequence Corresponding Guide Sequence chr4:3074163-3074183 SEQ ID NO:1561 chr4:3074430-3074450 SEQ ID NO:13 chr4:3074460-3074480 SEQ ID NO:1562 chr4:3074550-3074570 SEQ ID NO:1563 chr4:3074669-3074689 SEQ ID NO:1564 chr4:3074753-3074773 SEQ ID NO:1565 chr4:3074754-3074774 SEQ ID NO:1566 chr4:3072396-3072416 SEQ ID NO:1567 chr4:3072403-3072423 SEQ ID NO:1568 chr4:3072478-3072498 SEQ ID NO:1569 chr4:3073267-3073287 SEQ ID NO:1570 chr4:3073119-3073139 SEQ ID NO:1571 chr4:3073442-3073462 SEQ ID NO:1572 chr4:3073835-3073855 SEQ ID NO:1573 chr4:3073809-3073829 SEQ ID NO:1574 chr4:3073905-3073925 SEQ ID NO:1575 chr4:3073635-3073655 SEQ ID NO:1576

[0124] In some embodiments, the second guide sequence hybridizes over the entire length of a target sequence of Table 1. In some embodiments, the second guide sequence comprises a sequence of Table 1. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576.

[0125] In some embodiments, the second guide sequence targets chr4:3074430-3074450. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19. In some embodiments, the second guide sequence comprises SEQ ID NO:13. In some embodiments, the second guide sequence comprises at its 5’ end: at least one guanine and at least one additional nucleotide that do not hybridize to the target sequence and that prevent formation of a secondary structure, e.g., a stem loop. In some embodiments, the second guide sequence comprises two or more guanines at its 5’ end. In some embodiments, the two or more guanines at the 5’ end of the second guide sequence prevent formation of secondary structures. In some embodiments, the second guide sequence comprises 2, 3, 4, 5, or more than 5 guanines at its 5’ end. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:14-19. Scaffold sequence

[0126] In some embodiments, the polynucleotide of the present disclosure comprises a guide sequence, e.g., a first guide sequence or second guide sequence as described herein, and a scaffold sequence. In some embodiments, the scaffold sequence is capable of binding to a Cas protein. In some embodiments, the scaffold sequence is a tracrRNA sequence for a Cas protein. In some embodiments, the Cas protein is SpCas9. In some embodiments, the Cas protein is SaCas9.

[0127] In some embodiments, the scaffold sequence of the present disclosure provides improved stability and / or expression of the gRNA. In some embodiments, expression of the gRNA comprises transcription by a polymerase, e.g., an RNA Polymerase such as RNA Pol I, RNA Pol II, or RNA Pol III. In some embodiments, the scaffold sequence of the present disclosure provides improved CRISPR editing efficiency. In some embodiments, the scaffold sequence (i) does not comprise an early termination signal sequence; and / or (ii) comprises a stabilized secondary structure, thereby improving stability and / or expression of the gRNA.

[0128] In some embodiments, the scaffold sequence does not comprise an early termination signal sequence. A termination signal sequence is a nucleotide sequence that recognized by a polymerase, e.g., an RNA polymerase such as RNA Pol I, RNA Pol II, or RNA Pol III, to terminate transcription. In some embodiments, a scaffold sequence comprising an “early” termination signal sequence comprises the termination signal sequence within 10 nucleotides, within 15 nucleotides, within 20 nucleotides, within 25 nucleotides, or within 30 nucleotides of the 3’ end of the scaffold sequence. In some embodiments, a scaffold sequence comprising an “early” termination signal sequence comprises the termination signal sequence within 1 nucleotide, within 3 nucleotides, within 5 nucleotides, within 10 nucleotides, within 15 nucleotides, or within 20 nucleotides of the 5’ end of the scaffold sequence. In some embodiments, the early termination signal sequence comprises a stretch of at least 2, at least 3, at least 4, at least 5, or at least 6 consecutive thymine bases. In some embodiments, the early termination signal sequence comprises about 2 to about 8 consecutive thymine bases. In some embodiments, the early termination signal sequence comprises about 3 to about 7 consecutive thymine bases. In some embodiments, the early termination signal sequence comprises about 4 to about 6 consecutive thymine bases.

[0129] In some embodiments, a scaffold sequence that comprises an early termination signal sequence has lower expression as compared to a scaffold sequence that does not comprise an early termination signal sequence. In some embodiments, a conventional scaffold sequence for a Cas protein described herein, e.g., SaCas9, comprises an early termination signal sequence. In some embodiments, the conventional scaffold sequence is based on the wild- type tracrRNA sequence of the Cas protein. Conventional scaffold sequences for Cas proteins such as SaCas9 are described, e.g., in Ran et al., Nature 520:186-191, 2015. In some embodiments, the scaffold sequence of the present disclosure, which does not comprise the early termination signal sequence, is capable of binding to a Cas protein with substantially similar affinity as the conventional scaffold sequence for the Cas protein.

[0130] In some embodiments, the 5’ end of the conventional scaffold sequence comprises SEQ ID NO:1553. SEQ ID NO:1553 comprises four consecutive thymine bases at nucleotide positions 2-5. In some embodiments, the scaffold sequence comprises SEQ ID NO:1553, except the T at nucleotide position 2 is A, G, or C. In some embodiments, the scaffold sequence comprises SEQ ID NO:1553, except the T at nucleotide position 3 is A, G, or C. In some embodiments, the scaffold sequence comprises SEQ ID NO:1553, except the T at nucleotide position 4 is A, G, or C. In some embodiments, the scaffold sequence comprises SEQ ID NO:1553, except the T at nucleotide position 5 is A, G, or C.

[0131] In some embodiments, the scaffold sequence comprises SEQ ID NO:1553, except that at least one of nucleotide positions 2-5 is A, G, or C. In some embodiments, the scaffold sequence comprises SEQ ID NO:1553 with the following modifications at nucleotide positions 2-5 of SEQ ID NO:1553: position 2 is A, G, or C, and positions 3, 4, and 5 are each T; position 3 is A, G, or C, and positions 2, 4, and 5 are each T; position 4 is A, G, or C, and positions 2, 3, and 5 are each T; position 5 is A, G, or C, and positions 2, 3, and 4 are each T; positions 2 and 3 are each independently A, G, or C, and positions 4 and 5 are each T; positions 2 and 4 are each independently A, G, or C, and positions 3 and 5 are each T; positions 2 and 5 are each independently A, G, or C, and positions 3 and 4 are each T; positions 3 and 4 are each independently A, G, or C, and positions 2 and 5 are each T; positions 3 and 5 are each independently A, G, or C, and positions 2 and 4 are each T; positions 4 and 5 are each independently A, G, or C, and positions 2 and 3 are each T; positions 2, 3, and 4 are each independently A, G, or C, and position 5 is T; positions 2, 3, and 5 are each independently A, G, or C, and position 4 is T; positions 2, 4, and 5 are each independently A, G, or C, and position 3 is T; positions 3, 4, and 5 are each independently A, G, or C, and position 2 is T; or positions 2, 3, 4 and 5 are each independently A, G, or C.

[0132] In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:1554-1560. In some embodiments, at least one of nucleotide positions 2-5 of SEQ ID NOs:1554-1560 is A, G, or C.

[0133] In some embodiments, the scaffold sequence comprises SEQ ID NO:1553 comprising a modification as described herein, i.e., a modified SEQ ID NO:1553, and a downstream sequence capable of base pairing with the modified SEQ ID NO:1553. In some embodiments, the downstream sequence capable of base pairing with the modified SEQ ID NO:1553 is a reverse complement of SEQ ID NO:1553. In some embodiments, the base pairing forms a loop structure in the scaffold sequence. In some embodiments, the scaffold sequence comprises (A) SEQ ID NO:1553 comprising a modification as described herein; and (B) a downstream sequence that is a reverse complement of (A). In some embodiments, the scaffold sequence comprises (A) any one of SEQ ID NOs:1554-1560; and (B) a downstream sequence that is a reverse complement of (A). In some embodiments, the scaffold sequence comprises (A) any one of SEQ ID NOs:1554-1560, wherein at least one of nucleotide positions 2-5 of SEQ ID NO:1554-1560 is A, G, or C; and (B) a downstream sequence that is a reverse complement of (A). In some embodiments, (A) and (B) are paired together to form a loop. In some embodiments, (A) and (B) are separated by about 1 to about 10 nucleotides. In some embodiments, (A) and (B) are separated by about 2 to about 8 nucleotides. In some embodiments, (A) and (B) are separated by about 3 to about 6 nucleotides. In some embodiments, (A) and (B) are separated by about 4 to about 5 nucleotides. In some embodiments, (A) and (B) are separated by about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, or about 10 nucleotides. In some embodiments, (A) and (B) are directly adjacent to one another. As described herein, (A) and (B) may be complementary without necessarily having complementarity at each nucleotide. In some embodiments, (A) and (B) include one or more nucleotide mismatches, i.e., points at which hydrogen bonding does not occur. For example, complementary oligonucleotides can have at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of nucleotides hydrogen bonded. In some embodiments, (A) and (B) are fully complementary. In some embodiments, the complementarity between (A) and (B) comprises Watson-Crick base pairing, Hoogsteen base pairing, or a combination thereof.

[0134] In some embodiments, the scaffold sequence comprises a stabilized secondary structure. In some embodiments, the stabilized secondary structure comprises a sequence that promotes formation of the secondary structure. In some embodiments, the stabilized secondary structure comprises a sequence that improves stability of the secondary structure. In some embodiments, the stabilized secondary structure has improved binding affinity to its corresponding Cas protein, e.g., SaCas9. In some embodiments, the stabilized secondary structure comprises a stem loop.

[0135] In some embodiments, the stabilized secondary structure comprises a modification that promotes folding and / or improves stability of a stem loop sequence present in a conventional scaffold sequence for a Cas protein described herein, e.g., SaCas9. In some embodiments, the conventional scaffold sequence is based on the wild-type tracrRNA sequence of the Cas protein. In some embodiments, the stem loop sequence comprises “GAAA.” In some embodiments, the stem loop sequence comprises “AAAAT.” In some embodiments, the stem loop sequence comprises “ACTT.” In some embodiments, the scaffold sequence comprises at least 1, at least 2, or at least 3 stem loops, wherein each stem loop comprises a modification described herein. In some embodiments, the modification comprises one or more additional nucleotides that are incorporated into the stem loop sequence. In some embodiments, the modification increases the number of hydrogen bonding interactions in the stem loop, thereby promoting folding and / or improving stability of the stem loop. In some embodiments, the modification comprises adding the flanking nucleotides “CTG” and “CAG” to the stem loop sequence. In some embodiments, the modification comprises adding the flanking nucleotides “TAATT” and “AATTA” to the stem loop sequence.

[0136] In some embodiments, the stabilized secondary structure comprises a locked loop. In some embodiments, one or more stem loops of the conventional scaffold sequence is replaced with a locked loop. In some embodiments, the locked loop comprises a sequence that induces folding into a highly stabilized stem loop, e.g., with a melting temperature of at least 65°C, at least 66°C, at least 67°C, at least 68°C, at least 69°C, at least 70°C, or at least 71°C. An exemplary locked loop sequence is GGACTTCGGTCC (SEQ ID NO:1540). Locked loops are further described, e.g., in Riesenberg et al., Nat Commun 13:489, 2022. In some embodiments, the stabilized secondary structure comprises the sequence GGACTTCGGTCC (SEQ ID NO:1540).

[0137] In some embodiments, the stabilized secondary structure comprises the sequence CTGGAAACAG (SEQ ID NO:1541). In some embodiments, the stabilized secondary structure comprises the sequence TAATTGAAAAATTA (SEQ ID NO:1542). In some embodiments, the stabilized secondary structure comprises the sequence CTGAAAATCAG (SEQ ID NO:1543). In some embodiments, the stabilized secondary structure comprises the sequence TAATTAAAATAATTA (SEQ ID NO:1544). In some embodiments, the stabilized secondary structure comprises the sequence CTGACTTCAG (SEQ ID NO:1545). In some embodiments, the stabilized secondary structure comprises the sequence TAATTACTTAATTA (SEQ ID NO:1546).

[0138] In some embodiments, the stabilized secondary structure comprises the sequence CAGGAAACTG (SEQ ID NO:1547). In some embodiments, the stabilized secondary structure comprises the sequence AATTAGAAATAATT (SEQ ID NO:1548). In some embodiments, the stabilized secondary structure comprises the sequence CAGAAAATCTG (SEQ ID NO:1549). In some embodiments, the stabilized secondary structure comprises the sequence AATTAAAAATTAATT (SEQ ID NO:1550). In some embodiments, the stabilized secondary structure comprises the sequence CAGACTTCTG (SEQ ID NO:1551). In some embodiments, the stabilized secondary structure comprises the sequence AATTAACTTTAATT (SEQ ID NO:1552).

[0139] In some embodiments, the scaffold sequence comprises at least 1, at least 2, or at least 3 stem loops. In some embodiments, the scaffold sequence comprises at least three stem loops, wherein sequences of stem loops 1, 2, and 3 are according to Table 2. Table 2. Scaffold Sequence Stem Loops Combination # Stem loop 1 Stem loop 2 Stem loop 3 1a SEQ ID NO:1540 AAAAT ACTT 1b SEQ ID NO:1541 AAAAT ACTT 1c SEQ ID NO:1542 AAAAT ACTT 1d SEQ ID NO:1547 AAAAT ACTT 1e SEQ ID NO:1548 AAAAT ACTT 2a GAAA SEQ ID NO:1540 ACTT 2b GAAA SEQ ID NO:1543 ACTT 2c GAAA SEQ ID NO:1544 ACTT 2d GAAA SEQ ID NO:1549 ACTT 2e GAAA SEQ ID NO:1550 ACTT 3a GAAA AAAAT SEQ ID NO:1540 3b GAAA AAAAT SEQ ID NO:1545 3c GAAA AAAAT SEQ ID NO:1546 3d GAAA AAAAT SEQ ID NO:1551 3e GAAA AAAAT SEQ ID NO:1552 4a SEQ ID NO:1540 SEQ ID NO:1540 ACTT 4b SEQ ID NO:1541 SEQ ID NO:1543 ACTT 4c SEQ ID NO:1542 SEQ ID NO:1544 ACTT 4d SEQ ID NO:1547 SEQ ID NO:1549 ACTT 4e SEQ ID NO:1548 SEQ ID NO:1550 ACTT 5a SEQ ID NO:1540 AAAAT SEQ ID NO:1540 5b SEQ ID NO:1541 AAAAT SEQ ID NO:1545 5c SEQ ID NO:1542 AAAAT SEQ ID NO:1546 5d SEQ ID NO:1547 AAAAT SEQ ID NO:1551 5e SEQ ID NO:1548 AAAAT SEQ ID NO:1552 6a GAAA SEQ ID NO:1540 SEQ ID NO:1540 6b GAAA SEQ ID NO:1543 SEQ ID NO:1545 6c GAAA SEQ ID NO:1544 SEQ ID NO:1546 6d GAAA SEQ ID NO:1549 SEQ ID NO:1551 6e GAAA SEQ ID NO:1550 SEQ ID NO:1552 7a SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 7b SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 7c SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 7d SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 7e SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552

[0140] In some embodiments, the scaffold sequence (i) does not comprise an early termination signal sequence; and / or (ii) comprises a stabilized secondary structure. In some embodiments, the scaffold sequence comprises (a) SEQ ID NO:1553 at a 5’ end; and (b) at least three stem loops, wherein stem loops 1, 2, and 3 comprise any of combinations 1a-7e of Table 2. In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:28, 29, and 44-46. In some embodiments, each of SEQ ID NOs:28, 29, and 44-46 comprises SEQ ID NO:1553 at a 5’ end. In some embodiments, each of SEQ ID NOs:28, 29, and 44-46 comprises a stabilized secondary structure, e.g., a stem loop, as described herein.

[0141] In some embodiments, the scaffold sequence comprises (a) any one of SEQ ID NOs:1554-1560 at a 5’ end; and (b) at least three stem loops, wherein stem loops 1, 2, and 3 comprise the sequences GAAA, AAAAT, and ACTT, respectively. In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:21-27. In some embodiments, each of SEQ ID NOs:21-27 does not comprise an early termination signal sequence. In some embodiments, each of SEQ ID NOs:21-27 comprises three stem loops, wherein stem loops 1, 2, and 3 comprise the sequences GAAA, AAAAT, and ACTT, respectively.

[0142] In some embodiments, the scaffold sequence (i) does not comprise an early termination signal sequence; and (ii) comprises a stabilized secondary structure. In some embodiments, the scaffold sequence provides improved stability and / or expression of the gRNA as compared to a scaffold sequence characterized by only one of (i) and (ii) or a scaffold sequence characterized by neither of (i) and (ii).

[0143] In some embodiments, the scaffold sequence comprises (i) any one of SEQ ID NOs:1554-1560 at a 5’ end; and (ii) at least three stem loops, wherein stem loops 1, 2, and 3 comprise any of combinations 1a-7e of Table 2, e.g., according to Table 3: Table 3. Scaffold Sequences 5’ end sequence Stem loop 1 Stem loop 2 Stem loop 3 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:1554-1560 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:1554-1560 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:1554-1560 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:1554-1560 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:1554-1560 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:1554-1560 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:1554-1560 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:1554-1560 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552

[0144] In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:28- 43 or any one of SEQ ID NOs:47-95. In some embodiments, each of SEQ ID NOs:28-43 and 47-95 (i) does not comprise an early termination signal sequence and (ii) comprises a stabilized secondary structure, e.g., a stem loop, as described herein. In some embodiments, at least one of nucleotide positions 2-5 of each of SEQ ID NOs:28-43 and 47-95 is A, G, or C. In some embodiments, each of SEQ ID NOs:28-43 and 47-95 comprises at least three stem loops, wherein each stem loop independently comprises any one of SEQ ID NOs:1540-1552. In some embodiments, each of SEQ ID NOs:28-43 and 47-95 comprises at least three stem loops, wherein stem loop 1 comprises any one of SEQ ID NO:1540-1542, 1547, or 1548; stem loop 2 comprises any one of SEQ ID NOs:1540, 1543, 1544, 1549, or 1550; and / or stem loop 3 comprises any one of SEQ ID NOs:1540, 1545, 1546, 1551, or 1552.

[0145] In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:20- 95. In some embodiments, the scaffold sequence comprises any one of SEQ ID NOs:21-95. In some embodiments, the scaffold sequence comprises SEQ ID NO:20. In some embodiments, the scaffold sequence comprises SEQ ID NO:21. In some embodiments, the scaffold sequence comprises SEQ ID NO:22. In some embodiments, the scaffold sequence comprises SEQ ID NO:25. In some embodiments, the scaffold sequence comprises SEQ ID NO:28. In some embodiments, the scaffold sequence comprises SEQ ID NO:29. In some embodiments, the scaffold sequence comprises SEQ ID NO:30. In some embodiments, the scaffold sequence comprises SEQ ID NO:31. In some embodiments, the scaffold sequence comprises SEQ ID NO:44. In some embodiments, the scaffold sequence comprises SEQ ID NO:45. In some embodiments, the scaffold sequence comprises SEQ ID NO:46. In some embodiments, the scaffold sequence comprises SEQ ID NO:47. Polynucleotide comprising guide and scaffold sequences

[0146] In some embodiments, the polynucleotide of the present disclosure comprises a guide sequence described herein and a scaffold sequence described herein.

[0147] In some embodiments, the polynucleotide is a first polynucleotide, and the guide sequence is a first guide sequence. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence (i) does not comprise an early termination signal sequence; and / or (ii) comprises a stabilized secondary structure. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and nucleotide positions 2-5 of the scaffold sequence comprises at least one A, G, or C. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and a 5’ end of the scaffold sequence comprises any one of SEQ ID NOs:1554-1560. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises a stabilized secondary structure. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises at least three stem loops, wherein stem loops 1, 2, and 3 comprise any of combinations 1a-7e of Table 2. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises a sequence according to Table 3. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises any one of SEQ ID NOs:21-95. In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises any one of SEQ ID NOs:21, 22, 25, 28-31, and 44-47. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:1-12, and the scaffold sequence comprises any one of SEQ ID NOs:21-95. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:1-12, and the scaffold sequence comprises any one of SEQ ID NOs:21, 22, 25, 28-31, and 44-47.

[0148] In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2- 12, 13-19, and 1561-1576, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and the scaffold sequence is capable of binding a Cas protein, e.g., SpCas9 or SaCas9. In some embodiments, the scaffold sequence is capable of binding SaCas9. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and the scaffold sequence does not comprise an early termination signal sequence. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and the scaffold sequence comprises a stabilized secondary structure. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and the scaffold sequence (i) does not comprise an early termination signal sequence; and (ii) comprises a stabilized secondary structure. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and the scaffold sequence comprises any one of SEQ ID NOs:20-95.

[0149] In some embodiments, the polynucleotide is a second polynucleotide, and the guide sequence is a second guide sequence. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises SEQ ID NO:20. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence (i) does not comprise an early termination signal sequence; and / or (ii) comprises a stabilized secondary structure. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and nucleotide positions 2-5 of the scaffold sequence comprises at least one A, G, or C. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and a 5’ end of the scaffold sequence comprises any one of SEQ ID NOs:1554- 1560. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises a stabilized secondary structure. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346- 3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises at least three stem loops, wherein stem loops 1, 2, and 3 comprise any of combinations 1a-7e of Table 2. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises a sequence according to Table 3. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346- 3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises any one of SEQ ID NOs:21-95. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, e.g., any of the target sequences of Table 1, and the scaffold sequence comprises any one of SEQ ID NOs:21, 22, 25, 28-31, and 44-47. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence comprises any one of SEQ ID NOs:21-95. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence comprises any one of SEQ ID NOs:21, 22, 25, 28-31, and 44-47.

[0150] In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence is capable of binding a Cas protein, e.g., SpCas9 or SaCas9. In some embodiments, the scaffold sequence is capable of binding SaCas9. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence does not comprise an early termination signal sequence. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence comprises a stabilized secondary structure. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence (i) does not comprise an early termination signal sequence; and (ii) comprises a stabilized secondary structure. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, and the scaffold sequence comprises any one of SEQ ID NOs:20-95.

[0151] In some embodiments, the polynucleotide comprises: a guide sequence comprising any one of SEQ ID NOs:2-12, 13-19, and 1561-1576, any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and a scaffold sequence comprising any one of SEQ ID NOs:21- 95. In some embodiments, the polynucleotide comprises a sequence according to Table 4 or Table 5.

[0002] Table 4. Combinations of Guide and Scaffold Sequences Guide Sequence Scaffold Sequence (if the guide sequence comprises SEQ ID NO:2, the guide sequence does not comprise a 5’ guanine) 5’ end sequence Stem loop 1 Stem loop 2 Stem loop 3 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1547 SEQ ID NO:1549 ACTT

[0003] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1554 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1544 ACTT

[0004] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545

[0005] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 AAAAT SEQ ID NO:1540

[0006] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1555 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1550 ACTT

[0007] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551

[0008] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1556 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1542 AAAAT SEQ ID NO:1546

[0009] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1557 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA AAAAT SEQ ID NO:1545

[0010] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1558 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1540 AAAAT ACTT

[0011] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1548 AAAAT SEQ ID NO:1552

[0012] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1559 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1540 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1541 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1542 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1547 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1548 AAAAT ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA AAAAT SEQ ID NO:1551

[0013] any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1540 SEQ ID NO:1540 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1541 SEQ ID NO:1543 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1542 SEQ ID NO:1544 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1547 SEQ ID NO:1549 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1548 SEQ ID NO:1550 ACTT any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1540 AAAAT SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1541 AAAAT SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1542 AAAAT SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1547 AAAAT SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1548 AAAAT SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 GAAA SEQ ID NO:1550 SEQ ID NO:1552 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1540 SEQ ID NO:1540 SEQ ID NO:1540 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1541 SEQ ID NO:1543 SEQ ID NO:1545 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1542 SEQ ID NO:1544 SEQ ID NO:1546 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1547 SEQ ID NO:1549 SEQ ID NO:1551 any one of SEQ ID NOs:2-12, 13-19, and 1561-1576 SEQ ID NO:1560 SEQ ID NO:1548 SEQ ID NO:1550 SEQ ID NO:1552

[0014] Table 5. Combinations of Guide and Scaffold Sequences Guide Sequence Scaffold Sequence SEQ ID NO:2 (no 5’ guanine) any one of SEQ ID NOs:21-95 SEQ ID NO:3 any one of SEQ ID NOs:21-95 SEQ ID NO:4 any one of SEQ ID NOs:21-95 SEQ ID NO:5 any one of SEQ ID NOs:21-95 SEQ ID NO:6 any one of SEQ ID NOs:21-95 SEQ ID NO:7 any one of SEQ ID NOs:21-95 SEQ ID NO:8 any one of SEQ ID NOs:21-95 SEQ ID NO:9 any one of SEQ ID NOs:21-95 SEQ ID NO:10 any one of SEQ ID NOs:21-95 SEQ ID NO:11 any one of SEQ ID NOs:21-95 SEQ ID NO:12 any one of SEQ ID NOs:21-95 SEQ ID NO:13 any one of SEQ ID NOs:21-95 SEQ ID NO:14 any one of SEQ ID NOs:21-95 SEQ ID NO:15 any one of SEQ ID NOs:21-95 SEQ ID NO:16 any one of SEQ ID NOs:21-95 SEQ ID NO:17 any one of SEQ ID NOs:21-95 SEQ ID NO:18 any one of SEQ ID NOs:21-95 SEQ ID NO:19 any one of SEQ ID NOs:21-95 SEQ ID NO:1561 any one of SEQ ID NOs:21-95 SEQ ID NO:1562 any one of SEQ ID NOs:21-95 SEQ ID NO:1563 any one of SEQ ID NOs:21-95 SEQ ID NO:1564 any one of SEQ ID NOs:21-95 SEQ ID NO:1565 any one of SEQ ID NOs:21-95 SEQ ID NO:1566 any one of SEQ ID NOs:21-95 SEQ ID NO:1567 any one of SEQ ID NOs:21-95 SEQ ID NO:1568 any one of SEQ ID NOs:21-95 SEQ ID NO:1569 any one of SEQ ID NOs:21-95 SEQ ID NO:1571 any one of SEQ ID NOs:21-95 SEQ ID NO:1572 any one of SEQ ID NOs:21-95 SEQ ID NO:1573 any one of SEQ ID NOs:21-95 SEQ ID NO:1574 any one of SEQ ID NOs:21-95 SEQ ID NO:1575 any one of SEQ ID NOs:21-95 SEQ ID NO:1576 any one of SEQ ID NOs:21-95

[0152] In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:96- 1007. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:97-1007. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:97-171. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:172-247. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:97, 98, 101, 104-107, and 120-123. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:172-174, 177, 180-183, and 196-199. In some embodiments, the polynucleotide comprises SEQ ID NO:97. In some embodiments, the polynucleotide comprises SEQ ID NO:98. In some embodiments, the polynucleotide comprises SEQ ID NO:101. In some embodiments, the polynucleotide comprises SEQ ID NO:104. In some embodiments, the polynucleotide comprises SEQ ID NO:105. In some embodiments, the polynucleotide comprises SEQ ID NO:106. In some embodiments, the polynucleotide comprises SEQ ID NO:107. In some embodiments, the polynucleotide comprises SEQ ID NO:120. In some embodiments, the polynucleotide comprises SEQ ID NO:121. In some embodiments, the polynucleotide comprises SEQ ID NO:122. In some embodiments, the polynucleotide comprises SEQ ID NO:123. In some embodiments, the polynucleotide comprises SEQ ID NO:172. In some embodiments, the polynucleotide comprises SEQ ID NO:173. In some embodiments, the polynucleotide comprises SEQ ID NO:174. In some embodiments, the polynucleotide comprises SEQ ID NO:183. In some embodiments, the polynucleotide comprises SEQ ID NO:196. In some embodiments, the polynucleotide comprises SEQ ID NO:199. In some embodiments, the polynucleotide comprises SEQ ID NO:248. In some embodiments, the polynucleotide comprises SEQ ID NO:324. In some embodiments, the polynucleotide comprises SEQ ID NO:400. In some embodiments, the polynucleotide comprises SEQ ID NO:476. In some embodiments, the polynucleotide comprises SEQ ID NO:552. In some embodiments, the polynucleotide comprises SEQ ID NO:628. In some embodiments, the polynucleotide comprises SEQ ID NO:704. In some embodiments, the polynucleotide comprises SEQ ID NO:780. In some embodiments, the polynucleotide comprises SEQ ID NO:856. In some embodiments, the polynucleotide comprises SEQ ID NO:932. In some embodiments, the polynucleotide comprising any one of SEQ ID NOs:96- 1007 is a first polynucleotide, e.g., first gRNA.

[0153] In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:1008- 1539. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:1008- 1083. In some embodiments, the polynucleotide comprises any one of SEQ ID NOs:1008- 1010, 1013, 1016-1019, and 1032-1035. In some embodiments, the polynucleotide comprises SEQ ID NO:1008. In some embodiments, the polynucleotide comprises SEQ ID NO:1009. In some embodiments, the polynucleotide comprises SEQ ID NO:1010. In some embodiments, the polynucleotide comprises SEQ ID NO:1013. In some embodiments, the polynucleotide comprises SEQ ID NO:1016. In some embodiments, the polynucleotide comprises SEQ ID NO:1017. In some embodiments, the polynucleotide comprises SEQ ID NO:1018. In some embodiments, the polynucleotide comprises SEQ ID NO:1019. In some embodiments, the polynucleotide comprises SEQ ID NO:1032. In some embodiments, the polynucleotide comprises SEQ ID NO:1035. In some embodiments, the polynucleotide comprises SEQ ID NO:1084. In some embodiments, the polynucleotide comprises SEQ ID NO:1160. In some embodiments, the polynucleotide comprises SEQ ID NO:1236. In some embodiments, the polynucleotide comprises SEQ ID NO:1312. In some embodiments, the polynucleotide comprises SEQ ID NO:1388. In some embodiments, the polynucleotide comprises SEQ ID NO:1464. In some embodiments, the polynucleotide comprising any one of SEQ ID NOs:1008-1539 is a second polynucleotide, e.g., second gRNA.

[0154] In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1561 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1562 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1563 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1564 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1565 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1566 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1567 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1568 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1569 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1570 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1571 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1572 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1573 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1574 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1575 and any one of SEQ ID NOs:20-95. In some embodiments, the polynucleotide comprises a combination of SEQ ID NO:1576 and any one of SEQ ID NOs:20-95. As used herein, a combination of two sequences comprises the two sequences joined end-to-end (i.e., the 3’ end of the first indicated sequence joined directly to the 5’ end of the second indicated sequence). In some embodiments, the polynucleotide comprising the combination of any one of SEQ ID NOs:1561-1576 with any one of SEQ ID NOs:20-95 is a second polynucleotide, e.g., second gRNA. Expression System and Composition

[0155] In some embodiments, the disclosure provides an expression system comprising a nucleic acid sequence encoding one or more of the polynucleotides described herein. In some embodiments, the expression system comprises a first nucleic acid sequence encoding the first polynucleotide (e.g., first gRNA) described herein. In some embodiments, the expression system comprises a second nucleic acid sequence encoding the second polynucleotide (e.g., second gRNA) described herein. In some embodiments, the first polynucleotide comprises a first guide sequence and a first scaffold sequence as described herein. In some embodiments, the second polynucleotide comprises a second guide sequence and a second scaffold sequence as described herein. In some embodiments, the expression system further comprises a third nucleic acid sequence encoding an RNA-guided nuclease described herein. In some embodiments, the third nucleic acid sequence encodes a Cas protein. In some embodiments, the Cas protein is Cas9. In some embodiments, the Cas protein is capable of forming a complex with the first polynucleotide. In some embodiments, the Cas protein is capable of forming a complex with the second polynucleotide. In some embodiments, the Cas protein is capable of binding to the first scaffold sequence. In some embodiments, the Cas protein is capable of binding to the second scaffold sequence. In some embodiments, the Cas protein is SaCas9 or SpCas9. In some embodiments, the Cas protein is SaCas9.

[0156] In some embodiments, the disclosure provides a composition comprising an RNA- guided nuclease and one or more polynucleotides described herein. In some embodiments, the disclosure provides a composition comprising a Cas protein and one or more polynucleotides described herein. In some embodiments, the composition comprises a Cas protein and the first polynucleotide (e.g., first gRNA) described herein. In some embodiments, the composition comprises a Cas protein and the second polynucleotide (e.g., second gRNA) described herein. In some embodiments, the first polynucleotide comprises a first guide sequence and a first scaffold sequence. In some embodiments, the second polynucleotide comprises a second guide sequence and a second scaffold sequence. In some embodiments, the Cas protein is capable of binding to the first scaffold sequence, thereby forming a complex with the first polynucleotide. In some embodiments, the Cas protein is capable of binding to the second scaffold sequence, thereby forming a complex with the second polynucleotide. In some embodiments, the Cas protein is Cas9. In some embodiments, the Cas protein is SaCas9 or SpCas9. In some embodiments, the Cas protein is SaCas9. In some embodiments, the composition comprises a Cas protein and an expression system described herein.

[0157] In some embodiments, the first guide sequence targets human genome coordinates chr4:3078423-3078443 as described herein. In some embodiments, the first guide sequence comprises any one of SEQ ID NOs:1-12, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3072397-3074774 as described herein. In some embodiments, the second guide sequence targets a target sequence of Table 1. In some embodiments, the second guide sequence targets a sequence within human genome coordinates chr4:3074430-3074450 as described herein. In some embodiments, the second guide sequence comprises any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576. In some embodiments, the second guide sequence comprises SEQ ID NO:13. In some embodiments, the first guide sequence and / or the second guide sequence does not form a secondary structure. In some embodiments, the first guide sequence and / or the second guide sequence does not form a stem loop.

[0158] In some embodiments, the first and second scaffold sequences are each capable of binding to a Cas protein. In some embodiments, the Cas protein is Cas9, e.g., SaCas9 or SpCas9. In some embodiments, the first and second scaffold sequences are each capable of binding to the Cas protein of the composition. In some embodiments, the first scaffold sequence and the second scaffold sequence each independently (i) does not comprise an early termination signal sequence and (ii) comprises a stabilized secondary structure. In some embodiments, the early termination signal sequence comprises about 2 to about 8, or about 3 to about 7, or about 4 to about 6 consecutive thymine bases. In some embodiments, a 5’ end of the first scaffold sequence and / or the second scaffold sequence comprises any one of SEQ ID NOs:1554-1560. In some embodiments, the stabilized secondary structure comprises a locked loop. In some embodiments, the stabilized secondary structure comprises any one of SEQ ID NOs:1540-1552. In some embodiments, the first scaffold sequence and the second scaffold sequence each independently comprises at least three stem loops, wherein stem loops 1, 2, and 3 comprise any of combinations 1a-7e of Table 2. In some embodiments, the first scaffold sequence and the second scaffold sequence each independently comprises a 5’ end sequence and stem loop sequences as shown in Table 3. In some embodiments, the first scaffold sequence and the second scaffold sequence each independently comprises any one of SEQ ID NOs:20-95. In some embodiments, the first scaffold sequence and the second scaffold sequence are identical. In some embodiments, In some embodiments, the first scaffold sequence and the second scaffold sequence are different.

[0159] In some embodiments, the expression system comprises a first nucleic acid sequence encoding a first polynucleotide as described herein. In some embodiments, the composition comprises a first polynucleotide as described herein.

[0160] In some embodiments, the first polynucleotide encoded by the first nucleic acid sequence of the expression system, or the first polynucleotide of the composition comprises: (i) a first guide sequence targeting human genome coordinates chr4:3078423-3078443; and (ii) a first scaffold sequence comprising any one of SEQ ID NOs:21-95. In some embodiments, the first polynucleotide comprises: (i) a first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and optionally (ii) a first scaffold sequence.

[0161] In some embodiments, the first polynucleotide encoded by the first nucleic acid sequence of the expression system, or the first polynucleotide of the composition comprises: (i) a first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and (ii) a first scaffold sequence comprising any one of SEQ ID NOs:20-95. In some embodiments, the first polynucleotide comprises: (i) a first guide sequence comprising any one of SEQ ID NOs:1- 12; and (ii) a first scaffold sequence comprising any one of SEQ ID NOs:21-95.

[0162] In some embodiments, the expression system comprises a second nucleic acid sequence encoding a second polynucleotide as described herein. In some embodiments, the composition comprises a second polynucleotide as described herein.

[0163] In some embodiments, the second polynucleotide encoded by the second nucleic acid sequence of the expression system, or the second polynucleotide of the composition comprises: (i) a second guide sequence targeting a sequence within human genome coordinates chr4:3068346-3074862; and (ii) a second scaffold sequence comprising any one of SEQ ID NOs:21-95. In some embodiments, the second polynucleotide comprises: (i) a second guide sequence comprising any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576; and optionally (ii) a second scaffold sequence.

[0164] In some embodiments, the second polynucleotide encoded by the second nucleic acid sequence of the expression system, or the second polynucleotide of the composition comprises: (i) a second guide sequence comprising any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576; and (ii) a second scaffold sequence comprising any one of SEQ ID NOs:20-95.

[0165] In some embodiments, the expression system comprises (a) a first nucleic acid sequence encoding a first polynucleotide and (b) a second nucleic acid sequence encoding a second polynucleotide as described herein. In some embodiments, the composition comprises (a) a first polynucleotide and (b) a second polynucleotide as described herein. In some embodiments, the expression system further comprises a third nucleic acid sequence encoding a Cas protein. In some embodiments, the composition further comprises a Cas protein. In some embodiments, the Cas protein is SaCas9.

[0166] In some embodiments, (a) the first polynucleotide comprises (i) a first guide sequence targeting human genome coordinates chr4:3078423-3078443; and (ii) a first scaffold sequence comprising any one of SEQ ID NOs:21-95; and (b) the second polynucleotide comprises (i) a second guide sequence targeting a sequence within human genome coordinates chr4:3068346-3074862, and (ii) a second scaffold sequence comprising any one of SEQ ID NOs:21-95.

[0167] In some embodiments, (a) the first polynucleotide comprises (i) a first guide sequence targeting human genome coordinates chr4:3078423-3078443; and (ii) a first scaffold sequence comprising any one of SEQ ID NOs:21-95; and (b) the second polynucleotide comprises (i) a second guide sequence comprising any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576; and optionally (ii) a second scaffold sequence.

[0168] In some embodiments, (a) the first polynucleotide comprises (i) a first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and optionally (ii) a first scaffold sequence; and (b) the second polynucleotide comprises (i) a second guide sequence targeting a sequence within human genome coordinates chr4:3068346-3074862; and (ii) a second scaffold sequence comprising any one of SEQ ID NOs:21-95.

[0169] In some embodiments, (a) the first polynucleotide comprises (i) a first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and optionally (ii) a first scaffold sequence; and (b) the second polynucleotide comprises (i) a second guide sequence comprising any one of SEQ ID NOs:13-19 or any one of SEQ ID NOs:1561-1576; and optionally (ii) a second scaffold sequence.

[0170] In some embodiments, the first polynucleotide encoded by the first nucleic acid sequence of the expression vector, or the first polynucleotide of the composition comprises any one of SEQ ID NOs:96-1007. In some embodiments, first polynucleotide comprises any one of SEQ ID NOs:97-1007. In some embodiments, the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199. In some embodiments, the first polynucleotide comprises any one of SEQ ID NOs:97, 98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199. In some embodiments, the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932. In some embodiments, the first polynucleotide comprises any one of SEQ ID NOs:97, 98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

[0171] In some embodiments, the second polynucleotide encoded by the second nucleic acid sequence of the expression vector, or the second polynucleotide of the composition comprises any one of SEQ ID NOs:1008-1539. In some embodiments, the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035. In some embodiments, the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464. In some embodiments, the second polynucleotide comprises a combination of any one of SEQ ID NOs:1561-1576 with any one of SEQ ID NOs:20-95.

[0172] In some embodiments, the first and second polynucleotides encoded respectively by the first and second nucleic acid sequences of the expression system, or the first and second polynucleotides of the composition, comprise the sequences according to Table 6. Table 6. Combinations of First and Second Polynucleotides First Polynucleotide Second Polynucleotide Any one of SEQ ID NOs:96-1007 SEQ ID NO:1008 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1009 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1010 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1013 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1016 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1017 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1018 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1019 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1032 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1035 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1084 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1160 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1236 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1312 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1388 Any one of SEQ ID NOs:96-1007 SEQ ID NO:1464 SEQ ID NO:96 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:97 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:98 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:101 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:104 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:105 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:106 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:107 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:120 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:121 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:122 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:123 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:172 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:173 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:174 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:183 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:196 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:199 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:248 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:324 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:400 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:476 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:552 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:628 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:704 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:780 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:856 Any one of SEQ ID NOs:1008-1539 SEQ ID NO:932 Any one of SEQ ID NOs:1008-1539

[0173] In some embodiments, the first polynucleotide comprises SEQ ID NO:174 or SEQ ID NO:199, and the second polynucleotide comprises SEQ ID NO:1010 or SEQ ID NO:1013. In some embodiments, the first polynucleotide comprises SEQ ID NO:174, and the second polynucleotide comprises SEQ ID NO:1010. In some embodiments, the first polynucleotide comprises SEQ ID NO:174, and the second polynucleotide comprises SEQ ID NO:1013. In some embodiments, the first polynucleotide comprises SEQ ID NO:199, and the second polynucleotide comprises SEQ ID NO:1010. In some embodiments, the first polynucleotide comprises SEQ ID NO:199, and the second polynucleotide comprises SEQ ID NO:1013.

[0174] In some embodiments, the expression system comprises a vector. In some embodiments, the composition comprises a vector. In some embodiments, the first nucleic acid sequence, the second nucleic acid sequence, the third nucleic acid sequence, or any combination thereof, are on the vector. In some embodiments, the vector is an expression vector. In some embodiments, the vector is a viral vector. In some embodiments, the vector is a non-viral vector. In some embodiments, the vector is a bacterial expression vector. In some embodiments, the vector is a mammalian expression vector. In some embodiments, the vector is a human expression vector. Exemplary vectors are described herein.

[0175] In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a vector from retrovirus, adeno-associated virus, pox, baculovirus, vaccinia, herpes simplex, Epstein-Barr virus, adenovirus, geminivirus, or caulimovirus. In some embodiments, the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral (AAV) vector. Methods of introducing vectors, e.g., viral vectors, into cells are described herein. In some embodiments, viral transduction with adenoviral, AAV, and lentiviral vectors is a delivery method for in vivo gene therapy.

[0176] In some embodiments, the vector further comprises a regulatory element operably linked to the first polynucleotide, the second polynucleotide, the third polynucleotide, or combinations thereof. In some embodiments, the regulatory element comprises a promoter, an enhancer, a terminator, a 5’ UTR, a 3’ UTR, or combination thereof. Regulatory elements are described herein. In some embodiments, the regulatory element comprises a promoter. In some embodiments, the promoter is a bacterial promoter. In some embodiments, the promoter is a viral promoter. In some embodiments, the promoter is a mammalian promoter.

[0177] In some embodiments, the regulatory element comprises an RNA stabilizing sequence. In some embodiments, the RNA stabilizing sequence is located 3’ of the first, second, and / or third polynucleotide on the vector. In some embodiments, the RNA stabilizing sequence is a posttranscriptional regulatory element (PRE). In some embodiments, the RNA stabilizing sequence is a woodchuck hepatitis virus posttranscriptional regulatory element (WPRE).

[0178] In some embodiments, the expression system comprises the first nucleic acid sequence and the second nucleic acid sequence. In some embodiments, the first nucleic acid sequence and the second nucleic acid sequence are on a single vector. In some embodiments, the first nucleic acid sequence and the second nucleic acid sequence are on separate vectors. In some embodiments, the expression system comprises the first nucleic acid sequence and the third nucleic acid sequence. In some embodiments, the first nucleic acid sequence and the third nucleic acid sequence are on a single vector. In some embodiments, the first nucleic acid sequence and the third nucleic acid sequence are on separate vectors. In some embodiments, the expression system comprises the second nucleic acid sequence and the third nucleic acid sequence. In some embodiments, the second nucleic acid sequence and the third nucleic acid sequence are on a single vector. In some embodiments, the second nucleic acid sequence and the third nucleic acid sequence are on separate vectors. In some embodiments, each nucleic acid sequence of the expression system is operably linked to a distinct regulatory element. In some embodiments, one regulatory element is operably linked to more than one nucleic acid sequence of the expression system.

[0179] In some embodiments, the expression system comprises the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence. In some embodiments, the first nucleic acid sequence, the second nucleic acid sequence, and the third nucleic acid sequence are on a single vector. In some embodiments, the first and second nucleic acid sequences are on a first vector, and the third nucleic acid sequence is on a second vector. In some embodiments, the first and third nucleic acid sequences are on a first vector, and the second nucleic acid sequence is on a second vector. In some embodiments, the second and third nucleic acid sequences are on a first vector, and the first nucleic acid sequence is on a second vector. In some embodiments, each of the first, second, and third nucleic acid sequences is on a separate vector. In some embodiments, each nucleic acid sequence of the expression system is operably linked to a distinct regulatory element. In some embodiments, one regulatory element is operably linked to more than one nucleic acid sequence of the expression system.

[0180] In some embodiments, the composition comprises the Cas protein, the first polynucleotide, and the second polynucleotide described herein. In some embodiments, the composition comprises the Cas protein and the first polynucleotide described herein. In some embodiments, the composition comprises the Cas protein and the second polynucleotide described herein. In some embodiments, the first polynucleotide of the composition is encoded by a first nucleic acid sequence. In some embodiments, the second polynucleotide of the composition is encoded by a second nucleic acid sequence. In some embodiments, the first nucleic acid sequence and / or the second nucleic acid sequence is on a vector. Exemplary vectors are provided herein. In some embodiments, the vector is a viral vector. In some embodiments, the viral vector is a lentiviral vector, an adenoviral vector, or an adeno- associated viral vector. Cellular Delivery

[0181] In some embodiments, the disclosure provides a delivery particle comprising a polynucleotide described herein, an expression system described herein, a composition described herein, or combination thereof. In some embodiments, the polynucleotide is a first polynucleotide (e.g., first gRNA) described herein or a second polynucleotide (e.g., second gRNA) described herein. In some embodiments, the expression system comprises a first nucleic acid sequence encoding a first polynucleotide described herein; and / or a second nucleic acid sequence encoding a second polynucleotide described herein. In some embodiments, the composition comprises a first polynucleotide described herein; and / or a second polynucleotide described herein. In some embodiments, the first polynucleotide comprises a first guide sequence and a first scaffold sequence. In some embodiments, the second polynucleotide comprises a second guide sequence and a second scaffold sequence. In some embodiments, the expression system further comprises a third nucleic acid sequence encoding an RNA-guided nuclease described herein, e.g., a Cas protein such as SaCas9. In some embodiments, the composition further comprises an RNA-guided nuclease described herein, e.g., a Cas protein such as SaCas9.

[0182] Delivery particles for delivering biological components, e.g., CRISPR systems including the polynucleotides, expression systems, and / or compositions described herein, are known to one of ordinary skill in the art. Delivery particles may be in any form, including but not limited to: solid, semi-sold, emulsion, or colloidal particles. In some embodiments, the delivery particle is a lipid-based particle, a virus-like particle, a liposome, a micelle, a vesicle or microvesicle, an exosome, or a lipid nanoparticle. Delivery particles are further described, e.g., in US 2008 / 0234183, US 2011 / 0293703, US 2012 / 0251560, US 2013 / 0302401, US 2019 / 0167810, US 2020 / 0207833, US 5,543,158, US 5,855,913, US 5,895,309, US 6,007,845, and US 8,709,843.

[0183] In some embodiments, the polynucleotide, expression system, and / or composition described herein are comprised in a single delivery particle. In some embodiments, the polynucleotide, expression system, and / or composition described herein are comprised in multiple delivery particles. In some embodiments, the delivery particle is configured to deliver the polynucleotide, expression system, composition, or combination thereof into a cell.

[0184] In some embodiments, the disclosure provides a cell comprising a polynucleotide described herein, an expression system described herein, a composition described herein, a delivery particle described herein, or a combination thereof. In some embodiments, the polynucleotide is a first polynucleotide (e.g., first gRNA) described herein or a second polynucleotide (e.g., second gRNA) described herein. In some embodiments, the expression system comprises a first nucleic acid sequence encoding a first polynucleotide described herein; and / or a second nucleic acid sequence encoding a second polynucleotide described herein. In some embodiments, the composition comprises a first polynucleotide described herein; and / or a second polynucleotide described herein. In some embodiments, the first polynucleotide comprises a first guide sequence and a first scaffold sequence. In some embodiments, the second polynucleotide comprises a second guide sequence and a second scaffold sequence. In some embodiments, the expression system further comprises a third nucleic acid sequence encoding an RNA-guided nuclease described herein, e.g., a Cas protein such as SaCas9. In some embodiments, the composition further comprises an RNA-guided nuclease described herein, e.g., a Cas protein such as SaCas9.

[0185] In some embodiments, the cell is a bacterial cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is an animal cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is from an animal or human cell line. Examples of animal or human cells and cell lines include, but are not limited to, NSO, CHO, HT1080, H9, HepG2, MCF7, MDBK Jurkat, NIH3T3, PC12, BHK, EBX, EB14, EB24, EB26, EB66, Ebvl3, VERO, SP2 / 0, YB2 / 0, Y0, C127, L cell, COS (e.g., COS1 and COS7), QC1-3, VERO, PER.C6, HEK293, HeLa, and NT2. In some embodiments, the cell is a neuronal cell. In some embodiments, the cell is a neuronal precursor cell line. In some embodiments, the cell is from an animal or human cellular model for Huntington’s Disease (HD). In some embodiments, the cell is derived from a human subject having HD. In some embodiments, the cell is a fibroblast cell, e.g., derived from a human subject having HD. In some embodiments, the cell is a human stem cell, e.g., an induced pluripotent stem cell (iPSC), an embryonic stem cell (ESC), a tissue specific stem cell (e.g., neural stem cell), or a mesenchymal stem cell (MSC). In some embodiments, the cell is a neuronal cell differentiated from a stem cell described herein. Exemplary cells and cell lines for modeling HD are further described, e.g., in Szlachcic et al., Mol Neurosci 10:253, 2017; Hung et al., Mol Biol Cell 29(23):2809-2820, 2018; and Le Cann et al., Sci Rep 11:6934, 2021. Further exemplary cell lines for modeling HD include, but are not limited to, Huntingtin 150Q Stable PC12 Cell Line (Cat. No. T6018) from Applied Biological Materials Inc.; and cell lines GM04281, GM04282, AND GM23225 in the NIGMS Human Genetic Cell Repository, available from the Coriell Institute for Medical Research. Methods

[0186] In some embodiments, the disclosure provides a method of reducing CAG repeats in a mutant allele of a huntingtin (HTT) gene of a cell, comprising introducing to the cell: the polynucleotide, expression system, composition, or delivery particle described herein, or a combination thereof. In some embodiments, the disclosure provides a method of treating HD in a subject in need thereof, comprising administering to the subject: the polynucleotide, expression system, composition, or delivery particle described herein, or a combination thereof. In some embodiments, the subject comprises a heterozygous HD allele, i.e., an HTT gene comprising a mutant allele and a wild-type allele.

[0187] In some embodiments, the first polynucleotide (e.g., first gRNA), second polynucleotide (e.g., second gRNA), and Cas protein described herein are present in the cell following the introducing. In some embodiments, the first polynucleotide (e.g., first gRNA), second polynucleotide (e.g., second gRNA), and Cas protein described herein are present in the subject following the administering. In some embodiments, the Cas protein is SaCas9. In some embodiments, the first polynucleotide comprises a first guide sequence and first scaffold sequence as described herein. In some embodiments, the second polynucleotide comprises a second guide sequence and second scaffold sequence as described herein. In some embodiments, the first guide sequence targets a first target sequence, i.e., chr4:3078423-3078443 as described herein. In some embodiments, the second guide sequence targets a second target sequence, i.e., a sequence within human genome coordinates chr4:3068346-3074862 as described herein.

[0188] In some embodiments, the HTT gene of the cell and / or the subject comprises a mutant allele and a wild-type allele. In some embodiments, the mutant allele comprises a PAM (e.g., for SaCas9) adjacent to the first target sequence, and the wild-type allele does not comprise the PAM adjacent to the first target sequence. In some embodiments, a first complex comprising the first polynucleotide and the Cas protein, e.g., SaCas9, is guided to the first target sequence by the first guide sequence of the first polynucleotide. In some embodiments, the Cas protein, e.g., SaCas9, recognizes the PAM adjacent to the first target sequence in the mutant allele and cleaves the mutant allele at the first target sequence. In some embodiments, the Cas protein, e.g., SaCas9, does not cleave the wild-type allele at the first target sequence due to the lack of a PAM.

[0189] In some embodiments, both the wild-type allele and the mutant allele comprise a PAM (e.g., for SaCas9) adjacent to the second target sequence. In some embodiments, a second complex comprising the second polynucleotide and the Cas protein, e.g., SaCas9, is guided to the second target sequence by the second guide sequence of the second polynucleotide. In some embodiments, the Cas protein, e.g., SaCas9, recognizes the PAM adjacent to the second target sequence in both the wild-type and mutant alleles, and cleaves both the wild-type and mutant alleles at the second target sequence.

[0190] In some embodiments, the mutant allele is cleaved at two sites (i.e., the first target sequence and the second target sequence), thereby excising the region between the two sites. In some embodiments, the region between the two sites comprises the expanded CAG repeats associated with HD. In some embodiments, the CAG repeats region of the wild-type allele is not excised due to only one site of cleavage (i.e., the second target sequence). In some embodiments, the cleaved second target sequence in the wild-type allele is ligated and / or repaired, e.g., via native cellular repair pathways such as homology-directed repair or non- homologous end joining. See, e.g., FIG. 1.

[0191] In some embodiments, the mutant allele of the cell and / or the subject comprises greater than 27 CAG repeats prior to the introducing into the cell and / or the administering into the cell. In some embodiments, the mutant allele of the cell and / or the subject comprises greater than 35 CAG repeats prior to the introducing and / or the administering. In some embodiments, the mutant allele of the cell and / or the subject comprises fewer than 35 CAG repeats following the introducing and / or the administering. In some embodiments, the mutant allele of the cell and / or the subject comprises fewer than 27 CAG repeats following the introducing and / or the administering. In some embodiments, following the introducing and / or the administering, all CAG repeats of the mutant allele are removed. In some embodiments, following the introducing and / or the administering, exon 1 of the mutant allele is removed.

[0192] In some embodiments, the method provided herein reduces levels of RNA encoded by the mutant allele by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%. In some embodiments, the method reduces levels of RNA encoded by the mutant allele by at least 30%. In some embodiments, the method reduces levels of RNA encoded by the mutant allele by at least 50%. In some embodiments, the method reduces levels of protein expressed from the mutant allele by at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95%. In some embodiments, the method reduces levels of protein expressed from the mutant allele by at least 30%. In some embodiments, the method reduces levels of protein expressed from the mutant allele by at least 50%.

[0193] In some embodiments, the method provided herein does not reduce levels of RNA encoded by the wild-type allele and / or levels of protein expressed from the wild-type allele by more than 10%, more than 20%, more than 30%, more than 40%, more than 50%, more than 60%, or more than 70%. In some embodiments, the method does not reduce levels of RNA encoded by the wild-type allele and / or levels of protein expressed from the wild-type allele by more than 50%. In some embodiments, the method (1) reduces levels of RNA encoded by the mutant allele and / or levels of protein expressed from the mutant allele by at least 30%; and (2) does not reduce levels of RNA encoded by the wild-type allele and / or levels of protein expressed from the wild-type allele by more than 50%. In some embodiments, the method (1) reduces levels of RNA encoded by the mutant allele and / or levels of protein expressed from the mutant allele by at least 50%; and (2) does not reduce levels of RNA encoded by the wild-type allele and / or levels of protein expressed from the wild-type allele by more than 50%.

[0194] The entire contents of all publications, patents, and patent applications referenced herein are hereby incorporated herein by reference.

[0195] The specific examples included herein are for illustrative purposes only and are not to be considered as limiting to this disclosure. Moreover, the compositions and methods provided herein have been described in relation to certain embodiments thereof, and many details have been set forth for purposes of illustration. It will be apparent to those skilled in the art that the disclosure is susceptible to additional embodiments and that certain details described herein may be varied without departing from the basic principles of the disclosure. EMBODIMENTS

[0196] Embodiment 1. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence targets human genome coordinates chr4:3078423- 3078443, and the scaffold sequence comprises any one of SEQ ID NOs:21-95.

[0197] Embodiment 2. The polynucleotide of embodiment 1, wherein the guide sequence does not form a secondary structure.

[0198] Embodiment 3. The polynucleotide of embodiment 2, wherein the secondary structure is a stem loop.

[0199] Embodiment 4. The polynucleotide of any one of embodiments 1 to 3, wherein the guide sequence comprises any one of SEQ ID NOs:1-12.

[0200] Embodiment 5. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence comprises any one of SEQ ID NOs:2-12, and wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5' guanine.

[0201] Embodiment 6. The polynucleotide of embodiment 5, wherein the guide sequence comprises SEQ ID NO:2 and does not comprise a 5' guanine.

[0202] Embodiment 7. The polynucleotide of embodiment 5 or 6, wherein the scaffold sequence is capable of binding to a Cas protein.

[0203] Embodiment 8. The polynucleotide of embodiment 7, wherein the Cas protein is Cas9 from Staphylococcus aureus (SaCas9).

[0204] Embodiment 9. The polynucleotide of any one of embodiments 1 to 8, wherein the scaffold sequence does not comprise an early termination signal sequence.

[0205] Embodiment 10. The polynucleotide of embodiment 9, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

[0206] Embodiment 11. The polynucleotide of any one of embodiments 1 to 10, wherein the scaffold sequence comprises a stabilized secondary structure.

[0207] Embodiment 12. The polynucleotide of embodiment 11, wherein the stabilized secondary structure comprises a locked loop.

[0208] Embodiment 13. The polynucleotide of any one of embodiments 1 to 12, wherein the scaffold sequence comprises any one of SEQ ID NOs:20-95.

[0209] Embodiment 14. The polynucleotide of embodiment 13, wherein the scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

[0210] Embodiment 15. A polynucleotide comprising any one of SEQ ID NOs:97-1007.

[0211] Embodiment 16. The polynucleotide of embodiment 15, comprising any one of SEQ ID NOs:97, 98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199.

[0212] Embodiment 17. The polynucleotide of embodiment 15, comprising any one of SEQ ID NOs: 97, 98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

[0213] Embodiment 18. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein: (i) the guide sequence comprises any one of SEQ ID NOs:2-12, and the scaffold sequence comprises any one of SEQ ID NOs:20-95, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5' guanine; or (ii) the guide sequence comprises any one of SEQ ID NOs:1-12, and the scaffold sequence comprises any one of SEQ ID NOs:21-95.

[0214] Embodiment 19. An expression system comprising a nucleic acid sequence encoding the polynucleotide of embodiments 1-18.

[0215] Embodiment 20. The expression system of embodiment 19, wherein the nucleic acid sequence is a first nucleic acid sequence, the polynucleotide is a first polynucleotide, the guide sequence is a first guide sequence, and the scaffold sequence is a first scaffold sequence; and wherein the expression system further comprises a second nucleic acid sequence encoding a second polynucleotide, wherein the second polynucleotide comprises a second guide sequence and a second scaffold sequence, wherein the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862.

[0216] Embodiment 21. The expression system of embodiment 20, wherein the second guide sequence does not form a secondary structure.

[0217] Embodiment 22. The expression system of embodiment 21, wherein the secondary structure is a stem loop.

[0218] Embodiment 23. The expression system of any one of embodiments 20 to 22, wherein the second guide sequence comprises any one of SEQ ID NOs:13-19.

[0219] Embodiment 24. The expression system of embodiment 23, wherein the second guide sequence comprises SEQ ID NO:13.

[0220] Embodiment 25. The expression system of any one of embodiments 20 to 24, wherein the second scaffold sequence is capable of binding to a Cas protein.

[0221] Embodiment 26. The expression system of embodiment 25, wherein the Cas protein is SaCas9.

[0222] Embodiment 27. The expression system of any one of embodiments 20 to 26, wherein the second scaffold sequence does not comprise an early termination signal sequence.

[0223] Embodiment 28. The expression system of embodiment 27, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

[0224] Embodiment 29. The expression system of any one of embodiments 20 to 28, wherein the second scaffold sequence comprises a stabilized secondary structure.

[0225] Embodiment 30. The expression system of embodiment 29, wherein the stabilized secondary structure comprises a locked loop.

[0226] Embodiment 31. The expression system of any one of embodiments 20 to 30, wherein the second scaffold sequence comprises any one of SEQ ID NOs:20-95.

[0227] Embodiment 32. The expression system of any one of embodiments 20 to 30, wherein the second scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44- 47.

[0228] Embodiment 33. The expression system of any one of embodiments 20 to 32, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1539.

[0229] Embodiment 34. The expression system of embodiment 33, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032- 1035.

[0230] Embodiment 35. The expression system of embodiment 33, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0231] Embodiment 36. An expression system comprising: a first nucleic acid sequence encoding a first polynucleotide comprising any one of SEQ ID NOs:96-1007; and a second nucleic acid sequence encoding a second polynucleotide comprising any one of SEQ ID NOs:1008-1539.

[0232] Embodiment 37. The expression system of embodiment 36, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199.

[0233] Embodiment 38. The expression system of embodiment 36 or 37, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

[0234] Embodiment 39. The expression system of any one of embodiments 36 to 38, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035.

[0235] Embodiment 40. The expression system of any one of embodiments 36 to 39, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0236] Embodiment 41. The expression system of any one of embodiments 19 to 40, wherein the expression system comprises a vector.

[0237] Embodiment 42. The expression system of embodiment 41, wherein the vector is a viral vector.

[0238] Embodiment 43. The expression system of embodiment 42, wherein the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector.

[0239] Embodiment 44. The expression system of any one of embodiments 20 to 43, wherein the first nucleic acid sequence and the second nucleic acid sequence are on a single vector.

[0240] Embodiment 45. The expression system of any one of embodiments 20 to 43, wherein each of the first nucleic acid sequence and the second nucleic acid sequence is on a separate vector.

[0241] Embodiment 46. The expression system of embodiment 19, wherein the nucleic acid sequence is a first nucleic acid sequence, the polynucleotide is a first polynucleotide, and wherein the expression system further comprises a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide.

[0242] Embodiment 47. The expression system of embodiment 46, wherein the third nucleic acid sequence is on a separate vector from the first nucleic acid sequence .

[0243] Embodiment 48. The expression system of embodiment 46, wherein the first nucleic acid sequence and the third nucleic acid sequence are on a single vector.

[0244] Embodiment 49. The expression system of any one of embodiments 20 to 45, further comprising a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide and / or the second polynucleotide.

[0245] Embodiment 50. The expression system of embodiment 49, wherein the first, second, and third nucleic acid sequences are on a single vector.

[0246] Embodiment 51. The expression system of embodiment 49, wherein the first and second nucleic acid sequences are on a first vector, and the third nucleic acid sequence is on a second vector; or wherein the first and third nucleic acid sequences are on a first vector, and the second nucleic acid sequence is on a second vector; or wherein the second and third nucleic acid sequences are on a first vector, and the first nucleic acid sequence is on a second vector.

[0247] Embodiment 52. The expression system of embodiment 49, wherein each of the first, second, and third nucleic acid sequences is on a separate vector.

[0248] Embodiment 53. The expression system of any one of embodiments 46 to 52, wherein the Cas protein is SaCas9.

[0249] Embodiment 54. A composition comprising a Cas protein, and one or both of: a) a first polynucleotide comprising a first guide sequence, wherein: (i) the first guide sequence targets human genome coordinates chr4:3078423-3078443, and the first polynucleotide further comprises a first scaffold sequence comprising any one of SEQ ID NOs:21-95; or (ii) the first guide sequence comprises any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5' guanine; and b) a second polynucleotide comprising a second guide sequence, wherein: (i) the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, and the second polynucleotide further comprises a second scaffold sequence comprising any one of SEQ ID NOs:21-95; or (ii) the second guide sequence comprises any one of SEQ ID NOs:13-19.

[0250] Embodiment 55. The composition of embodiment 54, wherein the Cas protein is Cas9 from Staphylococcus aureus (SaCas9).

[0251] Embodiment 56. The composition of embodiment 54 or 55, wherein the first guide sequence and / or the second guide sequence does not form a secondary structure.

[0252] Embodiment 57. The composition of embodiment 56, wherein the secondary structure is a hairpin loop.

[0253] Embodiment 58. The composition of any one of embodiments 54 to 57, comprising both the first and second polynucleotides.

[0254] Embodiment 59. The composition of any one of embodiments 54 to 58, wherein the first polynucleotide comprises the first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5' guanine, and wherein the first polynucleotide further comprises a first scaffold sequence.

[0255] Embodiment 60. The composition of any one of embodiments 54 to 59, wherein the second polynucleotide comprises the second guide sequence comprising any one of SEQ ID NOs:13-19, and wherein the second polynucleotide further comprises a second scaffold sequence.

[0256] Embodiment 61. The composition of any one of embodiments 54 to 60, wherein the first and second scaffold sequences are each capable of binding to the Cas protein.

[0257] Embodiment 62. The composition of any one of embodiments 54 to 61, wherein the first scaffold sequence and the second scaffold sequence are identical.

[0258] Embodiment 63. The composition of any one of embodiments 54 to 61, wherein the first scaffold sequence and the second scaffold sequence are different.

[0259] Embodiment 64. The composition of any one of embodiments 54 to 63, wherein the first scaffold sequence and / or the second scaffold sequence does not comprise an early termination signal sequence.

[0260] Embodiment 65. The composition of embodiment 64, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

[0261] Embodiment 66. The composition of any one of embodiments 54 to 65, wherein the first scaffold sequence and / or the second scaffold sequence comprises a stabilized secondary structure.

[0262] Embodiment 67. The composition of embodiment 66, wherein the stabilized secondary structure comprises a locked hairpin loop.

[0263] Embodiment 68. The composition of any one of embodiments 54 to 67, wherein the first scaffold sequence and / or the second scaffold sequence comprises any one of SEQ ID NOs:20-96.

[0264] Embodiment 69. The composition of any one of embodiments 54 to 68, wherein the first scaffold sequence and / or the second scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

[0265] Embodiment 70. The composition of any one of embodiments 54 to 69, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-1007.

[0266] Embodiment 71. The composition of any one of embodiments 54 to 70, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172- 174, 177, 180-183, 196-199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

[0267] Embodiment 72. The composition of any one of embodiments 54 to 71, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1539.

[0268] Embodiment 73. The composition of any one of embodiments 54 to 72, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032-1035, 1084, 1160, 1236, 1312, 1388, and 1464.

[0269] Embodiment 74. The composition of any one of embodiments 54 to 73, wherein the first polynucleotide and the second polynucleotide are on a vector.

[0270] Embodiment 75. The composition of embodiment 74, wherein the vector is a viral vector.

[0271] Embodiment 76. The composition of embodiment 75, wherein the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector.

[0272] Embodiment 77. A delivery particle comprising the polynucleotide of any one of embodiments 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54 to 76, or combination thereof.

[0273] Embodiment 78. The delivery particle of embodiment 77, wherein the delivery particle comprises a lipid-based particle or a virus-like particle.

[0274] Embodiment 79. The delivery particle of embodiment 77, wherein the delivery particle comprises a liposome, a micelle, a vesicle, an exosome, or a lipid nanoparticle.

[0275] Embodiment 80. A cell comprising the polynucleotide of any one of embodiments 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

[0276] Embodiment 81. A method of reducing CAG repeats in a mutant allele of a huntingtin (HTT) gene of a cell, comprising introducing to the cell: the polynucleotide of any one of embodiments 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

[0277] Embodiment 82. The method of embodiment 81, wherein the mutant allele comprises greater than 35 CAG repeats prior to the introducing, and fewer than 35 CAG repeats following the introducing.

[0278] Embodiment 83. The method of embodiment 82, wherein, following the introducing, all CAG repeats of the mutant allele are removed.

[0279] Embodiment 84. The method of embodiment 83, wherein, following the introducing, exon 1 of the mutant allele is removed.

[0280] Embodiment 85. The method of any one of embodiments 81 to 84, wherein the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

[0281] Embodiment 86. The method of embodiment 85, wherein the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from the wild-type allele of the HTT gene by more than 50%.

[0282] Embodiment 87. A method of treating Huntington's Disease (HD) in a subject in need thereof, comprising administering to the subject: the polynucleotide of any one of embodiments 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

[0283] Embodiment 88. The method of embodiment 87, wherein the subject comprises greater than 35 CAG repeats in a mutant allele of a huntingtin (HTT) gene prior to the administering.

[0284] Embodiment 89. The method of embodiment 88, wherein the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

[0285] Embodiment 90. The method of embodiment 89, wherein the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

[0286] Embodiment 91. The method of any one of embodiments 88 to 90, wherein the subject comprises fewer than 35 CAG repeats in the mutant allele following the administering.

[0287] Embodiment 92. The method of embodiment 91, wherein, following the administering, all CAG repeats of the mutant allele are removed.

[0288] Embodiment 93. The method of embodiment 91 or 92, wherein, following the administering, exon 1 of the mutant allele is removed. EXAMPLES Example 1. Identification of Huntington’s Disease (HD) Single Nucleotide Polymorphisms (SNPs) and Allele-Specific Editing in Human Fibroblasts

[0289] To identify targetable SNPs in the huntingtin (HTT) gene present in significant patient populations amenable to allele-specific genome editing, an analysis of genome-wide associate study (GWAS) data from more than 10,000 HD patients was performed. A SNP, identified as rs3856973 in dbSNP, is present in heterozygosity in 42% HD patients, is present in intron 1 of the HTT gene, and is approximately 3.5 kb downstream from the CAG repeats region. The rs3856973 SNP is a G-A variant, with the guanine being present on the mutant allele and the adenine present on the wild-type allele of heterozygous HD patients. The guanine forms part of a PAM for SaCas9, TCGAGT, while the adenine does not have a SaCas9 PAM: TCAAGT. Thus, the PAM-containing mutant allele can be selectively edited with SaCas9.

[0290] To verify that SNP rs3856973 can be edited in an allele-specific manner via SaCas9, a CRISPR / Cas ribonucleoprotein (RNP) complex was electroporated into human HD patient fibroblasts having the rs3856973 SNP, or healthy human fibroblasts not having the rs3856973 SNP, using the NEON™ Transfection System. FIG. 2A shows a schematic of the assay that tested 7 different upstream gRNA sequences paired with a downstream gRNA targeting the SNP. A pair of gRNAs was introduced in each condition: a first gRNA that targets the mutant, PAM-containing allele described in Example 1 (shown in the figures as “gRNA 1”), and a second gRNA that targets both mutant and wild-type alleles in the promoter region of the HTT gene, i.e., upstream of exon 1 containing the CAG repeats region. Several second gRNA targets were tested, shown in FIG. 2A. Results are shown in FIG. 2B and indicate that the G allele was selectively edited with gRNA 1 when paired with the 7 different gRNAs targeting upstream of HTT exon 1, while the A allele was not edited, demonstrating allele-specific editing at SNP rs3856973 with SaCas9.

[0291] Further screening of upstream gRNAs that are located outside of the promoter was performed to avoid potential impact on wild-type allele expression. FIG. 3A shows a schematic of the assay that tested 10 different upstream gRNA sequences that are paired with gRNA1 targeting the SNP. Results are shown in FIG. 3B show the mutant allele excision bands on the gel. Example 2. Allele-Specific Reduction of HTT mRNA in HD Patient-Derived Fibroblasts

[0292] A further study was performed in HD patient-derived fibroblasts, in which a pair of selected gRNAs from FIG. 2A and FIG. 3A are introduced via electroporation: a first gRNA that targets the mutant, PAM-containing allele described in Example 1 (shown in the figures as “gRNA 1”), and a second gRNA that targets both mutant and wild-type alleles in the promoter region of the HTT gene, i.e., upstream of exon 1 containing the CAG repeats region. Several second gRNA targets were tested, shown in the figures as “gRNA11,” “gRNA25,” “gRNA27,” and “gRNA30.” Only the mutant allele is cleaved at the target sites of both gRNAs. See FIG. 2A for a schematic of the assay strategy.

[0293] Results are shown in FIGS. 4A and 4B. FIG. 4A shows that up to 50% of the mutant allele excision is observed using a quantitative ddPCR assay measuring the specific expected sequence upon excision when both the first and second gRNAs are introduced, while no excision is observed when only gRNA 1 is introduced or in a non-transfected control. FIG. 4B shows knockdown of greater than 70% of RNA from the mutant allele with minimal effect on the wild-type allele, measured by qPCR assays specifically designed to recognize a SNP that is located at HTT exon 50 and differs in mutant and wild-type alleles. The combination of gRNA 1 and gRNA 11 provided the highest mutant allele excision and mutant allele mRNA knockdown, while not significantly impacting the wild-type allele expression. Example 3. Viral Delivery of SaCas9 and gRNAs in HD Patient-Derived IPS-Neurons

[0294] SaCas9 and the combination of gRNA 1 and gRNA 11 was delivered into HD patient- derived IPS neurons via a lentiviral vector. FIG. 5A shows a schematic of the lentiviral vector design.

[0295] The results in FIG. 5B show low editing efficiency with gRNA 1 (~11%), therefore leading to overall low mutant allele correction as shown in FIG. 5C. This was in contrast to the results of Example 2, which shows high editing efficiency when synthetic gRNAs were delivered into the HD patient-derived fibroblast cells. Example 4. Engineering gRNA to Improve Efficiency

[0296] To increase the editing efficiency of gRNA 1, the secondary structure of gRNA 1 was examined. For gRNA expression from a viral vector, U6, a RNA Polymerase III (Pol III) promoter was used. U6 promoter ends with a guanine (G) for efficient initiation of the transcription, which leads to addition of a 5’ G to the transcribed gRNA1 sequence that matches to the target sequence. Secondary structure predictions showed that gRNA 1 including this 5’ G forms a stem loop in the spacer (i.e., sequence that hybridizes to the target) when expressed from a viral vector. See FIG. 6A, showing the predicted secondary structure of gRNA 1 as a synthetic gRNA, which does not contain the 5’ G and has an “open” spacer sequence; and FIG. 6B, showing the predicted secondary structure of gRNA 1 containing the 5’ G for expression from an adeno-associated viral (AAV) vector or lentiviral vector, which forms a stem loop in the spacer sequence that may inhibit binding of the spacer sequence with the target sequence. RNA folding was predicted using Geneious Prime RNA folding at 37 °C (Andronescu et al., Bioinformatics 23(13):i19–i28, 2007).

[0297] To verify whether the 5’ G indeed affects the editing efficiency, synthetic gRNA 1 and gRNA 11 were prepared with or without a 5’ G and tested for in vitro cleavage. The results are shown in FIG. 7E. “gRNA 1” corresponds to a synthetic gRNA 1 without a 5’ G; “gRNA 1 (+5’G)” corresponds to a synthetic gRNA 1 with a 5’ G; “gRNA 11” corresponds to a synthetic gRNA 11 without a 5’ G; and “gRNA 11 (+5’G)” corresponds to a synthetic gRNA 11 with a 5’ G. “gRNA DMD 16-58” and “gRNA EMX1-sg1” were positive control gRNAs. As shown in FIG. 6E, the DNA cleavage efficiency of gRNA 1 without the 5’ G is about 90%, and addition of the 5’ G reduced the DNA cleavage efficiency to less than 30%. A decrease in the efficiency was also observed when a 5’ G was added to gRNA 11, though to a lesser extent than gRNA 1.

[0298] FIGS. 6C and 6D show modifications to the spacer sequence of gRNA 1 that are predicted to disrupt the inhibitory stem loop and “open” the spacer sequence. FIG. 6C shows a mutation of the adenine directly following the 5’ G (A1) to uracil. FIG. 6D shows addition of a uracil directly following the 5’ G.

[0299] The scaffold sequence, i.e., region that binds to SaCas9, was also examined for modifications to improve expression and consequently editing efficiency. the tetraloop of the scaffold sequence contains a tract of consecutive thymines (T). Consecutive regions of 5-6 thymines (T5-6) is typically a strong termination signal for RNA Pol III, while 4 thymines (T4) can serve as premature termination signal and impact the gRNA transcription efficacy. See, e.g., Chen et al. 2016, Nucleic Acids Res 44(8):e75,.2016. FIG. 7A shows the sequence of gRNA 1 expression cassette with part of the U6 promoter region, including three loop structures in the scaffold region. The 5’ G described in Example 4 is circled. Two T4-6regions are boxed, one of which is present in the tetraloop, and the other is at the 3’ end.

[0300] FIG. 7B shows a non-limiting list of modifications that were made to the gRNA to improve expression and efficiency, including: (1) deletion of the 5’ G ( G);(2) mutation of T2 to G (T2G); (3) insertion of T between 5’ G and A1 (+5’T); (4) mutation of A1 to T (A1T); (5) insertion of a folding-inducing sequence that forms a highly stabilized loop (“locked loop”); (6) mutations of the third T in the T4 sequence of the tetraloop to A and its corresponding downstream paired nucleotide from A to T (shown in FIG. 7B as being 27 nucleotides downstream of the third T in the T4 sequence); and mutations of the fourth T in the T4 sequence to G and its corresponding downstream paired nucleotide from A to C (shown in FIG. 7B as being 25 nucleotides downstream of the fourth T in the T4sequence) (sc3); (7) mutations of the third T in the T4 sequence of the tetraloop to A and its corresponding downstream paired nucleotide from A to T (shown in FIG. 7B as being 27 nucleotides downstream of the third T in the T4sequence) (sc1); (8) mutation of the second T in the T4 sequence of the tetraloop to G and its corresponding downstream paired nucleotide from A to C (shown in FIG. 7B as being 29 nucleotides downstream of the second T in the T4sequence) (sc2); (9) insertion of a locked loop in the tetraloop; (10) insertion of a locked loop in loop 1; and (11) insertion of a locked loop in loop 2.

[0301] Modifications (1) to (5) above prevent formation of the inhibitory loop in the spacer sequence, modifications (6) to (8) above diminish the premature termination signal (T4), and modifications (9) to (11) above stabilize the scaffold sequence.

[0302] The individual modifications shown in FIG. 7B were made to gRNA 1, and the expression cassette shown in FIG. 5A was transfected into HD patient-derived IPS neurons (“iNeurons”).

[0303] FIG. 8A shows representative editing efficiency results with three “original,” unmodified, batches of gRNA 1, and gRNA 1 containing the G, T2G, A1T, +5’T, sc1, sc2, and sc3 modifications described above. An untreated (uninfected) sample was included as negative control. Mutant allele excision efficiency (%) was measured by ddPCR. As shown in FIG. 8A, removal of the 5’ G improved editing by at least 3-fold as compared to the original gRNA 1. Further, scaffold modifications disrupting the premature termination signal in the tetraloop (sc1, sc2, and sc3) improved editing by at least 4-fold as compared to the original gRNA 1.

[0304] FIG. 8B shows representative % indel results with the same modified and unmodified gRNA 1 as FIG. 8A. Similar improvements in % indels for gRNA 1 were observed with the G, sc1, sc2, and sc3 modifications. FIG. 8C shows the % indel results of gRNA 11 with the same modifications as in FIGS. 8A and 8B.

[0305] FIG. 9 shows representative % removal of mutant alleles in iNeurons with different combinations of optimized gRNAs. FIG. 10A shows % indel results with the same optimized gRNA combinations for gRNA1. FIG. 10B shows % indel results with the same optimized gRNA combinations for gRNA11.

[0306] Combinations of the different modifications described above in both gRNA 1 and gRNA 11, as shown in FIG. 11, are tested for further improvement to expression and consequently editing efficiency. Example 5. Optimized dual gRNA excision strategy achieves mutant HTT mRNA and protein reduction in HD patient-derived IPS-Neurons

[0307] FIG. 12A shows the fold change in HTT mRNA levels upon dual gRNA excision with the engineered guides and an optimized SaCas9 using lentiviral vector delivery in HD patient-derived IPS-Neurons. Measurements were obtained using allele specific SNP (rs362331) recognizing RT-qPCR assays that can distinguish wt and mutant HTT RNA. In this example the SaCas9 sequence is codon optimized, and SV40 NLS at both ends were replaced with Myc and bpSV40 NLS at 5`and 3` ends, respectively.

[0308] FIG. 12B shows the fold change of total and mutant HTT protein reduction achieved in HD patient-derived IPS-Neurons upon dual gRNA excision with the engineered guides and the optimized SaCas9 with lentivirus delivery. Gys1 refers to a guide RNA targeting the mouse glycogen synthase gene Gys1 as non-targeting control (Gumusgoz et al. 2021), and HTT to HTT targeting optimized dual gRNAs. Measurements were obtained using SMCxPRO immunoassay system. Data is normalized to untransduced cells. Example 6 AAV delivery of the optimized SaCas9-dual gRNA excision strategy in HD Patient-Derived IPS-Neurons

[0309] FIG. 13A shows the mutant HTT allele excision % in HD patient-derived iNeurons after treatment with escalating doses of AAVs (1E3, 1E4 and 1E5 AAV particles / cell) containing the optimized guides and the original or a codon-NLS optimized SaCas9. Results were obtained by ddPCR designed to detect the excision product. A plateau is reached at the 1E4 dose, most likely due to observed toxicity.

[0310] FIG. 13B Shows the corresponding efficiency of editing (indels) at single sites (g1 top, g11 bottom) associated with the dual excision shown in FIG. 14A. Measurements were obtained by NGS using primers flanking each single cut site. Example 7 In vivo AAV delivery of the optimized SaCas9-dual gRNA excision strategy in BACHD mouse brain

[0311] The optimized SaCas9 variant and the optimized gRNA 1 and 11 were tested in the brain of a mouse model for HD (BACHD mouse model). The striatum of BACHD mice was directly injected with AAV9 expressing the CRISPR components. Striatum of the mice was dissected for DNA extraction and analyzed.

[0312] FIG 15A shows that using the optimized SaCas9 variant resulted in an approximately 50 % increased mutant HTT allele excision compared to the original SaCas9, reaching to 9% measured by ddPCR designed to detect the excision product.

[0313] FIG. 15B shows the corresponding indel % at single gRNA target sites, reaching to 15% measured by NGS using primers flanking each single cut site.

Claims

WHAT IS CLAIMED IS:

1. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence targets human genome coordinates chr4:3078423-3078443, and the scaffold sequence comprises any one of SEQ ID NOs:21-95.

2. The polynucleotide of claim 1, wherein the guide sequence does not form a secondary structure.

3. The polynucleotide of claim 2, wherein the secondary structure is a stem loop.

4. The polynucleotide of any one of claims 1 to 3, wherein the guide sequence comprises any one of SEQ ID NOs:1-12.

5. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein the guide sequence comprises any one of SEQ ID NOs:2-12, and wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine.

6. The polynucleotide of claim 5, wherein the guide sequence comprises SEQ ID NO:2 and does not comprise a 5’ guanine.

7. The polynucleotide of claim 5 or 6, wherein the scaffold sequence is capable of binding to a Cas protein.

8. The polynucleotide of claim 7, wherein the Cas protein is Cas9 from Staphylococcus aureus (SaCas9).

9. The polynucleotide of any one of claims 1 to 8, wherein the scaffold sequence does not comprise an early termination signal sequence.

10. The polynucleotide of claim 9, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

11. The polynucleotide of any one of claims 1 to 10, wherein the scaffold sequence comprises a stabilized secondary structure.

12. The polynucleotide of claim 11, wherein the stabilized secondary structure comprises a locked loop.

13. The polynucleotide of any one of claims 1 to 12, wherein the scaffold sequence comprises any one of SEQ ID NOs:20-95.

14. The polynucleotide of claim 13, wherein the scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

15. A polynucleotide comprising any one of SEQ ID NOs:97-1007.

16. The polynucleotide of claim 15, comprising any one of SEQ ID NOs:97, 98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196-199.

17. The polynucleotide of claim 15, comprising any one of SEQ ID NOs: 97, 98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

18. A polynucleotide comprising a guide sequence and a scaffold sequence, wherein: (i) the guide sequence comprises any one of SEQ ID NOs:2-12, and the scaffold sequence comprises any one of SEQ ID NOs:20-95, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; or (ii) the guide sequence comprises any one of SEQ ID NOs:1-12, and the scaffold sequence comprises any one of SEQ ID NOs:21-95.

19. An expression system comprising a nucleic acid sequence encoding the polynucleotide of claims 1-18.

20. The expression system of claim 19, wherein the nucleic acid sequence is a first nucleic acid sequence, the polynucleotide is a first polynucleotide, the guide sequence is a first guide sequence, and the scaffold sequence is a first scaffold sequence; and wherein the expression system further comprises a second nucleic acid sequence encoding a second polynucleotide, wherein the second polynucleotide comprises a second guide sequence and a second scaffold sequence, wherein the second guide sequence targets a sequence within human genome coordinates chr4:3068346- 3074862.

21. The expression system of claim 20, wherein the second guide sequence does not form a secondary structure.

22. The expression system of claim 21, wherein the secondary structure is a stem loop.

23. The expression system of any one of claims 20 to 22, wherein the second guide sequence comprises any one of SEQ ID NOs:13-19.

24. The expression system of claim 23, wherein the second guide sequence comprises SEQ ID NO:

13.

25. The expression system of any one of claims 20 to 24, wherein the second scaffold sequence is capable of binding to a Cas protein.

26. The expression system of claim 25, wherein the Cas protein is SaCas9.

27. The expression system of any one of claims 20 to 26, wherein the second scaffold sequence does not comprise an early termination signal sequence.

28. The expression system of claim 27, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

29. The expression system of any one of claims 20 to 28, wherein the second scaffold sequence comprises a stabilized secondary structure.

30. The expression system of claim 29, wherein the stabilized secondary structure comprises a locked loop.

31. The expression system of any one of claims 20 to 30, wherein the second scaffold sequence comprises any one of SEQ ID NOs:20-95.

32. The expression system of any one of claims 20 to 30, wherein the second scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

33. The expression system of any one of claims 20 to 32, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1539.

34. The expression system of claim 33, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035.

35. The expression system of claim 33, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

36. An expression system comprising: a first nucleic acid sequence encoding a first polynucleotide comprising any one of SEQ ID NOs:96-1007; and a second nucleic acid sequence encoding a second polynucleotide comprising any one of SEQ ID NOs:1008-1539.

37. The expression system of claim 36, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 177, 180-183, and 196- 199.

38. The expression system of claim 36 or 37, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 183, 196, 199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

39. The expression system of any one of claims 36 to 38, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, and 1032-1035.

40. The expression system of any one of claims 36 to 39, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032, 1035, 1084, 1160, 1236, 1312, 1388, and 1464.

41. The expression system of any one of claims 19 to 40, wherein the expression system comprises a vector.

42. The expression system of claim 41, wherein the vector is a viral vector.

43. The expression system of claim 42, wherein the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector.

44. The expression system of any one of claims 20 to 43, wherein the first nucleic acid sequence and the second nucleic acid sequence are on a single vector.

45. The expression system of any one of claims 20 to 43, wherein each of the first nucleic acid sequence and the second nucleic acid sequence is on a separate vector.

46. The expression system of claim 19, wherein the nucleic acid sequence is a first nucleic acid sequence, the polynucleotide is a first polynucleotide, and wherein the expression system further comprises a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide.

47. The expression system of claim 46, wherein the third nucleic acid sequence is on a separate vector from the first nucleic acid sequence .

48. The expression system of claim 46, wherein the first nucleic acid sequence and the third nucleic acid sequence are on a single vector.

49. The expression system of any one of claims 20 to 45, further comprising a third nucleic acid sequence encoding a Cas protein capable of forming a complex with the first polynucleotide and / or the second polynucleotide.

50. The expression system of claim 49, wherein the first, second, and third nucleic acid sequences are on a single vector.

51. The expression system of claim 49, wherein the first and second nucleic acid sequences are on a first vector, and the third nucleic acid sequence is on a second vector; or wherein the first and third nucleic acid sequences are on a first vector, and the second nucleic acid sequence is on a second vector; or wherein the second and third nucleic acid sequences are on a first vector, and the first nucleic acid sequence is on a second vector.

52. The expression system of claim 49, wherein each of the first, second, and third nucleic acid sequences is on a separate vector.

53. The expression system of any one of claims 46 to 52, wherein the Cas protein is SaCas9.

54. A composition comprising a Cas protein, and one or both of: a) a first polynucleotide comprising a first guide sequence, wherein: (i) the first guide sequence targets human genome coordinates chr4:3078423- 3078443, and the first polynucleotide further comprises a first scaffold sequence comprising any one of SEQ ID NOs:21-95; or (ii) the first guide sequence comprises any one of SEQ ID NOs:2-12, wherein the guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine; and b) a second polynucleotide comprising a second guide sequence, wherein: (i) the second guide sequence targets a sequence within human genome coordinates chr4:3068346-3074862, and the second polynucleotide further comprises a second scaffold sequence comprising any one of SEQ ID NOs:21-95; or (ii) the second guide sequence comprises any one of SEQ ID NOs:13-19.

55. The composition of claim 54, wherein the Cas protein is Cas9 from Staphylococcus aureus (SaCas9).

56. The composition of claim 54 or 55, wherein the first guide sequence and / or the second guide sequence does not form a secondary structure.

57. The composition of claim 56, wherein the secondary structure is a hairpin loop.

58. The composition of any one of claims 54 to 57, comprising both the first and second polynucleotides.

59. The composition of any one of claims 54 to 58, wherein the first polynucleotide comprises the first guide sequence comprising any one of SEQ ID NOs:2-12, wherein the first guide sequence comprising SEQ ID NO:2 does not comprise a 5’ guanine, and wherein the first polynucleotide further comprises a first scaffold sequence.

60. The composition of any one of claims 54 to 59, wherein the second polynucleotide comprises the second guide sequence comprising any one of SEQ ID NOs:13-19, and wherein the second polynucleotide further comprises a second scaffold sequence.

61. The composition of any one of claims 54 to 60, wherein the first and second scaffold sequences are each capable of binding to the Cas protein.

62. The composition of any one of claims 54 to 61, wherein the first scaffold sequence and the second scaffold sequence are identical.

63. The composition of any one of claims 54 to 61, wherein the first scaffold sequence and the second scaffold sequence are different.

64. The composition of any one of claims 54 to 63, wherein the first scaffold sequence and / or the second scaffold sequence does not comprise an early termination signal sequence.

65. The composition of claim 64, wherein the early termination signal sequence comprises 4 to 6 consecutive thymine bases.

66. The composition of any one of claims 54 to 65, wherein the first scaffold sequence and / or the second scaffold sequence comprises a stabilized secondary structure.

67. The composition of claim 66, wherein the stabilized secondary structure comprises a locked hairpin loop.

68. The composition of any one of claims 54 to 67, wherein the first scaffold sequence and / or the second scaffold sequence comprises any one of SEQ ID NOs:20-96.

69. The composition of any one of claims 54 to 68, wherein the first scaffold sequence and / or the second scaffold sequence comprises any one of SEQ ID NOs:20-22, 25, 28-31, and 44-47.

70. The composition of any one of claims 54 to 69, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-1007.

71. The composition of any one of claims 54 to 70, wherein the first polynucleotide comprises any one of SEQ ID NOs:96-98, 101, 104-107, 120-123, 172-174, 177, 180- 183, 196-199, 248, 324, 400, 476, 552, 628, 704, 780, 856, and 932.

72. The composition of any one of claims 54 to 71, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1539.

73. The composition of any one of claims 54 to 72, wherein the second polynucleotide comprises any one of SEQ ID NOs:1008-1010, 1013, 1016-1019, 1032-1035, 1084, 1160, 1236, 1312, 1388, and 1464.

74. The composition of any one of claims 54 to 73, wherein the first polynucleotide and the second polynucleotide are on a vector.

75. The composition of claim 74, wherein the vector is a viral vector.

76. The composition of claim 75, wherein the viral vector is a lentiviral vector, an adenoviral vector, or an adeno-associated viral vector.

77. A delivery particle comprising the polynucleotide of any one of claims 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54 to 76, or combination thereof.

78. The delivery particle of claim 77, wherein the delivery particle comprises a lipid- based particle or a virus-like particle.

79. The delivery particle of claim 77, wherein the delivery particle comprises a liposome, a micelle, a vesicle, an exosome, or a lipid nanoparticle.

80. A cell comprising the polynucleotide of any one of claims 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

81. A method of reducing CAG repeats in a mutant allele of a huntingtin (HTT) gene of a cell, comprising introducing to the cell: the polynucleotide of any one of claims 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

82. The method of claim 81, wherein the mutant allele comprises greater than 35 CAG repeats prior to the introducing, and fewer than 35 CAG repeats following the introducing.

83. The method of claim 82, wherein, following the introducing, all CAG repeats of the mutant allele are removed.

84. The method of claim 83, wherein, following the introducing, exon 1 of the mutant allele is removed.

85. The method of any one of claims 81 to 84, wherein the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

86. The method of claim 85, wherein the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from the wild-type allele of the HTT gene by more than 50%.

87. A method of treating Huntington’s Disease (HD) in a subject in need thereof, comprising administering to the subject: the polynucleotide of any one of claims 1 to 18, the expression system of any one of claims 19 to 53, the composition of any one of claims 54-76, the delivery particle of any one of claims 77 to 79, or combination thereof.

88. The method of claim 87, wherein the subject comprises greater than 35 CAG repeats in a mutant allele of a huntingtin (HTT) gene prior to the administering.

89. The method of claim 88, wherein the method reduces levels of RNA encoded by and / or levels of protein expressed from the mutant allele by at least 30%, and wherein the method does not reduce levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

90. The method of claim 89, wherein the method reduces the levels of RNA encoded by and / or levels of protein expressed from by at least 50%, and wherein the method does not reduce the levels of RNA encoded by and / or levels of protein expressed from a wild-type allele of the HTT gene by more than 50%.

91. The method of any one of claims 88 to 90, wherein the subject comprises fewer than 35 CAG repeats in the mutant allele following the administering.

92. The method of claim 91, wherein, following the administering, all CAG repeats of the mutant allele are removed.

93. The method of claim 91 or 92, wherein, following the administering, exon 1 of the mutant allele is removed.

Citation Information

Patent Citations

  • Cell Penetrating Peptides

    US20080234183A1

  • Aminoalcohol lipidoids and uses thereof

    US20110293703A1

  • Conjugated lipomers and uses thereof

    US20120251560A1

  • Poly(beta-amino alcohols), their preparation, and uses thereof

    US20130302401A1

  • Exosomes comprising therapeutic polypeptides

    US20190167810A1