Products and methods for treatment of autosomal dominant genetic diseases involving diversified disease causing mutations
The Pseudomonas aeruginosa Type I CRISPR system with SNP-targeting guide RNA provides a mutation-independent solution for autosomal dominant diseases like RHO-adRP, achieving efficient and safe allele-specific ablation, addressing the limitations of existing CRISPR therapies.
Patent Information
- Application Number
- PCT/US2025/041271
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-08-08
- Publication Date
- 2026-02-12
AI Technical Summary
Current CRISPR-based therapies for autosomal dominant genetic diseases, such as retinitis pigmentosa, are mutation-specific and costly, with inefficiencies in large DNA deletion and risks of off-target editing, making them impractical for conditions with high genetic heterogeneity like RHO-adRP.
A mutation-independent CRISPR-Cas3 system using a Pseudomonas aeruginosa Type I CRISPR system with a guide RNA targeting a single-nucleotide polymorphism (SNP) for allele-specific ablation, achieving unidirectional large deletions in the target DNA to treat autosomal dominant diseases.
This approach effectively silences the disease-causing allele while preserving the wild-type allele, offering a one-therapy-fit-all solution for a broad patient population by leveraging common SNPs, reducing off-target risks and genetic heterogeneity challenges.
Smart Images

Figure IMGF000010_0001 
Figure IMGF000011_0001 
Figure IMGF000012_0001
Abstract
Description
PRODUCTS AND METHODS FOR TREATMENT OF AUTOSOMAL DOMINANT GENETIC DISEASES INVOLVING DIVERSIFIED DISEASE-CAUSING MUTATIONSCross-Reference to Related Applications
[0001] This application claims priority to Provisional Application No. 63 / 681 ,610, filed August 9, 2024, which is incorporated herein by reference in its entirety.Statement of Government Support
[0002] This invention was made with government support under GM 137833 awarded by the National Institutes of Health. The government has certain rights in the invention.Field
[0003] The disclosure relates to products and methods for treatment of autosomal dominant genetic diseases involving diversified disease-causing mutations. The products and methods use Type I CRISPR in a mutation-independent way to treat autosomal dominant genetic diseases such as autosomal dominant retinitis pigmentosa.Incorporation by Reference of the Sequence Listing
[0004] This application contains, as a separate part of disclosure, a Sequence Listing in computer-readable form (Filename: 70510_SeqListing.xml; 35,601 bytes; created August 7, 2025) which is incorporated by reference herein in its entirety.Background
[0005] Inherited retinal diseases (IRDs) are a diverse group of conditions causing vision loss due to abnormal development, dysfunction, or degeneration of photoreceptors or the retinal pigment epithelium. These diseases can be inherited in various patterns, including autosomal recessive, autosomal dominant, X-linked, mitochondrial, and digenic. To date, at least 280 genes associated with IRDs have been identified (RetNet: / / sph. uth.edu / Retnet / ). The most common IRD is retinitis pigmentosa (RP), a blinding disease caused by the progressive degeneration of rod-cone photoreceptors, affecting approximately 1 in 4,000 Americans [Ben-Yosef: Inherited Retinal Diseases. International journal of molecular sciences 2022, 23(21 )]. About 20% to 30% of RP cases are autosomal dominant (adRP), with RHO gene mutations accounting for 30-40% of these cases (RHO-adRP) [Zhen et al.: Rhodopsin-associated retinal dystrophy: Disease mechanisms and therapeutic strategies. Front Neurosci 2023, 17:1132179]. Currently, there is no cure for adRP. Significant progresshas been made in treating autosomal recessive disorders through gene supplementation, where exogenous genes are delivered to restore healthy biological function. This approach has led to a surge in clinical trials, encouraged by the approval of the first gene therapy for RPE65-linked retinal dystrophy. However, autosomal dominant (AD) disorders, driven by a single mutant allele on non-sex chromosomes, often do not respond to similar strategies. In AD disorders such as RHO-adRP, patients inherit one pathogenic allele causing a disease phenotype, often in a dominant-negative manner, alongside one normal allele. Treatment typically involves silencing the pathogenic allele in an allele-specific manner without affecting the wild-type allele.
[0006] CRISPR gene-editing approaches offer significant hope for inherited diseases including adRP. CRISPR-based precise repair methods, such as homologous-dependent repair, base editing, and prime editing, can correct disease-causing mutations using mutation-specific gRNA-directed CRISPR editing. However, these mutation-specific strategies are clinically impractical and prohibitively expensive for conditions with high genetic heterogeneity, like RHO-adRP, as each of the numerous causal mutations requires specific gRNA design and testing. Over 200 mutations in the RHO gene have been associated with RP (www.hgmd.cf.ac.uk / ac / ).
[0007] A mutation non-specific, ablate-and-replace strategy has been developed and demonstrated efficacy in preserving retinal function and structure in treating RHO-adRP [Wu etal., CRISPR genome surgery in a novel humanized model for autosomal dominant retinitis pigmentosa. Molecular therapy: the journal of the American Society of Gene Therapy 2022, 30(4):1407-1420; Tsai etal., Clustered Regularly Interspaced Short Palindromic Repeats- Based Genome Surgery for the Treatment of Autosomal Dominant Retinitis Pigmentosa. Ophthalmology 2018, 125(9):1421 -1430]. This strategy involves CRISPR-Cas9 ablation of both alleles and the introduction of healthy, CRISPR-Cas9-resistant RHO cDNA via in vivo delivery via adeno-associated virus (AAV). This strategy combines gene-editing and genesupplement therapies, thus inevitably share potential shortcomings of current FDA-approved gene supplement therapies for recessive diseases, such as variable expression levels and concerns about the sustainability of the rescue due to potential loss of the AAV episome, immunogenicity of AAV, and long-term transgene silencing [Shchaslyvyi etal. Current State of Human Gene Therapy: Approved Products and Vectors. Pharmaceuticals (Basel) 2023, 16(10)].
[0008] Recent proof-of-principle studies have demonstrated the potential of singlenucleotide polymorphism (SNP) Cas9 gene editing for treating dominant diseases, including Huntington’s Disease (HD) [Shin et al.: Permanent inactivation of Huntington's disease mutation by personalized allele-specific CRISPR / Cas9. Human molecular genetics 2016,25(20) :4566-4576; Monteys etal. CRISPR / Cas9 Editing of the Mutant Huntingtin Allele In vitro and In Vivo. Molecular therapy: The Journal of the American Society of Gene Therapy 2017, 25(1 ):12-23] and TGFBI (transforming growth factor p-induced) corneal dystrophies [Christie et al.: Mutation-Independent Allele-Specific Editing by CRISPR-Cas9, a Novel Approach to Treat Autosomal Dominant Disease. Molecular therapy: The Journal of the American Society of Gene Therapy 2020, 28(8) :1846-1857], Researchers identified common SNPs that create a NGG PAM site for SpCas9 binding, unique to the mutant chromosome haplotype, and designed an allele-specific sgRNA to target the mutant allele. A second sgRNA, either mutant allele-specific or common to both alleles located in intronic region, was paired to direct the Cas9 nuclease. Consequently, a large deletion was created by the sgRNA pair on the mutant allele to effectively silence its expression, while leaving the wildtype allele unaffected or creating a negligible indel in its noncoding region.
[0009] The above paired-Cas9 strategy is disease mutation agnostic. However, it suffers from two critical limitations. First, the large DNA deletion efficiency by paired guide RNAs can be subpar and varies significantly between different genomic loci, creating uncertainties in the therapeutic potential. Second, it is challenging to find two allele specific guides, so the second guide RNA commonly recognizes both alleles, causing a DNA double strand break in the WT allele, significantly increasing the chance of disruption of the WT gene and the frequency of genome instability and off-target editing risks.
[0010] There remains a need in the art for products and methods for treating AD disorders such as adRP.Summary
[0011] To overcome the barriers outlined above, this disclosure leverages a gene deletion system derived from Type I CRISPR-Cas, a CRISPR-Cas3 system (Fig. 1), to provide allelespecific but mutation-independent, solo gRNA ablation systems and methods for treating autosomal dominant diseases including, but not limited to, RHO-adRP (Fig. 2).
[0012] The disclosure provides products and methods for altering a target DNA containing an allele comprising a mutation associated with an autosomal dominant genetic disease, which method comprises introducing to the target DNA a Type I CRISPR system comprising: (a) a Type I CRISPR Cas 3 protein and Cascade complex and (b) a guide RNA that is complementary to the target DNA, wherein a CC protospacer adjacent motif (PAM) of the Type I CRISPR system is located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA, and wherein the Cas 3 protein translocates in thedirection along the target DNA to ablate the allele. The Type I CRISPR system may be a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system.
[0013] The disclosure also provides products and methods for altering a target DNA containing an allele comprising a mutation associated with an autosomal dominant genetic disease, which method comprises introducing to the target DNA a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system comprising: (a) a Pseudomonas aeruginosa (Pae) Type l-F CRISPR Cas 3 protein and Cascade complex; and (b) a guide RNA that is complementary to the target, wherein the Cas 3 protein translocates in the direction along the target DNA to ablate the allele.
[0014] In the methods, the target DNA may be in a cell and the introducing comprises introducing (a) and (b) into said cell.
[0015] In the methods, the Type I CRISPR Cas 3 protein induces cleavage of at least one strand of the target DNA at the PAM-proximal side of the binding site, thereby altering the target DNA in the cell.
[0016] The cleavage of one or both strands in the target DNA may result in a deletion of the target DNA. The deletion may be unidirectional. A deletion may comprise from about 500 nucleotides to about 100,000 nucleotides. A deletion comprise may comprise from about 5,000 nucleotides to about 20,000 nucleotides.
[0017] In the methods, the guide RNA and Cascade complex may be introduced into the cell as a Cascade ribonucleoprotein (RNP) complex.
[0018] In the methods, the guide RNA may be introduced into the cell as part of a first vector, a nucleic acid encoding the Type I CRISPR Cas 3 protein may be introduced into the cell as part of a second vector, and nucleic acids encoding the Cascade complex components may be introduced into the cell as part of one or more additional vectors. One or more of the vectors may be adeno-associated virus (AAV) vectors.
[0019] In the methods, the nucleic acids encoding the guide RNA and nucleic acids encoding the Type I CRISPR Cas 3 protein may be introduced into the cell as part of a single vector.
[0020] In the methods, the Type I CRISPR Cas 3 protein may be introduced into a cell by contacting the cell with an mRNA encoding the Type I Cas 3 protein.
[0021] In the methods, a CC protospacer adjacent motif (PAM) of the Pae Type l-F CRISPR system may be located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA.
[0022] In the methods, the target DNA may be genomic DNA. The target DNA may encode at least one gene product. The target DNA may encode a protein.
[0023] In the methods, a cell may be a eukaryotic cell. The cell may be a human cell. The human cell may be in situ.
[0024] In the methods, the Type I CRISPR Cas 3 protein may be encoded by a nucleic acid sequence that is codon optimized. The Type I Cas3 protein may be encoded by the nucleic acid sequence of SEQ ID NO: 18.
[0025] In the methods, the mutation may be a mutation in the rhodopsin (RHO) gene in a cell in a human subject and the alteration of the target DNA ablates the mutated RHO gene.
[0026] The disclosure provides products and methods for treating a subject who has an allele comprising a mutation associated with autosomal dominant retinitis pigmentosa (adRP), which method comprises altering the allele of the subject by introducing to the genomic DNA of the subject a Type I CRISPR system: (a) a Type I CRISPR Cas 3 protein and Cascade complex and (b) a guide RNA that is complementary to the target DNA, wherein a CC protospacer adjacent motif (PAM) of the Type I CRISPR system is located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA, and wherein the Cas 3 protein translocates in the direction along the target genomic DNA to ablate the allele. The Type I CRISPR system may be a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system. The Pseudomonas aeruginosa (Pae) Type l-F CRISPR system may comprise one or more hyper-activity Pae l-F variant Cas proteins.
[0027] The disclosure provides hyper-activity Pae l-F variant Cas proteins. A hyperactivity Pae l-F variant protein may be a Cas7 protein. A hyper-activity Pae l-F variant protein may be a Cas8 protein. The hyper-activity Pae l-F variant protein may be Cas7- E338R, Cas7-D234R, Cas7-D237R or Cas7-D68R. The hyper-activity Pae l-F protein may be Cas8-D76R or Cas8-E421 R. The hyper-activity Pae l-F variant Cas proteins may be used in CRISPR DNA editing systems and methods described herein, as well as in other Type I CRISPR DNA editing systems and methods known in the art.
[0028] The disclosure also provides products and methods for treating a subject who has an allele comprising a mutation associated with autosomal dominant retinitis pigmentosa (adRP), which method comprises altering the allele of the subject by introducing to the genomic DNA of the subject a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system comprising: (a) a Pseudomonas aeruginosa (Pae) Type l-F CRISPR Cas 3 protein and Cascade complex; and (b) a guide RNA that is complementary to the target DNA, wherein the Cas 3 protein translocates in the direction along the target genomic DNA to ablate the allele.
[0029] In the treatment methods, the genomic DNA may comprise a mutation in an allele of the Rhodopsin (RHO) gene and the adRP may be an RHO-adRP.
[0030] In the treatment methods, the guide RNA may, for example, recognize a singlenucleotide polymorphism at rs7984 in the 5' untranslated region of the RHO gene. 27. The guide RNA may, for example, recognize a single-nucleotide polymorphism at rs2855558 in the 3' untranslated region of the RHO gene. The guide RNA may, for example, recognize a single-nucleotide polymorphism at rs2410 in the 3' untranslated region of the RHO gene.
[0031] A Cascade complex may comprise Cas5 protein, Cas6 protein, Cas7 protein and Cas8 protein. The Cas3 protein may comprise the amino acid sequence of SEQ ID NO: 17. The Cas5 protein may comprise the amino acid sequence of SEQ ID NO: 9. The Cas6 protein may comprise the amino acid sequence of SEQ ID NO: 11 . The Cas7 protein may comprise the amino acid sequence of SEQ ID NO: 13. The Cas8 protein may comprise the amino acid sequence of SEQ ID NO: 15. As mentioned above, a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system Cascade complex may comprise one or more hyper-activity Pae l-F variant Cas proteins. A hyper-activity Pae l-F variant protein may be a Cas7 protein. A hyper-activity Pae l-F variant protein may be a Cas8 protein. The hyper-activity Pae l-F variant protein may be Cas7-E338R (SEQ ID NO: 19), Cas7-D234R (SEQ ID NO: 20), Cas7- D237R (SEQ ID NO: 21) or Cas7-D68R (SEQ ID NO: 22). The hyper-activity Pae l-F protein may be Cas8-D76R (SEQ ID NO: 24) or Cas8-E421 R (SEQ ID NO: 25).
[0032] Thus, the disclosure also provides illustrative Pae l-F variant proteins Cas7-E338R (SEQ ID NO: 19), Cas7-D234R (SEQ ID NQ:20), Cas7-D237R (SEQ ID NO: 21), Cas7- D68R (SEQ ID NO: 22), Cas8-D76R (SEQ ID NO: 24) and Cas8-E421 R (SEQ ID NO: 25).Brief Description of the Drawings
[0033] Fig. 1 shows Type I CRISPR-Cas introducing uni-directional large DNA deletions in the human genome. As an RNA-programmed DNA targeting system, Type I CRISPR employs a crRNA-guided multi-subunit DNA binding complex termed Cascade to recognize a PAM-flanked, crRNA-matched target site in the genome. Upon target binding, the Cascade complex undergoes confirmational changes to recruit the Cas3 enzyme in trans. Cas3 then initiates ATP-hydrolysis driven translocation and processively degrades target DNA over long distance.
[0034] Fig. 2A-C illustrates allele-specific but mutation-independent gene ablation therapy using type I CRISPR-Cas. A. If the dominant disease mutation is associated with a SNP in cis ( / .e., on the same allele / chromosome), Type I CRISPR will be programmed with a guide RNA that recognizes the SNP sequence but not the WT counterpart. The design leads to alarge deletion that destroys the mutation-containing gene region and thereby cures the adRP patients. B. If the dominant disease mutation and a SNP sequence are encoded on two different alleles / chromosomes ( / .e., in trans; disease mutation is associated with WT copy of this specific SNP region), type I CRISPR will be programmed with a guide RNA that recognizes the WT counterpart sequence but not the SNP, and therefore delete the mutation containing genomic region to cure the patients. C. For type I CRISPR-Cas to achieve allele differentiation using a SNP, the SNP can either be within the guide RNA sequence (underlined) close to the PAM or be within the PAM sequence (dashed box) itself. The version depicted here assumes the SNP is linked to disease mutation in cis.
[0035] Fig. 3A-C shows a proof of principle experiment demonstrating that Type I CRISPR achieves allele differentiation using the SNP. A. A diagram depicting the HEK293- WT / eGFP-SNP / tdTomato dual fluorescent reporter cell line used. This reporter line is positive for both GFP and TdTomato fluorescence. B. Guide RNA design for Type I CRISPR allele differentiation using SNP rs7984. Of note, the WT guide and SNP guide only differ by one nucleotide at the PAM-proximal “seed” region. C. Robust targeting and allele differentiation by Type I CRISPR-Cas demonstrated using reporter cell line and guides shown in A-B. The WT guide caused deletion into the eGFP gene, resulting in GFP-negative and tdTomato-positive cells. The SNP-targeting guide led to deletions into the tdTomato gene, causing tdTomato-negative and eGFP-positive cells.
[0036] Fig. 4A-B shows control experiments demonstrating that without successful allele differentiation, this therapeutic strategy will not work as both WT and diseased alleles would be destroyed. A. Design of third guide RNA (referred to as the common guide) that would target both alleles. It recognizes a site 13-nt away from the SNP site used in Fig. 3. B. Gene targeting experiment in the dual reporter line demonstrating that the common guide leads to high percentage of cells having both EGFP and tdTomato gene disrupted (dual negative), as well as both the GFP-negative / tdTomato-positive and GFP-positive / tdTomato-negative populations.
[0037] Fig. 5 shows the principle of Tn5-anchored-NGS profiling platform. This method has proven effective in comprehensively characterizing the unidirectional large DNA deletions generated by Type I CRISPR.
[0038] Fig. 6A-B shows Tn5-comprehensive profiling of genomic DNA deletions introduced at the RHO-WT-GFP locus by Type I CRISPR and the RHO-WT allele-targeting guide. A. Histogram showing the distribution of deletion lengths in 1 kb interval. B. Accumulative deletion length percentage. At this locus, 50% of the deletions are shorter than 2.3 kb, 90% of the deletions are shorter than 8.8 kb.
[0039] Fig. 7A-B shows Tn5-profiling of genomic DNA deletions introduced at the tdTomato locus by Type I CRISPR and the RHO-SNP-targeting guide. A. Histogram showing the distribution of deletion lengths in 1 kb interval. B. Accumulative deletion length percentage. At this locus, 50% of the deletions are shorter than 2.6 kb, 90% of the deletions are shorter than 8.9 kb.
[0040] Fig. 8A-B shows Tn5-profiling of genomic DNA deletions introduced at the endogenous RHO locus by Type I CRISPR and the RHO-WT-allele targeting guide. A. Histogram showing the distribution of deletion lengths in 1 kb interval. B. Accumulative deletion length percentage. At this locus, 50% of the deletions are shorter than 1 .7 kb, 90% of the deletions are shorter than 7.2 kb.
[0041] Fig. 9A-B shows Type I guide RNA design for additional common SNPs associated with human RHO gene. A. Guide RNA design for allele differentiation using SNP rs2855558. B. Guide RNA design for allele differentiation using SNP rs2410.
[0042] Fig. 10 is a bar graph showing the gene targeting / deletion activity of the wild-type Pae type l-F Cascade or a panel of its variants, delivered with WT Cas3 and a GFP targeting crRNA plasmid, into a HEK293- EGFP reporter cell line. Target / deletion activity is shown as the percentage of EGFP-negative cells in the total population, as measured by flow cytometry. Fold change in gene targeting activity relative to WT is indicated. The Pae l-F variant containing the Cas7-E338R mutation was used as the hyperactive construct for in vivo studies in the rabbit model.
[0043] Fig. 11 A-D shows in vivo Rho gene targeting / deletion in rabbit retina using dual- AAV delivered hyper-activity variant of Pae type l-F CRISPR-Cas. A. Design of a dual-AAV strategy to deliver Pae type l-F CRISPR-Cas components. A hyper-activity variant developed in Figure 10 (Cas7-E338R, depicted as Cas7*) was used. B. HEK293-WT / eGFP- SNP / tdTomato dual fluorescent reporter cell line (from Fig. 3) was transduced with dual-AAV expressing the hyper-activity Pae type l-F editor targeting the WT Rho allele. Plotted are the percentages of EGFP- / tdTomato+, EGFP+ / tdTomato-, and dual-negative cells in the total reporter population, as measured by flow cytometry. This system resulted in robust gene deletion, (over 40% of EGFP- / tdTomato+ cells) with exceptional allele specificity (background level of EGFP- / tdTomato+ cells). C. Diagram of sub-retinal injection of AAV in a transgenic rabbit model. One allele of the endogenous rabbit Rho gene has its exon 1 (light grey) replaced with human RHO exon 1 (dark grey). The target sequence and flanking PAM utilized by the Pae type l-F editor is given in the middle, with Cas3 translocation direction indicated. D. Genomic deletions at the humanized Rho allele in rabbit retina detected after AAV injection using an anchored Tn5 NGS method [Dolan et al.: Introducing a Spectrum ofLong-Range Genomic Deletions in Human Embryonic Stem Cells Using Type I CRISPR- Cas. Mol Cell. 2019, 74(5):936-950.e5]. Each line represents a deletion event identified in a Tn5 sequencing read. X-axis gives the relative location of start and end of each deletion, with PAM set as position +1 (as labeled in panel C). The humanized Rho allele is aligned above as a reference for positional alignment.Detailed Description
[0044] Type I CRISPR is the most widespread and diverse type of naturally existing CRISPR systems, accounting for over half of all identified prokaryotic CRISPR systems. Type I CRISPR is further classified into eight subclasses (l-A through l-F, l-Fv, and l-G) based on cas gene composition. Table 1 below shows the Type I CRISPR subclasses, the bacterial source of each subclass, its PubMed reference number (PMID), its PAM and its spacer length.Table 1
[0045] The Type I CRISPR DNA targeting machinery consists of two modules: an RNA- guided Cascade complex that recognizes a dsDNA target sequence, and a helicase- nuclease fusion enzyme Cas3 that is recruited in trans to the Cascade-bound target and translocates in the 3’-to-5’ direction to processively degrade large stretches of DNA [Sinkunas etal.: Cas3 is a single-stranded DNA nuclease and ATP-dependent helicase in the CRISPR / Cas immune system. EMBO J 2011 , 30(7):1335-1342], Unlike the commonly used CRISPR-Cas9 that creates a site-specific double-strand DNA break, Type I CRISPR creates kb-scale large deletions in the human genome [Dolan et al., supra] (Fig. 1). This “DNA shredder” tool is contemplated by the disclosure herein to be superior to the paired Cas9-sgRNA tools in that it creates large deletions with higher efficiencies and requires only one CRISPR binding site. It is contemplated herein that Type I CRISPR technology is particularly useful for removing disease-causing genomic loci such as genes with dominant negative mutations, toxic repeat expansions, large parasitic sequence, or integrated viral genome.
[0046] The illustrative RHO-adRP treatment strategy herein makes use of common benign and heterozygous SNPs present near the RHO gene in RHO-adRP patients. One such example is SNP rs7984 located in the 5’ untranslated region of the RHO gene that shows nearly equal allele frequencies in global population (Global minor allele frequency 0.47165, reported by dbSNP) (Fig. 2). The disclosure provides two illustrative versions of Type I CRISPR reagents, each of which can selectively recognize either allele of this SNP locus while ignoring the other allele (Fig. 3B). This allele differentiation effect is achieved by the SNP lying within the PAM sequence or the PAM-proximal segment of the guide RNA. Upon allele recognition by Type I CRISPR, the processive nuclease Cas3 will start unidirectional translocation from this SNP site and remove kb-sized genome fragments towards the PAM-proximal direction. Therefore, regardless of the Rho mutation being in cis or in trans with that SNP, there is a Type I CRISPR strategy to ablate the gain-of-function disease allele while leaving the non-diseased allele untouched (Fig. 2). This gene editingoutcome is contemplated to effectively cure the disease because the RHO gene is haplo- sufficient, and RHO+ / null individuals exhibit no phenotypes. Given the long-range nature of the deletions (a few hundred nucleotides to kilobases in size) caused by Type I CRISPR, this approach can address a panel of patient mutations that occur in the RHO gene, making it a one-therapy-fit-all to benefit a broad patient population. Also, the heterozygous rate for SNP rs7984 is close to 50%, meaning that the two guide RNA designs provided herein together cover about half of RHO-adRP patients. By utilizing additional common benign SNPs with high heterozygous rate, such as rs2855558 and rs2410 (Fig. 9), an even larger patient population can be treated. Lastly, this strategy is contemplated herein to be widely applicable to other single-gene dominant genetic diseases with diversified disease-causing mutations.
[0047] Table 2 lists other benign SNPs in the RHO gene, the SNP position on human chromosome 3, the variation, the variant type, the SNP ID and functional class.Table 2
[0048] To design a Type I CRISPR guide to achieve allele-specific but mutationindependent deletion, a common benign SNP located within or close to the gene of interestcan be leveraged. To differentiate between a mutant allele and the wild-type allele (Fig. 2A and B), the designer identifies a short PAM sequence either encompassing the SNP (scenario 2 in Fig. 2C) or located very close to it (scenario 1 in Fig. 2C, i.e., SNP is in the PAM-proximal half of the guide sequence).
[0049] Based on available PAM options, the designer chooses a specific subtype of Type I CRISPR-Cas system to use. Table 1 is a summary of some representative PAM preferences and guide lengths of Type I CRISPR subtypes l-A through l-G. Once the appropriate subtype(s) are picked, the guide RNA sequences recognizing the PAM-adjacent region can be designed. For scenario 1 , the allele-differentiating SNP should be within the guide, excluding the +6, +12, +18, +24, +30 nts positions and preferably in the PAM- proximal-half of guide. For scenario 2, the SNP should be within the short PAM’s consensus nucleotide positions.
[0050] Given that the Type I CRISPR systems are known to fire Cas3 helicase-nuclease either unidirectionally, from the CRISPR-complementary site toward the PAM, or bidirectionally, the overall Type I PAM / guide design should ensure the processive Cas3 nuclease will travel towards the target gene’s body to ablate the disease-associated allele. “Ablate” herein means that enough of the disease-associated allele is deleted to prevent its disease-causing activity.
[0051] The disclosure provides DNAs encoding illustrative guide RNAs that target the Rho gene as follows.DNA encoding Rho-WT Guide (set out in Fig 3 and SEQ ID NO: 1) GTGGCTGCTCCCACCCAAGAATGCTGCGAAGGDNA encoding Rho-SNP Guide (set out in Fig3 and SEQ ID NO: 2) GCGGCTGCTCCCACCCAAGAATGCTGCGAAGGDNA encoding Rho-common Guide (set out in Fig 4 and SEQ ID NO: 3) ACCCAAGAATGCTGCGAAGGCCTGAGCTCAGCDNA encoding Rho-rs2855558 WT Guide (set out in Fig 9 and SEQ ID NO: 4) AATGAGGGTGAGATTGGGCCTGGGGTCTCACCDNA encoding Rho-rs2855558 SNP Guide (set out in Fig 9 and SEQ ID NO: 5) AGTGAGGGTGAGATTGGGCCTGGGGTCTCACCDNA encoding Rho-rs2410 WT Guide (set out in Fig 9 and SEQ ID NO: 6) AAGGCCAGCGGGATGTGTGCCCCTCCTCCTCCDNA encoding Rho-rs2410 SNP Guide (set out in Fig 9 and SEQ ID NO: 7) GAGGCCAGCGGGATGTGTGCCCCTCCTCCTCC
[0052] The Pae-IF CRISPR repeat sequence follows.Pae-IF-CRISPR repeat (SEQ ID NO: 8) GTTCACTGCCGTATAGGCAGCTAAGAAA
[0053] Definitions
[0054] As used herein, a “nucleic acid” or a “nucleic acid sequence” refers to a polymer or oligomer of pyrimidine and / or purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (See Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982)). The present technology contemplates any deoxyribonucleotide, ribonucleotide, or peptide nucleic acid component, and any chemical variants thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases, and the like. The polymers or oligomers may be heterogenous or homogenous in composition, and may be isolated from naturally occurring sources or may be artificially or synthetically produced. In addition, the nucleic acids may be DNA or RNA, or a mixture thereof, and may exist permanently or transitionally in single-stranded or double-stranded form, including homoduplex, heteroduplex, and hybrid states. A nucleic acid or nucleic acid sequence may comprise other kinds of nucleic acid structures such as, for instance, a DNA / RNA helix, peptide nucleic acid (PNA), morpholino nucleic acid (see, e.g., Braasch and Corey, Biochemistry, 41 (14): 4503-4510 (2002)) and U.S. Patent 5,034,506), locked nucleic acid (LNA; see Wahlestedt etal., Proc. Natl. Acad. Sci. U.S.A., 97: 5633-5638 (2000)), cyclohexenyl nucleic acids (see Wang, J. Am. Chem. Soc., 122: 8595-18602 (2000)), and / or a ribozyme. Hence, the term “nucleic acid” or “nucleic acid sequence” may also encompass a chain comprising non-natural nucleotides, modified nucleotides, and / or non- nucleotide building blocks that can exhibit the same function as natural nucleotides e.g., “nucleotide analogs”); further, the term “’’nucleic acid” or nucleic acid sequence” as used herein refers to an oligonucleotide, nucleotide or polynucleotide, and fragments or portions thereof, and to DNA or RNA of genomic or synthetic origin, which may be single or double-stranded, and represent the sense or antisense strand. The terms “nucleic acid,” “polynucleotide,” “nucleotide sequence,” and “oligonucleotide” are used interchangeably. They refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof.
[0055] The terms “complementary” and “complementarity” refer to the ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid by either traditional Watson-Crick base-paring or other non-traditional types of pairing. The degree of complementarity between two nucleic acids can be indicated by the percentage of nucleotides in a nucleic acid which can form hydrogen bonds e.g., Watson-Crick base pairing) with a second nucleic acid {e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary). Two nucleic acids are “perfectly complementary” if all the contiguous nucleotides of a nucleic acid will hydrogenbond with the same number of contiguous nucleotides in a second nucleic acid. Two nucleic acids are “substantially complementary” if the degree of complementarity between the two nucleic acids is at least 60% {e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%. 97%, 98%, 99%, or 100%) over a region of at least 8 nucleotides e.g., 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 30, 35, 40, 45, 50, or more nucleotides), or if the two nucleic acids hybridize under at least moderate, preferably high, stringency conditions. Exemplary moderate stringency conditions include overnight incubation at 37° C in a solution comprising 20% formamide, 5xSSC (150 mM NaCI, 15 mM trisodium citrate), 50 mM sodium phosphate (pH 7.6), 5xDenhardt’s solution, 10% dextran sulfate, and 20 mg / ml denatured sheared salmon sperm DNA, followed by washing the filters in 1 xSSC at about 37-50° C, or substantially similar conditions, e.g., the moderately stringent conditions described in Sambrook et al., infra. High stringency conditions are conditions that use, for example (1 ) low ionic strength and high temperature for washing, such as 0.015 M sodium chloride / 0.0015 M sodium citrate / 0.1 % sodium dodecyl sulfate (SDS) at 50° C, (2) employ a denaturing agent during hybridization, such as formamide, for example, 50% (v / v) formamide with 0.1% bovine serum albumin (BSA) / 0.1% Ficoll / 0.1 % polyvinylpyrrolidone (PVP) / 50 mM sodium phosphate buffer at pH 6.5 with 750 mM sodium chloride and 75 mM sodium citrate at 42° C, or (3) employ 50% formamide, 5xSSC (0.75 M NaCI, 0.075 M sodium citrate), 50 mM sodium phosphate (pH 6.8), 0.1% sodium pyrophosphate, 5xDenhardt’s solution, sonicated salmon sperm DNA (50 pg / ml), 0.1% SDS, and 10% dextran sulfate at 42° C, with washes at (i) 42° C in 0.2xSSC, (ii) 55° C in 50% formamide, and (iii) 55° C in 0.1 xSSC (preferably in combination with EDTA). Additional details and an explanation of stringency of hybridization reactions are provided in, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Press, Cold Spring Harbor, N.Y. (2001); and Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and John Wiley & Sons, New York (1994).
[0056] As used herein, the term “percent sequence identity” refers to the percentage of nucleotides or nucleotide analogs in a nucleic acid, or amino acids in an polypeptide, that is identical with the corresponding nucleotides or amino acids in a reference molecule after aligning the two sequences. In case a nucleic acid according to the technology is longer than a reference, additional nucleotides in the nucleic acid, that do not align with the reference, are not taken into account for determining sequence identity. Methods and computer programs for alignment are well known in the art, including BLAST, Align 2, and FASTA.
[0057] As used herein, the term “hybridization” is used in reference to the pairing of complementary nucleic acids. Hybridization and the strength of hybridization e.g., the strength of the association between the nucleic acids) is influenced by such factors as thedegree of complementary between the nucleic acids, stringency of the conditions involved, and the Tm of the formed hybrid. Hybridization methods involve the annealing of one nucleic acid to another, complementary nucleic acid, e.g., a nucleic acid having a complementary nucleotide sequence. The ability of two polymers of nucleic acid containing complementary sequences to find each other and “anneal” or “hybridize” through base pairing interaction is a well-recognized phenomenon. The initial observations of the “hybridization” process by Marmur and Lane, Proc. Natl. Acad. Sci. USA, 46: 453 (1960) and Doty et al., Proc. Natl. Acad. Sci. USA, 46: 461 (1960), have been followed by the refinement of this process into an essential tool of modern biology. For example, hybridization and washing conditions are now well known and exemplified in Sambrook etal., supra. The conditions of temperature and ionic strength determine the “stringency” of the hybridization.
[0058] As used herein, a “double-stranded nucleic acid” may be a portion of a nucleic acid, a region of a longer nucleic acid, or an entire nucleic acid. A “double-stranded nucleic acid” may be, e.g., without limitation, a double-stranded DNA, a double-stranded RNA, a double-stranded DNA / RNA hybrid, etc. A single-stranded nucleic acid having secondary structure e.g., base-paired secondary structure) and / or higher order structure e.g., a stemloop structure) may also be considered a “double-stranded nucleic acid.” For example, triplex structures are considered to be “double-stranded.” Any base-paired nucleic acid is a “double-stranded nucleic acid.”
[0059] The term “gene” refers to a DNA that comprises control and coding sequences necessary for the production of an RNA having a non-coding function {e.g., a ribosomal or transfer RNA), a polypeptide, or a precursor of any of the foregoing. The RNA or polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence so long as the desired activity or function is retained. Thus, a “gene” refers to a DNA or RNA, or portion thereof, that encodes a polypeptide or an RNA chain that has functional role to play in an organism. For the purpose of this disclosure, it may be considered that genes include regions that regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.
[0060] The term “wild type” refers to a gene or a gene product that has the characteristics of that gene or gene product when isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designated the “normal” or “wild-type” form of the gene. In contrast, the term “engineered,”“modified,” “mutant,” or “polymorphic” refers to a gene or gene product that displays modifications in sequence and or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product.
[0061] As used herein, the term “variant” refers to the exhibition of qualities that have a pattern that deviates from what occurs in nature. A variant may also be a mutant.
[0062] The terms “non-naturally occurring,” “engineered,” and “synthetic” are used interchangeably and indicate the involvement of the hand of man. The terms, when referring to nucleic acid molecules or polypeptides mean that the nucleic acid molecule or the polypeptide is at least substantially free from at least one other component with which they are naturally associated in nature and as found in nature.
[0063] The terms “peptide,” “polypeptide,” and “protein” are used interchangeably herein, and refer to a polymeric form of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides having modified peptide backbones.
[0064] An amino acid “replacement” or “substitution” refers to the replacement of any one amino acid at a given position or residue by another amino acid at the same position or residue within a polypeptide. Amino acids are broadly grouped as “aromatic” or “aliphatic.” An aromatic amino acid includes an aromatic ring. Examples of “aromatic” amino acids include histidine (H or His), phenylalanine (F or Phe), tyrosine (Y or Tyr), and tryptophan (W or Trp). Non- aromatic amino acids are broadly grouped as “aliphatic.” Examples of “aliphatic” amino acids include glycine (G or Gly), alanine (A or Ala), valine (V or Vai), leucine (L or Leu), isoleucine (I or He), methionine (M or Met), serine (S or Ser), threonine (T or Thr), cysteine (C or Cys), proline (P or Pro), glutamic acid (E or Glu), aspartic acid (A or Asp), asparagine (N or Asn), glutamine (Q or Gin), lysine (K or Lys), and arginine (R or Arg). The amino acid replacement or substitution can be conservative, semi-conservative, or nonconservative. The phrase “conservative amino acid substitution” or “conservative mutation” refers to the replacement of one amino acid by another amino acid with a common property. A functional way to define common properties between individual amino acids is to analyze the normalized frequencies of amino acid changes between corresponding proteins of homologous organisms (Schulz and Schirmer, Principles of Protein Structure, Springer- Verlag, New York (1979)). According to such analyses, groups of amino acids may be defined where amino acids within a group exchange preferentially with each other, and therefore resemble each other most in their impact on the overall protein structure (Schulzand Schirmer, supra). Examples of conservative amino acid substitutions include substitutions of amino acids within the sub-groups described above, for example, lysine for arginine and vice versa such that a positive charge may be maintained, glutamic acid for aspartic acid and vice versa such that a negative charge may be maintained, serine for threonine such that a free -OH can be maintained, and glutamine for asparagine such that a free -NH2 can be maintained. “Semi-conservative mutations” include amino acid substitutions of amino acids within the same groups listed above, but not within the same sub-group. For example, the substitution of aspartic acid for asparagine, or asparagine for lysine, involves amino acids within the same group, but different sub-groups. “Nonconservative mutations” involve amino acid substitutions between different groups, for example, lysine for tryptophan, or phenylalanine for serine, etc.
[0065] “Binding” as used herein e.g., with reference to an DNA-binding domain of a polypeptide) refers to a non-covalent interaction between macromolecules e.g., between a protein and a nucleic acid). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” {e.g., when a molecule X is said to interact with a molecule Y, it is meant the molecule X binds to molecule Y in a non-covalent manner). Not all components of a binding interaction need be sequence-specific {e.g., contacts with phosphate residues in a DNA backbone), but some portions of a binding interaction may be sequence-specific. Binding interactions are generally characterized by a dissociation constant (Kd) of less than 10-6M, less than 10-7M, less than 10-8M, less than 10-9M, less than 10-1° M, less than 10-11M, less than 10-12M, less than 10-13M, less than 10-14M, or less than 10-15M. “Affinity” refers to the strength of binding, increased binding affinity being correlated with a lower Kd.
[0066] By “binding domain” it is meant a protein domain that is able to bind non-covalently to another molecule. A binding domain can bind to, for example, a DNA molecule (a DNA- binding protein), an RNA molecule (an RNA-binding protein) and / or a protein molecule (a protein binding protein). In the case of a protein domain-binding protein, it can bind to itself (to form homodimers, homotrimers, etc.) and / or it can bind to one or more molecules of a different protein or proteins.
[0067] A “vector” or “expression vector” is a replicon, such as plasmid, phage, virus, or cosmid, to which another DNA segment, e.g., an “insert,” may be attached or incorporated so as to bring about the replication of the attached segment in a cell.
[0068] A cell has been “genetically modified,” “transformed,” or “transfected” by exogenous DNA, e.g., a recombinant expression vector, when such DNA has been introduced inside the cell. The presence of exogenous DNA results in permanent or transientgenetic change. The transforming DNA may or may not be integrated (covalently linked) into the genome of the cell. In prokaryotes, yeast, and mammalian cells for example, the transforming DNA may be maintained on an episomal element such as a plasmid. With respect to eukaryotic cells, a stably transformed cell is one in which the transforming DNA has become integrated into a chromosome so that it is inherited by daughter cells through chromosome replication. This stability is demonstrated by the ability of the eukaryotic cell to establish cell lines or clones that comprise a population of daughter cells containing the transforming DNA. A “clone” is a population of cells derived from a single cell or common ancestor by mitosis. A “cell line” is a clone of a primary cell that is capable of stable growth in vitro for many generations.
[0069] "Recombinant," as used herein, means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, polymerase chain reaction (PCR) and / or ligation steps resulting in a construct having a structural coding or non-coding sequence distinguishable from endogenous nucleic acids found in natural systems. DNAs encoding polypeptides can be assembled from cDNA fragments or from a series of synthetic oligonucleotides, to provide a synthetic nucleic acid which is capable of being expressed from a recombinant transcriptional unit contained in a cell or in a cell-free transcription and translation system. Genomic DNA comprising the relevant sequences can also be used in the formation of a recombinant gene or transcriptional unit. Non-translated DNA may be present 5' or 3' from the open reading frame, as long as such DNA does not interfere with manipulation or expression of the coding regions, and may indeed act to modulate production of a desired product by various mechanisms). Alternatively, DNAs encoding RNA e.g., DNA-targeting RNA) that are not translated may also be considered recombinant.Thus, the term "recombinant" nucleic acid refers to one which is not naturally occurring, e.g., is made by the artificial combination of two otherwise separated segments of nucleic acids through human intervention. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques. Such is usually done to replace a codon with a codon encoding the same amino acid, a conservative amino acid, or a non-conservative amino acid. Alternatively, it is performed to join together nucleic acid segments of desired functions to generate a desired combination of functions. This artificial combination is often accomplished by either chemical synthesis means, or by the artificial manipulation of isolated segments of nucleic acids, e.g., by genetic engineering techniques. When a recombinant polynucleotide encodes a polypeptide, the sequence of the encoded polypeptide can be naturally occurring ("wild type") or can be a variant e.g., a mutant) of the naturally occurring sequence. Thus, the term "recombinant" polypeptide does notnecessarily refer to a polypeptide whose sequence does not naturally occur. Instead, a "recombinant" polypeptide is encoded by a recombinant DNA, but the sequence of the polypeptide can be naturally occurring ("wild type") or non-naturally occurring (e.g., a variant, a mutant, etc.). Thus, a "recombinant" polypeptide is the result of human intervention but may be a naturally occurring amino acid sequence.
[0070] A “subject” or “patient” may be human or non-human and may include, for example, animal strains or species used as “model systems” for research purposes, such as a mouse model. Likewise, patient may include either adults, juveniles e.g., children), or infants. Moreover, patient may mean any living organism, preferably a mammal e.g., humans and non-humans) that may benefit from the administration of compositions contemplated herein. Examples of mammals include, but are not limited to, any member of the Mammalian class: humans, non-human primates such as chimpanzees, and other apes and monkey species; farm animals such as cattle, horses, sheep, goats, swine; domestic animals such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice, and guinea pigs, and the like. Examples of non-mammals include, but are not limited to, birds, fish, and the like. A mammal may be a human.
[0071] The term “contacting” as used herein refers to bring or put in contact, to be in or come into contact. The term “contact” as used herein refers to a state or condition of touching or of immediate or local proximity. Contacting to a target destination, such as, but not limited to, an organ, tissue, cell, or tumor, may occur by any means of administration known to the skilled artisan.
[0072] As used herein, the terms “providing,” “administering,” and “introducing,” are used interchangeably herein and refer to the placement of the proteins or systems of the disclosure into a subject by a method or route which results in at least partial localization to a desired site. Administration can use any appropriate route which results in delivery to a desired location in the subject.
[0073] CRISPR / Cas
[0074] In bacteria and archaea, Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR)-CRISPR associated (Cas) (CRISPR-Cas) systems provide immunity by incorporating fragments of invading phage, virus, and plasmid DNA into CRISPR loci and using corresponding CRISPR RNAs (“crRNAs”) to guide the degradation of homologous DNAs. Transcription of a CRISPR locus produces a “pre-crRNA,” which is processed to yield crRNAs containing spacer-repeat fragments that guide effector nuclease complexes to cleave dsDNAs complementary to the spacer. As mentioned in the Background section above, several different types of CRISPR systems are known, {e.g., Type I, Type II, or TypeIII), and classified based on the Cas protein type and the use of a proto-spacer-adjacent motif (PAM) for selection of proto-spacers in invading DNA. Type I systems are the most widespread and diversified type and are classified into eight subtypes (l-A through l-F, l-Fv and l-G).
[0075] Engineering CRISPR / Cas systems for use in eukaryotic cells involves reconstitution of a CRISPR / Cas complex.
[0076] The system disclosed herein comprises an engineered Type I CRISPR-Cas system, and / or one or more nucleic acids encoding the engineered CRISPR-Cas system, wherein the engineered CRISPR-Cas system comprises: (a) Cas3; (b) Cascade complex; and (c) a guide RNA (gRNA), wherein the gRNA is configured to specifically hybridize to (“recognize”) a portion of a target nucleic acid.
[0077] Target recognition by the gRNA and Cascade complex results in a conformational change which facilitates recruitment of Cas3. Cas3 may comprise a single protein unit which contains helicase and nuclease domains. After target validation by the Cascade complex, Cas3 nicks the strand of DNA that is looped out by the R-loop formed by the Cascade complex. Cas3 then uses its helicase / nuclease activity to processively degrade substrate nucleic acids, moving in a 3’ to 5’ direction.
[0078] An illustrative Pseudomonas aeruginosa (Pae) Type l-F CRISPR system provided herein comprises Cas3 as well as Cascade complex proteins including one or more of: Cas5, Cas6 Cas7 and Cas8. The system may comprise a Cas5, Cas6, Cas7 or Cas8, which is an engineered protein or fusion protein as disclosed herein. The system may comprise two, three or all of Cas5, Cas6, Cas7 and Cas8. The Cas proteins may be wild-type proteins or other variants thereof, for example those disclosed in International Application No. PCT / US2022 / 031091 . Illustrative hyper-activity Cas7 and Cas8 proteins are provided herein. The system may comprise two, three or all of Cas5, Cas6, Cas7 and Cas8, all of which are an engineered protein or fusion protein as disclosed herein.
[0079] The one or more nucleic acids encoding the engineered CRISPR-Cas system may be any nucleic acid including DNA, RNA, or combinations thereof. The one or more nucleic acids may comprise one or more messenger RNAs, one or more vectors, or any combination thereof.
[0080] The Cas3 and the other Cas proteins, e.g., the engineered Cas protein(s) or fusion protein(s) may be encoded by a single nucleic acid e.g., a single vector). The Cas3 and the other Cas proteins, e.g., the engineered Cas protein(s) or fusion protein(s) may be encoded by different nucleic acids e.g., multiple mRNAs or two or more vectors).
[0081] Engineering the system for use in eukaryotic cells may involve codon-optimization or other modification (e.g., to include an appropriate nuclear localization signal (NLS) or purification tag). It will be appreciated that changing native codons to those most frequently used in mammals allows for maximum expression of the system proteins in mammalian cells e.g., human cells). Such modified nucleic acids are commonly described in the art as “codon-optimized,” or as utilizing “mammalian-preferred” or “human-preferred” codons. The nucleic acid is considered codon-optimized herein if at least about 60% e.g., 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 98%) of the codons encoded therein are mammalian preferred codons.
[0082] Illustrative Pae Type l-F CRISPR Cas proteins and illustrative DNAs encoding the Cas proteins include, but are not limited to, the following. “NLS” is a nuclear localization signal. “HA tag” is a hemagglutinin tag.
[0083] Pae-cas5 protein with NLS and HA tag (SEQ ID NO: 9)MSVTDPEALLLLPRLSIQNANAISSPLTWGFPSPGAFTGFVHALQRRVGISLDIELDGVGIVC HRFEAQISQPAGKRTKVFNLTRNPLNRDGSTAAIVEEGRAHLEVSLLLGVHGDGLDDHPAQ EIARQVQEQAGAMRLAGGSILPWCNERFPAPNAELLMLGGSDEQRRKNQRRLTRRLLPGF ALVSREALLQQHLETLRTTLPEATTLDALLDLCRINFEPPATSSEEEASPPDAAWQVRDKPG WLVPIPAGYNALSPLYLPGEVRNARDRETPLRFVENLFGLGEWLSPHRVAALSDLLWYHHA EPDKGLYRWSTPRFVEHAIAGSVGYPYDVPDYAGSYPEFPKKKRKV
[0084] Human codon optimized DNA encoding Pae-cas5 with NLS and HA tag (SEQ ID NO: 10)ATGAGCGTGACCGATCCTGAAGCTCTGCTGCTGCTCCCCAGACTGAGCATCCAGAACG CCAACGCCATCAGCAGCCCTCTGACATGGGGATTTCCAAGTCCTGGCGCCTTCACCGG ATTTGTGCACGCCCTGCAGAGAAGAGTGGGCATCAGCCTGGACATCGAGCTGGATGG CGTGGGCATCGTGTGCCACAGATTTGAGGCCCAGATCTCTCAGCCTGCCGGCAAGAGA ACAAAGGTGTTCAACCTGACCAGAAATCCCCTGAACCGCGACGGATCTACAGCCGCCA TTGTGGAAGAAGGCAGAGCCCACCTGGAAGTGTCACTGCTGCTTGGAGTTCACGGCGA CGGCCTGGATGATCACCCTGCTCAAGAGATCGCCAGACAGGTGCAAGAACAGGCTGG CGCCATGAGACTTGCCGGCGGATCTATTCTGCCCTGGTGCAACGAGAGATTCCCCGCT CCTAATGCCGAACTGCTGATGCTCGGAGGCTCCGATGAACAGCGGAGAAAGAACCAGC GGAGACTGACCCGTAGACTGCTGCCTGGATTTGCTCTGGTGTCCAGAGAAGCCCTGCT CCAGCAGCACCTGGAAACCCTGAGAACCACACTGCCTGAGGCCACCACACTGGATGCA CTGCTGGACCTGTGCCGGATCAACTTTGAGCCTCCTGCCACCAGCAGCGAGGAAGAAG CCTCTCCTCCTGACGCTGCTTGGCAAGTGCGAGATAAGCCTGGATGGCTGGTGCCTAT TCCTGCCGGCTACAATGCCCTGTCTCCTCTGTATCTGCCTGGCGAAGTGCGGAACGCCAGAGACAGAGAGACACCCCTGAGATTCGTGGAAAACCTGTTCGGCCTCGGCGAGTGGCTGTCTCCACATAGAGTTGCCGCTCTGAGCGACCTGCTGTGGTATCACCATGCCGAGCCTGACAAGGGCCTGTACAGATGGTCTACCCCACGCTTCGTGGAACACGCCATTGCCGGCTCTGTGGGCTACCCCTACGATGTGCCTGATTACGCCGGCAGCTACCCTGAGTTCCCCAAGAAAAAGCGGAAAGTGTGA
[0085] Pae-cas6 protein with NLS and HA tag (SEQ ID NO: 11 )MDHYLDIRLRPDPEFPPAQLMSVLFGKLHQALVAQGGDRIGVSFPDLDESRSRLGERLRIHASADDLRALLARPWLEGLRDHLQFGEPAVVPHPTPYRQVSRVQAKSNPERLRRRLMRRHDLSEEEARKRIPDTVARALDLPFVTLRSQSTGQHFRLFIRHGPLQVTAEEGGFTCYGLSKGGFVPWFGSVGYPYDVPDYAGSYPEFPKKKRKV
[0086] Human codon optimized DNA encoding Pae-cas6 with NLS and HA tag (SEQ IDNO: 12)ATGGACCACTACCTGGACATCAGACTGCGGCCCGATCCTGAGTTTCCTCCAGCTCAGCTGATGAGCGTGCTGTTCGGCAAACTGCATCAGGCCCTGGTTGCCCAAGGCGGCGATAGAATCGGAGTGTCATTCCCCGACCTGGACGAGAGCAGATCTAGACTGGGCGAGAGACTGAGAATCCACGCCAGCGCCGATGATCTGAGAGCCCTGCTTGCTAGACCTTGGCTGGAAGGCCTGAGGGACCATCTGCAGTTTGGAGAACCTGCCGTGGTGCCCCATCCTACACCTTACAGACAGGTGTCCAGAGTGCAGGCCAAGAGCAACCCCGAAAGACTGCGGAGAAGGCTGATGCGGAGACACGACCTGTCTGAGGAAGAGGCCCGGAAGAGAATCCCCGACACAGTGGCTAGAGCCCTGGACCTGCCTTTCGTGACACTGAGAAGCCAGAGCACCGGCCAGCACTTCCGGCTGTTTATCAGACACGGCCCTCTGCAAGTGACCGCCGAAGAAGGCGGCTTTACCTGTTACGGCCTGAGCAAAGGCGGATTCGTGCCCTGGTTTGGCTCTGTGGGCTACCCCTACGATGTGCCTGATTACGCCGGCAGCTACCCTGAGTTCCCCAAGAAAAAGCGGAAAGTGTGA
[0087] Pae-cas7 protein with NLS and HA tag (SEQ ID NO: 13)
[0088] MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKDRDPAKLDASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTVATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGLRDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILDKGDKKGQKSKTLYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVEPYGSVTSQGKAYRQPKQKLDFYTLLDNWVLRDEAPAVEQQHYVIANLIRGGVFGEAEEKGSVGYPYDVPDYAGSYPEFPKKKRKV*
[0089] Human codon optimized DNA encoding Pae-cas7 with NLS and HA tag (SEQ ID NO: 14)ATGAGCAAGCCCATCCTGAGCACAGCCAGCGTGCTGGCCTTTGAGAGAAAGCTGGACCCCTCTGACGCCCTGATGTCTGCTGGTGCTTGGGCTCAGAGGGATGCCTCTCAAGAATGGCCTGCTGTGACCGTGCGGGAAAAGTCTGTGCGGGGCACCATCAGCAACCGGCTGAAAACAAAGGACAGGGACCCCGCCAAACTGGACGCCTCTATCCAGTCTCCTAACCTGCAGACCGTGGACGTGGCCAATCTGCCTTCCGATGCCGACACACTGAAAGTGCGGTTCACCCTGAGAGTGCTCGGCGGAGCTGGAACACCTAGCGCCTGTAATGATGCCGCCTACCGGGATAAGCTGCTGCAGACAGTGGCCACCTACGTGAACGATCAGGGCTTTGCCGAGCTGGCCAGAAGATACGCCCACAACCTGGCCAACGCCAGATTCCTGTGGCGGAATAGAGTGGGAGCCGAGGCTGTGGAAGTGCGGATCAACCACATCAGACAGGGCGAAGTGGCCAGAGCTTGGAGATTTGATGCCCTGGCCATCGGCCTGAGAGACTTTAAGGCCGACGCCGAACTGGATGCTCTGGCCGAACTGATTGCCAGCGGCCTGTCTGGATCTGGACACGTGCTGCTGGAAGTGGTGGCCTTTGCCAGAATCGGCGACGGCCAAGAAGTGTTCCCTAGCCAAGAGCTGATCCTGGACAAGGGCGACAAGAAGGGCCAGAAGTCTAAGACCCTGTACAGCGTGCGGGATGCCGCCGCTATTCACTCTCAGAAGATCGGAAACGCCCTGCGGACCATCGACACCTGGTATCCCGATGAGGATGGCCTGGGACCTATCGCCGTGGAACCTTACGGCAGCGTGACATCTCAGGGCAAAGCCTACAGACAGCCTAAGCAGAAGCTGGACTTCTACACCCTGCTGGACAACTGGGTGCTGAGAGATGAAGCCCCTGCTGTGGAACAGCAGCACTACGTGATCGCCAACCTGATCAGAGGCGGCGTGTTCGGAGAGGCCGAGGAAAAAGGCTCTGTGGGCTACCCCTACGATGTGCCTGATTACGCCGGCAGCTACCCTGAGTTCCCCAAGAAAAAGCGGAAAGTGTGA
[0090] Pae-cas8 protein with NLS and HA tag (SEQ ID NO: 15)
[0091] MTSPLPTPTWQELRQFIESFIQERLQGKLDKLQPDEDDKRQTLLATHRREAWLADAARRVGQLQLVTHTLKPIHPDARGSNLHSLPQAPGQPGLAGSHELGDRLVSDVVGNAAALDVFKFLSLQYQGKNLLNWLTEDSAEALQALSDNAEQAREWRQAFIGITTVKGAPASHSLAKQLYFPLPGSGYHLLAPLFPTSLVHHVHALLREARFGDAAKAAREARSRQESWPHGFSEYPNLAIQKFGGTKPQNISQLNNERRGENWLLPSLPPNWQRQNVNAPMRHSSVFEHDFGRTPEVSRLTRTLQRFLAKTVHNNLAIRQRRAQLVAQICDEALQYAARLRELEPGWSATPGCQLHDAEQLWLDPLRAQTDETFLQRRLRGDWPAEVGNRFANWLNRAVSSDSQILGSPEAAQWSQELSKELTMFKEILEDERDGSVGYPYDVPDYAGSYPEFPKKKRKV*
[0092] Human codon optimized DNA encoding Pae-cas8 with NLS and HA tag (SEQ IDNO: 16)ATGACAAGCCCTCTGCCTACACCTACCTGGCAAGAGCTGCGGCAGTTCATCGAGAGCTTCATCCAAGAGCGGCTGCAGGGCAAGCTGGACAAGCTGCAGCCTGACGAGGACGACAAGAGACAGACACTGCTGGCCACACACAGAAGAGAAGCCTGGCTGGCCGATGCCGCTAGAAGAGTTGGACAGCTCCAGCTGGTCACCCACACACTGAAGCCTATTCACCCCGATGCCAGAGGCAGCAACCTGCATTCTCTGCCTCAAGCTCCAGGCCAGCCTGGACTGGCTGGATCTCATGAGCTGGGCGATAGACTGGTGTCCGACGTGGTGGGAAATGCCGCTGCTCTGGACGTGTTCAAGTTCCTGAGCCTGCAGTACCAGGGCAAGAACCTGCTGAACTGGCTGACCGAGGATAGCGCCGAAGCTCTGCAGGCCCTGTCTGATAATGCCGAGCAGGCTAGAGAATGGCGGCAGGCCTTTATCGGCATCACCACAGTGAAAGGCGCCCCTGCCTCTCACTCTCTGGCCAAGCAGCTGTACTTTCCCCTGCCTGGCTCTGGCTACCATCTGCTGGCTCCTCTGTTTCCCACAAGCCTGGTGCATCATGTGCACGCCCTGCTGAGAGAGGCCAGATTTGGCGACGCTGCCAAAGCCGCCAGAGAAGCCAGAAGCAGACAAGAGTCTTGGCCCCACGGCTTCAGCGAGTACCCTAATCTGGCCATCCAGAAGTTCGGCGGCACCAAGCCTCAGAACATCAGCCAGCTGAACAACGAGCGGAGAGGCGAGAATTGGCTGCTGCCAAGCCTGCCTCCTAACTGGCAGAGACAGAACGTGAACGCCCCTATGAGACACAGCAGCGTGTTCGAGCACGACTTCGGCAGAACACCCGAGGTGTCCAGACTGACAAGAACCCTGCAGAGATTCCTGGCCAAGACCGTGCACAACAACCTGGCCATCAGACAGAGAAGGGCCCAGCTGGTGGCCCAGATTTGTGATGAGGCCCTGCAGTATGCCGCCAGACTGAGAGAATTGGAGCCAGGCTGGAGCGCCACACCTGGTTGTCAACTGCATGACGCCGAACAGCTGTGGCTGGACCCTCTGAGAGCCCAGACCGATGAGACATTCCTGCAGCGAAGGCTGAGAGGCGATTGGCCAGCCGAAGTGGGCAACAGATTTGCTAATTGGCTGAACCGGGCCGTGTCCAGCGATTCTCAGATCCTGGGATCTCCTGAGGCCGCTCAGTGGTCCCAAGAGCTGAGCAAAGAACTGACCATGTTCAAAGAGATTCTCGAGGACGAGCGGGACGGCTCTGTGGGCTACCCCTACGATGTGCCTGATTACGCCGGCAGCTACCCTGAGTTCCCCAAGAAAAAGCGGAAAGTGTGA
[0093] Pae-cas2-cas3 protein with NLS and HA tag (SEQ ID NO: 17)MNILLVSQCEKRALSETRRILDQFAERRGERTWQTPITQAGLDTLRRLLKKSARRNTAVACHWIRGRDHSELLWIVGDASRFNAQGAVPTNRTCRDILRKEDENDWHSAEDIRLLTVMAALFHDIGKASQAFQAKLRNRGKPMADAYRHEWVSLRLFEAFVGPGSSDEDWLRRLADKRETGDAWLSQLARDDRQSAPPGPFQKSRLPPLAQAVGWLIVSHHRLPNGDHRGSASLARLPAPIQSQWCGARDADAKEKAACWQFPHGLPFASAHWRARTALCAQSMLERPGLLARGPALLHDSYVMHVSRLILMLADHHYSSLPADSRLGDPNFPLHANTDRDSGKLKQRLDEHLLGVALHSRKLAGTLPRLERQLPRLARHKGFTRRVEQPRFRWQDKAYDCAMACREQAMEHGFFGLNLASTGCGKTLANGRILYALADPQRGARFSIALGLRSLTLQTGQAYRERLGLGDDDLAILVGGSAARELFEKQQERLERSGSESAQELLAENSHVHFAGTLEDGPLREWLGRNSAGNRLLQAPILACTIDHLMPASESLRGGHQIAPLLRLMTSDLVLDEVDDFDIDDLPALSRLVHWAGLFGSRVLLSSATLPPALVQGLFEAYRSGREIFQRHRGAPGRATEIRCAWFDEFSSQSSAHGAVTSFSEAHATFVAQRLAKLEQLPPRRQAQLCTVHAAGEARPALCRELAGQMNTWMADLHRCHHTEHQGRRISFGLLRLANIEPLIELAQAILAQGAPEGLHVHLCVYHSRHPLLVRSAIERQLDELLKRSDDDAAALFARPTLAKALQASTERDHLFVVLASPVAEVGRDHDYDWAIVEPSSMRSIIQLAGRIRRHRSGFSGEANLYLLSRNIRSLEGQNPAFQRPGFETPDFPLDSHDLHDLLDPALLARIDASPRIVEPFPLFPRSRLVDLEHRRLRALMLADDPPSSLLGVPLWWQTPASLSGALQTSQPFRAGAKERCYALLPDEDDEERLHFSRYEEGTWSNQDNLLRNLDLTYGPRIQTWGTVNYREELV AMAGREDLDLRQCAMRYGEVRLRENTQGWSYHPYLGFKKYNLGSVGYPYDVPDYAGSYPEFPKKKRKV
[0094] Human codon optimized DNA encoding Pae-cas2-cas3 with NLS and HA tag (SEQ ID NO: 18)ATGAACATCCTGCTGGTGTCCCAGTGCGAGAAGAGAGCCCTGAGCGAGACAAGACGGATCCTGGATCAGTTCGCCGAGCGGAGAGGCGAGAGAACATGGCAGACACCTATCACACAGGCCGGACTGGACACCCTGCGGAGACTGCTGAAGAAGTCCGCCAGACGGAATACCGCCGTGGCCTGTCACTGGATCAGAGGCAGAGATCACTCCGAGCTGCTGTGGATCGTGGGCGACGCCTCTAGATTCAATGCTCAGGGCGCCGTGCCTACCAACAGAACCTGCAGAGACATCCTGCGGAAAGAGGACGAGAACGACTGGCACAGCGCCGAGGATATCAGGCTGCTGACAGTGATGGCCGCTCTGTTCCACGATATCGGCAAAGCCAGCCAGGCCTTCCAGGCCAAGCTGAGAAATAGAGGCAAGCCCATGGCCGACGCCTACAGACATGAATGGGTGTCCCTGAGACTGTTCGAGGCCTTTGTTGGCCCTGGCAGCTCCGATGAAGATTGGCTGAGAAGGCTGGCCGACAAGAGAGAAACAGGCGACGCTTGGCTGTCTCAGCTGGCCAGAGATGACAGACAGTCTGCCCCTCCTGGACCTTTCCAGAAGTCCAGATTGCCTCCTCTGGCTCAGGCCGTTGGCTGGCTGATTGTGTCTCACCACAGACTGCCCAACGGCGACCATAGAGGCTCTGCTTCTCTGGCTAGACTGCCCGCTCCTATCCAGTCTCAGTGGTGCGGAGCTAGAGATGCCGACGCCAAAGAGAAAGCCGCCTGCTGGCAGTTTCCTCACGGCCTGCCTTTTGCCAGCGCTCATTGGAGAGCCAGAACAGCCCTGTGTGCCCAGAGCATGCTGGAAAGACCTGGACTGCTGGCTAGAGGCCCTGCTCTGCTGCACGATAGCTACGTGATGCACGTGTCCCGGCTGATCCTGATGCTGGCCGATCACCACTACTCTAGCCTGCCTGCCGATAGCAGACTGGGCGACCCTAATTTTCCCCTGCACGCCAACACCGACAGAGACAGCGGAAAGCTGAAGCAGCGGCTGGACGAACATCTGCTTGGAGTGGCCCTGCACTCCAGAAAGCTGGCTGGAACACTGCCCAGACTGGAACGCCAACTGCCTAGGCTGGCAAGACACAAGGGCTTCACCAGAAGAGTGGAACAGCCCCGGTTCCGGTGGCAGGATAAGGCCTATGATTGCGCCATGGCCTGTAGAGAACAGGCTATGGAACACGGCTTCTTCGGCCTGAATCTGGCCTCTACCGGATGCGGCAAAACCCTGGCCAATGGCAGAATCCTGTACGCCCTGGCTGACCCTCAAAGAGGCGCCAGATTTTCTATCGCCCTGGGCCTGAGAAGCCTGACACTGCAGACAGGCCAGGCCTACAGAGAGAGACTCGGCCTGGGAGATGACGACCTGGCCATTCTCGTTGGAGGATCTGCCGCCAGAGAGCTGTTTGAGAAGCAGCAAGAGAGACTGGAAAGATCCGGCAGCGAGAGCGCCCAAGAACTGCTCGCCGAAAATTCCCACGTGCACTTCGCCGGCACACTGGAAGATGGACCTCTGAGAGAGTGGCTGGGCAGAAACAGCGCCGGCAATAGACTGCTGCAGGCCCCTATTCTGGCCTGCACCATCGATCATCTGATGCCCGCCAGCGAGTCTCTGAGAGGCGGACATCAAATTGCCCCTCTGCTGCGGCTGATGACCAGCGATCTGGTGCTGGATGAGGTGGACGACTTCGACATCGACGACCTGCCTGCTCTGTCCAGACTGGTGCATTGG GCTGGCCTGTTTGGCAGCAGAGTGCTGCTGTCTAGCGCCACACTTCCTCCAGCTCTGG TGCAGGGACTGTTTGAAGCCTACAGAAGCGGCAGAGAGATCTTCCAGAGACACAGAGG CGCTCCAGGCAGAGCCACCGAGATTAGATGCGCTTGGTTTGACGAGTTCAGCAGCCAG TCTAGCGCACATGGCGCCGTGACAAGCTTTTCTGAAGCCCACGCCACCTTCGTGGCCC AGAGACTGGCTAAACTCGAGCAGCTGCCACCTCGGAGACAGGCCCAACTGTGTACAGT TCATGCCGCTGGCGAAGCCAGACCTGCACTGTGCAGAGAACTGGCCGGACAGATGAA CACCTGGATGGCCGATCTGCACAGATGCCACCACACCGAGCACCAGGGCAGAAGAAT CTCTTTCGGCCTGCTGAGGCTGGCCAACATCGAGCCTCTGATTGAACTGGCCCAGGCC ATCCTTGCTCAAGGCGCTCCTGAAGGACTGCATGTGCACCTGTGCGTGTACCACAGCA GACACCCTCTGCTCGTCAGAAGCGCCATCGAGAGACAGCTGGACGAGCTGCTCAAGA GAAGCGACGATGATGCCGCCGCACTGTTCGCCAGACCAACACTTGCTAAAGCCCTGCA GGCCTCCACCGAGAGGGATCACCTGTTTGTGGTGCTGGCCTCTCCTGTGGCCGAAGT GGGAAGAGATCACGACTACGACTGGGCCATTGTGGAACCCAGCAGCATGCGGAGCAT CATCCAGCTGGCCGGCAGAATCAGAAGGCACAGATCTGGCTTTAGCGGCGAGGCCAA CCTGTACCTGCTGAGCCGGAATATCAGATCCCTGGAAGGACAGAACCCCGCCTTTCAG AGGCCCGGCTTTGAGACACCTGACTTCCCACTGGACTCCCACGACCTGCACGATCTGC TTGATCCAGCTCTCCTGGCCAGAATCGACGCTAGCCCCAGAATCGTCGAGCCCTTTCC ACTGTTCCCTAGAAGCAGGCTGGTGGACCTGGAACACAGACGGCTGAGAGCCCTCATG CTGGCTGACGATCCTCCATCTTCTCTGCTGGGAGTGCCACTCTGGTGGCAGACTCCAG CTTCTTTGTCTGGCGCCCTGCAGACCAGCCAGCCTTTTAGAGCTGGCGCCAAAGAACG GTGCTACGCCCTGCTGCCAGACGAGGACGATGAGGAAAGACTGCACTTCAGCAGATAC GAGGAAGGCACCTGGTCCAACCAGGACAACCTGCTGCGGAACCTGGACCTGACATAC GGCCCTAGAATCCAGACCTGGGGCACCGTGAACTACCGCGAAGAACTGGTGGCCATG GCTGGCAGAGAGGACCTGGATCTGAGACAGTGCGCCATGAGATACGGGGAAGTGCGG CTGCGCGAGAATACCCAAGGCTGGTCTTATCACCCCTACCTGGGCTTCAAGAAGTACA ACCTGGGCTCTGTGGGCTACCCCTACGATGTGCCTGATTACGCCGGCAGCTACCCTGA GTTCCCCAAGAAAAAGCGGAAAGTGTGA
[0095] Illustrative Pae Type l-F CRISPR Cas variant proteins include, but are not limited to, the following hyper-activity variant proteins.
[0096] Pae-cas7 E338R protein with NLS and HA tag (SEQ ID NO: 19)MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKD RDPAKLDASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTVATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGLRDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILDKGDKKGQKSKTLYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVEPYGSVTSQGKAYRQPKQKLDFYTLLDNWVLRDEAPAVEQQHYVIANLIRGGVFGRAEEKGSVGYPYDVPDYAGSYPEFPKKKRKV*
[0097] Pae-cas7 D234R protein with NLS and HA tag (SEQ ID NO: 20)MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKDRDPAKLDASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTVATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGLRDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILRKGDKKGQKSKTLYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVEPYGSVTSQGKAYRQPKQKLDFYTLLDNWVLRDEAPAVEQQHYVIANLIRGGVFGEAEEKGSVGYPYDVPDYAGSYPEFPKKKRKV
[0098] Pae-cas7 D237R protein with NLS and HA tag (SEQ ID NO: 21 )MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKDRDPAKLDASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTVATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGLRDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILDKGRKKGQKSKTLYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVEPYGSVTSQGKAYRQPKQKLDFYTLLDNWVLRDEAPAVEQQHYVIANLIRGGVFGEAEEKGSVGYPYDVPDYAGSYPEFPKKKRKV
[0099] Pae-cas7 D68R protein with NLS and HA tag (SEQ ID NO: 22)MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKDRDPAKLRASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTVATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGLRDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILDKGDKKGQKSKTLYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVEPYGSVTSQGKAYRQPKQKLDFYTLLDNWVLRDEAPAVEQQHYVIANLIRGGVFGEAEEKGSVGYPYDVPDYAGSYPEFPKKKRKV
[0100] Pae-cas8 D76R protein with NLS and HA tag (SEQ ID NO: 24)MTSPLPTPTWQELRQFIESFIQERLQGKLDKLQPDEDDKRQTLLATHRREAWLADAARRVGQLQLVTHTLKPIHPRARGSNLHSLPQAPGQPGLAGSHELGDRLVSDVVGNAAALDVFKFLSLQYQGKNLLNWLTEDSAEALQALSDNAEQAREWRQAFIGITTVKGAPASHSLAKQLYFPLPGSGYHLLAPLFPTSLVHHVHALLREARFGDAAKAAREARSRQESWPHGFSEYPNLAIQKFGGTKPQNISQLNNERRGENWLLPSLPPNWQRQNVNAPMRHSSVFEHDFGRTPEVSRLTRTLQRFLAKTVHNNLAIRQRRAQLVAQICDEALQYAARLRELEPGWSATPGCQLHDAEQLWLDPLRAQTDETFLQRRLRGDWPAEVGNRFANWLNRAVSSDSQILGSPEAAQWSQELSKELTMFKEILEDERDGSVGYPYDVPDYAGSYPEFPKKKRKV
[0101] Pae-cas8 E421 R protein with NLS and HA tag (SEQ ID NO: 25)MTSPLPTPTWQELRQFIESFIQERLQGKLDKLQPDEDDKRQTLLATHRREAWLADAARRVG QLQLVTHTLKPIHPDARGSNLHSLPQAPGQPGLAGSHELGDRLVSDVVGNAAALDVFKFLS LQYQGKNLLNWLTEDSAEALQALSDNAEQAREWRQAFIGITTVKGAPASHSLAKQLYFPLP GSGYHLLAPLFPTSLVHHVHALLREARFGDAAKAAREARSRQESWPHGFSEYPNLAIQKFG GTKPQNISQLNNERRGENWLLPSLPPNWQRQNVNAPMRHSSVFEHDFGRTPEVSRLTRT LQRFLAKTVHNNLAIRQRRAQLVAQICDEALQYAARLRELEPGWSATPGCQLHDAEQLWLD PLRAQTDETFLQRRLRGDWPAEVGNRFANWLNRAVSSDSQILGSPEAAQWSQELSKRLT MFKEILEDERDGSVGYPYDVPDYAGSYPEFPKKKRKV
[0102] The disclosure also provides the following variant protein.
[0103] Pae-cas7 E283R protein with NLS and HA tag (SEQ ID NO: 23)MSKPILSTASVLAFERKLDPSDALMSAGAWAQRDASQEWPAVTVREKSVRGTISNRLKTKD RDPAKLDASIQSPNLQTVDVANLPSDADTLKVRFTLRVLGGAGTPSACNDAAYRDKLLQTV ATYVNDQGFAELARRYAHNLANARFLWRNRVGAEAVEVRINHIRQGEVARAWRFDALAIGL RDFKADAELDALAELIASGLSGSGHVLLEVVAFARIGDGQEVFPSQELILDKGDKKGQKSKT LYSVRDAAAIHSQKIGNALRTIDTWYPDEDGLGPIAVRPYGSVTSQGKAYRQPKQKLDFYTL LDNWVLRDEAPAVEQQHYVIANLIRGGVFGEAEEKGSVGYPYDVPDYAGSYPEFPKKKRK V
[0104] The system disclosed herein comprises a guide RNA (gRNA), wherein the gRNA is configured to specifically hybridize to a target nucleic acid. The “guide RNA” or “gRNA” comprises a guide sequence that determines the binding specificity of the CRISPR-Cas complex. The terms “guide sequence,” “guide,” and “spacer” are used interchangeably herein and refer to the nucleotide sequence within a gRNA that specifies the target site.
[0105] The gRNA may be encoded by the same or different nucleic acid as any of the Cas3 and the other Cas proteins, e.g., the engineered Cas protein(s) or fusion protein(s). For example, a single vector may encode any or all of the gRNA, the Cas3, and the Cascade proteins, e.g., the engineered Cas protein(s) or fusion protein(s).
[0106] The terms “target DNA,” “target nucleic acid,” “target sequence,” and “target site” are used interchangeably herein to refer to a polynucleotide (nucleic acid, gene, chromosome, genome, etc.) to which a guide sequence e.g., in a guide RNA) is designed to have complementarity, wherein hybridization between the target DNA and a guide sequence promotes the formation of a CRISPR / Cas complex. The target DNA and guide sequenceneed not exhibit complete complementarity, provided that there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex.
[0107] The strand of the target DNA that is complementary to and hybridizes with the guide RNA is referred to as the “complementary strand” and the strand of the target DNA that is complementary to the “complementary strand” (and is therefore not complementary to the guide RNA) is referred to as the “noncomplementary strand” or “non-complementary strand.”
[0108] The target nucleic acid includes a protospacer adjacent motif (PAM). A PAM may be 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleotides in length. A PAM may be between 2-6 nucleotides in length. The PAM may be 2 nucleotides in length. The PAM is “adjacent to” the target nucleic acid sequence in that it typically immediately precedes the target sequence, that is, is 5’ of the target site.
[0109] PAM sequences are typically specific to the particular Cas endonuclease being used in the CRISPR / Cas complex and the species from which it was derived. For example, Type l-F CRISPR-Cas3 elements typically are active in a genome which comprises a protospacer adjacent motif (PAM) comprising the nucleic acid sequence CC located adjacent to the target genomic DNA sequence. PAM sequences and methods of determining PAM sequences for specific Cas proteins are known in the art. The gRNA or portion thereof ( / .e., the guide sequence) that hybridizes to a target nucleic acid may be between any length.
[0110] The guide sequence of the gRNA does not need to be completely complementary to the target site. The guide sequence of the gRNA may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the target site. The gRNA sequence may be at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or at least 100% complementary to the 3’ end of the target site (e.g., the last 5, 6, 7, 8, 9, or 10 nucleotides of the 3’ end of the target site
[0111] A gRNA herein is a non-naturally occurring gRNA.
[0112] Any element of any suitable CRISPR / Cas gene editing system known in the art can be employed in the systems and methods described herein, as understood by the ordinarily skilled person.
[0113] Conventional viral and non-viral based gene transfer methods can be used to introduce nucleic acids encoding components of the present system into cells, tissues, or a subject. Such methods can be used to administer nucleic acids encoding components of the present system to cells in culture, or in a host organism. Non-viral vector delivery systemsinclude DNA plasmids, cosmids, RNA e.g., a transcript of a vector described herein), and a nucleic acid complexed with a delivery vehicle.
[0114] Viral vector delivery systems include DNA and RNA viruses, which have either episomal or integrated genomes after delivery to the cell. A variety of viral constructs may be used to deliver the present system and / or components to the cells, tissues, and / or a subject. Viral vectors include, for example, retroviral, lentiviral, adenoviral, adeno-associated and herpes simplex viral vectors. Nonlimiting examples of such recombinant viruses include recombinant adeno-associated virus (AAV), recombinant adenoviruses, recombinant lentiviruses, recombinant retroviruses, recombinant herpes simplex viruses, recombinant poxviruses, phages, etc. See, e.g., Ausubel et al., Current Protocols in Molecular Biology, John Wiley & Sons, New York, 1989; Kay, M. A., et al., 2001 Nat. Medic. 7(1 ) :33-40; and Walther W. and Stein II., 2000 Drugs, 60(2): 249-71 .
[0115] Drug selection strategies may be adopted for positively selecting for cells comprising the nucleic acids encoding the present system or components thereof.
[0116] The present disclosure also provides for DNA segments encoding the proteins and nucleic acids disclosed herein, vectors containing these segments and cells containing the vectors. The vectors may be used to propagate the segment in an appropriate cell and / or to allow expression from the segment e.g., an expression vector). The person of ordinary skill in the art would be aware of the various vectors available for propagation and expression of a nucleic acid.
[0117] To construct cells that express the present system, expression vectors for stable or transient expression of the present system may be constructed via conventional methods and introduced into cells. For example, nucleic acids encoding the components of the present system may be cloned into a suitable expression vector, such as a plasmid or a viral vector in operable linkage to a suitable promoter. The selection of expression vectors / plasmids / viral vectors should be suitable for integration and replication in eukaryotic cells.
[0118] Vectors of the present disclosure can drive the expression of one or more nucleic acids in mammalian cells using a mammalian expression vector. Examples of mammalian expression vectors include pCDM8 (Seed, Nature (1987) 329:840, incorporated herein by reference) and pMT2PC (Kaufman, et al., EMBO J. (1987) 6:187, incorporated herein by reference). When used in mammalian cells, the expression vector's control functions are typically provided by one or more regulatory elements. For example, commonly used promoters are derived from polyoma, adenovirus 2, cytomegalovirus, simian virus 40, and others disclosed herein and known in the art. For other suitable expression systems for bothprokaryotic and eukaryotic cells see, e.g., Chapters 16 and 17 of Sambrook, et al., MOLECULAR CLONING: A LABORATORY MANUAL., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., incorporated herein by reference.
[0119] Vectors of the present disclosure can comprise any of a number of promoters known to the art, wherein the promoter is constitutive, regulatable or inducible, cell type specific, tissue-specific, or species specific. In addition to the sequence sufficient to direct transcription, a promoter can also include other regulatory elements that are involved in modulating transcription e.g., enhancers, Kozak sequences and introns). Many promoter / regulatory sequences useful for driving constitutive expression of a gene are available in the art and include, but are not limited to, for example, CMV (cytomegalovirus promoter), EF1a (human elongation factor 1 alpha promoter), SV40 (simian vacuolating virus 40 promoter), PGK (mammalian phosphoglycerate kinase promoter), Ubc (human ubiquitin C promoter), human beta-actin promoter, rodent beta-actin promoter, CBh (chicken betaactin promoter), CAG (hybrid promoter contains CMV enhancer, chicken beta actin promoter, and rabbit beta-globin splice acceptor), TRE (Tetracycline response element promoter), H1 (human polymerase III RNA promoter), U6 (human U6 small nuclear promoter), and the like. Additional promoters that can be used for expression of the components of the present system, include, without limitation, cytomegalovirus (CMV) intermediate early promoter, a viral LTR such as the Rous sarcoma virus LTR, HIV-LTR, HTLV-1 LTR, Maloney murine leukemia virus (MMLV) LTR, myeloproliferative sarcoma virus (MPSV) LTR, spleen focus-forming virus (SFFV) LTR, the simian virus 40 (SV40) early promoter, herpes simplex tk virus promoter, elongation factor 1 -alpha (EF1 -a) promoter with or without the EF1 -a intron. Additional promoters include any constitutively active promoter. Alternatively, any regulatable promoter may be used, such that its expression can be modulated within a cell. Thus, as one example, protein coding genes may be driven by an EF1 alpha promoter. As another example, crRNA expression may be driven by a U6 promoter.
[0120] Moreover, inducible expression can be accomplished by placing the nucleic acid encoding such a molecule under the control of an inducible promoter / regulatory sequence. Promoters well known in the art can be induced in response to inducing agents such as metals, glucocorticoids, tetracycline, hormones, and the like, are also contemplated for use herein. Thus, it will be appreciated that the present disclosure includes the use of any promoter / regulatory sequence known in the art that is capable of driving expression of the desired protein operably linked thereto.
[0121] The vectors of the present disclosure may direct the expression of the nucleic acid in a particular cell type e.g., tissue-specific regulatory elements are used to express thenucleic acid). Such regulatory elements include promoters that may be tissue specific or cell specific. The term “tissue specific” as it applies to a promoter refers to a promoter that is capable of directing selective expression of a nucleic acid of interest in a specific type of tissue {e.g., seeds) in the relative absence of expression of the same nucleic acid of interest in a different type of tissue. The term “cell type specific” as applied to a promoter refers to a promoter that is capable of directing selective expression of a nucleic acid of interest in a specific type of cell in the relative absence of expression of the same nucleic acid of interest in a different type of cell within the same tissue. The term “cell type specific” when applied to a promoter also means a promoter capable of promoting selective expression of a nucleic acid of interest in a region within a single tissue. Cell type specificity of a promoter may be assessed using methods well known in the art, e.g., immunohistochemical staining.
[0122] Additionally, the vector may contain, for example, some or all of the following: a selectable marker gene, such as the neomycin gene for selection of stable or transient transfectants in host cells; enhancer / promoter sequences from the immediate early gene of human CMV for high levels of transcription; transcription termination and RNA processing signals from SV40 for mRNA stability; 5’-and 3’-untranslated regions for mRNA stability and translation efficiency from highly-expressed genes like a-globin or f-globin; SV40 polyoma origins of replication and ColE1 for proper episomal replication; internal ribosome binding sites (IRESes), versatile multiple cloning sites; T7 and SP6 RNA promoters for in vitro transcription of sense and antisense RNA; a “suicide switch” or “suicide gene” which when triggered causes cells carrying the vector to die e.g., HSV thymidine kinase, an inducible caspase such as iCasp9), and reporter gene for assessing expression.
[0123] When introduced into a cell, the vectors may be maintained as an autonomously replicating sequence or extrachromosomal element or may be integrated into host DNA.
[0124] The present system or components thereof e.g., Cas proteins or fusion proteins as described herein) may be delivered to a cell by any suitable means. The system may be delivered in vivo. Vectors according to the present disclosure can be transformed, transfected, or otherwise introduced into a wide variety of cells. Numerous methods of transfection are known to the ordinarily skilled artisan, for example, lipofectamine, calcium phosphate co-precipitation, electroporation, DEAE-dextran treatment, microinjection, viral infection, and other methods known in the art. Transduction refers to entry of a virus into the cell and expression {e.g., transcription and / or translation) of nucleic acid delivered by the viral vector genome. In the case of a recombinant vector, “transduction” generally refers to entry of the recombinant viral vector into the cell and expression of a nucleic acid of interest delivered by the vector genome.
[0125] Any of the vectors comprising a nucleic acid that encodes the components of the present system is also within the scope of the present disclosure. Such a vector may be delivered into cells by a suitable method. Methods of delivering vectors to cells are well known in the art and may include DNA or RNA electroporation, transfection reagents such as liposomes or nanoparticles to delivery DNA or RNA; delivery of DNA, RNA, or protein by mechanical deformation (see, e.g., Sharei etal. Proc. Natl. Acad. Sci. USA (2013) 110(6): 2082-2087, incorporated herein by reference); or viral transduction. Vectors may be delivered to cells by viral transduction. Nucleic acids can be delivered as part of a larger construct, such as a plasmid or viral vector, or directly, e.g., by electroporation, lipid vesicles, viral transporters, microinjection, and biolistics (high-speed particle bombardment). The construct(s) or the nucleic acid(s) encoding the components of the present system may be a DNA molecule. The nucleic acid(s) encoding the components of the present system maya be a DNA vector and may be electroporated to cells. The nucleic acid(s) encoding the components of the present system may an RNA molecule, which may be electroporated to cells. Delivery of DNA may be by adeno-associated virus (AAV).
[0126] Additionally, delivery vehicles such as nanoparticle- and lipid-based mRNA or protein delivery systems can be used. Further examples of delivery vehicles ribonucleoprotein (RNP) complexes, gene gun, hydrodynamic, electroporation or nucleofection microinjection, and biolistics. Various gene delivery methods are discussed in detail by Nayerossadat et al. (Adv Biomed Res. 2012; 1 : 27) and Ibraheem et al. (Int J Pharm. 2014 Jan 1 ;459(1-2):70-83), incorporated herein by reference. Delivery may be by lipid nanoparticles (LNP). Delivery may be by virus-like particles (VLP).
[0127] One or more components of the system may be introduced into a cell as a ribonucleoprotein (RNP) complex. The term “ribonucleoprotein complex,” as used herein, refers to a complex of ribonucleic acid (RNA) and RNA-binding protein(s). In the context of CRISPR-Cas systems, an RNP complex typically comprises Cascade proteins e.g., Cas5, Cas6, Cas7, and Cas8) in complex with a gRNA. RNPs may be assembled in vitro and can be delivered directly to cells using standard electroporation, cationic lipids, gold nanoparticles, or other transfection techniques (see, e.g., Kim et al., Genome Res., 24: 1012-1019 (2014); Zuris et al., Nat. BiotechnoL, 33: 73-80 (2015); and Mout et al., ACS Nano., 11 : 2452-2458 (2017)).
[0128] The systems may further comprise components in addition to those listed, including, but not limited to, sequence tags, protein markers or marker proteins, spacers, capture sequences, and the like.
[0129] Methods
[0130] Disclosed herein are methods for utilizing the disclosed proteins and systems. The descriptions and illustrative examples provided above for the proteins and systems are applicable to the methods described herein.
[0131] The disclosure provides a method of altering a target nucleic acid such as a target DNA. The phrase “altering a target DNA,” as used herein, refers to modifying at least one physical feature of a DNA of interest. DNA alterations may include, for example, single or double strand DNA breaks as well as deletions that affect the structural integrity of the DNA.
[0132] The methods comprise contacting a target nucleic acid such as a target DNA with a system disclosed herein or a composition comprising the system.
[0133] The methods introduce a single strand or double strand break in the target DNA. In this respect, the disclosed systems may direct cleavage of one or both strands of a target DNA, such as within a target genomic DNA.
[0134] Altering a DNA may comprise a deletion of nucleotides. If the deletion occurs in one direction from the PAM binding site, the deletion is a unidirectional deletion. If the deletion encompasses nucleotides on either side of the PAM binding site, it is a bidirectional deletion. The systems provided herein may produce unidirectional DNA deletions. The systems introduce a deletion without prominent off-target activity.
[0135] The deletion of the DNA may be of any size. For example, the deletion of the DNA may comprise from about 500 nucleotides to about 100,000 nucleotides (e.g., about 1 ,000, 5,000, 10,000, or 50,000 nucleotides, or a range defined by any two of the foregoing values). The deletion of the DNA may comprise from about 5,000 nucleotides to about 20,000 nucleotides (e.g., about 6,000, 6,500, 7,000, 7,500, 8,000, 8,500, 9,000, 9,500, 10,000, 10,500, 11 ,000, 11 ,500, 12,000, 12,500, 13,000, 13,500, 14,000, 14,500, 15,000, 15,500, 16,000, 16,500, 17,000, 17,500, 18,000, 18,500, 19,000, or 19,500 nucleotides, or a range defined by any two of the foregoing values).
[0136] Contacting target DNA herein comprises introducing a CRISPR system into a cell. As described above, the system may be introduced into cells by methods known in the art. As an example, to introduce Type I CRISPR systems into the retina of subjects, adeno- associated virus (AAV) may be used as a delivery vector including but not limited to AAV serotypes AAV5, AAV9 and AAV44.9 (doi.org / 10.1016 / j.ymthe.2020.04.002). The cell may be a mammalian cell. The cell may be a human cell.
[0137] Introducing the system into a cell comprises administering the system to a subject. The subject may be a human subject. The administering may comprise in vivo administration.
[0138] The target nucleic acid may be genomic DNA. The term “genomic,” as used herein, refers to a nucleic acid (e.g., a gene or locus) that is located on a chromosome in a cell.
[0139] The target nucleic acid may encode a gene product. The term “gene product,” as used herein, refers to any biochemical product resulting from expression of a gene. Gene products may be RNA or protein. RNA gene products include non-coding RNA, such as tRNA, rRNA, micro RNA (miRNA), and small interfering RNA (siRNA), and coding RNA, such as messenger RNA (mRNA). Gene products may be a protein or peptide.
[0140] The target DNA may comprise a “disease-associated mutation” and / or “disease- associated” gene. The term “disease-associated gene,” refers to any gene or polynucleotide whose gene products are expressed at an abnormal level or in an abnormal form in cells obtained from a disease-affected individual as compared with tissues or cells obtained from an individual not affected by the disease. A disease-associated gene may be expressed at an abnormally high level or at an abnormally low level, where the altered expression correlates with the occurrence and / or progression of the disease. A disease-associated gene also refers to a gene, the mutation or genetic variation of which is directly responsible or is in linkage disequilibrium with a gene(s) that is responsible for the etiology of a disease. The RHO gene is an example of a gene associated with a “single gene” or “monogenic” disease.
[0141] The method of altering a target DNA is used to delete nucleic acids from a target DNA in a cell by cleaving the target DNA and allowing the cell to repair the cleaved DNA in the absence of an exogenously provided donor nucleic acid molecule. Deletion of a nucleic acid target in this manner can be used in a variety of applications, including but not limited to, ablating an allele and / or a gene comprising a disease-associated mutation.
[0142] The components of the present systems or fusion proteins cells may be administered with a pharmaceutically acceptable carrier or excipient as a pharmaceutical composition. The components of the present system may be mixed, individually or in any combination, with a pharmaceutically acceptable carrier to form pharmaceutical compositions, which are also within the scope of the present disclosure.
[0143] An effective amount of the components of the present system or compositions as described herein is administered. Within the context of the present disclosure, the term “effective amount” refers to that quantity of the components of the system such that recruitment of one or more effector domains and deletion is achieved. When utilized as a method of treatment, the effective amount may depend on the condition being treated, the severity of the condition, the individual patient parameters including age, physical condition, size, gender and weight, the duration of the treatment, the nature of concurrent therapy (ifany), the specific route of administration and like factors within the knowledge and expertise of the health practitioner. The subject may be a human.
[0144] In the context of the present disclosure insofar as it relates to any of the disease conditions recited herein, the terms “treat,” “treatment,” and the like mean to relieve or alleviate at least one symptom associated with such condition, or to slow, stop or reverse the progression of such condition. Within the meaning of the present disclosure, the term “treat” also may mean to delay the onset {e.g., the period prior to clinical manifestation of a disease) and / or reduce the risk of developing a disease.
[0145] The phrase “pharmaceutically acceptable,” as used in connection with compositions and / or cells of the present disclosure, refers to molecular entities and other ingredients of such compositions that are physiologically tolerable and do not typically produce untoward reactions when administered to a subject {e.g., a mammal, a human). Preferably, as used herein, the term “pharmaceutically acceptable” means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in mammals, and more particularly in humans. “Acceptable” means that the carrier is compatible with the active ingredient of the composition {e.g., the nucleic acids, vectors, cells, or therapeutic antibodies) and does not negatively affect the subject to which the composition(s) are administered. Any of the pharmaceutical compositions and / or cells to be used in the present methods can comprise pharmaceutically acceptable carriers, excipients, or stabilizers in the form of lyophilized formations or aqueous solutions.
[0146] Pharmaceutically acceptable carriers, including buffers, are well known in the art, and may comprise phosphate, citrate, and other organic acids; antioxidants including ascorbic acid and methionine; preservatives; low molecular weight polypeptides; proteins, such as serum albumin, gelatin, or immunoglobulins; amino acids; hydrophobic polymers; monosaccharides; disaccharides; and other carbohydrates; metal complexes; and / or nonionic surfactants. See, e.g., Remington: The Science and Practice of Pharmacy 20th Ed. (2000) Lippincott Williams and Wilkins, Ed. K. E. Hoover.
[0147] Kits
[0148] The disclosure further provides kits containing one or more reagents or other components useful, necessary, or sufficient for practicing any of the methods described herein. For example, kits may include the disclosed engineered Cas proteins, fusion proteins, Cascade complex components and / or guide RNAs, vectors, compositions, transfection or administration reagents, negative and positive control samples e.g., cells, template DNA), cells, containers housing one or more components e.g., microcentrifugetubes, boxes), detectable labels, detection and analysis instruments, software, instructions, and the like.
[0149] Other terminology and disclosure
[0150] This entire document is intended to be read as a unified disclosure, and it should be understood that all combinations of features described herein are contemplated, even if the combination of features is not found together in the same sentence, or paragraph, or section of this document. The disclosure also includes, for instance, all embodiments of the disclosure narrower in scope in any way than the embodiments specifically mentioned. With respect to aspects of the disclosure described as a genus, all individual species are considered separate aspects of the disclosure.
[0151] As used herein and in the appended claims, the singular forms "a," "and," and "the" mean "one or more" unless the context unambiguously requires a more restricted meaning. It is further noted that the claims may be drafted to exclude any element, e.g., any optional element. As such, this sentence is intended to serve as antecedent basis for use of such exclusive terminology as "solely," "only" and the like in connection with the recitation of claim elements, or use of a "negative" limitation.
[0152] Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integer or step. When used herein the term "comprising" can be substituted with the term "containing" or "including" or sometimes when used herein with the term "having." When used herein, "consisting of" excludes any element, step, or ingredient not specified in the claim. When used herein, "consisting essentially of" does not exclude materials or steps that do not materially affect the basic and novel characteristics of the subject matter of a claim.
[0153] As used herein, “may,” “may comprise,” “may be,” “can,” “can comprise” and “can be” all indicate something envisaged by the inventors that is functional and available as part of the subject matter provided.
[0154] When a range of values is provided herein, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the statedrange. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the disclosure.
[0155] The term "about" or "approximately" as used herein means within 20%, preferably within 10%, and more preferably within 5% of a given value or range. It includes, however, also the concrete number, e.g., about 10 includes 10.
[0156] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure.
[0157] All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials for the purpose for which the publications are cited. To the extent the material incorporated by reference contradicts or is inconsistent with this specification, the specification will supersede any such material.Examples
[0158] A better understanding of the disclosure and of its advantages will be obtained from the following examples, offered for illustrative purposes only. The examples are not intended to limit the scope of the disclosure since variations and modifications will occur to those skilled in the art. Accordingly, only such limitations as appear in the claims should be placed on the invention.Materials and Methods Used in the Examples
[0159] Cell culture
[0160] HEK293T-Rho-EGFP-tdTomato reporter cells were cultured in DMEM / F12 (Gibco) supplemented with 10% FBS in a tissue culture incubator at 37°C with 5% CO2.
[0161] Plasmid transfection and estimation of gene targeting efficiency
[0162] CRISPR-Cas3 plasmid transfection was conducted using jetPrime Transfection Reagent per manufacturer’s instructions. HEK293T cells were seeded one day before transfection at 1 .5x105cells per well of a 24-well plate respectively. For each transfection, we used 1 pL jetPrime reagent, and a total of 500 ng crispr-cas plasmids (50, 92.5, 50, 162.5, 95 and 50 ng of Cas2-3, Cas5, Cas6, Cas7, Cas8 and CRISPR plasmids, respectively). To monitor genome targeting efficiency, cells were analyzed by flow cytometry 4-5 days post transfection.
[0163] Tn5 tagmentation-based NGS library construction for deletion size profiling
[0164] Tn5 library is constructed as described previously [Dolan et al, supra]. Briefly, Tn5 transposase was purified and loaded with one pre-annealed oligo pair ME-A / ME-rev as described [Picelli et al. Tn5 transposase and tagmentation procedures for massively scaled sequencing projects. Genome Res 2014, 24(12):2033-2040]. Tagmentation was performed in 10 mM T ris pH8.5, 5mM MgCh and 50% DMF using 300 ng of genomic DNA and 1 .4 pg of loaded Tn5, in a total volume of 40 pL. After 7 min incubation at 55°C, tagmentation reactions were purified using Zymo DNA Clean and Concentrator colums. For NGS library construction, 1ststep PCR amplification was carried out using Q5 DNA Polymerase for 15 cycles with a Tn5 adaptor oligo and a gene specific oligo, and then purified with Ampure beads. 2ndstep of nested-PCR was done for another 15 cycles using Q5 with a Tn5 adaptor oligo and nested gene specific oligo with partial Illumina adaptor sequence. After Ampure beads purification, the 3rdstep PCR was carried out for 10 cycles with Illumina index primers. The final NGS libraries were size selected on a 1 .5% agarose gel between 500 and 750 bp. The purified library is sequenced on Illumina MiSeq with a 300-cycle kit.
[0165] Bioinformatic analysis of the Tn5-NGS datasets
[0166] MiSeq sequencing reads were first subjected to adapter trimming using cutadapt 1 .8.1 [Martin M: Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnetjournal; Vol 17, No 1 : Next Generation Sequencing Data Analysis 201 1] to trim off the adapter sequence from the ends of reads in case of read-through. Reads were then quality trimmed using Trimmomatic v0.33 [Bolger et al. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics 2014, 30(15) :2114-2120] filter settings "TRAILING:3 SLIDINGWINDOW:4:15 MINLEN:10", and then aligned to a defined window (~1 OOkb) of the human genome spanning the region of interest. Alignment was performed using nucmer 4.0.0beta2 [Kurtz et al. Versatile and open software for comparing large genomes. Genome biology 2004, 5(2):R12] with a minimum match length of 10 and minimum cluster length of 20. Nucmer alignments were then filtered using an in-house python program [Dolan et al., supra]. Python and bash programs were subsequently used to extract, and plot read counts and locations.
[0167] Construction of a HEK293-Rho-EGFP / SNP-tdTomato dual reporter cell line
[0168] WT HEK293T cells were individualized with TrypLE Express (Gibco), washed once with DMEM supplemented with 10% FBS and resuspended in Neon buffer R to a concentration of 2x106cells / mL. 3 pg SpyCas9 protein was assembled with 0.6 pg of AAVS / -targeting sgRNA, the resulting Cas9 RNP was then mixed with 0.5 pg pSmart- AAVS1 -EF1 alpha- EGFP-RHO-WT plasmid, 0.5 pg pSmart-AAVS7-EF1 alpha-tc / Tomato-RHO-SNP and approximately 105cells in buffer R in a total volume of 10 pL. The mixture was electroporated with a 10 pL Neon tip (Invitrogen, 1575V, 10ms, 3 pulses), and plated in one well of a 24-well tissue culture plate containing 500 pL DMEM / F12 supplemented with 10% FBS. The cells were cultured for 2 weeks after electroporation, with daily medium change. EGFP and tdTomato dual positive cells were then sorted into 96 well plates at 1 cell per well via FACS. The sorted single cells were expanded, correctly targeted clones were identified by genomic junction PCRs.
[0169] Creation of hyper-activity variants of the Pae type l-F CRISPR-Cas gene editor
[0170] Plasmids encoding cascade subunits with arginine (Arg) substitutions of the selected Glu / Asp residues were generated by Q5 site-directed mutagenesis (NEB). All plasmid clones were sequence confirmed by Sanger sequencing.
[0171] Dual-AAV packaging and transduction
[0172] AAV packaging is done by OBiO Tech, Inc using a proprietary GC03 capsid. For transduction of HEK293 reporter cell, cells were washed with 1xPBS, dissociated with TrypLE Express, and neutralized with growth media. These cells were then pelleted and resuspended in media to a final concentration of 1 x 103cells per 100 pL. Dual-AAV viruses expressing Cascade and Cas3 / CRISPR were then added to 100 pL cells at an MOI of 1 x 107each. Transduced cells were analyzed by flow cytometry with a ZE5 Cell Analyzer (BioRad) 6 days after transduction.
[0173] AAV subretinal injection of rabbit
[0174] Packaged AAVs expressing Cascade and Cas3 / CRISPR were mixed 1 :1 and diluted to 2x101° vg / mL each. 30 pL of diluted virus was then subjected to sub-retinal injection into one eye of the humanized rabbit model. Rabbits were sacrificed 14 days after injection. Genomic DNA from retina within the injection bleb was extracted and used for an anchored Tn5-based NGS profiling method.Example 1
[0175] Generating a dual reporter cell line model for the allele differentiation test.
[0176] To provide proof-of-principle data that Type I CRISPR strategy can enable efficient gene targeting and allele differentiation, a HEK293T dual reporter cell line was generated that carries the SNP rs7984 and nearby partial RHO gene (Fig. 3A, HEK293T-Rho-wt- EGFP / SNP-tdTomato line). Two 98 bp DNA fragments were synthesized that correspond to partial human RHO gene including the 5’ UTR sequence containing the SNP position. Thesetwo fragments differ by only one base pair at the SNP site, with one fragment containing the WT base (A) and the other the SNP base (G). The fragment with WT base was cloned 3’ of an EGFP reporter construct and the fragment with SNP base was cloned 3’ of a tdTomato reporter construct. The two resulting reporters were then knocked-in at the safe harbor AAVS1 locus of the HEK293T cell’s genome into two different alleles. After single cell isolation and genotyping validation, the final reporter line HEK293T-Rho-wt-EGFP / SNP- tdTomato was obtained, which is positive for both eGFP and TdTomato fluorescence. Targeting of the GFP reporter leads to GFP-negative and tdTomato-positive cells, whereas targeting of the tdTomato reporter results in GFP-positive and tdTomato-negative cells. Targeting of both alleles leads to dual negative cells.Example 2
[0177] Design of Type I CRISPR guides that recognize either the RHO-SNP or the RHO-WT allele.
[0178] A Type l-F CRISPR system from Pseudomonas aeruginosa (Pae) was then leveraged to achieve SNP-based allele differentiation. Pae l-F CRISPR is known to recognize a CC dinucleotide PAM positioned next to the guide RNA-complementary sequence. A CC PAM is located just 2 bp away from the SNP rs7984 (Fig. 3B). Accordingly, Type l-F guide RNAs were designed that place this SNP at the PAM-proximal position 2 in the target sequence, which is in the “seed” region of guide that is critical for CRISPR targeting specificity. Given that Cas3 is known to translocate unidirectionally towards the PAM-side of the CRISPR binding site, this design ensured that DNA deletion occurred towards the RHO gene body. Two Type l-F guide RNAs were designed that differ by 1 nt at the SNP position, one fully complementary to the SNP-containing sequence and the other completely matching the WT counterpart (Fig. 3B).Example 3
[0179] Type I CRISPR achieving robust RHO sequence targeting in an allele-specific manner.
[0180] To test if Type I CRISPR can achieve allele differentiation through a single-nt SNP, a plasmid mix encoding the Pae Type l-F CRISPR-Cas components was transfected into the dual reporter cell line and eGFP and tdTomato disruption was examined by FACS. When a WT RHO-targeting guide was used, robust knockdown of GFP but not tdTomato fluorescence was observed, as evident by the formation of GFP- / tdTomato+ cells in 26% of the total population. There were no tdTomato- or dual- negative cells observed abovebackground level, indicative of robust WT-Rho-GFP allele-specific targeting by the Pae l-F CRISPR (Fig. 3C). Conversely, when the RHO-SNP-targeting guide was used, a 34% GFP+ tdTomato- cell population was observed, with no GFP- or dual- negative cells observed above background (Fig. 3C). This result suggests effective and Rho-SNP-tdTomato allelespecific targeting by the Pae l-F CRISPR. Collectively, the cell line level data demonstrated that the Type I CRISPR targeting / deletion machinery can distinguish between the RHO SNP and its WT counterpart that differ by just a single base pair. As a control experiment, a third guide targeted to a common RHO gene sequence located 13-bp away from the target site of the two previously used guides (Fig. 4A) was used. As expected, this common guide is capable of targeting both the Rho-WT-eGFP and SNP-tdTomato alleles, resulting in simultaneous disruption of GFP and tdTomato fluorescence in the cell population (Fig. 4B).Example 4
[0181] Profiling the large deletion lengths created by Type I CRISPR.
[0182] T o comprehensively define the lengths of the DN A deletions generated by T ype ICRISPR a Tn5-anchored-NGS (next-gen sequencing) profiling method was used (Fig. 5). This profiling method leverages Tn5 transposase tagmentation of the genomic DNA and locus-specific nested primers to capture unidirectional deletions generated by Type I CRISPR. DNA lesions at both the AAVS1 -reporter locus (Figs. 6-7) and the endogenous RHO gene locus (Fig. 8) were profiled. In line with prior reports, Type I CRISPR-induced deletions exhibit heterogeneity in terms of the deletion endpoints and overall deletion lengths. As shown in Figs. 6 and 7, nearly 50% of all the DNA deletions are shorter than 2.3 kb for the Rho-WT-GFP locus and less than 2.6 kb for the RHO- SNP-tdTomato locus. A similar trend was confirmed at the endogenous RHO gene locus of the same reporter cell line, which is also targeted by Type l-F CRISPR via the WT-RHO-targeting guide (Fig. 8). Tn5-anchored-NGS profiling revealed that about half of the DNA deletions are shorter than 1.7 kb.
[0183] Importantly, over 90% of the deletions are within 8.8 kb at the reporter loci and less than 7.2 kb at the endogenous RHO locus (Figs. 6-8). Given that the human RHO gene is ~7.7 kb in size and is separated from the immediate downstream H1-8 gene by a ~7 kb intergenic region, over 90% of the DNA deletions created by Type I CRISPR targeting are within the RHO gene body and therefore impose minimal safety risks of deleting into downstream genes.
[0184] Other Type I CRISPR-Cas orthologs with less Cas3 processivity can be exploited to further constrain the deletion length if needed to apply the Type I CRISPR therapeutic strategy to other autosomal dominant diseases and their gene loci.Example 5
[0185] Exploiting additional common and benign SNPs associated with RHO gene to cure more adRP patients
[0186] By utilizing additional common benign SNPs with high heterozygous rate, an even larger adRP patient population can be cured. For instance, rs2855558 and rs2410 are both located in the 3’ UTR of human RHO gene. Two CC dinucleotide PAMs for the Pae l-F CRISPR-Cas system can be utilized that place these SNP positions in the PAM-proximal end of the guide RNA, at the +2 and +1 position, respectively (Fig. 9) so that Cas3 translocation fires the processive nuclease towards the RHO gene body and causes large deletion-mediated destruction of the gain-of-function pathogenic allele. Meanwhile, the SNP discrimination power of Type I system leaves the wt RHO allele intact.Example 6
[0187] Developing a hyper-activity variant of Pae type l-F gene editor through structural guided protein engineering
[0188] To maximize the clinical efficacy of Pae type l-F CRISPR-Cas, structure-guided protein engineering was applied to enhance its activity. Based on the available structure of Pae type l-F CRISPR-Cas [Rollins eta! . Structure Reveals a Mechanism of CRISPR-RNA-Guided Nuclease Recruitment and Anti-CRISPR Viral Mimicry. Molecular cell 2019, 74(1 ):132-142 e135], negatively charged residues within 8 A of the target-strand DNA were targeted and individually mutated to positively charged arginine to enhance the DNA binding capacity of Cascade. As a result, seven point mutations were created in either the Cas7 or Cas8 subunit of Cascade. When plasmids encoding each mutated Cas subunit were transfected along with other Pae l-F components and a GFP-targeting crRNA plasmid into a HEK293-EGFP reporter cell line, increased targeting activity was observed for six out of the seven variants tested (Fig. 10). Five of these variants tested, Cas7- E338R (SEQ ID NO: 19), Cas7-D234R (SEQ ID NO: 20), Cas7-D237R (SEQ ID NO: 21), Cas7- D68R (SEQ ID NO: 22), Cas8-D76R (SEQ ID NO: 24), and Cas8-E421 R (SEQ ID NO: 25), yielded a 6.8-, 5.0-, 6.5-, 6.8- and 4.5- fold increase in gene targeting efficiency, respectively (Fig. 10). These are defined herein as “hyper-activity” Pae l-F variants and Cas7-E338R (SEQ ID NO: 25), the most potent variant, was selected for subsequent in vivo studies.Example 7
[0189] Dual-AAV-mediated in vivo targeting of the human Rho gene
[0190] AAV is widely used for in vivo gene delivery in clinical settings [Wang et ai:. Adeno- associated virus as a delivery vector for gene therapy of human diseases. Signal Transduct Target Ther 2024, 9(1 ):78]. Given the ~4.7 kb packing limit of a single AAV vector, a dualAAV strategy (Fig. 11 A) was designed to deliver all components of the Pae type l-F CRISPR-Cas system. One AAV carries the cas5, cas6, cas7 and cas8 genes, linked via 2A self-cleaving peptides, while the second AAV contained Cas3 and a short CRISPR array (Fig. 11 A). To target the wild-type human RHO SNP sequence, the corresponding CRISPR guide and all other components of the hyper-activity Pae type l-F editor (Fig .11C) were packaged into AAVs pseudo typed with a GC03 capsid, which has been shown to efficiently transduce rabbit retina (OBiO Tech Inc. proprietary capsid). The activity of packaged dual- AAV was first validated in the HEK293-WT / eGFP-SNP / tdTomato dual fluorescent reporter cell line. As shown in Fig. 11 B, dual AAV transduction resulted in robust disruption of GFP (which is linked with WT Rho sequence), with no increase above background in tdTomato- negative cells (tdTomato is linked to the SNP RHO sequence), indicating both high gene targeting / deletion efficiency and allele differentiation.
[0191] The sub-retinal injection of these dual-AAVs into the eye of a transgenic rabbit model was then performed, in which one allele of the endogenous rabbit Rho gene has its exon 1 replaced with human RHO exon 1 (Fig. 11C). An anchored Tn5-based NGS profiling method [Dolan etal., supra] revealed successful genomic deletions at the humanized RHO locus from the injected rabbit retina (Fig. 11 D). The deletions range from -600 bp to -8 kb in size, starting from PAM-CRISPR target site and extending into the RHO gene body, to disrupt expression of the mutant RHO allele linked with the SNP in adRP patients. By contrast, no gene editing or deletion events were detected in the un-injected eye control from the same animals. These results validate the approach for in vivo gene editing therapy.
Claims
Claims\Ne claim:1 . A method of altering a target DNA containing an allele comprising a mutation associated with an autosomal dominant genetic disease, which method comprises introducing to the target DNA a Type I CRISPR system comprising:(a) a Type I CRISPR Cas 3 protein and Cascade complex and(b) a guide RNA that is complementary to the target DNA, wherein a CC protospacer adjacent motif (PAM) of the Type I CRISPR system is located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA, and wherein the Cas 3 protein translocates in a direction along the target DNA to ablate the allele.
2. A method of altering a target DNA containing an allele comprising a mutation associated with an autosomal dominant genetic disease, which method comprises introducing to the target DNA a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system comprising:(a) a Pseudomonas aeruginosa (Pae) Type l-F CRISPR Cas 3 protein and Cascade complex; and(b) a guide RNA that is complementary to the target, wherein the Cas 3 protein translocates in a direction along the target DNA to ablate the allele.
3. The method of claim 1 or 2, wherein the target DNA is in a cell and said introducing comprises introducing (a) and (b) into said cell.
4. The method of claim 3, wherein the method is conducted under conditions such that the guide RNA binds to the target DNA in the cell, and the Type I CRISPR Cas 3 protein induces cleavage of at least one strand of the target DNA at the PAM-proximal side of the binding site, thereby altering the target DNA in the cell.
5. The method of claim 4, wherein the cleavage of one or both strands in the target DNA results in a deletion of the target DNA.
6. The method of claim 5, wherein the deletion is unidirectional.
7. The method of claim 5 or 6, wherein the deletion comprises from about 500 nucleotides to about 100,000 nucleotides.
8. The method of claim 5 or 6, wherein the deletion comprises from about 5,000 nucleotides to about 20,000 nucleotides.
9. The method of claim 3, wherein the guide RNA and Cascade complex are introduced into the cell as a Cascade ribonucleoprotein (RNP) complex.
10. The method of claim 3, wherein the guide RNA is introduced into the cell as part of a first vector, a nucleic acid encoding the Type I CRISPR Cas 3 protein is introduced into the cell as part of a second vector, and nucleic acids encoding the Cascade complex components are introduced into the cell as part of one or more additional vectors.11 . The method of claim 3, wherein the nucleic acids encoding the guide RNA and nucleic acids encoding the Type I CRISPR Cas 3 protein are introduced into the cell as part of a single vector.
12. The method of claim 3, wherein the Type I CRISPR Cas 3 protein is introduced into a cell by contacting the cell with an mRNA encoding the Type I Cas 3 protein.
13. The method of claim 2, wherein a CC protospacer adjacent motif (PAM) of the Pae Type l-F CRISPR system is located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA.
14. The method of any preceding claim, wherein said target DNA is genomic DNA.
15. The method of claim 14, wherein the target DNA encodes at least one gene product.
16. The method of claim 15, wherein the target DNA encodes a protein.
17. The method of any preceding claim, wherein the cell is a eukaryotic cell.
18. The method of claim 17, wherein the cell is a human cell.
19. The method of claim 18, wherein the human cell is in situ.
20. The method of claim 10 or 11 , wherein Type I CRISPR Cas 3 protein is encoded by a nucleic acid sequence that is codon optimized.21 . The method of claim 20, wherein the Type I Cas3 protein is encoded by the nucleic acid sequence of SEQ ID NO: 18.
22. The method of any preceding claim wherein the mutation is a mutation in a rhodopsin (RHO) allele in a cell in a human subject and the alteration of the target DNA ablates the mutated RHO allele.
23. A method of treating a subject who has an allele comprising a mutation associated with autosomal dominant retinitis pigmentosa (adRP), which method comprises altering a genomic DNA of the subject by introducing to the genomic DNA of the subject a Type I CRISPR system comprising:(a) a Type I CRISPR Cas 3 protein and Cascade complex and(b) a guide RNA that is complementary to the target DNA, wherein a CC protospacer adjacent motif (PAM) of the Type I CRISPR system is located adjacent to or encompasses a single-nucleotide polymorphism (SNP) in the target DNA, and wherein the Cas 3 protein translocates in a direction along the target genomic DNA to ablate the allele.
24. A method of treating a subject who has an allele comprising a mutation associated with autosomal dominant retinitis pigmentosa (adRP), which method comprises altering a genomic DNA of the subject by introducing to the genomic DNA of the subject a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system comprising:(a) a Pseudomonas aeruginosa (Pae) Type l-F CRISPR Cas 3 protein and Cascade complex; and(b) a guide RNA that is complementary to the target DNA, wherein the Cas 3 protein translocates in a direction along the target genomic DNA to ablate the allele.
25. The method of claim 23 or 24 wherein the genomic DNA comprises a mutation in an allele of the Rhodopsin (RHO) gene and the adRP is an RHO-adRP.
26. The method of claim 25 wherein the synthetic guide RNA recognizes a singlenucleotide polymorphism at rs7984 in the 5' untranslated region of the RHO gene.
27. The method of claim 25 wherein the synthetic guide RNA recognizes a singlenucleotide polymorphism at rs2855558 in the 3' untranslated region of the RHO gene.
28. The method of claim 25 wherein the synthetic guide RNA recognizes a singlenucleotide polymorphism at rs2410 in the 3' untranslated region of the RHO gene.
29. The method of claim 2, 9 or 24 wherein the Cascade complex comprises Cas5 protein, Cas6 protein, Cas7 protein and Cas8 protein.
30. The method of claim 29 wherein the Cas3 protein comprises the amino acid sequence of SEQ ID NO: 17, the Cas5 protein comprises the amino acid sequence of SEQ ID NO: 9, the Cas6 protein comprises the amino acid sequence of SEQ ID NO: 1 1 , the Cas7 protein comprises the amino acid sequence of SEQ ID NO: 13 and the Cas8 protein comprises the amino acid sequence of SEQ ID NO: 15.31 . The method of claim 1 or 23 wherein the Type I CRISPR system is a Pseudomonas aeruginosa (Pae) Type l-F CRISPR system.
32. The method of claim 31 wherein the Pseudomonas aeruginosa (Pae) Type l-F CRISPR system comprises a hyper-activity Pae l-F variant Cas protein.
33. The method of claim 32 wherein the hyper-activity Pae l-F variant protein is a Cas7 protein.
34. The method of claim 33 wherein the hyper-activity Pae l-F variant protein is a Cas8 protein.
35. The method of claim 32 wherein the hyper-activity Pae l-F variant protein is Cas7- E338R (SEQ ID NO: 19), Cas7-D234R (SEQ ID NQ:20), Cas7-D237R (SEQ ID NO: 21 ), Cas7-D68R (SEQ ID NO: 22), Cas8-D76R (SEQ ID NO: 24) or Cas8-E421 R (SEQ ID NO: 25).
36. A Pae l-F variant protein Cas7-E338R (SEQ ID NO: 19), Cas7-D234R (SEQ ID NO:20), Cas7-D237R (SEQ ID NO: 21), Cas7-D68R (SEQ ID NO: 22), Cas8-D76R (SEQ ID NO: 24) or Cas8-E421 R (SEQ ID NO: 25).
37. Use of one or more of the Pae l-F variant proteins of claim 36 for DNA editing.
38. The method of claim 10 wherein one or more of the vector is an adeno- associated virus (AAV) vector.