Engineered CAS9 endonucleases with enhanced editing efficiency
Mutant Cas9 proteins with Keapl degron mutations address CRISPR limitations by enhancing stability and efficiency, achieving improved gene editing and regulation with reduced off-target effects.
Patent Information
- Application Number
- PCT/US2025/037645
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Existing CRISPR-Cas9 systems face limitations such as specific PAM sequence restrictions, DNA damage, off-target effects, and instability of sgRNAs, with mammalian hosts' response to bacterial Cas9 proteins being unclear, necessitating improved compositions and methods for gene therapies.
Mutant Cas9 proteins with mutations in the Keapl degron sequence, such as alanine substitutions, to enhance stability and efficiency by suppressing Keapl-mediated degradation, thereby increasing CRISPR efficiency and promoting cell growth.
The mutant Cas9 proteins exhibit enhanced CRISPR-mediated gene editing and regulation capabilities, with increased half-life and reduced off-target effects, facilitating improved gene knock-out and knock-in efficiencies and gene expression modulation.
Smart Images

Figure US2025037645_22012026_PF_FP_ABST
Abstract
Description
ENGINEERED CAS9 ENDONUCLEASES WITH ENHANCED EDITING EFFICIENCYCROSS REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of United States Provisional Patent Application Serial No. 63 / 671 ,846, filed July 16, 2024, the disclosure of which is incorporated herein by reference in its entirety.GOVERNMENT INTEREST
[0002] This invention was made with government support under Grant No. R21 CA270967 awarded by the National Institutes of Health. The government has certain rights in the invention.REFERENCE TO SEQUENCE LISTING XML SUBMITTED ELECTRONICALLY
[0003] The content of the Sequence Listing XML filed using Patent Center as an XML file (Name: 4210_0545WO.xml; Size: 169,996 bytes; and Date of Creation: July 14, 2025) is incorporated herein by reference in its entirety.TECHNICAL FIELD
[0004] The presently disclosed subject matter relates to clustered regulatory interspaced short palindromic repeats (CRISPR)-associated protein (Cas) mutant proteins, such as mutant Cas9 proteins that contain one or more mutations in the Kelch-like ECH-associated protein 1 (Keapl ) degron sequence of the Cas9 protein as compared to the Cas9 protein that is free of mutations in the Keapl degron sequence. The presently disclosed subject matter further relates to nucleic acids encoding the mutant Cas9 proteins and to related compositions, as well as to methods of enhancing CRISPR efficiency, suppressing Keapl activity, and promoting cell growth.BACKGROUND
[0005] CRISPR and Cas proteins provide an adaptive immune defense against phage for many bacteria. Cas proteins, especially Streptococcuspyogenes Cas9 (SpCas9) and Staphylococcus aureus Cas9 (SaCas9), have also been repurposed to provide a potent and convenient mammalian genome-editing tool (1 ,2). Cas protein-driven gene editing has been widely used in biomedical research as a loss-of-function approach to study the function of genes (3) and enhancer elements (4) and to screen for new drug targets (5,6). More recently gene editing has been increasingly investigated for the treatment of human diseases (see (7) for review), including Barth Syndrome (8), Duchenne Muscular Dystrophy (9,10), blindness (11 ), deafness (12), cancer (13), removing HIV infection (14), transthyretin amyloidosis (15), heterozygous familial hypercholesterolemia and atherosclerotic cardiovascular diseases (16), and many others. Notably, Cas-protein driven gene editing has also been used to generate murine models by replacing traditional gene knockout and knockin approaches to mimic human genetic diseases (17), and more importantly, has demonstrated the ability to correct disease mutation(s) in mouse (9,17-19) and dog (20). Further, Cas-induced gene editing has also been initiated in human clinical trials, including edited T cells for enhanced immune-therapies (7) to treat p-thalassemia (NCT03655678) and sickle cell disorders (NCT03745287). As a result, the FDA has approved the first therapeutic use of CRISPR to correct the genetic defect causing sickle cell disease.
[0006] Although CRISPR is the source of many breakthroughs in both biomedical research and disease therapy, technical limitations still exist. These limitations include the restrictions imposed by specific PAM (protospacer adjacent motif) sequences for Cas9 recognition, unexpected DNA damage caused by Cas9 (21 ), off-target effects caused by Cas9, instability of the sgRNAs in cells, unwanted sustained activity of Cas9 and other issues. To overcome these drawbacks, extensive efforts have been devoted to further improving and broadening the power and applications of Cas9 in both biomedical research and clinical usages. These efforts include but are not limited to finding new Cas9 proteins with altered and expanded PAM sequence recognition (22), additional prokaryotic endonucleases such as Cpf 1 (23), CasY (24) and CasX (25) or even the eukaryotic endonuclease Fanzor (26), engineered Cas9 enzymes with improved fidelity (27-29), base-editors to rewrite the genome by editing DNA or RNA bases (30), dCas9- fusions (catalytic-dead Cas9) to modulate epigenome (31 ,32), natural Cas9 inhibitors (33), controlled Cas by small molecules (34) and many others. In addition, manipulating DNA damage responses by inhibiting NHEJ (35) via suppressing DNAPK (36) or enhancing HDR by small molecules (37), or modifying gRNAs by engineering a tRNA-gRNA architecture (38) all lead to improved and enhanced CRISPR techniques.
[0007] As Cas9 was originally identified in prokaryotes and is not naturally present in mammalian cells, whether and how mammalian hosts recognize and respond to the bacterial Cas9 protein remains unknown. Accordingly, there is an ongoing need to provide addresses this issue and to provide additional compositions and methods for applying CRISPR in gene therapies.SUMMARY
[0008] This summary lists several embodiments of the presently disclosed subject matter, and in many cases lists variations and permutations of these embodiments. This summary is merely exemplary of the numerous and varied embodiments. Mention of one or more representative features of a given embodiment is likewise exemplary. Such an embodiment can typically exist with or without the feature(s) mentioned; likewise, those features can be applied to other embodiments of the presently disclosed subject matter, whether listed in this summary or not. To avoid excessive repetition, this Summary does not list or suggest all possible combinations of such features.
[0009] In some embodiments, the presently disclosed subject matter provides a mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence of a Cas9 protein, wherein said mutant Cas9 protein has a half-life greater than the half-life of the corresponding Cas9 protein that does not comprise the one or more mutations in the Keapl degron sequence.
[0010] In some embodiments, the Cas9 protein is a Streptococcus pyogenes Cas9 protein comprising a sequence of SEQ ID NO: 1 or a protein having a sequence having at least 90% homology to SEQ ID NO: 1 . In some embodiments, the Keapl degron sequence is the ETGE sequence at amino acids 1068 to 1071 of SEQ ID NO: 1 . In some embodiments, the one or moremutations comprise an alanine substitution at position 1069 or at position 1070 of SEQ ID NO: 1 . In some embodiments, the one or more mutations comprise an alanine substitution at position 1069 and at position 1070 of SEQ ID NO: 1.
[0011] In some embodiments, the Cas9 protein is a Staphylococcus aureus Cas9 of SEQ ID NO: 2 or a protein having at least 90% homology to SEQ ID NO: 2. In some embodiments, the Keapl degron sequence is the ETGE-like sequence at amino acids 860 to 863 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprise an alanine substitution at position 860 of SEQ ID NO: 2, a glutamic acid substitution at position 861 of SEQ ID NO: 2, and an alanine substitution at postion 863 of SEQ ID NO: 2.
[0012] In some embodiments, the mutant Cas9 protein further comprises one or more mutations outside of the Keapl degron sequence of the Cas9 protein, optionally wherein the one or more mutations outside of the Keapl degron sequence are conservative and / or nonconservative substitutions. In some embodiments, the one or more mutations outside of the Keapl degron sequence of the Cas9 protein comprise an alanine at position 10 of SEQ ID NO: 1 and an alanine at position 840 of SEQ ID NO: 1 .
[0013] In some embodiments, the presently disclosed subject matter provides an isolated nucleic acid sequence comprising a sequence that encodes the mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence of a Cas9 protein, wherein said mutant Cas9 protein has a half-life greater than the half-life of the corresponding Cas9 protein that does not comprise the one or more mutations in the Keapl degron sequence.
[0014] In some embodiments, the presently disclosed subject matter provides a construct comprising the nucleic acid sequence that encodes the mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence of a Cas9 protein. In some embodiments, the construct further comprises a promoter operably linked to the nucleic acid sequence encoding the mutant Cas9 protein. In some embodiments, the construct further comprises a sequence encoding a marker gene. In some embodiments, the construct comprises a viral based vector or a non-viral based vector.
[0015] In some embodiments, the presently disclosed subject matter provides a composition comprising the construct that comprises the nucleic acid sequence that encodes the mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence of a Cas9 protein and a pharmaceutically acceptable carrier.
[0016] In some embodiments, the presently disclosed subject matter provides a cell line comprising an isolated nucleic acid sequence that encodes the mutant Cas9 protein or construct comprising a nucleic acid that encodes the mutant Cas9 protein.
[0017] In some embodiments, the presently disclosed subject matter provides a method of enhancing CRISPR efficiency comprising administering one or more guide RNAs (gRNAs) to a cell comprising a mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence to thereby increase or improve CRISPR efficiency. In some embodiments, the enhanced CRISPR efficiency increases CRIPSR-mediated knock-out efficiency. In some embodiments, the enhanced CRISPR efficiency increases CRISPR- mediated knock-in efficiency.
[0018] In some embodiments, the cell has decreased expression of Keapl . In some embodiments, the cell comprises a Keapl mutant, and wherein presence of the Keapl mutant leads to decreased degradation of any Cas9 protein or mutant Cas9 in said cell. In some embodiments, the Keapl mutant is a Keapl protein having the mutation of G333C relative to wild-type Keapl as set forth in SEQ ID NO: 3.
[0019] In some embodiments, the presently disclosed subject matter provides a CRISPR-Cas9 system for use in editing a gene, wherein the CRISPR-Cas9 system comprises (i) a guide RNA; and (ii) a mutant Cas9 protein, wherein the mutant Cas9 protein comprises one or more mutations in a Keapl degron sequence of a Cas9 protein.
[0020] In some embodiments, the presently disclosed subject matter provides a method of editing a gene, wherein the method comprises introducing a mutant Cas9 comprising one or more mutations in a Keapl degron sequence of a Cas9 protein into a cell or tissue, and introducing one or more guide RNAs into said cell or tissue.
[0021] In some embodiments, the presently disclosed subject matter provide a method of suppressing Keapl activity comprising administering a protein or nucleic acid that blocks the Keapl - Kelch domain responsible for Keapl binding to its substrates. In some embodiments, the protein that blocks the Keapl - Kelch domain comprises an antibody specific to the Kelch domain. In some embodiments, the nucleic acid that blocks the Keapl - Kelch domain is a nucleic acid that mimics a Kelch domain substrate. In some embodiments, the Kelch domain substrate is a Keapl degron sequence of a Cas9 protein.
[0022] In some embodiments, the suppressed Keapl activity is the ability to ubiquitinate the Cas9 protein. In some embodiments, suppression of Keapl activity increases the half-life of the Cas9 protein.
[0023] In some embodiments, the presently disclosed subject matter provides a method of promoting cell growth comprising administering an inhibitor of Keapl and NRF2 binding in amounts sufficient to promote cell growth. In some embodiments, the inhibitor is a mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence.
[0024] Accordingly, it is an object of the presently disclosed subject matter to provide mutant Cas9 proteins that comprise one or more mutations in a Keapl degron sequence of the Cas9 protein, as well as related nucleic acids and constructs, compositions, cell lines, and methods.
[0025] These and other objects are achieved in whole or in part by the presently disclosed subject matter. Further, objects of the presently disclosed subject matter having been stated above, other objects and advantages of the presently disclosed subject matter will become apparent to those skilled in the art after a study of the following description, Drawings and Examples.BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The presently disclosed subject matter can be better understood by referring to the following figures. The components in the figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the presently disclosed subject matter (often schematically). Thedrawings are not intended to limit the scope of this presently disclosed subject matter, which is set forth with particularity in the claims as appended or as subsequently amended, but merely to clarify and exemplify the presently disclosed subject matter.
[0027] Figures 1A-1G demonstrate that Kelch-like ECH-associated protein 1 (Keapl ) interacts with Streptococcus pyogenes Gas 9 (SpCas9) in an ETGE (SEQ ID NO: 18) degron-dependent manner. Fig. 1A is a schematic drawing of SpCas9 protein domain structures and shows the location of the Keapl ETGE (SEQ ID NO: 18) binding motif degron at amino acids 1068 to 1071 in the amino acid sequence of SpCas9 (SEQ ID NO:1 ). Fig. 1 B and Fig. 1C are images of immunoblot (IB) analysis of whole cell lysates (WCL) and Flag- immunoprecipitants (IP) derived from HEK293 cells transfected with the indicated DNA constructs: empty vector (EV), wild-type (WT) Flag-SpCas9 and a Flag-SpCas9 construct for a mutant SpCas9 with an EAAE (SEQ IE NO: 19) mutation in place of the ETGE (SEQ ID NO: 18) degron of WT SpCas9. Fig. 1 D is a schematic drawing showing a structural simulation by PyMOL indicating that the ETGE (SEQ ID NO: 18)-containing motif in SpCas9 (PDB: 4oo8; SEQ ID NO: 1 ) has a similar structure topology as the “ETGE” (SEQ ID NO: 18)-containing motif in NRF2 (PDB: 2flu; SEQ ID NO: 16). Fig. 1 E is an image of a Coomassie blue-stained SDS-PAGE gel showing Flag-IP (arrow) derived from HEK293T cells transfected with EV (left lane) or WT- SpCas9 (right lane). Fig. 1 F is a volcano plot indicating that multiple E3 ubiquitin ligases were identified from the SpCas9 interactome. Fig. 1 G presents a mass spectrometry spectrum indicating identification of a Keapl peptide (LLYAVGGFDGTNR, amino acids 471 to 483 of SEQ ID NO: 3) from the SpCas9 interactome.
[0028] Figures 2A-2J illustrate that Keapl targets SpCas9 for ubiquitination and degradation. Fig. 2A and Fig. 2B show images of the IB analysis of WCL derived from HEK293T cells transfected with WT Flag- SpCas9 or a Flag-SpCas9 for a mutant SpCas9 containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif. Fig. 2C shows an image of the IB analysis of WCL derived from control or endogenous Keapl - depleted HEK293 cells by shRNAs transfected with WT or mutant Flag-SpCas9 constructs, where the mutant Flag-SpCas9 construct is for a SpCas9 comprising a EAAE (SEQ ID NO: 19) mutation in place of the WT Keapl binding motif. Fig. 2D shows an image of an IB analysis of WCL derived from control or endogenous Keapl -depleted HBE cells by shRNAs transfected with either WT or mutant spCas9 constructs for 48 hours, where the mutant spCas9 contains a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif. Fig. 2E is an image of an IB analysis of WCL derived from control or endogenous Keapl -depleted HEK293 cells by shRNAs transfected with the indicated doses of WT-SpCas9 construct (0, 0.5, 1 or 2 micrograms (pg)) for 48 hours. Fig. 2F shows an image of an IB analysis of WCL derived from control or endogenous Keapl -depleted HEK293 cells by shRNAs transfected with the indicated Flag-SpCas9 and HA-Keap1 (resistant to shKeapI ) constructs for 48 hours. Fig. 2G shows an image of an IB analysis of WCL derived from HEK293 cells transfected with either WT- or mutant-SpCas9 constructs for 48 hours, where the mutant SpCas9 construct is for a SpCas9 containing EAAE (SEQ ID NO: 19) in place of the WT Keapl binding motif. Where indicated (by + above lane), 10 micromolar (pM) MG132 was added to the cell culture media overnight before cell collection. Fig. 2H is an image of an IB analysis of Ni-NTA pulldowns and WCL from HEK293 cells transfected with the indicated DNA constructs for 48 hours. 10 pM MG132 was added to the cell culture media overnight before cell collection. Fig. 21 and Fig. 2J show images of the IB analysis of WCL derived from HEK293 cells transfected with either WT or mutant SpCas9 constructs (where the mutant SpCas9 construct is for a SpCas9 containing EAAE (SEQ ID NO: 19) in place of the WT Keapl binding motif) and treated with the indicated compounds. As indicated, DMSO, CDDO (50 nanomolar (nM)) or tBHQ (10 pM) was added to culture medium for 16 hours before cell collection.
[0029] Figures 3A-3H demonstrates that Keapl targets SaCas9 and Fanzor for degradation. Fig. 3A and Fig. 3E depict sequence alignments for potential ETGE (SEQ ID NO: 18)-like motifs from SaCas9 and Fanzor to SpCas9. Sequences in Fig. 3A are (top) amino acids 1066 to 1073 of SEQ ID NO: 1 ; (middle) amino acids 857 to 864 of SEQ ID NO: 2; and (bottom)SEQ ID NO: 18. Sequences in Fig. 3E are (top) amino acids 1066 to 1073 of SEQ ID NO: 1 ; (middle) amino acids 293 to 300 and 408 to 514 of SEQ ID NO: 15; and (bottom) SEQ ID NO: 18. Fig. 3B and Fig. 3D show IB analyses of WCL from HEK293 cells transfected with HA-SaCas9 (WT or a mutant SaCas9 containing a AEGA (SEQ ID NO: 20) mutation in place of the WT binding motif at amino acids 859 to 862 of SaCas9 (SEQ ID NO: 2)) and the indicated amounts of Flag-Keap1 for 48 hours. Fig. 3C shows an IB analysis of Flag-IP and WCL from HEK293 cells transfected with indicated HA-SaCas9 constructs with Flag-Keap1 for 48 hours. The HA-SaCas9 constructs included those with mutant Keapl binding motifs in place of the binding motif at amino acids 859 to 862 of SaCas9 where the mutant Keapl binding motifs have a sequence selected from DSGD (SEQ ID NO: 21 ), AEGA (SEQ ID NO: 20), or EAAE (SEQ ID NO: 19). Fig. 3F shows an image of the IB analysis of WCL from HEK293T cells transfected with Flag-Fanzor with increasing doses of HA-Keap1 constructs) for 48 hours. Fig. 3G shows IB analysis of Flag-IPs and WCL from HEK293T cells transfected with indicated HA-Keap1 and the indicated Flag-Fanzor constructs (WT or corresponding to mutant proteins where the amino acid sequence at 295 to 298 or at 410 to 413 of the WT Fanzor protein is replaced by the sequence EAAE (SEQ ID NO: 19)) for 48 hours. Fig. 3H shows IB analyses of WCL from HEK293T cells transfected with indicated Flag-Fanzor constructs (WT or constructs for mutant proteins where the amino acid sequence at 295 to 298 or 410 to 413 of the WT Fanzor protein is replace by the sequence EAAE (SEQ ID NO: 19)) with or without HA-Keap1 for 48 hours.
[0030] Figures 4A-4T demonstrate that engineered Keapl degron- mutated Cas9 variants display enhanced Cas9-mediated genome editing ability in vitro. Fig. 4A is a schematic showing the sgRNA sequence targeting the human AAVS1 locus used in CRISPR-mediated knockout efficiency tests and the guide RNA sequence used in this study to introduce a Kpn1 site for knockin efficacy tests. Sequences shown correspond, from top to bottom, to the sequences of SEQ ID NOs: 25-27. Fig. 4B shows IB analysis of WCL derived from HEK293T cells transfected with Flag-SpCas9 with or without HA- Keapl constructs for 72 hours. Fig. 4C, Fig. 4E, Fig. 4F, Fig. 4H, and Fig.4Q show representative DNA agarose gel images of standard T7E1 assays derived from HEK293T cells transfected with indicated Cas9 (WT SpCas9 or a mutant SpCas 9 containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO:1 ) and sgRNA performed 3-day post-transfection. Indel% was calculated and presented to indicate target gene editing efficiency. Fig. 4D shows IB analysis of WCL derived from control or endogenous Keapl -depleted HEK293T cells transfected with indicated Cas9 constructs (WT SpCas9 or a mutant containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO: 1 ) and EMX1 - sgRNAs. Where indicated, lentiviruses for shscramble or shKeap1 -11 were used to infect HEK293 cells for 24 hrs followed by 72 hrs selection with 1 mg / mL puromycin to eliminate non-infected cells. Fig. 4G shows IB analysis of WCL derived from HEK293T cells transfected with EV (empty vector), WT- SpCas9-Flag or a mutant SpCas9-Flag (corresponding to a mutant SpCas9 containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO: 1 ) and AAVS1 - sgRNAs for 72 hrs. Fig. 41 shows (left) IB analysis of WCL derived from HEK293T cells transfected with EV (empty vector), WT-SpCas9-Flag or mutant SpCas9-Flag (for a mutant SpCas9 with an EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO:1 ) and EMX1 -sgRNAs for 72 hrs; and (Right) representative DNA agarose gel images of standard T7E1 assays derived from HEK293T cells transfected with indicated Cas9 and sgRNA performed 3-day posttransfection. Indel% was calculated and presented to indicate target gene editing efficiency. Fig. 4J, Fig. 4K, and Fig. 4L show (left) IB analysis of WCL derived from HEK293T (Fig. 4J), HBE (Fig. 4K) or BPH1 (Fig. 4L) cells transfected with EV (empty vector), WT-SpCas9-Flag or mutant SpCas9-Flag (for a mutant SpCas9 containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO: 1 ) and EMX1 -sgRNAs for 72 hrs; and (right) calculations of EMX-1 gene editing efficiency derived from DNA sanger sequencing of PCR products of targeted EMX-1 genomic regions from HEK293T cells transfected with indicated Cas9and sgRNA performed 3-day post-transfection. n=2 (biological duplicates). *p<0.05 (one-way ANOVA test). Fig. 4M is a schematic diagram showing a test of CRISPRki efficiency by introducing a new Sca-I site in the mouse Rosa26 locus. Fig. 4N shows IB analysis of WCL derived from MEF cells transfected with WT or mutant SpCas9-Flag constructs for 48 hrs. The mutant SpCas9-Flag corresponds to a mutant SpCas9 protein containing a EAAE (SEQ ID NO: 19) sequence in place of the WT Keapl binding motif at amino acids 1068 to 1071 of SEQ ID NO:1. Fig. 40 shows a representative DNA agarose gel image of Sca-I digestion of PCR products using genomic DNA from cells transfected with the WT or mutant SpCas9-Flag constructs from Fig. 4N to measure CRISPRki efficiency. Fig. 4P shows an image of IB analysis of WCL derived from HEK293T cells transfected with EV (empty vector), WT-SaCas9-Flag or mutant SaCas9-Flag and hDMD-sgRNAs for 72 hrs. The mutant SaCas9-Flag construct is for a mutant SaCas9 protein containing a AEGA (SEQ ID NO: 20) at amino acids 859-862. Fig. 4R shows calculations of hDMD gene editing efficiency derived from DNA sanger sequencing of PCR products of targeted hDMD genomic regions from HEK293T cells transfected with EV, WT SaCas9-flag or mutant SaCas9-Flag and sgRNA performed 3-day post-transfection. n=2 (biological duplicates). *p<0.05 (one-way ANOVA test). The mutant SaCas9-Flag construct corresponds to a mutant SaCas9 protein containing a AEGA (SEQ ID NO: 20) sequence in place of amino acids 859-862 of SEQ ID NO: 2. Fig. 4S and Fig. 4T show (left) IB analysis of WCL derived from BPH1 cells transfected with EV (empty vector), WT or mutant SaCas9-HA and indicated sgRNAs for 48 hrs; and (right) calculations of indicated gene editing efficiency derived from DNA sanger sequencing of PCR products of targeted genomic regions from BPH1 cells transfected with indicated Cas9 and sgRNA performed 3-day posttransfection. The mutant SaCas9-Ha construct corresponds to a mutant SaCas9 protein containing a AEGA (SEQ ID NO: 20) sequence in place of amino acids 859-862 of SEQ ID NO:2. n=2 (biological duplicates). *p<0.05 (one-way ANOVA test).
[0031] Figures 5A-5O demonstrate that dCas9-EAAE (SEQ ID NO: 19)- fusions exert enhanced gene regulation ability in cells. Fig. 5A is a schematicillustration of dCas9-p300 fusion proteins bringing the p300 transcription activator to a given site determined by sgRNA in regulating expression of genes of interest. Fig. 5B shows IB analyses of cytoplasm or chromatin fractions from HEK293T cells transfected with WT- or EAAE (SEQ ID NO: 19)- dCas9-p300 constructs for 48 hours. Fig. 5C is a graph showing the results of fluorescent reporter assays from HEK293T cells transfected with WT- dCas9-p300 or EAAE (SEQ ID NO: 19)-dCas9-p300 constructs with a fluorescence reporter indicating that dSpCas9-EAAE (SEQ ID NO: 19)-p300 fusion proteins display a increase in activating targeted exogenous gene expression compared with dSpCas9-WT-p300. Fluorescent signals were detected by FACS analyses against GFP. The * indicates p<0.05 (one-way ANOVA test). Fig. 5D is a graph showing RT-PCR analyses of expression of indicated mRNA targets from HEK293T cells transfected with WT-dCas9-p300 or EAAE (SEQ ID NO: 19)-dCas9-p300 with sgRNAs targeting either OCT4 or IL1RN, which indicates that dSpCas9-EAAE (SEQ ID NO: 19)-p300 fusion proteins display a significantly increased efficiency in activating targeted endogenous gene expression compared to dSpCas9-WT-p300. N=3 (biological triplicates). Fig. 5E is an SDS-PAGE gel showing IB analysis of chromatin fractions of HEK293T cells transfected with indicated WT- or EAAE (SEQ ID NO: 19)-dCas9-p300 constructs for 48 hours. Where indicated, 200 |ig / mL CHX was added to culture and cells were harvested at the indicated time periods. Quantifications of HA-dCas9-p300 protein half-life is presented in Fig. 5F. Fig. 5G shows an IB analysis of WT- or EAAE (SEQ ID NO: 19)- dCas9-CRAB expression levels in either cytoplasm or chromatin fractions from 48 hour transfected HEK293T cells. Fig. 5H is an image showing IB analysis of chromatin fractions of HEK293T cells transfected with indicated WT- or EAAE (SEQ ID NO: 19)-HA-dCas9-CRAB constructs for 48 hours. Where indicated, 200 p.g / mL CHX was added to culture and cells were harvested at indicated time periods. Fig. 51 provides a plot of the chromatin fractions subjected to IB analysis as described for Fig. 5H and quantified. Figs. 5J, 5K, 5L, 5M, 5N and 50 each have two panels with the lower panels showing IB analyses of WCL from HEK293T cells transfected with indicateddCas9-CRAB and sgRNAs, and the upper panels showing bar graphs of RT- PCR analyses of mRNAs from cells in lower panels expressing either WT- or EAAE (SEQ ID NO: 19)-dCas9-CRAB for indicated target gene expression. Notably, U6 snRNA is used as an internal normalization control. n=3 (biological triplicates) * p<0.05 (one-way ANOVA test).
[0032] Figures 6A-6P demonstrate that SpCas9 stabilizes endogenous Keapl substrates by competitively binding Keapl to modulate cellular functions. Fig. 6A shows an IB analysis of WCL derived from HEK293 cells transfected with Flag-NRF2 and increasing doses of Flag-SpCas9 constructs for 48 hours. Fig. 6B shows an IB analysis of WCL derived from HEK293 cells transfected with Flag-NRF2 and WT-Cas9 or EAAE (SEQ ID NO: 19)-Cas9 constructs for 48 hours. Fig. 6C shows an IB analysis of WCL and HA-IPs derived from HEK293 cells transfected with HA-Keap1 , Flag-NRF2, or Flag- SpCas9 constructs for 48 hours. Fig. 6D shows an IB analysis of WCL derived from HEK293 cells transfected with HA-SpCas9 with increasing doses of Flag- NRF2 constructs for 48 hours. Fig. 6E shows an IB analysis of WCL and Flap- IPs derived from HEK293 cells transfected with Flag-PALB2, HA-Keap1 and increasing doses of HA-SpCas9 constructs for 48 hours. Fig. 6F shows an IB analysis of WCL derived from HEK293 cells transfected with Flag-PALB2 and increasing doses of HA-SpCas9 constructs for 48 hours. Fig. 6G shows an IB analysis of WCL derived from HEK293 cells transfected with HA-SpCas9 and increasing doses of Flag-PALB2 constructs for 48 hours. Fig. 6H shows a proposed model of SpCas9 expression to modulate cellular function. SpCas9 binds and titrates Keapl substrates via the ETGE (SEQ ID NO: 18) binding motif, leading to stabilization of Keapl substrates including NRF2 and subsequently alters cellular signalling. Fig. 61 shows an IB analysis of WCL derived from HEK293 cells transfected with Flag-NRF2 with the indicated HA- Keapl WT or mutant constructs for 48 hours. Fig. 6J shows an IB analysis of WCL derived from HEK293 cells transfected with Flag-SpCas9 with the indicated HA-Keap1 WT or mutant constructs for 48 hours. Fig. 6K is a schematic diagram showing cancer patient derived Kelch domain mutations in Keapl (SEQ ID NO: 3). Fig. 6L shows an IB analysis of WCL from HEK293T cells transfected with Flag-NRF2 with the indicated HA-Keap1 WTor mutant constructs. Fig. 6M shows an IB analyses of WCL from HEK293T cells transfected with Flag-SpCas9 with the indicated HA-Keap1 WT or mutant constructs. Fig. 6N shows a table summarizing various Keapl mutants with impaired function in either degrading NRF2 or Cas9. Fig. 60 shows an IB analysis of WCL from HEK293T cells transfected with Flag-SpCas9 with the indicated HA-Keap1 WT or mutant constructs and EMX1 -sgRNA for 72 hours. Fig. 6P shows a representative DNA agarose gel for standard T7E1 assays using cells from Fig. 60.
[0033] Figures 7A-7B show that sgRNA does not affect Cas9 binding to Keapl or Cas9 expression. Fig. 7A shows an IB analysis of HA-IP and WCL from HEK293 cells transfected with Flag-SpCas9, HA-Keap1 , and the indicated amount of sgAAVS plasmids (0, 2, or 4 pg) for 48 hours. Fig. 7B shows an IB analysis of WCL from HEK293 cells transfected with 0.5, 1 , 1 .5, or 2 pg sgAAVS plasmids with or without CMV-Flag-SpCas9.
[0034] Figures 8A-8O demonstrate that Keapl targets Cas9 for ubiquitination and degradation. Fig. 8A is a schematic drawing showing Keapl protein domain structures. Keapl -Kelch (including the sequence at amino acids 327 to 624 of SEQ ID NO: 3) and Keapl -AKelch (including the sequence at amino acids 1 to 286 of SEQ ID NO: 3) Fig. 8B shows an IB analysis of Flag-IP and WCL from HEK293 cells transfected with the indicated Keapl and Cas9 constructs. Fig. 8C and Fig. 8D show IB analyses of WCL from HEK293 cells transfected with the indicated DNA constructs. Fig. 9D includes analysis of WCL from HEK293 cells transfected with a construct corresponding to a SpCas9 mutant containing the sequence EAAE (SEQ ID NO: 19) in place of the WT Keapl binding motif. Fig. 8E shows an IB analysis of WCL from indicated shscramble or shKeap1 -H520 cells transfected with either WT-Cas9-Flag or EAAE (SEQ ID NO: 19)-Cas9-Flag for 48 hours. Fig. 8F is a graph showing RT-PCR analysis of mRNA levels of transfected Flag- SpCas9-WT plasmids from HEK293 cells at 48 hours. Fig. 8G and Fig. 8H show IB analyses of WCL from HEK293 cells transfected with the indicated Flag-Cas9 and / or HA-Keap1 constructs. Where indicated, MG132 (10 mM, overnight) or bortezomib (10 mM, overnight) was added to cell culture mediaprior to cell collection. Fig. 81 shows an IB analysis of Ni-NTA pulldown and WCL from HEK293 cells transfected with indicated Flag-Cas9 and His-Ub plasmids for 48 hours. Fig. 8J and Fig. 8K show IB analyses of WCL from HEK293T cells transfected with Flag-SpCas9 (200 ng; WT or a mutant with a EAAE (SEQ ID NO:19) sequence in place of the WT Keapl binding motif) or HA-Keap1 for 48 hours. As indicated, 200 pg / mL cycloheximide (CHX) was added to cell culture and cells were collected at indicated time periods after CHX addition. Fig. 8L shows an IB analysis of WCL from HEK293 cells transfected with Flag-SpCas9 and treated with CDDO (5 mM) or tBHQ (200 mM) for 3.5 hours before cell collection. Fig. 8M and Fig. 8N show IB analyses of WCL from HEK293T cells transfected with indicated Cas9 (WT or EAAE (SEQ ID NO: 19) mutant) or Keapl plasmids. As indicated, cells were treated with CDDO (50 nM), tBHQ (10 mM) or sulforaphane (Sulf, 10 mM) for 16 hours before cell collection. Fig. 80 shows an IB analysis of Flag-IP and WCL from HEK293 cells transfected with HA-Keap1 and indicated Flag-Cas9 (WT or EAAE (SEQ ID NO:19) mutant). For 48 hours. Where indicated, cells were treated with tBHQ (10 pM) for 16 hours before cell collection.
[0035] Figures 9A-9G demonstrate that Keapl targets SaCas9 for ubiquitination and degradation. Figs. 9A shows IB analysis of WCL from HEK293 cells transfected with indicated Flag-SpCas9 plasmids with increasing doses (0, 1 , or 2 pg) of HA-Keap1 for 48 hours. Fig. 9B shows an IB analysis of WCL from HEK293 cells transfected with the indicated HA- SaCas9 plasmids (WT or mutants with a DSGD (SEQ ID NO: 21 ), AEGA (SEQ ID NO: 20), or EAAE (SEQ ID NO: 19) sequence in the Keapl binding motif) at 24 or 48 hours. Fig. 9C and Fig. 9D show IB analyses of WCL from HEK293 cells transfected with indicated HA-SaCas9 plasmids (WT or mutants with a EAAE (SEQ ID NO: 19 or AEGA (SEQ ID NO: 20 sequence in the Keapl binding motif) for 48 hours. Fig. 9E shows IB analysis of Ni-NTA pulldown and WCL from HEK293 cells transfected with HA-SaCas9 with indicated His-Ub- only plasmids for 48 hours. As indicated, cells were treated with 10 mM MG132 overnight before cell collection. Fig. 9F shows IB analysis of Ni-NTA pulldown and WCL from HEK293 cells transfected with HA-SaCas9, Flag-Keapl with indicated His-Ub-only plasmids for 48 hours. As indicated, cells were treated with 10 mM MG 132 overnight before cell collection. Fig. 9G is a schematic drawing showing E3 ligase Keapl recognition of the ETGE (SEQ ID NO: 18) motif in WT Cas9 proteins to target the mutants for ubiquitination and proteasomal degradation, leading to reduced gene editing ability.
[0036] Figures 10A-10J show engineered Cas9 mutants that evade Keapl recognition display enhanced gene editing ability in cells. Fig. 10A shows an image of a DNA gel with T7E1 assays showing increased AAVS1 gene deletion efficiency in shKeapI HEK293 cells. Fig. 10B is an image showing a DNA gel with T7E1 assays showing increased AAVS1 gene deletion efficiency in EAAE (SEQ ID NO:19)-SpCas9 mutant expressing HEK293 cells compared with WT-Cas9 expressing cells. Fig. 10C is an image of a DNA gel with T7E1 assays (by Kpnl digestion) showing increased Kpnl knockin efficiency in HEK293 cells expressing EAAE (SEQ ID NO: 19)-Cas9 compared to cells expressing WT-SpCas9. Fig. 10D is a IB analysis of WCL from HEK293T cells transfected with the indicated Flag-SpCas9 plasmids (WT or mutant SpCas9 containing EAAE (SEQ ID NO:19) in the Keapl binding motif) for 72 hours. Fig. 10E (top) shows a portion of a primer sequence (i.e., the CACTAGGGACAGGTACCGTGACAGAAA; nucleotides 15-41 of the sequence of SEQ ID NO: 25) and (bottom) a DNA agarose gel image from qPCR using knockin specific primers with genomic DNA showed increased AAVS1 knockin efficiency in EAAE (SEQ ID NO: 19)-SpCas9 expressing HEK293 cells compared with WT-SpCas9 expressing cells. Fig. 10F shows a presentative DNA PAGE gel image from T7E1 assays (by Kpnl digestion) showing increased Kpnl site knockin efficiency in Keapl depleted HEK293 cells expressing WT-Cas9 but not EAAE (SEQ ID NO: 19)-Cas9. Fig. 10G and Fig. 10H, top panels, show IP analyses of WCL from HEK293 cells transfected with indicated HA-SaCas9 constructs (WT or a mutant comprising a AEGA (SEQ ID NO: 20) in the WT Keapl binding motif) and sgRNAs. Fig. 10G and Fig. 10H, bottom panels, show representative DNA gel images from T7E1 assays showing increased deletion efficiency of the indicated gene targets in mutated-SaCas9 expressing cells compared with WT-SaCsa9 expressing cells. Fig. 101 and Fig. 10J, left panels, show IB analysis of WCLderived from cells transfected with EV (empty vector), WT-, or AEGA(SEQ ID NO: 20)-Ha-SaCas9 and FANCF-sgRNA (Fig. 101) or EMX1 -sgRNA(Fig. 10J) for 72 hours. Fig. 101 and Fig. 10J, right, show TIDE analyses of FANCF(Fig. 101) or EMX-1 (Fig. 10 J) gene editing efficiency derived from DNA sanger sequencing of PCR products targeted FANCF or EMX- 1 genomic regions from left panels. n=2 (biological duplicates), *p<0.05 (one-way ANOVA test).
[0037] Figures 11A-11 E show engineered SaCas9 mutants that evade Keapl recognition display enhanced gene editing ability in cells but minimal effects in a DMD animal model. Fig. 11A is a DNA gel image from T7E1 assays showing increased DMD gene deletion efficiency from C2C12 cells electroporated with the indicated HA-SaCas9 constructs (WT or a mutant with an AEGA (SEQ ID NO: 20) in place of the WT Keapl binding motif) and DMD guide RNAs. Fig. 11 B is a schematic diagram shwoing the testing protocol for in vivo gene editing ability of either SaCas9-WT or SaCas9-AEGA (SEQ ID NO: 20) using the DMD murine model. Fig. 11C, left panel, shows an IB analysis of WCL derived from tibialis anterior tissue harvested from DMD mice infected with the indicated adeno-viral SaCas9 viruses (WT or a mutant with AEGA (SEQ ID NO: 20) in place of the WT Keapl binding motif) 8-weeks postinfection. Fig. 11 C, right panel, shows representative end-point PCR analyses of DMD gene deletion efficiency in tibialis anterior tissue. Fig. 11 D shows (top panel) representative end-point PCR analyses of DMD gene deletion efficiency in mouse liver tissue; (middle panel) IB analysis of SaCas9-HA expression from mouse liver tissue; and (bottom panel) RT-PCR analyses of SaCas9 mRNA expression levels from mouse liver tissue. Data from individual mice infected with WT saCas9 viruses or mutant saCas9 virus with an AEGA (SEQ ID NO: 20) in place of the WT Keapl binding motif). Fig. 11 E shows (top panel) representative end-point PCR analyses of DMD gene deletion efficiency in mouse heart tissue; (middle panel) IB analysis of SaCas9-HA expression from mouse heart tissue; and (bottom panel) RT-PCR analyses of SaCas9 mRNA expression levels from mouse heart tissue. Data from individual mice infected with WT saCas9 viruses or mutant saCas9 virus with an AEGA (SEQ ID NO: 20) in place of the WT Keapl binding motif).
[0038] Figures 12A-12I demonstrate that the eSpCas9-EAAE (SEQ ID NO: 19) mutant displays improved gene editing but less off-target effects in cells. Fig. 12A shows an IB analyses of WCL from HEK293T cells transfected with indicated Flag-SpCas9 or Flag-eSpCas9 constructs with EMX1 sgRNA for 72 hours. Fig. 12B is a graph showing the calculations of EMX1 on-target gene editing efficiency derived from DNA sanger sequencing of PCR products of targeted EXM1 genomic regions from HEK293T cells transfected with indicated Cas9 and sgRNA performed 3-day post-transfection. n=2 (biological triplicates). *p<0.05 (one-way ANOVA test). Fig. 12C, Fig. 12D, Fig. 12E, Fig. 12F, Fig. 12G, Fig. 12H, and Fig. 121 are graphs showing the calculations of indicated EMX1 off-target gene editing efficiency (Fig. 12C, Off-target 2; Fig. 12D, Off-target 3; Fig. 12E, Off-target 4; Fig. 12F, Off-target 5; Fig. 12G, Off- target 6; Fig. 12H, Off-target 7; and Fig. 121, Off-target 8) derived from DNA sanger sequencing of PCR products of targeted EXM1 genomic regions from HEK293T cells transfected with indicated Cas9 and sgRNA performed 3-day post-transfection. n=2 (biological triplicates). *p<0.05 (one-way ANOVA test).
[0039] Figures 13A-13L demonstrate that EAAE (SEQ ID NO: 19)-dCas9 fusions exert enhanced epigenome editing ability. Fig. 13A shows IB analysis of WCL from HEK293T cells transfected with Flag-dSpCas9 with increasing doses of HA-Keap1 for 48 hours. Fig. 13B shows an IB analyses of WCL from HEK293 cells transfected with HA-dSpCas9-p300 fusion with increasing doses of Flag-Keap1 for 48 hours. Fig. 13C shows an IB analyses of WCL from HEK293 cells transfected with either HA-dSpCas9-WT-p300 or HA- dSpCas9-EAAE (SEQ ID NO: 19)-p300 fusion for 48 hours. Fig. 13D shows an IB analyses of WCL from HEK293 cells transfected with either HA- dSpCas9-WT-p300 or HA-dSpCas9-EAAE (SEQ ID NO: 19)-p300 fusion for indicated days. Fig. 13E and Fig. 131 show IB analyses of WCL from HEK293T cells transfected with indicated dCas9-fusions for 48 hours. Where indicated, cells were treated with 200 mg / mL CHX and collected at indicated time points post-CHX additions. Fig. 13F shows IB analyses of WCL from HEK293 cells transfected with either HA-dSpCas9-WT-VPR, HA-dSpCas9- WT-FKBP, HA-dSpCas9-EAAE (SEQ ID NO: 19)-VPR, or HA-dSpCas9-EAAE (SEQ ID NO: 19)-FKBP fusion for indicated days. Fig. 12G is a graph showingfluorescent reporter assay results indicating that dSpCas9-EAAE (SEQ ID NO: 19)-p300 fusion proteins display a 1 .9-fold increase in activating targeted exogenous gene expression compared with dSpCas9-WT-p300 at day 3 posttransfection. Fluorescent signals were detected by FACS analyses against GFP. Fig. 13H is a graph showing fluorescent reporter assay results indicating that dSpCas9-EAAE (SEQ ID NO: 19)-Suntag fusion proteins display a 3.1 - fold increase in activating targeted exogenous gene expression compared with dSpCas9-WT-p300. Fluorescent signals were detected by FACS analyses against GFP. Fig. 13J, Fig. 13K, and Fig. 13L show RT-PCR analyses of mRNAs levels from cells in lower panels expressing either WT- or EAAE (SEQ ID NO: 19)-dCas9-CRAB for indicated target gene expression. Notably, Actin is used as an internal normalization control. n=3 (biological triplicates). *p<0.05 (one-way ANOVA test).DETAILED DESCRIPTION
[0040] As a potent and convenient genome-editing tool, Cas9 has been widely used in biomedical research and evaluated for gene therapy approaches to treat human diseases. Distinctly engineered Cas9s, dCas9s and additional related prokaryotic endonucleases have been identified. However, as these bacterial enzymes are not naturally present in mammalian cells, whether and how bacterial Cas9 proteins are regulated by mammalian hosts remains poorly understood. As disclosed herein, it has now been shown that Keapl acts as a host endogenous E3 ligase that targets Cas9 / dCas9 / Fanzor for ubiquitination and degradation and that Cas9 containing mutations in the ETGE (SEQ ID NO: 18, also amino acids 1068 to 1071 of SEQ ID NO: 1 ) sequence of the Keapl binding motif of Cas9, (which mutants are also referred to herein as Cas9-“ETGE” (SEQ ID NO: 18) mutants) are capable of evading Keapl recognition display enhanced gene editing ability in cells. Moreover, corresponding dCas9 fusion mutants (referred to herein as dCas9-“ETGE” (SEQ ID NO: 18)-fusion mutants) exert extended protein half-life on chromatin, leading to improved CRISPRa and CRISPRi efficacy. In addition, Cas9 binding to Keapl also impairs Keapl function by competing with Keapl substrates or binding partners, whileengineered Cas9 mutants show less perturbation of Keapl biology. Thus, the presently disclosed subject matter reveals Cas9 regulation that is specific to mammalian systems and provides new Cas9 designs not only with enhanced gene regulatory capacity but also with minimal effects on disrupting endogenous Keapl signaling.
[0041] The presently disclosed subject matter will now be described more fully. The presently disclosed subject matter can, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein below and in the accompanying Examples. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the embodiments to those skilled in the art.
[0042] All references listed herein, including but not limited to all patents, patent applications and publications thereof, and scientific journal articles, are incorporated herein by reference in their entireties to the extent that they supplement, explain, provide a background for, or teach methodology, techniques, and / or compositions employed herein.
[0043] It is to be understood that the disclosed method and compositions are not limited to specific synthetic methods, specific analytical techniques, or to particular reagents unless otherwise specified, and, as such, may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.
[0044] Disclosed are materials, compositions, and components that can be used for, can be used in conjunction with, can be used in preparation for, or are products of the disclosed method and compositions. These and other materials are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these materials are disclosed that while specific reference of each various individual and collective combinations and permutation of these compounds may not be explicitly disclosed, each is specifically contemplated and described herein. For example, if a mutant Cas9 protein is disclosed and discussed and a number of modifications that can be made to a number of molecules including the mutant Cas9 protein are discussed, each and every combination and permutation of the mutant Cas9protein and the modifications that are possible are specifically contemplated unless specifically indicated to the contrary. Thus, if a class of molecules A, B, and C are disclosed as well as a class of molecules D, E, and F and an example of a combination molecule, A-D is disclosed, then even if each is not individually recited, each is individually and collectively contemplated. Thus, is this example, each of the combinations A-E, A-F, B-D, B-E, B-F, C-D, C-E, and C-F are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A- D. Likewise, any subset or combination of these is also specifically contemplated and disclosed. Thus, for example, the sub-group of A-E, B-F, and C-E are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A- D. This concept applies to all aspects of this application including, but not limited to, steps in methods of making and using the disclosed compositions. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific embodiment or combination of embodiments of the disclosed methods, and that each such combination is specifically contemplated and should be considered disclosed.I. Definitions
[0045] It is understood that the disclosed method and compositions are not limited to the particular methodology, protocols, and reagents described as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention which will be limited only by the appended claims.
[0046] While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.
[0047] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniquesemployed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject matter.
[0048] In describing the presently disclosed subject matter, it will be understood that a number of techniques and steps are disclosed. Each of these has individual benefit and each can also be used in conjunction with one or more, or in some cases all, of the other disclosed techniques.
[0049] Accordingly, for the sake of clarity, this description will refrain from repeating every possible combination of the individual steps in an unnecessary fashion. Nevertheless, the specification and claims should be read with the understanding that such combinations are entirely within the scope of the invention and the claims.
[0050] It must be noted that as used herein and in the appended claims, the singular forms "a", "an”, and "the" include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to "a mutant Cas9 protein" includes a plurality of such mutant Cas9 proteins, reference to "the mutant Cas9 protein" is a reference to one or more mutant Cas9 proteins and equivalents thereof known to those skilled in the art, and so forth.
[0051] “Optional” or “optionally” means that the subsequently described event, circumstance, or material may or may not occur or be present, and that the description includes instances where the event, circumstance, or material occurs or is present and instances where it does not occur or is not present.
[0052] Ranges may be expressed herein as from "about" one particular value, and / or to "about" another particular value. When such a range is expressed, also specifically contemplated and considered disclosed is the range from the one particular value and / or to the other particular value unless the context specifically indicates otherwise. Similarly, when values are expressed as approximations, by use of the antecedent “about,” it will be understood that the particular value forms another, specifically contemplated embodiment that should be considered disclosed unless the contextspecifically indicates otherwise. For example, term “about,” when referring to a value or to an amount of a composition, dose, sequence identity (e.g., when comparing two or more nucleotide or amino acid sequences), mass, weight, temperature, time, volume, concentration, percentage, etc., is meant to encompass variations of in some embodiments ±20%, in some embodiments ±10%, in some embodiments ±5%, in some embodiments ±1 %, in some embodiments ±0.5%, and in some embodiments ±0.1 % from the specified amount, as such variations are appropriate to perform the disclosed methods or employ the disclosed compositions. It will be further understood that the endpoints of each of the ranges are significant both in relation to the other endpoint, and independently of the other endpoint unless the context specifically indicates otherwise. Finally, it should be understood that all of the individual values and sub-ranges of values contained within an explicitly disclosed range are also specifically contemplated and should be considered disclosed unless the context specifically indicates otherwise. For example, 1 to 5 includes 1 , 1.5, 2, 2.75, 3, 3.90, 4, and 5. The foregoing applies regardless of whether in particular cases some or all of these embodiments are explicitly disclosed.
[0053] The term “comprising”, which is synonymous with “including”, “containing”, or “characterized by”, is inclusive or open-ended and does not exclude additional, unrecited elements and / or method steps. “Comprising” is a term of art that means that the named elements and / or steps are present, but that other elements and / or steps can be added and still fall within the scope of the relevant subject matter.
[0054] The phrase “consisting essentially of” limits the scope of the related disclosure or claim to the specified materials and / or steps, plus those that do not materially affect the basic and novel characteristic(s) of the disclosed and / or claimed subject matter. For example, a pharmaceutical composition can “consist essentially of” a pharmaceutically active agent or a plurality of pharmaceutically acitive agents, which means that the recited pharmaceutically active agent(s) is / are the only pharmaceutically active agent(s) present in the pharmaceutical composition. It is noted, however, that carriers, excipients, and / or other inactive agents can and likely would bepresent in such a pharmaceutical compostion and are encompassed within the nature of the phrase “consisting essentially of.”
[0055] As used herein, the phrase “consisting of” excludes any element, step, or ingredient not specifically recited. It is noted that, when the phrase “consists of” appears in a clause of the body of a claim, rather than immediately following the preamble, it limits only the element set forth in that clause, other elements are not excluded from the claim as a whole.
[0056] With respect to the terms “comprising”, “consisting of”, and “consisting essentially of”, where one of these three terms is used herein, the presently disclosed and claimed subject matter can include the use of either of the other two terms. For example, a composition that in some embodiments comprises a given active agent also in some embodiments can consist essentially of that same active agent, and indeed can in some embodiments consist of that same active agent.
[0057] As used herein, the term “and / or” when used in the context of a listing of entities, refers to the entities being present singly or in combination. Thus, for example, the phrase “A, B, C, and / or D” includes A, B, C, and D individually, but also includes any and all combinations and subcombinations of A, B, C, and D.
[0058] The term “gene” refers broadly to any segment of DNA associated with a biological function. A gene can comprise sequences including but not limited to a coding sequence, a promoter region, a cis-regulatory sequence, a non-expressed DNA segment that is a specific recognition sequence for regulatory proteins, a non-expressed DNA segment that contributes to gene expression, a DNA segment designed to have desired parameters, or combinations thereof. A gene can be obtained by a variety of methods, including cloning from a biological sample, synthesis based on known or predicted sequence information, and recombinant derivation of an existing sequence.
[0059] As is understood in the art, a gene comprises a coding strand and a non-coding strand. As used herein, the terms “coding strand”, “coding sequence” and “sense strand” are used interchangeably, and refer to a nucleic acid sequence that has the same sequence of nucleotides as an mRNA fromwhich the gene product is translated. As is also understood in the art, when the coding strand and / or sense strand is used to refer to a DNA molecule, the coding / sense strand includes thymidine residues instead of the uridine residues found in the corresponding mRNA. Additionally, when used to refer to a DNA molecule, the coding / sense strand can also include additional elements not found in the mRNA including, but not limited to promoters, enhancers, and introns. Similarly, the terms “template strand” and “antisense strand” are used interchangeably and refer to a nucleic acid sequence that is complementary to the coding / sense strand.
[0060] Similarly, all genes, gene names, and gene products disclosed herein are intended to correspond to homologs from any species for which the compositions and methods disclosed herein are applicable. Thus, the terms include, but are not limited to genes and gene products from humans and mice. It is understood that when a gene or gene product from a particular species is disclosed, this disclosure is intended to be exemplary only, and is not to be interpreted as a limitation unless the context in which it appears clearly indicates. Also encompassed are any and all nucleotide sequences that encode the disclosed amino acid sequences, including but not limited to those disclosed in the corresponding GENBANK® entries.
[0061] The term “gene expression” generally refers to the cellular processes by which a biologically active polypeptide is produced from a DNA sequence and exhibits a biological activity in a cell. As such, gene expression involves the processes of transcription and translation, but also involves post- transcriptional and post-translational processes that can influence a biological activity of a gene or gene product. These processes include, but are not limited to RNA syntheses, processing, and transport, as well as polypeptide synthesis, transport, and post-translational modification of polypeptides. Additionally, processes that affect protein-protein interactions within the cell can also affect gene expression as defined herein.
[0062] The terms "modulate" or “alter” are used interchangeably and refer to a change in the expression level of a gene, or a level of RNA molecule or equivalent RNA molecules encoding one or more proteins or protein subunits, or activity of one or more proteins or protein subunits is up regulated or downregulated, such that expression, level, or activity is greater than or less than that observed in the absence of the modulator. For example, the terms "modulate" and / or “alter” can mean "inhibit" or “suppress”, but the use of the words "modulate" and / or “alter” are not limited to this definition.
[0063] As used herein, the terms "inhibit", “suppress”, “repress”, “downregulate”, “loss of function", “block of function", and grammatical variants thereof are used interchangeably and refer to an activity whereby gene expression (e.g., a level of an RNA encoding one or more gene products) is reduced below that observed in the absence of a composition of the presently disclosed subject matter. In some embodiments, inhibition results in a decrease in the steady state level of a target RNA. By way of example and not limitation, histone methyltransferases, such as G9a, can suppress transcription of a number of genes below that observed in the absence of histone methyltransferases.
[0064] The term "RNA" refers to a molecule comprising at least one ribonucleotide residue. By "ribonucleotide" is meant a nucleotide with a hydroxyl group at the 2' position of a D-ribofuranose moiety. The terms encompass double stranded RNA, single stranded RNA, RNAs with both double stranded and single stranded regions, isolated RNA such as partially purified RNA, essentially pure RNA, synthetic RNA, recombinantly produced RNA, as well as altered RNA, or analog RNA, that differs from naturally occurring RNA by the addition, deletion, substitution, and / or alteration of one or more nucleotides. Such alterations can include addition of non-nucleotide material, for example at one or more nucleotides of the RNA. Nucleotides in the RNA molecules of the presently disclosed subject matter can also comprise non-standard nucleotides, such as non-naturally occurring nucleotides or chemically synthesized nucleotides or deoxynucleotides. These altered RNAs can be referred to as analogs or analogs of a naturally occurring RNA.
[0065] The term “transcription factor” generally refers to a protein that modulates gene expression, such as by interaction with the cis-regulatory element and / or cellular components for transcription, including RNA Polymerase, Transcription Associated Factors (TAFs), chromatin-remodelingproteins, reverse tet-responsive transcriptional activator, and any other relevant protein that impacts gene transcription.
[0066] The term “promoter” defines a region within a gene that is positioned 5' to a coding region of a same gene and functions to direct transcription of the coding region. The promoter region includes a transcriptional start site and at least one cis-regulatory element. The term “promoter” also includes functional portions of a promoter region, wherein the functional portion is sufficient for gene transcription. To determine nucleotide sequences that are functional, the expression of a reporter gene is assayed when variably placed under the direction of a promoter region fragment.
[0067] The term “degron” can be defined as a portion of a protein that is important in regulation of protein degradation rates. Known degrons include short amino acid sequences, structural motifs and exposed amino acids located anywhere in the protein. Some proteins can even contain multiple degrons.
[0068] The terms “active”, “functional” and “physiological”, as used for example in “enzymatically active”, “functional chromatin” and “physiologically accurate”, and variations thereof, refer to the states of genes, regulatory components, chromatin, etc. that are reflective of the dynamic states of each as they exists naturally, or in vivo, in contrast to static or non-active states of each. Measurements, detections or screenings based on the active, functional and / or physiologically relevant states of biological indicators can be useful in elucidating a mechanism, or defining a disease state or phenotype, as it occurs naturally. This is in contrast to measurements taken based on static concentrations or quantities of a biological indicator that are not reflective of level of activity or function thereof.
[0069] As used herein “protein half-life” refers to the time it takes for half of the amount of a specific protein in a biological system to be degraded or eliminated. It is a measure of the stability of the protein and is influenced by factors such as cellular environment, post-translational modifications, and interactions with other molecules, including enzymes like ubiquitin ligases that target proteins for degradation.
[0070] As used herein, the terms “antibody" and “antibodies” refer to proteins comprising one or more polypeptides substantially encoded by immunoglobulin genes or fragments of immunoglobulin genes. The presently disclosed subject matter also includes functional equivalents of the antibodies of the presently disclosed subject matter. As used herein, the phrase “functional equivalent” as it refers to an antibody refers to a molecule that has binding characteristics that are comparable to those of a given antibody. In some embodiments, chimerized, humanized, and single chain antibodies, as well as fragments thereof, are considered functional equivalents of the corresponding antibodies upon which they are based. In some embodiments, the presently disclosed subject matter provides methods for identifying, characterizing and / or developing disease-related components of a genespecific chromatin regulatory protein complex, wherein one or more antibodies can be used directly, or in assays related thereto, in the identification, characterization and / or isolation of such components.
[0071] The term "substantially identical”, as used herein to describe a degree of similarity between nucleotide sequences, peptide sequences and / or amino acid sequences refers to two or more sequences that have in one embodiment at least about least 60%, in another embodiment at least about 70%, in another embodiment at least about 80%, in another embodiment at least about 85%, in another embodiment at least about 90%, in another embodiment at least about 91%, in another embodiment at least about 92%, in another embodiment at least about 93%, in another embodiment at least about 94%, in another embodiment at least about 95%, in another embodiment at least about 96%, in another embodiment at least about 97%, in another embodiment at least about 98%, in another embodiment at least about 99%, in another embodiment about 90% to about 99%, and in another embodiment about 95% to about 99% nucleotide identity, when compared and aligned for maximum correspondence, as measured using a sequence comparison algorithm or by visual inspection.
[0072] The terms “variant” and “derivative” as described herein can refer to both nucleic acid sequences and protein sequences. It is understood that one way to define any known variants and derivatives or those that might arise, ofthe disclosed nucleic acid sequences and proteins herein is through defining the variants and derivatives in terms of homology to specific known sequences. Specifically disclosed are variants of the genes and proteins herein disclosed which have at least, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, or 99 percent homology to the stated sequence. Those of skill in the art readily understand how to determine the homology of two proteins or nucleic acids, such as genes. For example, the homology can be calculated after aligning the two sequences so that the homology is at its highest level.
[0073] Another way of calculating homology can be performed by published algorithms. Optimal alignment of sequences for comparison may be conducted by the local homology algorithm of Smith and Waterman Adv. Appl. Math. 2: 482 (1981 ), by the homology alignment algorithm of Needleman and Wunsch, J. MoL Biol. 48: 443 (1970), by the search for similarity method of Pearson and Lipman, Proc. Natl. Acad. Sci. U.S.A. 85: 2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wl), or by inspection.
[0074] The same types of homology can be obtained for nucleic acids by for example the algorithms disclosed in Zuker, M. Science 244:48-52, 1989, Jaeger et al. Proc. Natl. Acad. Sci. USA 86:7706-7710, 1989, Jaeger et al. Methods Enzymol. 183:281 -306, 1989 which are herein incorporated by reference for at least material related to nucleic acid alignment.
[0075] For example, as used herein, a sequence recited as having a particular percent homology to another sequence refers to sequences that have the recited homology as calculated by any one or more of the calculation methods described above. For example, a first sequence has 80 percent homology, as defined herein, to a second sequence if the first sequence is calculated to have 80 percent homology to the second sequence using the Zuker calculation method even if the first sequence does not have 80 percent homology to the second sequence as calculated by any of the other calculation methods. As another example, a first sequence has 80 percent homology, as defined herein, to a second sequence if the first sequence iscalculated to have 80 percent homology to the second sequence using both the Zuker calculation method and the Pearson and Lipman calculation method even if the first sequence does not have 80 percent homology to the second sequence as calculated by the Smith and Waterman calculation method, the Needleman and Wunsch calculation method, the Jaeger calculation methods, or any of the other calculation methods. As yet another example, a first sequence has 80 percent homology, as defined herein, to a second sequence if the first sequence is calculated to have 80 percent homology to the second sequence using each of calculation methods (although, in practice, the different calculation methods will often result in different calculated homology percentages).II. CAS9 MutantsA. Proteins
[0076] In some embodiments, the presently disclosed subject matter provides a mutant Cas9 protein. The mutant Cas9 protein can comprise one or more mutations in the Keapl degron sequence of the Cas9 protein. The Keapl degron sequence of Cas9 protein is a sequence targeted by Keapl , an E3 ligase that can act as a Cas9 suppressor. Mutating the Keapl degron sequence of Cas9 can prevent Keapl from suppressing Cas9. In some embodiments, the mutant Cas9 protein has a half-life greater than the half-life of the corresponding Cas9 protein that does not comprise (i.e., is free of) the one or more mutations in the Keapl degron sequence. In some embodiments, the mutant Cas9 protein has a half-life greater than the half-life of the corresponding wild-type (WT) Cas9 protein.
[0077] In some embodiments, the Cas9 protein (i.e., the Cas9 protein that is free of one or more mutations in the Keapl degron sequence) is a wild-type (WT) Streptococcus pyogenes Cas9 (SpCas9) protein. WT Cas9 protein has the sequence found in UNIPROT Q99ZW2 and shown below:MDKKYS IGLD IGTNSVGWAV ITDEYKVP SK KFKVLGNTDR HSIKKNLIGA LLFDSGETAE ATRLKRTARR RYTRRKNRIC YLQEIFSNEM AKVDDSFFHR LEESFLVEED KKHERHP IFGNIVDEVAYHE KYPTIYHLRK KLVDSTDKAD LRLIYLALAHMIKFRGHFLI EGDLNPDNSD VDKLFIQLVQ TYNQLFEENPINASGVDAKA ILSARLSKSR RLENLIAQLP GEKKNGLFGNLIALSLGLTP NFKSNFDLAE DAKLQLSKDT YDDDLDNLLAQIGDQYADLF LAAKNLSDAI LLSDILRVNT EITKAPLSASMIKRYDEHHQ DLTLLKALVR QQLPEKYKEI FFDQSKNGYAGYIDGGASQE EFYKFIKP IL EKMDGTEELL VKLNREDLLRKQRTFDNGSI PHQIHLGELH AILRRQEDFY PFLKDNREKIEKILTFRIPY YVGPLARGNS RFAWMTRKSE ETITPWNFEEVVDKGASAQS FIERMTNFDK NLPNEKVLPK HSLLYEYFTVYNELTKVKYV TEGMRKPAFL SGEQKKAIVD LLFKTNRKVTVKQLKEDYFK KIECFDSVEI SGVEDRFNAS LGTYHDLLKIIKDKDFLDNE ENEDILEDIV LTLTLFEDRE MIEERLKTYAHLFDDKVMKQ LKRRRYTGWG RLSRKLINGI RDKQSGKTILDFLKSDGFAN RNFMQLIHDD SLTFKEDIQK AQVSGQGDSLHEHIANLAGS PAIKKGILQT VKVVDELVKV MGRHKPENIVIEMARENQTT QKGQKNSRER MKRIEEGIKE LGSQILKEHPVENTQLQNEK LYLYYLQNGR DMYVDQELDI NRLSDYDVDHIVPQSFLKDD S IDNKVLTRS DKNRGKSDNV P SEEVVKKMKNYWRQLLNAK LITQRKFDNL TKAERGGLSE LDKAGFIKRQLVETRQITKH VAQILDSRMN TKYDENDKLI REVKVITLKSKLVSDFRKDF QFYKVREINN YHHAHDAYLN AVVGTALIKKYPKLESEFVY GDYKVYDVRK MIAKSEQEIG KATAKYFFYSNIMNFFKTEI TLANGEIRKR PLIETNGETG EIVWDKGRDFATVRKVLSMP QVNIVKKTEV QTGGFSKES I LPKRNSDKLIARKKDWDPKK YGGFDSPTVA YSVLVVAKVE KGKSKKLKSVKELLGITIME RSSFEKNP ID FLEAKGYKEV KKDLI IKLPKYSLFELENGR KRMLASAGEL QKGNELALP S KYVNFLYLASHYEKLKGSPE DNEQKQLFVE QHKHYLDEI I EQI SEFSKRVILADANLDKV LSAYNKHRDK P IREQAENI I HLFTLTNLGAPAAFKYFDTT IDRKRYTSTK EVLDATLIHQ S ITGLYETRIDLSQLGGD ( SEQ ID NO : 1 ) •
[0078] In some embodiments, the Cas9 protein (i.e., the Cas9 protein that does not comprise the one or more mutations in the Keapl degraon sequence) is a Cas9 comprising or consisting of the sequence of SEQ ID NO: 1 or a sequence having at least about 90% homology to the sequence of SEQ ID NO: 1. In some embodiments, the sequence has at least about 91 % homology, at least about 92% homology, at least about 93% homology, at least about 94% homology, at least about 95% homlogy, at least about 96% homlogy, at least about 97% homlogy, at least about 98% homology, or at least about 99% homology to the sequence of SEQ ID NO: 1 .
[0079] In some embodiments, the Keapl degron sequence comprises an ETGE (SEQ ID NO: 18) sequence. An ETGE (SEQ ID NO: 18)- or ETGE (SEQ ID NO: 18)-like motif is present in most Cas9 species.
[0080] In some embodiments, the Keapl degron sequence is the ETGE (SEQ ID NO: 18) sequence at amino acids 1068 to 1071 of SEQ ID NO: 1 . In some embodiments, the Keapl degron sequence is located in the SpCas9 Ruvlll domain that regulates Cas9 enzyme activity.
[0081] In some embodiments, the one or more mutations of the presently disclosed mutant Cas9 protein comprise an alanine substitution at position 1069 or position 1070 as compared to SEQ ID NO: 1 . In some embodiments, the one or more mutations comprise an alanine substitution at both of positions 1069 and 1070 as compared to SEQ ID NO: 1 . Stated another way, the presently disclosed mutant SpCas9 can comprise the sequence of SEQ ID NO: 1 , i.e., the sequence of WT SpCas9, except comprising one or both mutations selected from T 1069A and G1070A.
[0082] In some embodiments, the mutant Cas9 protein comprises a Keapl degron sequence that comprises the sequence EAAE (SEQ ID NO: 19). Insome embodiments, the mutant Cas9 protein can have a sequence of SEQ ID NO: 1 except that the amino acids at positions 1068 to 1071 of SEQ ID NO: 1 are replaced by the sequence EAAE (SEQ ID NO: 19). This mutant can also be referred to herein as “SpCas9-EAAE (SEQ ID NO: 19)”, where reference to SEQ ID NO: 19 refers to the four amino acid mutated sequence in the Keapl degron sequence of SpCas9, and not the full mutant protein sequence.
[0083] Also disclosed herein are mutant Cas9 proteins comprising one or more mutations in a Keapl degron sequence of Cas9, wherein the Cas9 is a Staphylococcus aureus Cas9 (SaCas9). The WT SaCas9 can have the sequence found in UNIPROT J7RUA5 and shown below:MKRNYILGLD IGITSVGYGI IDYETRDVIDAGVRLFKEAN VENNEGRRSK RGARRLKRRRRHRIQRVKKL LFDYNLLTDH SELSGINPYEARVKGLSQKL SEEEFSAALL HLAKRRGVHNVNEVEEDTGN ELSTKEQI SR NSKALEEKYVAELQLERLKK DGEVRGSINR FKTSDYVKEAKQLLKVQKAY HQLDQSFIDT YIDLLETRRTYYEGPGEGSP FGWKDIKEWY EMLMGHCTYFPEELRSVKYA YNADLYNALN DLNNLVITRDENEKLEYYEK FQI IENVFKQ KKKPTLKQIAKEILVNEEDI KGYRVTSTGK PEFTNLKVYHDIKDITARKE I IENAELLDQ IAKILTIYQSSEDIQEELTN LNSELTQEEI EQI SNLKGYTGTHNLSLKAI NLILDELWHT NDNQIAIFNRLKLVPKKVDL SQQKEIPTTL VDDFILSPVVKRSFIQSIKV INAI IKKYGL PNDI I IELAREKNSKDAQKM INEMQKRNRQ TNERIEEI IRTTGKENAKYL IEKIKLHDMQ EGKCLYSLEAIPLEDLLNNP FNYEVDHI IP RSVSFDNSFNNKVLVKQEEN SKKGNRTPFQ YLSSSDSKISYETFKKHILN LAKGKGRI SK TKKEYLLEERDINRFSVQKD FINRNLVDTR YATRGLMNLLRSYFRVNNLD VKVKSINGGF TSFLRRKWKFKKERNKGYKH HAEDALI IAN ADFIFKEWKKLDKAKKVMEN QMFEEKQAES MPEIETEQEYKEIFITPHQI KHIKDFKDYK YSHRVDKKPNRELINDTLYS TRKDDKGNTL IVNNLNGLYDKDNDKLKKLI NKSPEKLLMY HHDPQTYQKLKLIMEQYGDE KNPLYKYYEE TGNYLTKYSKKDNGPVIKKI KYYGNKLNAH LDITDDYPNSRNKVVKLSLK PYRFDVYLDN GVYKFVTVKNLDVIKKENYY EVNSKCYEEA KKLKKISNQAEFIASFYNND LIKINGELYR VIGVNNDLLNRIEVNMIDIT YREYLENMND KRPPRI IKTIASKTQSIKKY STDILGNLYE VKSKKHPQI I KKG ( SEQ ID NO : 2 ) .
[0084] In some embodiments, the Cas9 protein (i.e., the Cas9 protein that does not comprise the one or more mutations in the Keapl degron sequence) is a Cas9 having the sequence of SEQ ID NO: 2 or a sequence having at least about 90% homology to the sequence of SEQ ID NO: 2. In some embodiments, the sequence has at least about 91 % homology, at least about 92% homology, at least about 93% homology, at least about 94% homology, at least about 95% homlogy, at least about 96% homlogy, at least about 97% homlogy, at least about 98% homology, or at least about 99% homology to the sequence of SEQ ID NO: 2.
[0085] In some embodiments, the Keapl degron sequence is the ETGE (SEQ ID NO:18)-like sequence at positions 860 to 863 of SEQ ID NO: 2.
[0086] In some embodiments, the one or more mutations are at one or more of positions 860, 861 , and 863 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprise an alanine substitution at position 860 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprise a glutamic acid substiution at position 861 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprise an alanine substitution at position 863 of SEQ ID NO: 2. Stated another way, the mutant Cas9 protein can comprise the sequence of be SEQ ID NO: 2, i.e., the sequence of WT SaCas9, with one or more mutations selected from E860A, T861 E, and N863A. In some embodiments, the one or more substitutions comprise two or three substitutions selected from an alanine substitution at position 860 of SEQ ID NO: 2, a glutamic acid substitution at position 861 of SEQ ID NO: 2; and an alanine substitution at position 863 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprises all three mutations (i.e., all three of E860A, T861 E, and N863A).
[0087] In some embodiments, the mutant Cas9 comprises a Keapl degron sequence that comprises the sequence AEGA (SEQ ID NO: 20). In some embodiments, the mutant Cas9 protein can have a sequence of SEQ ID NO: 2 except that the amino acids at positions 860 to 863 of SEQ ID NO: 2 are replaced by AEGA (SEQ ID NO: 20). This mutant can also be referred to herein as “SaCas9-AEGA (SEQ ID NO: 20)”, where reference to SEQ ID NO: 20 refers to the four amino acid sequence mutation in the Keapl degron sequence of SaCas9, and not the full mutant protein sequence. In some aspects, the EAAE (SEQ ID NO: 19) mutations of SpCas9 are not as effective in the SaCas9 in increasing half-life and / or gene editing ability as AEGA (SEQ ID NO: 20).
[0088] In some aspects, the disclosed mutant Cas9 proteins can have one or more mutations outside of the Keapl degron sequence of the Cas9 protein, e.g., relatative to a wild-type (WT) Cas9, such as SpCas9 or SaCas9. For example, in some embodiments, the mutant Cas9 can be based on the sequence of another mutant Cas9 protein known in the art, e.g., a dCas9. Stated another way, the presently disclosed mutant Cas9 protein can comprise one or more mutations in a Keapl degron sequence relative to a“parent” Cas9 protein while the “parent” Cas9 protein can be a WT Cas9 protein or another mutant Cas9 protein that contains one or more (e.g., 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10) mutations relative to a WT Cas9 where said one or more mutations of the “parent” mutant Cas9 protein are mutations outside a Keapl degron sequence. In some embodiments, the parent Cas9 protein of the presently disclosed mutants is a dCas9 protein that contains one or more mutations in an endonucelase domain relative to its corresponding WT Cas9 protein. The one or more mutations in the endonucelase domain are mutations that do not affect the ability of the protein to bind a target DNA, but can remove its ability to cleave the target DNA. Thus, a dCas9 protein (or a mutant Cas9 protein of the presently disclosed subject matter that is based on a “parent” dCas9 protein) can be used in CRISPR interference (CRISPRi) and CRISPR activation (CRISPRa). In some embodiments, the presently disclosed mutant Cas9 protein is based on dSpCas9 (which comprises a D10A and a H840A mutation relative to WT SpCas9 (i.e., SEQ ID NO: 1 ). In some embodiments, the presently disclosed mutant Cas9 protein comprises one or more mutations in a Keapl degron sequence of a Cas9 protein and further comprises one or more mutations outside of the Keapl degron sequence of the Cas9 protein, wherein the Cas9 protein has a sequence of SEQ ID NO: 1 , where the one or more mutations outside of the Keapl degron sequence of the Cas9 protein comprise an alanine at position 10 of SEQ ID NO: 1 and an alanine at position 840 of SEQ ID NO: 1 , and wherein the one or more mutations in the Keapl degron sequence comprise an alanine substitution at position 1069 and / or at position 1070 of SEQ ID NO: 1 . In some embodiments, the mutant Cas9 protein has a sequence of SEQ ID NO: 1 except comprising an alanine at each of positions 10, 840, 1069, and 1070 (i.e., the mutant Cas9 protein of the presently disclosed subject matter can comprise the amino acid sequence of SEQ ID NO: 1 with the following mutations: D10A, H840A, T1069A, and G1070A).
[0089] The one or more mutations outside of the Keapl degron sequence can be conservative or nonconservative substitutions. In some embodiments, the one or more mutations outside the Keapl degron are mutations that do not alter one or more function of the mutant Cas9 protein. For example, oneor more mutations outside of the Keapl degron sequence would not negatively affect the mutant Cas9 protein’s ability to work in a CRISPR assay (e.g., a CRISPRi or CRISPRa assay).
[0090] The replacement of one amino acid residue with another that is biologically and / or chemically similar is known to those skilled in the art as a conservative substitution. For example, a conservative substitution would be replacing one hydrophobic residue for another, or one polar residue for another. The substitutions include combinations such as, for example, Gly, Ala; Vai, lie, Leu; Asp, Glu; Asn, Gin; Ser, Thr; Lys, Arg; and Phe, Tyr. Such conservatively substituted variations of each explicitly disclosed sequence are included within the mosaic polypeptides provided herein.
[0091] Substitutional or deletional mutagenesis can be employed to insert sites for N-glycosylation (Asn-X-Thr / Ser) or O-glycosylation (Ser or Thr). Deletions of cysteine or other labile residues also can be desirable. Deletions or substitutions of potential proteolysis sites, e.g. Arg, is accomplished for example by deleting one of the basic residues or substituting one by glutaminyl or histidyl residues.
[0092] Certain post-translational derivatizations are the result of the action of recombinant host cells on the expressed polypeptide. Glutaminyl and asparaginyl residues are frequently post-translationally deamidated to the corresponding glutamyl and asparyl residues. Alternatively, these residues are deamidated under mildly acidic conditions. Other post-translational modifications include hydroxylation of proline and lysine, phosphorylation of hydroxyl groups of seryl or threonyl residues, methylation of the o-amino groups of lysine, arginine, and histidine side chains (T. E. Creighton, Proteins: Structure and Molecular Properties, W.H. Freeman & Co., San Francisco pp 79-86
[1983] ), acetylation of the N-terminal amine and, in some instances, amidation of the C-terminal carboxyl.
[0093] It is understood that one way to define the variants and derivatives of the disclosed mutant Cas9 proteins herein is through defining the variants and derivatives in terms of homology / identity to specific known sequences. For example, the sequences of wild type Cas9s are known (see, for example, SEQ ID NOs: 1 and 2). Specifically disclosed are variants of these hereindisclosed which have at least, 70% or 75% or 80% or 85% or 90% or 95% homology to the full length or a fragment of wild type sequence. Wherein a sequence is said to have at least about 70% sequence identity, it is understood to also have at least about 75%, 80%, 85%, 90%, 92%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity.B. Nucleic Acids
[0094] In some embodiments, the presently disclosed subject matter provides a nucleic acid sequence comprising a sequence that encodes one or more of the mutant Cas9 proteins disclosed herein.
[0095] Thus, in some embodiments, the presently disclosed subject matter provides an isolated nucleic acid encoding a mutant Cas9 protein that comprises one or more mutations in the Keapl degron sequence of the Cas9 protein and which has a half-life greater than the half-life of the corresponding Cas9 protein that does not comprise the one or more mutations in the Keapl degron sequence. In some embodiments, the isolated nucleic acid encodes a mutant Cas9 protein that that comprises one or more mutations in a Keapl degron sequence of a Cas9 protein, wherein the Cas9 protein comprises or consists of the sequence of SEQ ID NO: 1 or SEQ ID NO: 2, or a sequence having at least 90%, 91 %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homology to the sequence of SEQ ID NO: 1 or SEQ ID NO: 2. In some embodiments, the one or more mutations provide a mutant Cas9 protein that comprises a Keapl degron sequence comprising the sequence EAAE (SEQ ID NO: 19) or AEGA (SEQ ID NO: 20).
[0096] In some embodiments, the isolated nucleic acid endodes a mutant Cas9 protein that comprises the sequence of SEQ ID NO: 1 or a protein having a sequence having at least 90% homology to SEQ ID NO: 1 , wherein the sequence of SEQ ID NO: 1 or the sequence having at least 90% homology thereto is mutated to comprise one or more mutations in the Keapl degron sequence. In some embodiments, the one or more mutations are at one or more of amino acids 1068 to 1071 of SEQ ID NO: 1 . In some embodiments, the one or more mutations are an alanine substitution at one or both of position 1069 and position 1070 of SEQ ID NO: 1 . In some embodiments, the one ormore mutations further comprise one or more mutations outside the Keapl degron sequence, e.g., an alanine substitution at position 10 of SEQ ID NO: 1 and an alanine substitution at position 840 of SEQ ID NO: 1 .
[0097] In some embodiments, the isolated nucleic acid endodes a mutant Cas9 protein that comprises the sequence of SEQ ID NO: 2 or a protein having a sequence having at least 90% homology to SEQ ID NO: 2, wherein the sequence of SEQ ID NO: 2 or the sequence having at least 90% homology thereto is mutated to comprise one or more mutations in the Keapl degron sequence. In some embodiments, the one or more mutations comprise an alanine substitution at position 860 of SEQ ID NO: 2, a glutamic acid substitution at position 861 of SEQ ID NO: 2, and / or an alanine substitution at position 863 of SEQ ID NO: 2. In some embodiments, the one or more mutations comprise alanine substitution at postion 860, a glutamic acid substitution at position 861 , and an alanine substitution at position 863 of SEQ ID NO: 2, such that the nucleic acid encodes a mutant Cas9 protein that comprises the sequence AEGA (SEQ ID NO: 20) in place of the wild-type Keapl degron sequence of SEQ ID NO: 2.
[0098] In some embodiments, disclosed herein is an isolated nucleic acid encoding a mutant Cas9 protein wherein the mutant Cas9 protein comprises a variant Cas9 protein comprising an amino acid sequence that has at least 95% sequence identity to the amino acid sequence of SEQ ID NO: 1 , wherein the amino acid corresponding to SEQ ID NO: 1 is mutated in the mutant Cas9 protein, and wherein the mutant Cas9 protein binds to a guide RNA and a target DNA. In some aspects, the sequence at amino acids 1068 to 1071 of SEQ ID NO: 1 is mutated. In some aspects, the mutated sequence at amino acids 1068 to 1071 is EAAE (SEQ ID NO: 19) (e.g., in place of the ETGE (SEQ ID NO: 18) at positions 1068 to 1071 of SEQ ID NO: 1 ).
[0099] Also disclosed are isolated nucleic acid encoding a mutant Cas9 protein, wherein the mutant Cas9 protein comprises a variant Cas9 protein comprising an amino acid sequence that has at least 95% sequence identity to the amino acid sequence of SEQ ID NO: 2, wherein the amino acid corresponding to SEQ ID NO: 2 is mutated in the mutant Cas9 protein, and wherein the mutant Cas9 protein binds to a guide RNA and a target DNA. Insome aspects, the amino acid sequence at 860 to 863 of SEQ ID NO: 2 is mutated. In some aspects, the one or more mutations are at amino acids 860 to 863 of SEQ ID NO: 2. In some aspects, mutated sequence at amino acids 860 to 863 is AEGA (SEQ ID NO: 20).
[0100] Also disclosed are constructs comprising one or more of the nucleic acid sequences disclosed herein. The term “vector” can be used interchangeably with “construct.”
[0101] In some aspects, disclosed are nucleic acid sequences and constructs comprising a sequence that encodes one or more of the mutant Cas9 proteins disclosed herein operably linked to a promoter.
[0102] In some aspects, disclosed are nucleic acid sequences and constructs comprising a sequence that encodes one or more of the mutant Cas9 proteins disclosed herein and further comprising a sequence encoding a marker gene. A marker gene can be any gene that encodes a protein that can be used for detection or purification.C. Compositions Comprising Nucleic Acid Sequences or Vectors
[0103] Also disclosed are composition comprising the nucleic acid sequences or vectors described herein. For example, disclosed are compositions comprising vectors, wherein the vectors comprise any of the mutated Cas9 proteins or nucleic acid sequences disclosed herein.
[0104] In some aspects, the compositions further comprise a pharmaceutically acceptable carrier. Phamaceutically acceptable carriers include, for example, water, saline, buffered aqueous solutions (e.g., phosphate or citrate buffered solutions). In some embodiments, the carrier can include a carrier suitable for depot of an isolated nucleic acid or vector, such as a cationic or polycationic carrier (e.g., protamine, spermidine, poly-L- lysine, poly-arginine, etc.). Carriers can also include lipids and lipid-based nanoparticles.Viral and Non-Viral Vectors
[0105] There are a number of compositions and methods which can be used to deliver nucleic acids to cells, either in vitro or in vivo. These methodsand compositions can largely be broken down into two classes: viral based delivery systems and non-viral based delivery systems. For example, the nucleic acids can be delivered through a number of direct delivery systems such as, electroporation, lipofection, calcium phosphate precipitation, plasmids, viral vectors, viral nucleic acids, phage nucleic acids, phages, cosmids, or via transfer of genetic material in cells or carriers such as cationic liposomes. Appropriate means for transfection, including viral vectors, chemical transfectants, or physico-mechanical methods such as electroporation and direct diffusion of DNA, are described by, for example, Wolff, J. A., et al., Science, 247, 1465-1468, (1990); and Wolff, J. A. Nature, 352, 815-818, (1991 ). Such methods are well known in the art and readily adaptable for use with the compositions and methods described herein. In certain cases, the methods will be modified to specifically function with large DNA molecules. Further, these methods can be used to target certain diseases and cell populations by using the targeting characteristics of the carrier.
[0106] Expression vectors can be any nucleotide construction used to deliver genes or gene fragments into cells (e.g., a plasmid), or as part of a general strategy to deliver genes or gene fragments, e.g., as part of recombinant retrovirus or adenovirus (Ram et al. Cancer Res. 53:83-88, (1993)). For example, disclosed herein are expression vectors comprising a nucleic acid sequence capable of encoding one or more of the disclosed mutated Cas9 proteins operably linked to a control element.
[0107] The “control elements” present in an expression vector are those non-translated regions of the vector-enhancers, promoters, 5’ and 3’ untranslated regions-which interact with host cellular proteins to carry out transcription and translation. Such elements can vary in their strength and specificity. Depending on the vector system and host utilized, any number of suitable transcription and translation elements, including constitutive and inducible promoters, can be used. For example, when cloning in bacterial systems, inducible promoters such as the hybrid lacZ promoter of the pBLUESCRIPT phagemid (Stratagene, La Jolla, California, United States of America) or pSPORTI plasmid (Gibco BRL, Gaithersburg, Maryland, UnitedStates of America) and the like can be used. In mammalian cell systems, promoters from mammalian genes or from mammalian viruses can be used. If it is necessary to generate a cell line that contains multiple copies of the sequence encoding a polypeptide, vectors based on SV40 or EBV can be advantageously used with an appropriate selectable marker.
[0108] Promoters for controlling transcription from vectors in mammalian host cells can be obtained from various sources, for example, the genomes of viruses such as polyoma, Simian Virus 40 (SV40), adenovirus, retroviruses, hepatitis-B virus and cytomegalovirus, or from heterologous mammalian promoters (e.g., beta actin promoter). The early and late promoters of the SV40 virus can be obtained as an SV40 restriction fragment, which also contains the SV40 viral origin of replication (Fiers et aL, Nature, 273: 113 (1978)). The immediate early promoter of the human cytomegalovirus can be obtained as a Hindlll E restriction fragment (Greenway, P.J. et al., Gene 18: 355-360 (1982)). Additionally, promoters from the host cell or related species can also be used.
[0109] The term “enhancer” generally refers to a sequence of DNA that functions at no fixed distance from the transcription start site and can be either 5’ (Laimins, L. et al., Proc. Natl. Acad. Sci. 78: 993 (1981 )) or 3’ (Lusky, M.L., et al., Mol. Cell Bio. 3: 1108 (1983)) to the transcription unit. Furthermore, enhancers can be within an intron (Banerji, J.L. et al., Cell 33: 729 (1983)) as well as within the coding sequence itself (Osborne, T.F., et al., Mol. Cell Bio. 4: 1293 (1984)). They are usually between 10 and 300 bp in length, and they function in cis. Enhancers function to increase transcription from nearby promoters. Enhancers also often contain response elements that mediate the regulation of transcription. Promoters can also contain response elements that mediate the regulation of transcription. Enhancers often determine the regulation of expression of a gene. While many enhancer sequences are now known from mammalian genes (globin, elastase, albumin, a-fetoprotein and insulin), typically an enhancer from a eukaryotic cell virus can be used for general expression. Non-limiting examples include the SV40 enhancer on the late side of the replication origin (bp 100-270), the cytomegalovirus earlypromoter enhancer, the polyoma enhancer on the late side of the replication origin, and adenovirus enhancers.
[0110] The promoter or enhancer can be specifically activated either by light or specific chemical events which trigger their function. Systems can be regulated by reagents such as tetracycline and dexamethasone. There are also ways to enhance viral vector gene expression by exposure to irradiation, such as gamma irradiation, or alkylating chemotherapy drugs.
[0111] Optionally, the promoter or enhancer region can act as a constitutive promoter or enhancer to maximize expression of the polynucleotides of the invention. In certain constructs, the promoter or enhancer region can be active in all eukaryotic cell types, even if it is only expressed in a particular type of cell at a particular time. An exemplary promoter of this type is the CMV promoter (650 bases). Other exemplary promoters include, but are not limited to, SV40 promoters, cytomegalovirus (full length promoter), and retroviral vector LTR.
[0112] Expression vectors used in eukaryotic host cells (yeast, fungi, insect, plant, animal, human or nucleated cells) can also contain sequences for the termination of transcription which can affect mRNA expression. These regions are transcribed as polyadenylated segments in the untranslated portion of the mRNA encoding tissue factor protein. The 3’ untranslated regions also include transcription termination sites. In some embodiments, the transcription unit also contains a polyadenylation region. One benefit of this region is that it increases the likelihood that the transcribed unit will be processed and transported like mRNA. The identification and use of polyadenylation signals in expression constructs is well established. In some embodiments, homologous polyadenylation signals can be used in the transgene constructs. In certain transcription units, the polyadenylation region is derived from the SV40 early polyadenylation signal and consists of about 400 bases.
[0113] The expression vectors can include a nucleic acid sequence encoding a marker product. This marker product is used to determine if the gene has been delivered to the cell and once delivered is being expressed. Exemplary marker genes include the E. coli lacZ gene, which encodes B- galactosidase, and the gene encoding the green fluorescent protein.
[0114] In some embodiments, the marker can be a selectable marker. Examples of suitable selectable markers for mammalian cells include, for instance, dihydrofolate reductase (DHFR), thymidine kinase, neomycin, neomycin analog G418, hydromycin, and puromycin. When such selectable markers are successfully transferred into a mammalian host cell, the transformed mammalian host cell can survive if placed under selective pressure. There are two widely used distinct categories of selective regimes. The first category is based on a cell’s metabolism and the use of a mutant cell line which lacks the ability to grow independent of a supplemented media. Two examples are CHO DHFR-cells and mouse LTK-cells. These cells lack the ability to grow without the addition of such nutrients as thymidine or hypoxanthine. Because these cells lack certain genes necessary for a complete nucleotide synthesis pathway, they cannot survive unless the missing nucleotides are provided in a supplemented media. An alternative to supplementing the media is to introduce an intact DHFR or TK gene into cells lacking the respective genes, thus altering their growth requirements. Individual cells which were not transformed with the DHFR or TK gene will not be capable of survival in non-supplemented media.
[0115] The second category is dominant selection which refers to a selection scheme used in any cell type and does not require the use of a mutant cell line. These schemes typically use a drug to arrest growth of a host cell. Those cells which have a novel gene would express a protein conveying drug resistance and would survive the selection. Examples of such dominant selection use the drugs neomycin, (Southern P. and Berg, P., J. Molec. AppL Genet. 1 : 327 (1982)), mycophenolic acid, (Mulligan, R.C. and Berg, P. Science 209: 1422 (1980)) or hygromycin, (Sugden, B. et al., Mol. Cell. Biol. 5: 410-413 (1985)). The three examples employ bacterial genes under eukaryotic control to convey resistance to the appropriate drug G418 or neomycin (geneticin), xgpt (mycophenolic acid) or hygromycin, respectively. Others include the neomycin analog G418 and puramycin.
[0116] As used herein, the terms “plasmid” and “viral vector” refer to agents that transport the disclosed nucleic acids, such as a nucleic acid sequence capable of encoding one or more of the disclosed mutant peptides into the cellwithout degradation and include a promoter yielding expression of the gene in the cells into which it is delivered. In some embodiments the nucleic acid sequences disclosed herein are derived from either a virus or a retrovirus. Viral vectors include, for example, Adenovirus, Adeno-associated virus, Herpes virus, Vaccinia virus, Polio virus, AIDS virus, neuronal trophic virus, Sindbis and other RNA viruses, including these viruses with the HIV backbone. Other examples include any viral families which share the properties of these viruses which make them suitable for use as vectors. Retroviruses include Murine Maloney Leukemia virus, MMLV, and retroviruses that express the desirable properties of MMLV as a vector. Retroviral vectors are able to carry a larger genetic payload, i.e., a transgene or marker gene, than other viral vectors, and for this reason are a commonly used vector. However, in some embodiments, they are not as useful in non-proliferating cells. Adenovirus vectors are relatively stable and easy to work with, have high titers, can be delivered in aerosol formulation, and can transfect non-dividing cells. Pox viral vectors are large and have several sites for inserting genes, they are thermostable and can be stored at room temperature. In some embodiments, the viral vector is a viral vector which has been engineered so as to suppress the immune response of the host organism, elicited by the viral antigens. Exemplary vectors of this type can carry coding regions for Interleukin 8 or 10.
[0117] In some embodiments, viral vectors can have higher transaction abilities (i.e., ability to introduce genes) than chemical or physical methods of introducing genes into cells. Typically, viral vectors contain, nonstructural early genes, structural late genes, an RNA polymerase III transcript, inverted terminal repeats necessary for replication and encapsidation, and promoters to control the transcription and replication of the viral genome. When engineered as vectors, viruses typically have one or more of the early genes removed and a gene or gene / promoter cassette is inserted into the viral genome in place of the removed viral DNA. Constructs of this type can carry up to about 8 kb of foreign genetic material. The necessary functions of the removed early genes are typically supplied by cell lines which have been engineered to express the gene products of the early genes in trans.
[0118] Retroviral vectors, in general, are described by Verma, I.M., (Retroviral vectors for gene transfer. In Microbiology, Amer. Soc. for Microbiology, pp. 229-232, Washington, (1985)), which is hereby incorporated by reference in its entirety. Examples of methods for using retroviral vectors for gene therapy are described in U.S. Patent Nos. 4,868,116 and 4,980,286; PCT applications WO 90 / 02806 and WO 89 / 07136; and Mulligan, (Science 260:926-932 (1993)); the teachings of which are incorporated herein by reference in their entirety for their teaching of methods for using retroviral vectors for gene therapy.
[0119] A retrovirus is essentially a package which has packed into it nucleic acid cargo. The nucleic acid cargo carries with it a packaging signal, which ensures that the replicated daughter molecules will be efficiently packaged within the package coat. In addition to the package signal, there are a number of molecules which are needed in cis, for the replication, and packaging of the replicated virus. Typically a retroviral genome contains the gag, pol, and env genes which are involved in the making of the protein coat. It is the gag, pol, and env genes which are typically replaced by the foreign DNA that it is to be transferred to the target cell. Retrovirus vectors typically contain a packaging signal for incorporation into the package coat, a sequence which signals the start of the gag transcription unit, elements needed for reverse transcription, including a primer binding site to bind the tRNA primer of reverse transcription, terminal repeat sequences that guide the switch of RNA strands during DNA synthesis, a purine rich sequence 5' to the 3' LTR that serves as the priming site for the synthesis of the second strand of DNA synthesis, and specific sequences near the ends of the LTRs that provide for the insertion of the DNA state of the retrovirus to insert into the host genome. This amount of nucleic acid is sufficient for the delivery of one or more genes depending on the size of each transcript. In some embodiments, either positive or negative selectable markers are included along with other genes in the insert.
[0120] Since the replication machinery and packaging proteins in most retroviral vectors have been removed (gag, pol, and env), the vectors are typically generated by placing them into a packaging cell line. A packaging cell line is a cell line which has been transfected or transformed with aretrovirus that contains the replication and packaging machinery but lacks any packaging signal. When the vector carrying the DNA of choice is transfected into these cell lines, the vector containing the gene of interest is replicated and packaged into new retroviral particles, by the machinery provided in cis by the helper cell. The genomes for the machinery are not packaged because they lack the necessary signals.
[0121] The construction of replication-defective adenoviruses has been described (Berkner etal., J. Virology 61 :1213-1220 (1987); Massie et al., Mol. Cell. Biol. 6:2872-2883 (1986); Haj-Ahmad ef al., J. Virology 57:267-274 (1986); Davidson et al., J. Virology 61 :1226-1239 (1987); Zhang, “Generation and identification of recombinant adenovirus by liposome-mediated transfection and PCR analysis” BioTechniques 15:868-872 (1993)). The benefit of the use of these viruses as vectors is that they are limited in the extent to which they can spread to other cell types, since they can replicate within an initial infected cell but are unable to form new infectious viral particles. Recombinant adenoviruses have been shown to achieve high efficiency gene transfer after direct, in vivo delivery to airway epithelium, hepatocytes, vascular endothelium, CNS parenchyma and / or a number of other tissue sites (see Morsy, J. Clin. Invest. 92:1580-1586 (1993); Kirshenbaum, J. Clin. Invest. 92:381 -387 (1993); Roessler, J. Clin. Invest. 92:1085-1092 (1993); Moullier, Nature Genetics 4:154-159 (1993); La Salle, Science 259:988-990 (1993); Gomez-Foix, J. Biol. Chem. 267:25129-25134 (1992); Rich, Human Gene Therapy 4:461 -476 (1993); Zabner, Nature Genetics 6:75-83 (1994); Guzman, Circulation Research 73:1201 -1207 (1993); Bout, Human Gene Therapy 5:3-10 (1994); Zabner, Cell 75:207-216 (1993); Caillaud, Eur. J. Neuroscience 5:1287-1291 (1993); and Ragot, J. Gen. Virology 74:501 -507 (1993)) the teachings of which are incorporated herein by reference in their entirety for their teaching of methods for using retroviral vectors for gene therapy). Recombinant adenoviruses achieve gene transduction by binding to specific cell surface receptors, after which the virus is internalized by receptor-mediated endocytosis, in the same manner as wild type or replication-defective adenovirus (Chardonnet and Dales, Virology 40:462-477 (1970); Brown and Burlingham, J. Virology 12:386-396 (1973);Svensson and Persson, J. Virology 55:442-449 (1985); Seth, et al., J. Virol. 51 :650-655 (1984); Seth, et al., Mol. Cell. Biol., 4:1528-1533 (1984); Varga et a!., J. Virology 65:6061 -6070 (1991 ); Wickham eta!., Cell 73:309-319 (1993)).
[0122] A viral vector can be one based on an adenovirus which has had the E1 gene removed and these virons are generated in a cell line such as the human 293 cell line. Optionally, both the E1 and E3 genes are removed from the adenovirus genome.
[0123] Another type of viral vector that can be used to introduce the polynucleotides of the presently disclosed subject matter into a cell is based on an adeno-associated virus (AAV). This defective parvovirus can infect many cell types and is nonpathogenic to humans. AAV type vectors can transport about 4 to 5 kb and wild type AAV is known to stably insert into chromosome 19. An exemplary embodiment of this type of vector is the P4.1 C vector produced by Avigen (San Francisco, California, United States of America) which can contain the herpes simplex virus thymidine kinase gene, HSV-tk, or a marker gene, such as the gene encoding the green fluorescent protein, GFP.
[0124] In another type of AAV virus, the AAV contains a pair of inverted terminal repeats (ITRs) which flank at least one cassette containing a promoter which directs cell-specific expression operably linked to a heterologous gene. “Heterologous” in this context refers to any nucleotide sequence or gene which is not native to the AAV or B19 parvovirus. Typically the AAV and B19 coding regions have been deleted, resulting in a safe, noncytotoxic vector. The AAV ITRs, or modifications thereof, confer infectivity and site-specific integration, but not cytotoxicity, and the promoter directs cellspecific expression. United States Patent No. 6,261 ,834 is herein incorporated by reference in its entirety for material related to the AAV vector.
[0125] The inserted genes in viral and retroviral vectors usually contain promoters, or enhancers to help control the expression of the desired gene product. A promoter is generally a sequence or sequences of DNA that function when in a relatively fixed location in regard to the transcription start site. A promoter contains core elements required for basic interaction of RNApolymerase and transcription factors, and can contain upstream elements and response elements.
[0126] Other useful systems include, for example, replicating and host- restricted non-replicating vaccinia virus vectors. In addition, the disclosed nucleic acid sequences can be delivered to a target cell in a non-nucleic acid based system. For example, the disclosed polynucleotides can be delivered through electroporation, or through lipofection, or through calcium phosphate precipitation. The delivery mechanism chosen will depend in part on the type of cell targeted and whether the delivery is occurring for example in vivo or in vitro.
[0127] Thus, the compositions can comprise, in addition to the disclosed expression vectors, lipids such as liposomes, such as cationic liposomes (e.g., DOTMA, DOPE, DC-cholesterol) or anionic liposomes. Liposomes can further comprise proteins to facilitate targeting a particular cell, if desired. Administration of a composition comprising a peptide and a cationic liposome can be administered to the blood, to a target organ, or inhaled into the respiratory tract to target cells of the respiratory tract. For example, a composition comprising a peptide or nucleic acid sequence described herein and a cationic liposome can be administered to a subjects lung cells. Regarding liposomes, see, e.g., Brigham et al. (Am. J. Resp. Cell. Mol. Biol. 1 :95-100 (1989)); Feigner et al. (Proc. Natl. Acad. Sci USA 84:7413-7417 (1987)); and U.S. Patent No. 4,897,355. Furthermore, the compound can be administered as a component of a microcapsule that can be targeted to specific cell types, such as macrophages, or where the diffusion of the compound or delivery of the compound from the microcapsule is designed for a specific rate or dosage.D. Cells
[0128] Disclosed are cell lines comprising any of the disclosed nucleic acid sequences, mutant Cas9 proteins, or constructs. The disclosed cell lines can be used, for example, for CRISPR experiments. The cells can be bacterial (e.g., E.coli) or mammalian (e.g., human) cells.
[0129] In some aspects, the cell lines can have any of the disclosed nucleic acid sequences integrated into their genome. In some aspects, the cell lines can express one or more of the disclosed mutant Cas9 proteins (e.g., SpCas9 (SEQ ID NO:1 ) comprising one or more mutations selected from T1069A and G1070A, or SaCas9 (SEQ ID NO:2) comprising one or more mutations selected from E860A, T861 E, and N863A) constitutively or under a regulatable expression system.E. CRISPR-Cas9 System
[0130] In some embodiments, the presently disclosed subject matter provides a CRISPR / Cas9 system for use in editing a gene. The system can comprise (i) a guide RNA (gRNA); and (ii) a mutant Cas9 protein as disclosed herein (e.g., SpCas9 (SEQ ID NO:1 ) comprising one or more mutations selected from T1069A and G1070A, or SaCas9 (SEQ ID NO:2) comprising one or more mutations selected from E860A, T861 E, and N863A). In some embodiments, the gRNA is a sgRNA.III. MethodsA. Method of Enhancing CRISPR
[0131] The methods and compositions disclosed herein can utilize CRISPR / CRISPR-associated (Gas) systems or components of such systems to modify a genome within a cell. CRISPR / Cas systems include transcripts and other elements involved in the expression of, or directing the activity of, Gas genes. The methods and compositions disclosed herein employ CRISPR / Cas systems by utilizing CRISPR complexes (comprising a guide RNA (gRNA) complexed with a Cas protein) for site-directed cleavage of nucleic acids. Accordingly, in some embodiments, the presently disclosed subject matter provides a method of editing a gene, wherein the method comprises introducing a mutant Cas9 protein into a cell or tissue and introdugin one or more guide RNAs into said cell or tissue. For example, the method disclosed herein can use the disclosed mutant Cas9 proteins to enhance CRISPR systems.
[0132] Thus, in some aspects, disclosed are methods of enhancing CRISPR efficiency comprising administering one or more gRNAs to a cell comprising any of the mutant Cas9 proteins disclosed herein. Guide RNAs program Cas9 nucleases to cut at a specific genomic location. The design of an effective guide RNA is a component to achieving efficient gene knockout via CRISPR.
[0133] In some aspects, the enhanced CRISPR efficiency is enhanced CRISPR knock out efficiency. In some aspects, the enhanced CRISPR efficiency is enhanced CRISPR knock in efficiency.
[0134] In some aspects, the mutant Cas9 half-life can be increased compared to wild type Cas9 or the Cas9 protein that does not comprises a mutant Keapl degron sequence. Wild type Cas9 can be any known wild type Cas9 sequence, for example, SpCas9 (SEQ ID NO:1 ) or SaCas9 (SEQ ID NO:2). The mutant Gas 9 can comprise SaCas9 comprising one or both of the mutations T1069A and G1070A or SaCas9 comprising one, two or three of the mutations selected from the group comprising E860A, T861 E, and N863A.
[0135] In some aspects, the cell has decreased Keapl activity. The decreased Keapl activity can be because the wild type Cas9 protein present in the cells binds and suppresses Keapl from targeting its bona fide substrates for degradation.
[0136] In some aspects, the cell has a decreased expression or activity of Keapl . The decreased expression or activity of Keapl can be because the cells are genetically modified to have decreased Keapl expression or because they can be treated with an agent that reduces Keapl affinity with its substrates.
[0137] Also disclosed are methods of enhancing CRISPR efficiency comprising administering one or more gRNAs to a cell comprising a Keapl mutant, wherein the Keapl mutant is unable to promote degradation of Cas9. In some aspects, the cell can further comprise any one of the disclosed mutant Cas9 proteins.
[0138] In some aspects, the Keapl mutant can be wild type Keapl having the mutation of G333C. In some aspects, wild type Keapl can have the sequence found in UNIPROT Q14145 and shown below:MQPDPRPSGA GACCRFLPLQ SQCPEGAGDAVMYASTECKA EVTPSQHGNR TFSYTLEDHTKQAFGIMNEL RLSQQLCDVT LQVKYQDAPAAQFMAHKVVL ASSSPVFKAM FTNGLREQGMEVVSIEGIHP KVMERLIEFA YTASISMGEKCVLHVMNGAV MYQIDSVVRA CSDFLVQQLDPSNAIGIANF AEQIGCVELH QRAREYIYMHFGEVAKQEEF FNLSHCQLVT LISRDDLNVRCESEVFHACI NWVKYDCEQR RFYVQALLRAVRCHSLTPNF LQMQLQKCEI LQSDSRCKDYLVKIFEELTL HKPTQVMPCR APKVGRLIYTAGGYFRQSLS YLEAYNPSDG TWLRLADLQVPRSGLAGCVV GGLLYAVGGR NNSPDGNTDSSALDCYNPMT NQWSPCAPMS VPRNRIGVGVIDGHIYAVGG SHGCIHHNSV ERYEPERDEWHLVAPMLTRR IGVGVAVLNR LLYAVGGFDGTNRLNSAECY YPERNEWRMI TAMNTIRSGAGVCVLHNCIY AAGGYDGQDQ LNSVERYDVETETWTFVAPM KHRRSALGIT VHQGRIYVLGGYDGHTFLDS VECYDPDTDT WSEVTRMTSGRSGVGVAVTM EPCRKQIDQQ NCTC ( SEQ IDNO : 3 ) .B. Method of Suppressing Keapl Activity
[0139] Disclosed are methods of suppressing Keapl activity comprising administering a protein or nucleic acid that blocks the Keapl - Kelch domain responsible for Keapl binding to its substrates.
[0140] In some aspects, a protein that blocks the Keapl - Kelch domain is an antibody or chemical compound specific to the Kelch domain. For example, see e.g. Jiang et al. J Enzyme Inhib Med Chem. 2018 Dec;33(1 ):833-841 and Abed et al. Acta Pharma Sinica B. 2015 July;5(4):285- 299. In some aspects, a nucleic acid that blocks the Keapl - Kelch domain is a nucleic acid that mimics a Kelch domain substrate. For example, see e.g. Guntas et al. Protein Eng Des Sei. 2016 Jan; 29(1 ): 1-9.
[0141] In some aspects, a Kelch domain substrate is a Keapl degron sequence of Cas9. In some aspects, the suppressed Keapl activity is the ability to ubiquitinate Cas9. Thus, in some aspects, the half-life of Cas9 can be increased.C. Method of Promoting Cell Growth
[0142] Disclosed are methods of promoting cell growth comprising administering an inhibitor of Keapl and NRF2 binding. In some aspects, the inhibitor can be Cas9. The presence of Cas9 can allow Keapl to bind Cas9 which frees up NRF2. In some aspects, NRF2 inhibitors can be used. For example, an anti-NRF2 antibody can be used as an NRF2 inhibitor. The use of NRF2 inhibitors can overcome Cas9 expression induced cellular behavior changes in certain settings.IV. Kits
[0143] The materials described above as well as other materials can be packaged together in any suitable combination as a kit useful for performing, or aiding in the performance of, the disclosed method. It is useful if the kit components in a given kit are designed and adapted for use together in the disclosed method. For example disclosed are kits for enhancing CRISPR efficiency, the kit comprising one or more of the mutated Cas9 proteins disclosed herein or the nucleic acid sequences or constructs that encode oneor more of the mutated Cas9 proteins disclosed herein. The kits also can contain cells that comprise a nucleic acid sequence that encodes one or more of the mutated Cas9 proteins disclosed herein or cell lines that express one or more of the mutated Cas9 proteins disclosed herein.
[0144] While some embodiments of the disclosure have been described by way of illustration, it will be apparent that the disclosure can be put into practice with many modifications, variations and adaptations, and with the use of numerous equivalents or alternative solutions that are within the scope of persons skilled in the art, without departing from the spirit of the disclosure or exceeding the scope of the claims.
[0145] All publications, patents, and patent applications are herein incorporated by reference in their entirety to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference in its entirety.Examples
[0146] The examples presented herein represent certain embodiments of the present disclosure. However, it is to be understood that these examples are for illustration purposes only and do not intend, nor should any be construed, to be wholly definitive as to conditions and scope of this disclosure. The examples were carried out using standard techniques, which are known and routine to those of skill in the art, except where otherwise described in detail.EXAMPLE 1Keapl is an endogenous spCas9 binding protein
[0147] Mammalian-specific Cas9 regulation can provide guidance to engineering Cas9 variants with enhanced gene modulating ability, including improved robustness and / or fidelity. This knowledge can also help in understanding how this bacterial protein can interfere with mammalian signaling networks. Cas9 has been found to undergo SUMOylation in mammalian cells, which stabilizes Cas9 proteins and enhances DNA binding (49). A close examination of the SpCas9 protein sequence (SEQ ID NO: 1 )led to the discovery of a canonical E3 ligase Keapl degron sequence ETGE (SEQ ID NO: 18), which corresponds to amino acids 1068 to 1071 of SEQ ID NO: 1 ), located in its RuvC III domain, needed for SpCas9 enzyme activity (Figure 1A). Interestingly, a similar motif is present in most Cas9 species, CasX, Cas12, Cas13a and even the eukaryotic enzyme Fanzor (26). See Table 1 , below, which lists several Cas9 and other Cas or Cas-like proteins (with Uniprot numbers and SEQ ID Nos). The potential KEAP1 binding motif of each protein (with flanking amino acid pairs), along with its domain location within the protein are also provided. The amino acid residues shown in bold font in the subsequences shown in the “KEAP1 binding motif” column represent the potential 4 amino acid binding motif within each protein, while the underlined amino acid residues in this potential binding motif indicate an alternative “mutant” amino acid residue in comparison to canonical E3 degron the KEAP1 binding motif of SpCas9.Table 1. Cas9 and other proteins with the SpCas9 E3 ligase Keapl degron sequence or a similar sequence.
[0148] Mutating the Keapl degron of SpCas9 led to reduced SpCas9 binding to Keapl in human cells (Figures 1B, 1C), suggesting that Keapl is largely reliant upon the degron sequence for SpCas9 binding. Notably, co- expression of sgRNAs did not affect SpCas9 binding to Keapl (Figure 7A), nor SpCas9 protein abundance (Figure 7B). That the native ETGE (amino acids 1066 to 1073 of SEQ ID NO:1 ) degron sequence is the key motif in SpCas9 for Keapl recognition is further supported by a structure simulationdemonstrating that structurally the ETGE (SEQ ID NO: 18) motif in SpCas9 resembles the conformation of the ETGE (SEQ ID NO: 18) domain in a well- characterized Keapl substrate NRF2 (50) (Figure 1 D). The room mean square deviation (RMSD) calculated between the SpCas9 ETGEIV (amino acids 1068 to 1073 of SEQ ID NO: 1 ) and the NRF2 ETGEFL (amino acids 79 to 84 of SEQ ID NO: 16) sequence was 1.078 A, which further confirms the structural similarity between the two motifs. Further, the SpCas9 interactome in HEK293T cells was established by proteomics (Figure 1 E, 1F) and unique Keapl peptides were found (Figure 1G), along with other E3-related proteins (see Table 2, below) as potential SpCas9 interacting partners. Together, these data indicate Keapl is an endogenous mammalian protein interacting with SpCas9.
[0149] Table 2. E3 or E3-related Proteins from Flag-SpCas9 interactome from HEK293T cells.EXAMPLE 2 Keapl targets SpCas9 for ubiquitination and degradation
[0150] Given that Keapl binds and earmarks protein substrates for destruction, such as NRF2 (nuclear factor erythroid-related factor 2) (51 ), or binds and suppresses its binding partner’s function, such as PALB2 (partner and localizer of BRCA2) (52), whether and how Keapl controls SpCas9 function was examined. It was found that SpCas9 bound the Keapl -Kelch domain (Figure 8A and Figure 8B) responsible for Keapl binding to its substrates (52). As a result, Keapl targeted SpCas9 for degradation in a Keapl dose-dependent (Figure 2A and Figure 8C) and SpGas9 “ETGE” (SEQ ID NO: 18) in a degron-dependent manner in cells (Figure 2B and Figure 8D). More importantly, depletion of endogenous Keapl led to accumulated WT-SpCas9, but not EAAE (SEQ ID NO: 19)-SpCas9 proteins in cells (Figure 2C, Figure 2D, and Figure 8E). Interestingly, increased SpCas9 expression upon Keapl depletion was observed even at conditions where the amount of SpCas9 used was significantly greater than amounts typically suggested for gene editing (54) (Figure 2E). Keapl depletion did not significantly affect mRNA levels of these transfected SpCas9s in cells (Figures 8F and 8G). Together, these data suggest that depletion of endogenous Keapl is likely to cause SpCas9 protein accumulation. Notably, re-introducingan shKeapI -resistant version of WT-Keap1 reduced expression of SpCas9 (Figure 2F), alleviaing concerns for off-target effects derived from shKeapI . To further support a role of proteasome-mediated SpCas9 protein stability control, it was found that blocking the 26S proteasome by either MG132 or bortezomib increased SpCas9 protein levels in a Cas9 “ETGE” (SEQ ID NO:18) degron-dependent manner (Figure 2G and Figures 8G and 8H). Given that poly-ubiquitination occurs via eight distinctive ubiquitin linkages that determine the fate of poly-ubiquitinated proteins (55), SpCas9 was mainly modified by K48-linked poly-ubiquitination in cells (Figure 8I) and Keapl robustly promoted K48-linked ubiquitination of SpCas9 (Figure 2H) to induce its degradation. As such, ectopic Keapl expression significantly shortened WT-SpCas9 protein half-life (Figure 8J), while SpCas9-EAAE (SEQ ID NO:19) displayed an extended protein half-life (Figure 8K), presumably due to its inability to be recognized and degraded by Keapl .
[0151] In addition to Keapl knockdown, inactivating Keapl by pharmacological inhibitors including Keapl cysteine modifying agents tertbutyl hydroxyquinone (tBHQ) and sulforaphane, or chemicals blocking Keapl binding to cullin 3 such as ODDO (2-cyano-3,12-dioxooleana-1 ,9(1 1 )-dien-28- oic acid), similarly led to accumulation of WT, but not EAAE (SEQ ID NO: 19)- SpCas9 proteins in cells (Figure 21 and Figures 8L and 8M). Moreover, tBHQ treatment blocked Keapl -mediated WT-SpCas9 degradation (Figure 2J and Figure 8N) presumably by attenuating Keapl binding to SpCas9 (Figure 80). These data further support the notion that SpCas9 is a Keapl substrate.EXAMPLE 3Keapl targets Cas9s from other bacterial species and Fanzorl for degradation
[0152] Given that the “ETGE” (SEQ ID NO: 18) motif is present in multiple species of Cas9s (see Table 5, above) used for gene editing, it was determined whether protein stability of other Cas9s is also regulated by Keapl . To this end, it was found that, similar to WT-SpCas9, versions of SpCas9s engineered for higher fidelity (including HF1 -SpCas9 (28) and eSpCas9 (27))were also subjected to Keapl -mediated degradation (Figure 9A). Moreover, S. aureus Cas9 (SaCas9) with an “ETGN” (SEQ ID NO: 22, which also corresponds to amino acids 860 to 863 of SEQ ID NO: 2) motif (Figure 3A), that has been widely used in in vivo CRISPR-mediated gene editing due to its smaller size than SpCas9 for adenovirus packaging, could be targeted by Keapl for degradation as well (Figure 3B). Mutating the first conserved Keapl degron “ETGN” (SEQ ID NO: 18) to either “DSGD” (SEQ ID NO: 21 ), “AEGA” (SEQ ID NO: 20) or “EAAE” (SEQ ID NO: 19) efficiently reduced its binding to Keapl (Figure 3C). Notably, both DSGD (SEQ ID NO: 21 )- and EAAE (SEQ ID NO: 19)-SaCas9 mutants displayed reduced protein levels in cells, while the AEGA (SEQ ID NO: 20)-SaCas9 mutant showed a comparable expression to WT-SaCas9 (Figures 9B-9D). AEGA (SEQ ID NO: 20)- SaCas9 was further investigated. Consistent with AEGA (SEQ ID NO: 20)- SaCas9 being deficient in binding Keapl (Figure 3C), this mutant was resistant to Keapl -mediated degradation in cells (Figure 3D). Like SpCas9, Keapl also robustly promoted a K48-linked SaCas9 poly-ubiquitination in cells (Figures 9E and 9F), which further supports the notion that Keapl targets SaCas9 for ubiquitination and destruction (Figure 9G). Recently, an eukaryotic programmable RNA-guided endonuclease Fanzor was discovered as an OMEGA system with genome engineering capacity (26). Fanzor (S. punctatus) proteins also contain two “ETGE” (SEQ ID NO: 18)-like motifs, including “ETVE” (SEQ ID NO: 23) and “ETGS” (SEQ ID NO: 24) (Figure 3E). In cells, Fanzor proteins were degraded by Keapl in a Keapl dose-dependent manner (Figure 3F). Mutating the “ETVE” (SEQ ID NO: 23) but not the “ETGS” (SEQ ID NO: 24) motif to “EAAE” (SEQ ID NO: 19) led to reduced Fanzor binding to Keapl (Figure 3G) and subsequent resistance to Keapl -mediated degradation in cells (Figure 3H). Together, these data suggest that Keapl can function as a mammalian E3 ligase recognizing and targeting Cas9 and Fanzor for degradation in cells.EXAMPLE 4Engineered Cas9 mutants evade Keapl recognition and exert enhanced gene editing ability in vitro
[0153] To examine if evading Keapl -mediated Cas9 protein degradation leads to enhanced gene editing efficiency, SpCas9 and SaCas9 were used for in cell and in vivo gene editing tests, respectively, due to their size differences. To test the in cell gene editing efficiency, sgRNAs and donor DNA were engineered to target the safe human endogenous AAVS1 locus to measure both gene knockout (CRISPRko) and knockin (CRISPRki) efficiency (Figure 4A). Consistent with SpCas9 being degraded by Keapl (Figure 2A), introducing exogenous Keapl reduced SpCas9 protein levels (Figure 4B) and subsequently attenuated Cas9-mediated AAVS1 editing efficiency in HEK293T cells using T7E1 assays for indel measurements (Figure 4C). On the other hand, depletion of endogenous Keapl in HEK293T cells resulted in protein accumulation of WT-, but not EAAE (SEQ ID NO: 19)-SpCas9 (Figure 4D). As a result, Keapl depletion led to increased EMX1 and AAVS1 editing efficiency in WT-Cas9, but not in EAAE (SEQ ID NO: 19)-Cas9 expressing cells (Figure 4E and Figure 10A), which positively correlated with SpCas9 protein levels. Consistent with Keapl being a gene editing suppressor, pharmacological Keapl inhibition by CDDO increased SpCas9-mediated EMX1 gene editing in HEK293T cells (Figure 4F).
[0154] When expressed to a comparable level to WT-SpCas9, EAAE (SEQ ID NO: 19)-SpCas9 displayed an enhanced AAVS1 editing efficiency at either the 2-day (Figure 10B) or 3-day period post-transfection (Figures 4G and 4H). Similarly enhanced gene editing ability was also observed using EMX1- sgRNA in either E7T1 assays (Figure 41) or DNA sanger sequencing coupled with TIDE analyses (Figure 4J) that are commonly used to analyze indels generated by CRISPR. In addition to HEK293T cells, it was also observed that, compared to WT-SpCas9, EAAE (SEQ ID NO: 19)-spCas9 generated more indels guided by E / WX7-sgRNA in HBE (human bronchial epithelial) (Figure 4K) or BPH1 (benign prostatic hyperplasia) cells (Figure 4L). Together these data support that EAAE (SEQ ID NO: 19)-SpCas9 displays an enhanced CRISPRko efficiency than WT-SpCas9.
[0155] To examine effects of SpCas9 mutants on CRISPRki efficiency, a guide DNA was used to introduce a Kpnl site into the human AAVS1 locus (Figure 4A). Notably, EAAE (SEQ ID NO: 19)-SpCas9, when expressed to a similar level as WT-SpCas9, also displayed an increased AAVS1 knockin efficiency (Figure 10C) evidenced by Kpnl digestion. To further reinforce this conclusion, unique primers targeting only Kpnl inserted genomic DNA regions were designed to measure knockin efficiency by qPCR. This approach also indicated that EAAE (SEQ ID NO: 19)-spCas9 displayed increased knockin efficiency compared with WT-SpCas9 (Figures 10D and 10E). Moreover, depletion of endogenous Keapl resulted in increased AAVS1 indel generation and knockin efficiencies in WT-, but not EAAE (SEQ ID NO: 19)-SpCas9 expressing cells (Figure 10F). To further reinforce this conclusion, sgRNAs targeting the mouse Rosa26 locus and a donor DNA template (Figure 4M) were used to measure the knockin efficiency in MEFs (mouse embryonic fibroblasts). To this end, it was found that when WT and EAAE (SEQ ID NO: 19)-SpCas9 were expressed to a comparable protein level (Figure 4N), EAAE (SEQ ID NO: 19)-SpCas9 exerted a significantly increased ability to facilitate CRISPRki (Figure 40).
[0156] For SaCas9, an AEGA (SEQ ID NO: 20)-SaCas9 mutant bypassing Keapl -mediated binding and degradation was identified (Figures 3C and 3D) and an examination was undertaken to study whether evading the Keapl negative regulation facilitates SaCas9-mediated gene editing. It was found that AEGA (SEQ ID NO: 20)-SaCas9 exerted an increased gene editing efficiency in HEK293T cells using sgRNAs targeting either DMD (Duchenne muscular dystrophy) (Figures 4P-4R), FANCF(FA complementation group F) (Figures 10G and 10H) or EMX1 (empty spiracles homeobox 1 ) gene (Figures 101 and 10J). Considering that in clinic CRISPR-mediated gene editing is more applicable to primary cells, the effects of SaCas9-AEGA (SEQ ID NO: 20) were compared with those of SaCas9-WT in BPH1 cells and it was observed that AEGA (SEQ ID NO: 20) exerted a significantly increased indel generation ability using sgRNAs either targeting endogenous EMX-1 locus (Figure 4S) or the DMD locus (Figure 4T). Together, these data support thenotion that SpCas9-EAAE (SEQ ID NO: 19) or SaCas9-AEGA (SEQ ID NO: 20) mutant evading Keapl recognition displays enhanced gene editing ability.
[0157] To further examine whether AEGA (SEQ ID NO: 20)-SaCas9 displays enhanced gene editing ability in vivo, an established DMD mouse model (9) was used. First, an increased dystrophin exon deletion efficiency was observed in C2C12 mouse myoblast cells using AEGA (SEQ ID NO: 20)- SaCas9, compared with WT-SaCas9 (Figure 11 A). Additionally, efficacy of AAV-SaCas9-WT or AAV-SaCas9-AEGA (SEQ ID NO: 20) to remove the mutated exon 23 from the dystrophin gene in live DMD mice was evaluated. At 8-weeks post-injection, muscle and other tissues were harvested to examine exon deletion efficiency (Figure 11 B). Although less AEGA (SEQ ID NO: 20)-SaCas9 proteins were detected than WT-SaCas9 in distinct organs from these animals (Figures 11C-11 E: 0.52-fold, 0.37-fold, and 0.88-fold compared with WT-SaCas9 in Figures 11C-11 E, respectively), a slightly increased dystrophin gene deletion efficiency was observed in tibialis anterior (Figure 11C), as well as a comparable dystrophin gene deletion efficiency in liver (Figure 11 D) and heart (Figure 11 E). Cumulatively, these data indicate that Cas9-mutants evading Keapl -mediated degradation exert enhanced gene editing ability in cells, while effects on in vivo gene editing remain to be further determined.
[0158] A potential concern for increased protein half-life of Cas9 variants by evading Keapl recognition is the build-up of off-target effects. Previously characterized off-target sites for sgEMX-1 (27) were used to test this possibility. By DNA sanger sequencing coupled TIDE analysis, it was observed that, although EAAE (SEQ ID NO: 19)-SpCas9 improved on-target gene editing efficiency compared with WT-SpCas9 (Figures 12A and 12B), it also increased the off-target frequency at 7 characterized off-target regions tested (Figures 12C-12I). eSpCas9 has been reported with improved on- target specificity (27).Thus, an eSpCas9-EAAE (SEQ ID NO: 19) mutant was engineered and it was found that, compared with SpCas9, eSpCas9 indeed reduced off-target effects on all 7 known off-target sites tested (Figures 12C- 121). Moreover, compared with EAAE (SEQ ID NO: 19)-SpCas9, EAAE (SEQ ID NO: 19)-eSpCas9 significantly decreased 6 out of 7 off-target indels(Figures 12C-12I). Compared with WT-eSpCas9, EAAE (SEQ ID NO: 19)- eSpCas9 did not generate more indels on 2 off-target sites (Figures 12C-12I). Together, these data suggest that the eSpCas9-EAAE (SEQ ID NO: 19) variant is a relatively “safer” candidate not only evading Keapl recognition with improved gene editing ability, but also with less off-target effects.EXAMPLE 5Engineered EAAE-dCas9-fusions exert increased CRISPRa or CRISPRi ability
[0159] In addition to Cas9-mediated gene editing, catalytic-dead versions of Cas9 (dCas9) have been widely used in epigenetic regulation, taking advantage of an accurate anchoring of dCas9 (and its fusion partners) to a specific chromatin locus (56). Given that the identified Cas9 mutants herein display extended protein half-lives (Figure 8J), the ability of EAAE (SEQ ID NO: 19)-dSpCas9 to enhance dCas9-governed transcriptional modulation was examined. To this end, like SpCas9, dSpCas9 was also subjected to Keapl -mediated protein degradation control (Figure 13A). A series of dSpCas9 fusions with distinct well-defined epigenetic regulators was created, including dCas9-VPR (a tripartite fusion of 3 heterologous viral transactivation domains), dCas9-p300 (E1 A binding protein p300), and and dCas9- SunTagxI O for transcription activation, as well as dCas9-KRAB (kruppel associated box) for transcription inactivation (57). The fusions were designed to have the same dSpCas9-sequence and promoter. A similar set of dCas9- fusion mutants was created with the “EAAE (SEQ ID NO: 19)-dCas9” mutation (Figure 5A). WT-dCas9-p300 was subject to Keapl -mediated degradation (Figure 13B). Protein expression of EAAE (SEQ ID NO: 19)-dCas9-p300 was comparable to WT-dCas9-p300 in cells (Figure 13C), regardless of transfection periods (Figure 13D) and cellular locations (either cytoplasm or chromatin, Figure 5B). Compared with WT-dCas9-p300, EAAE (SEQ ID NO: 19)-dCas9-p300 exerted an increased ability to activate transcription of either an exogenous fluorescent reporter (58) (Figure 5C), or endogenous genes such as OCT4 (POU5F1, POU class5 homoboxl) (5.1 -fold) and ILFIN {interleukin 1 receptor antagonist) (13.7-fold) (Figure 5D). This increase waslargely due to EAAE (SEQ ID NO: 19)-dCas9-p300 proteins remaining on chromatin for a longer duration compared to dCas9-WT, evidenced by an extended EAAE (SEQ ID NO: 19)-dCas9-p300 protein half-life on chromatin (Figures 5E and 5F). Notably, cytoplasmic EAAE (SEQ ID NO: 19)-dCas9- p300 also exerted an extended protein half-life (Figure 13E). Similarly, even the EAAE (SEQ ID NO: 19)-dCas9-VPR fusion and EAAE (SEQ ID NO: 19)- dCas9-Suntag fusion showed less expression to their WT counterparts in cells (Figure 13F), mutant versions displayed an enhanced ability to increase transcription of an exogenous fluorescent reporter (Figures 13G and 13H). Together, these data suggest that EAAE (SEQ ID NO: 19)-dCas9 fusions with epigenetic activators evading Keapl -mediated degradation display increased CRISPRa ability in vitro.
[0160] To investigate if evading Keapl -mediated dCas9 degradation also increases CRISPRi efficiency, WT- or EAAE (SEQ ID NO: 19)-dCas9-KRAB fusions were engineered with comparable expression levels in both cytoplasm and on chromatin (Figure 5G). Compared with WT-dCas9-KRAB, EAAE (SEQ ID NO: 19)-dCas9-KRAB showed an extended protein half-life in cytoplasm (Figure 131) and, more importantly, on chromatin (Figures 5H and 51). As a result, EAAE (SEQ ID NO: 19)-dCas9-KRAB exerted a significantly enhanced ability to suppress expression of endogenous genes including BRCA (Figure 5J), CANX (calnexin) (Figure 5K), SYVN (synoviolin 1) (Figure 5L), BLM (BLM RecQ like helicase) (Figure 5M), CHX1 (IRE2, immediate early response 2) (Figure 5N) and ERK1 (mitogen activated protein kinase 3) (Figure 50) using U6 or actin as a normalizatoin control (Figures 13J-13L). These data suggest EAAE (SEQ ID NO: 19)-dCas9 fusions also enhance CRISPRi ability in vitro.EXAMPLE 6 Cas9 is a Keapl inhibitor
[0161] Considering that Cas9 proteins do not naturally exist in mammalian cells, it was believed that introducing bacterial Cas9 proteins into mammalian systems could trigger immune responses and perturb mammalian signaling. Indeed, anti-Cas9 antibodies or Cas9 immune-reactive T cells were detectedin human serum (59-62) but not in the eye (63). Whether and how Cas9 specifically modulates intrinsic cellular programs remains unclear. It was found that expression of SpCas9 led to stabilization of Keapl substrates including NRF2 and IKKp (64) (Figure 6A) in an “ETGE” (SEQ ID NO: 18) degron-dependent manner (Figure 6B). Without being bound to any one theory, this appeared largely due to the SpCas9 proteins sequestering Keapl from its endogenous substrates (such as NRF2) (Figure 6C). On the other hand, increasing NRF2 protein levels did not significantly affect SpCas9 protein levels (Figure 6D). In addition to Keapl substrates, SpCas9 binding to Keapl also blocked binding of other Keapl binding partners, including PALB2 (52) (Figure 6E). Expression of SpCas9 did not affect PALB2 protein stability (Figure 6F) and vice versa (Figure 6G). Together, these data suggest that SpCas9 functions as a Keapl inhibitor, expression of which titrates Keapl away from endogenous Keapl substrates and Keapl binding partners, leading to stabilization of Keapl substrates and attenuated Keapl -mediated regulation of cellular processes (Figure 6H). On the other hand, EAAE (SEQ ID NO: 19)-SpCas9 does not inhibit Keapl function due to its inability to interact with Keapl , and thus can have minimal effects on Keapl biology.
[0162] Keapl mutations are frequently observed in human cancers, especially lung cancer, where these cancerous Keapl mutants are deficient in degrading oncogenic substrates including NRF2 (46). To examine if cancer associated Keapl mutants are also deficient in degrading Cas9, CRISPR- mediated gene editing was assessed to determine if higher editing activity could be observed in Keapl -mutant tumors. To this end, Keapl -C273S and C288S inactivation mutants (to mimic electrophiles-stimulated states) that are deficient in degrading NRF2 (65) (Figure 61) were still capable of degrading SpCas9 (Figure 6J). Given that these two critical cysteines are modified in response to oxidative and electrophilic stresses to suppress Keapl -mediated degradation of NRF2 (65), these data suggest that Keapl can utilize distinct mechanisms to target NRF2 and SpCas9 for degradation. Possible Keapl cancer-associated mutations that are deficient in degrading Cas9 were sought to further understand how possible environmental cues can modulate Keapl - mediated Cas9 protein stability control. From a panel of commonly observedKeapl -Kelch domain cancerous mutations, including G423V, D422N and G333C (Figure 6K), it was found that while all of these Keapl mutants were deficient in degrading NRF2 (Figure 6L), only the G333C-Keap1 mutant was unable to degrade SpCas9 in cells (Figures 6M and 6N). As a result, ectopic expression of WT- or G423V-Keap1 , but not G333C-Keap1 (Figure 60), significantly impaired SpCas9-guided AAVS1 editing (Figure 6P). These data suggest that not all cells expressing Keapl mutants associated with cancer and NRF2 accumulation can necessarily result in enhanced CRISPR efficiency, and that effects of Keapl mutants on NRF2 signaling and Cas9 / dCas9 mediated gene modulation can be separated. Thus, Keapl expression, but not Keapl somatic mutation nor activity, could be a more faithful marker to predict gene therapy efficacy by CRISPR-Cas9 / dCas9.Discussion of Examples
[0163] The mammalian E3 ligase Keapl earmarks Gas9s and dCas9s for ubiquitination and thus leads to degradation that reduces CRISPR efficiency. Engineered Cas9s and dCas9s with mutated Keapl degrons that evade Keapl -mediated suppression display improved Cas9 and dCas9 protein stability, as well as gene editing and epigenome editing efficiency, respectively, in cells. Without being bound to any one theory, the minimal effects of AEGA (SEQ ID NO: 20)-SaCas9 in improving DMD gene editing efficacy compared with WT-SaCas9 in the DMD mouse model could be due to lower mutated Cas9 expression levels, or other unknown reasons. Nonetheless, evading Keapl -mediated protein degredation leads to sustained dCas9 deposition on chromatin, leading to enhanced dCas9-mediated transcriptional regulation. These results provide a rationale to utilize engineered dCas9 variants in further improving CRISPRi and CRISPRa capability.
[0164] In addition to bacterial Cas9 proteins, Keapl also targets the eukaryotic enzyme Fanzor for degradation. These data suggest that mammalian Keapl can serve as a regulator for these bacterial / eukaryotic endonucleases bearing potential Keapl degrons as a general host defense mechanism.
[0165] On the other hand, Cas9 also functions as a Keapl suppressor in cells by competing with bona fide Keapl substrates or binding proteins to inactivate Keapl function. Inducible or constitutive Cas9 expression in various mouse tissues did not lead to morphological and breeding abnormalities (66); however, two reports demonstrate that Cas9 can induce a cellular toxic effect by triggering p53-dependent DNA damage response, which indirectly enriches p53-mutated cells (67,68). Considering the tight connection of p53 mutations with cancer, these studies could connect Cas9 activity with tumorigenesis. Notably, the Keapl / NRF2 signaling exerts context-dependent roles in tumorigenesis (69). For example, NRF2 is an oncogene in lung cancer (70) and breast cancer (71 ), while it functions as a tumor suppressor by exerting chemo-preventive activities or in established Nrf2 knockout murine models in skin (72), stomach (73), colon (74) and bladder (75) cancer. Given that SpCas9 perturbs the Keap1 / NRF2 signaling to cause NRF2 accumulation, how Cas9 affects cell behaviors could be determined by the pathophysiological settings. Elevated NRF2 protein levels have also been reported to contribute to resistance to anti-cancer drugs and cause metabolic reprogramming (76), suggesting the application of the CRISPR-Cas9 system in non-biased screens for drug resistant genes or metabolic vulnerability would be undertaken cautiously with considerations from the impact of Cas9 on NRF2 activity. To this end, Cas9 is not the only protein known to affect the Keapl / NRF2 signaling. A handful of mammalian proteins have been shown to compete with NRF2 to bind Keapl , leading to NRF2 accumulation, including HBXIP (71 ), CDK20 (77), P62 / sqstm1 (78) and p21 (79). Thus, it is plausible that in addition to Keap1 / NRF2, expression of the bacterial Cas9 proteins in mammalian cells can also affect other mammalian signaling networks, through which Cas9 modulates cellular functions and behaviors. This knowledge will provide a further guide to improve the safe applications of CRISPR-Cas9 technique in gene therapies in humans.EXAMPLE 7Materials and MethodsMaterials
[0166] MG132 (S2619), cycloheximide (S6611 ), bortezomib (S1013) and sulforaphane (S5771 ) were purchased from Selleckchem. Puromycin (P8833), anti-Flag agarose beads (A-2220), anti-HA agarose beads (A-2095) glutathione agarose beads (G4510) Tris Base (648311 ), NaCI (S9888) and the Keapl inhibitors tBHQ (112941 ) and CDDO (SMB00376) were purchased from Sigma-Aldrich. Nickel-NTA agarose beads (H-350-5) were purchased from Goldbio. Polyethylenimine (PEI, 23966) was purchased from Polysciences. NP-40 (AAJ61055AP) was purchased from Thermo Fisher Scientific.Antibodies
[0167] All antibodies were used at a 1 :2000 dilution in TBST (1 x Tris- Buffered Saline, 0.1 % Tween20) buffer with 5% non-fat milk for western blotting. Anti-HA antibody (3724), anti-Keap1 antibody (8047), anti-IKKp- antibody (8943), anti-rabbit IgG, HRP-linked antibody (7074) and anti-mouse IgG, HRP-linked antibody (7076) were obtained from Cell Signaling Technology. Anti-NRF2 antibody (ab62352) was obtained from Abeam. Anti- p27-antibody (sc1641 ) and anti-GFP-antibody (sc4304) were obtained from Santa Cruz. Polyclonal anti-Flag antibody (F-2425), monoclonal anti-Flag antibody (F-3165, clone M2) and anti-Tubulin antibody (T-5168) were obtained from Sigma-Aldrich.Cell culture and transfection
[0168] HEK293, HEK293T, MEF (mouse embryonic fibroblast), BPH1 , HeLa, HBE (human bronchial epithelial cell) and H520 (lung cancer) cells were cultured in DMEM medium (Gibco 11965092) supplemented with 10% FBS (Gibco 16000044), 100 units of penicillin and 100 mg / ml streptomycin (Gibco 15140122). HEK293, HEK293T and HeLa cells were purchased from UNC Lineberger TCF (tissue culture facility). MEF cells were obtained from Dr. Wenyi Wei Lab (BIDMC, Harvard Medical School). HBE (transformed by Bmi-1 / hTERT as described in (39)) and H520 (human epithelial lung squamous cell carcinoma, as ATCC-HTB-182) cells were obtained from Dr. Gang Greg Wang Lab (Duke University). BPH1 cells (a benign hyperplastic prostatic epithelial cell line, as Sigma-Aldrich SCC256 ) were obtained from Dr.Haojie Huang Lab (Mayo Clinic). Cell transfection was performed by polyethylenimine (PEI) as described previously (40,41 ). Briefly, 100 ng to 3 pg indicated DNA plasmids were mixed with PEI (1 :3 ratio) in Opti-MEM medium and vortexed, followed by incubation at room temperature (RT) for 15 min before gently dropping into cell culture media. Packaging of lentiviral shKeapI viruses and subsequent infection of various cell lines to generate stable endogenous Keapl depleted cells were performed according to the protocols described previously (42,43). Briefly, 2.5 pg shKeapI plasmid, 1 .25 pg VSVG and 1 .25 pg A8.9 plasmids were co-transfected into HEK293T cells in 10 cm dishes to package virus. Medium was refreshed 24 hr post-transfection. 10 mL medium from each transfection was collected each day for two consecutive days. Collected virus-containing media was pooled and filtered through 0.45 pM filter. Indicated cells were infected with indicated virus in the presence of polybrene for 24 hr followed by a recovery period of 24 hr with fresh medium. Then, cells were selected for 3 days using 1 pg / mL puromycin to eliminate non-infected cells.Plasmids construction
[0169] The plasmids pLenti-V2-Flag-SpCas9 (52961 ), eSpCas9 (71814), HF1 -SpCas9 (92012) and Flag-PALB2 (71114) were purchased from Addgene. The parental HA-SaCas9 was generated in the Dr. Charles Gersbach Lab at Duke University (44). Flag-Fanzor1 was cloned into the pLenti-GFP-puro vector (Addgene 17448) by Age1 and Xho1 sites. HA- dSpCas9-WT-CRAB HA-dSpCas9-WT-p300, HA-dSpCas9-WT-VPR, HA- dSpCas9-WT-FKBP and HA-dSpCas9-WT-suntagx10 were generated by the Nate Hathaway in UNC at Chapel Hill (45). All the Cas9 related mutations were generated using the QuikChange XL Site-Directed Mutagenesis Kit (200516, Agilent Technologies) according to the manufacturer’s instructions. HA-Keap1 was cloned into pCDNA3.0-HA vector with BamH1 and Xhol sites. HA-Keap1 -C151 S, C273S and C288S mutants were generated using the QuikChange XL Site-Directed Mutagenesis Kit (200516, Agilent Technologies) according to the manufacturer’s instructions. HA-Keap1 -G423V, D422N, G333C, Kelch and AKelach were cloned into pCDNA3.0-HA vector withBamH1 and Xhol using Flag-Keap1 plasmid as the template (46). Flag-NRF2 was cloned into pCDN3.0 backbone with BamH1 and Xho1 sites. shRNA vectors for depleting endogenous Keapl including shKeap1 -7, -11 , and -83 were generated in a previous study (46). sgRNAs for CRISPR-Cas9 mediated deletion of targeted gene were generated by cloning the annealed sgRNA oligos into Bbsl-digested phU6-sg vector (addgene 53188) as previously described (41 ). sgRNAs for CRISPR-SaCas9 mediated deletion of target genes were generated by cloning the annealed sgRNA oligos into Bbsl- digested phU6- sg vector for SaCas9 (pDO240-pZDonor-SaCas9-hU6-gRNA, from Dr. Charles Gersbach Lab, Duke University). Primers used for all vector constructions are listed in Table 3. The vector sequence for pDO240-pZDonor- SaCas9-hU6-gRNA is listed in Table 4.Table 3. Primers used for vector construction.Table 4. Vector sequence for pDO240-pZDonor-SaCas9-hU6-gRNA.Immunoblot and immunoprecipitations analyses
[0170] Cells were washed with sterile 1xPBS and lysed in EBC buffer (50 mM Tris pH 7.5, 120 mM NaCI, 0.5% NP-40) supplemented with protease inhibitors (B14002, Selleckchem) and phosphatase inhibitors (B15002, Selleckchem). The protein concentrations of whole cell lysates were measured by NanoDrop OneC (Thermo Fisher Scientific) using the Bio-Rad protein Bradford assay reagent (5000006, Bio-Rad) as described previously (40). Equal amounts of whole cell lysates (WCLs) were resolved by SDS-PAGE and immunoblotted with the indicated antibodies. For immunoprecipitations analysis, 1 mg WCLs were incubated with indicated antibody-conjugated agarose beads for 3-4 hr at 4 °C with rotations. The immuno-complexes were washed three times with NETN buffer (20 mM Tris, pH 8.0, 100 mM NaCI, 1 mM EDTA and 0.5% NP-40), resuspended in 3xSDS sample buffers, and boiled at 95 °C on heat block before being resolved by SDS-PAGE and immunoblotted with the indicated antibodies.RMSD measurements
[0171] The structure for SpCas9’s ‘ETGE’ (SEQ ID NO: 18) motif was obtained from PDB: 2flu and the NRF2’s ‘ETGE’ (SEQ ID NO: 18) motif was obtained from PDB: 4oo8. RMSD calculations were made using an in-house python script which takes the average distance between optimal rigid body superpositions of each protein’s Co backbone carbons. Calculations were made using the ‘ETGE’ (SEQ ID NO: 18) motif and the following three C- terminal and N-terminal flanking amino acids. Each structure was then compared to their respective full-length AlphFold2 generated structures; these were then compared in a similar fashion wherein similar values were obtained.Mass spectrometry analyses to identify Flag-SpCas9 interacting proteins
[0172] EV (empty vector) and Flag-SpCas9 constructs were transfected into HEK293T cells, respectively. 48 hrs post-transfection, cells were harvested with EBC buffer (50 mM Tris pH 7.5, 120 mM NaCI, 0.5% NP-40) with PI (protease inhibitor cocktail) and PPI (protein phosphatase inhibitor cocktail). 3 mg of WCL was used for Flag-M2 agarose beads mediated immunoprecipitations with gentle rotations at 4°C for 4 hrs. Flag- immunoprecipitants were washed with NETN buffer (20 mM Tris, pH 8.0, 100 mM NaCI, 1 mM EDTA and 0.5% NP-40) for 4 times. Proteins were diluted using 8 M urea (U4883, Sigma-Aldrich) to 100 pg / pL and then subjected to FASP trypsin digestion protocol. Briefly, protines were reduced using 50 mM DTT (A39255, Thermo Fisher Scientific) for 15 min at 65°C and diluted with 200 pL of 8 M urea. Then sarnies were transferred into 30K MWCO spin filter (14-558-349, Thermo Fisher Scientific) for centrifuging at 10,000 x g for 30min at room temperature. The proteins were washed twice using 200 pL of 8 M urea at 10,000 x g for 30 min at room temperature. Then, proteins were alkylated using 100 pL of 15 mM 2-chloroacetamide (148415000, Thermo Fisher Scientific) diluted in 8 M urea for 20 min in the dark at room temperature. The spin filter was washed twice at 10,000 x g for 20 min at room temperature. Then, the buffer was exchanged with 50 mM ammonium bicarbonate (ABC) pH 8.0 at 10,000 x g for 15 min at room temperature. 100 pL of 50 mM ABC was added into the spin filter with 2.5 pg of trypsin (V511 C, Promega). Samples were trypsinized overnight at 37°C for 18 hr. Following trypsinization, peptides were recovered in a new receiver tube by centrifugation at 10,000 x g for 15 min. Peptides were eluted twice using 50 pL 0.5% TFA in water at 10,000 x g for 10 min. Samples were then concentrated to 100 pL using a concentrator sold under the tradename SAVANT™ SPD131 DDA SpeedVac Concentrator (Thermo Fisher Scientific) followed by C18 column desalting (89870, Thermo Fisher Scientific). Samples were then concentrated and resolubilized in 100 pL of LC-Optima MS-grade water using SAVANT™ SPD131 DDA SpeedVac Concentrator. Ethyl acetate extraction followed by SAVANT™ SPD131 DDA SpeedVac Concentrator was performed to remove residual detergents. QFP (quantitative peptide assays and standards) (23290, Thermo Fisher Scientific) was performed for peptide quantification. Detailed liquid chromatography-mass spectrometry / mass spectrometry methods and data filtering methods were described previously (47). Data was searched using MaxQuant (version 1 .6.6.7), and all statistical analyses were done in Perseus (version 1.6.3.4). The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the PRIDE (48) partner repository with the dataset identifier PXD050935.T7E1 gene editing assays
[0173] 3-day post-transfection with indicated DNA constructs, cells were washed with sterile 1 xPBS and collected by trypsinization and centrifugation. Genomic DNA was extracted from cell pellets using a solution sold under the tradename QUICKEXTRACT™ DNA Extraction Solution (Biosearch technology SS000035-D2) following manufacturer’s instructions. End-pointPCR was performed using 10 pL genomic DNA as the template with indicated primers in the presence of high-fidelity Taq DNA polymerase (M7122, Promega). PCR products were verified by DNA electrophoresis in 1 xTAE buffer and purified by PCR cleanup kits (BS664, Bio Basic, Inc). 600-800 ng PCR products from control or indicated Cas9-expressing cells were mixed, denatured and then annealed to generate small indels in hybridization products. 1 pLT7 Endonuclease I (NEB M0689) was added and the reactions were incubated at 37°C for 1 hr. Digested PCR products were resolved by 2% TAE DNA agarose gel electrophoresis. Indels were calculated using the intensities of T7 digested bands normalized to the sum of both full-length and digested bands quantified by Imaged.Sanger sequencing and TIDE analyses for gene knockout efficacy
[0174] 3-day post-transfection with indicated DNA constructs, genomic DNA was extracted using QUICKEXTRACT™ DNA Extraction Solution following manufacturer’s instructions. A region of -500 bp flanking the target site was amplified by PCR with indicated primer pairs. PCR amplicons were purified and subjected to Sanger sequencing by GENEWIZ, Inc. TIDE analyses (see online at tide.nki.nl) was performed to calculate the total gene editing efficiency. Biological duplicates were used to generate error bars. Statistical significance was determined by one-way ANOVA tests. Primers used for all sanger sequencing were listed in the Table 5.
[0175] Table 5. Primers used for gene editing testingCRISPRa (dCas9-fusion CRISPR assays) and Flow Cytometry
[0176] 500,000 cells per well were plated in a 6-well plate and 24 hrs later, cells were transfected with 3 pg DNA and 9 pL PEI mixed in 200 pL Opti-MEM medium. For experiments using p300 fusion plasmids, 1.5 pg dCas9-p300, 1 pg sgRNA, and 0.5 pg tre3g-GFP plasmids were transfected into cells. 16- hours post-transfection, the medium was refreshed. For CRISPRa assays targeting endogenous genes, 2 pg dCas9-p300 plasmid and 1 pg sgRNA vector (0.25 pg per gRNA) were used. The medium was refreshed 16-hour post-transfection. Alternatively, 100,000 cells per well were plated into a 12- well plate and transfected with 1 pg DNA and 3 pL PEI mixed in 100 pL Opti- MEM medium. For each transfection, 0.35 pg sgRNA, 0.15 pg tre3g-BFPvector, and 0.5 pg dCas9 plasmid were used (except for dCas9-SunTag, where 0.35 pg dCas9 and 0.15 pg scFv-VP64 plasmids were used). The indicated time course experiments started 24 hours post-media change. To quantify the BFP or GFP expression, cells were dissociated with 0.05% Trypsin for 8-minutes, quenched by media with 10% FBS and spun down at 1 ,700 RPM for 5 minutes. The pellet was washed with 1 X PBS and resuspended in FACS buffer (1 mM EDTA and 0.2% BSA in 1 X PBS). Flow cytometry was performed with the Thermo Fisher Attune NxT and the data was analyzed with the Flow Jo Software.RNA Extraction and qRT-PCR
[0177] 3 days post-transfection, total RNA was extracted using RNA extraction kits (BS584, Bio Basic) following the manufacturer’s instructions. Briefly, cells were washed with sterile 1 xPBS and resuspended in RLT solution and passed through QIAshredder columns for homogenization (QIAGEN 79656). Equal volume of 7-% ethanol was added and mixed. Mixtures were transferred to EZ-10 columns with 2 ml_ collection tubes and spun at 6,000xg for 1 min. 500 uL RW and RPE solutions were added to the columns, respectively, to clean up the RNA. Total RNA was collected in a RNase-free 1 .5 mL microtube with 50 pL RNase-free water with centrifugation at 8,000xg for 1 min. The RNA concentrations were determined by Nanodrop OnecSpectrophotometer (Thermo Fisher Scientific). cDNA was synthesis using a kit sold under the tradename ISCRIPT™ cDNA Synthesis Kit (1708890, BioRad) following manufacturer’s instructions. qRT-PCR was performed with QuantStudio 6 Flex Real-Time PCR System (Thermo Fisher Scientific) using iTaq universal SYBR green supermix (1725124, Bio-Rad). Briefly, the PCR reaction was composed of 4 pL cDNA template (20 X dilution after cDNA synthesis) , 5 pL iTaq Mix and 1 pL 5 pM primer mix in a 10 pL volume. The comparative Ct method was used to calculate fold changes in gene expression, which was normalized to U6 snRNA as a house-keeping gene. Biological triplicates were used to generate error bar. Statistical significance was determined by one-way ANOVA tests. Primers used for testing gene expression were listed in the Table 6.
[0178] Table 6. Primers used for qRT-PCR.DMD mouse model study
[0179] ITR-containing plasmids, AAV-2XgRNA (9), AAV-WT-SaCas9 (9) and AAV-AEGA-SaCas9 were generated in Dr. Charles Gersbach's Lab (Duke University). Intact ITRs were verified by Smal digest on all vectors before AAV production. AAV9 was produced by the Duke University Viral Vector Core and titers were measured by qPCR with a plasmid standard curve. Neonatal 2- day-old (P2) mdx mice (Jackson Labs) were administered AAV by intravenous injection through the facial vein with 40ul AAV vector per mouse (7.2e11 vg AAV-2XgRNA and 2.2e11 vg AAV-WT-SaCas9 or AAV-AEGA-SaCas9. At 8 weeks post-injection, mice were euthanized and tibialis anterior muscle tissue was collected. Protein analysis with western blot and genomic DNA analysis with endpoint PCR were performed as described previously (9). In brief, protein lysate was loaded onto NuPAGE Bis-Tris gels (Invitrogen) with MOPS buffer (Invitrogen), transferred to nitrocellulose membranes, and blocked overnight. Blots were probed with anti-HA (abeam) for Cas9 detection, MANDYs8 (Sigma D8168) for dystrophin detection, and rabbit anti-GAPDH (Cell Signaling), followed by mouse or rabbit horseradish peroxidase- conjugated secondary antibodies (Santa Cruz) and visualized using Western- C ECL substrate (Biorad) on a ChemiDoc chemiluminescence system (Biorad). In brief, genomic DNA was isolated using the DNeasy kit (Qiagen) and endpoint PCR was performed with primers flanking the Cas9 cut sites (5‘- TACACTAACACGCATATTTG (SEQ ID NO: 121 ), and 5‘- CATTGCATCCATGTCTGACT (SEQ ID NO: 122)) to amplify a 1638bp unedited region or the expected 467bp deletion product. PCR products were electrophoresed in a 1 % agarose gel and viewed on a BioRad GelDoc imager.Statistics
[0180] Differences between control and experimental conditions were evaluated by Student's t test or one-way ANNOVA. These analyses were performed using the SPSS 11.5 Statistical Software and P < 0.05 considered statistically significant.REFERENCES
[0181] All references listed in the instant disclosure, including but not limited to all patents, patent applications and publications thereof, scientific journal articles, and database entries are incorporated herein by reference in their entireties to the extent that they supplement, explain, provide a background for, and / or teach methodology, techniques, and / or compositions employed herein. The discussion of the references is intended merely to summarize the assertions made by their authors. No admission is made that any reference (or a portion of any reference) is relevant prior art. Applicants reserve the right to challenge the accuracy and pertinence of any cited reference.
[0182] The references cited herein are as follows:1 . Mali, P., Yang, L., Esvelt, K.M., Aach, J., Guell, M., DiCarlo, J.E., Norville, J.E. and Church, G.M. (2013) RNA-guided human genome engineering via Cas9. Science, 339, 823-826.2. Cong, L., Ran, F.A., Cox, D., Lin, S., Barretto, R., Habib, N., Hsu, P.D., Wu, X., Jiang, W., Marraffini, L.A. et al. (2013) Multiplex genome engineering using CRISPR / Cas systems. Science, 339, 819-823.3. Chen, S., Sanjana, N.E., Zheng, K., Shalem, O., Lee, K., Shi, X., Scott, D.A., Song, J., Pan, J.Q., Weissleder, R. et al. (2015) Genome-wide CRISPR screen in a mouse model of tumor growth and metastasis. Cell, 160, 1246-1260.4. Korkmaz, G., Lopes, R., Ugalde, A.P., Nevedomskaya, E., Han, R., Myacheva, K., Zwart, W., Elkon, R. and Agami, R. (2016) Functional genetic screens for enhancer elements in the human genome using CRISPR-Cas9. Nat Biotechnol, 34, 192-198.Huang, A., Garraway, L.A., Ashworth, A. and Weber, B. (2020) Synthetic lethality as an engine for cancer drug target discovery. Nat Rev Drug Discov, 19, 23-38. Kurata, M., Yamamoto, K., Moriarity, B.S., Kitagawa, M. and Largaespada, D.A. (2018) CRISPR / Cas9 library screening for drug target discovery. J Hum Genet, 63, 179-186. Xiao-Jie, L., Hui-Ying, X., Zun-Ping, K., Jin-Lian, C. and Li-Juan, J. (2015) CRISPR-Cas9: a new and promising player in gene therapy. J Med Genet, 52, 289-296. Wang, G., McCain, M.L., Yang, L., He, A., Pasqualini, F.S., Agarwal, A., Yuan, H., Jiang, D., Zhang, D., Zangi, L. et al. (2014) Modeling the mitochondrial cardiomyopathy of Barth syndrome with induced pluripotent stem cell and heart-on-chip technologies. Nat Med, 20, 616-623. Nelson, C.E., Hakim, C.H., Ousterout, D.G., Thakore, P.L, Moreb, E.A., Castellanos Rivera, R.M., Madhavan, S., Pan, X., Ran, F.A., Yan, W.X. et al. (2016) In vivo genome editing improves muscle function in a mouse model of Duchenne muscular dystrophy. Science, 351 , 403-407. Tabebordbar, M., Zhu, K., Cheng, J.K.W., Chew, W.L., Widrick, J. J., Yan, W.X., Maesner, C., Wu, E.Y., Xiao, R., Ran, F.A. etal. (2016) In vivo gene editing in dystrophic mouse muscle and muscle stem cells. Science, 351 , 407-41 1 . Kim, K., Park, S.W., Kim, J.H., Lee, S.H., Kim, D., Koo, T., Kim, K.E., Kim, J.H. and Kim, J.S. (2017) Genome surgery using Cas9 ribonucleoproteins for the treatment of age-related macular degeneration. Genome Res, 27, 419-426. Gao, X., Tao, Y., Lamas, V., Huang, M., Yeh, W.H., Pan, B., Hu, Y.J., Hu, J.H., Thompson, D.B., Shu, Y. et al. (2018) Treatment of autosomal dominant hearing loss by in vivo delivery of genome editing agents. Nature, 553, 217-221 . Chen, Z.H., Yu, Y.P., Zuo, Z.H., Nelson, J.B., Michalopoulos, G.K., Monga, S., Liu, S., Tseng, G. and Luo, J.H. (2017) Targeting genomic rearrangements in tumor cells through Cas9-mediated insertion of a suicide gene. Nat Biotechnol, 35, 543-550.Kaminski, R., Bella, R., Yin, C., Otte, J., Ferrante, P., Gendelman, H.E., Li, H., Booze, R., Gordon, J., Hu, W. et al. (2016) Excision of HIV-1 DNA by gene editing: a proof-of-concept in vivo study. Gene Ther, 23, 696. Gillmore, J.D., Gane, E., Taubel, J., Kao, J., Fontana, M., Maitland, M.L., Seitzer, J., O'Connell, D., Walsh, K.R., Wood, K. et al. (2021 ) CRISPR- Cas9 In Vivo Gene Editing for Transthyretin Amyloidosis. N Engl J Med, 385, 493-502. Lee, R.G., Mazzola, A.M., Braun, M.C., Platt, C., Vafai, S.B., Kathiresan, S., Rohde, E., Bellinger, A.M. and Khera, A.V. (2023) Efficacy and Safety of an Investigational Single-Course CRISPR Base-Editing Therapy Targeting PCSK9 in Nonhuman Primate and Mouse Models. Circulation, 147, 242-253. Yin, H., Xue, W., Chen, S., Bogorad, R.L., Benedetti, E., Grompe, M., Koteliansky, V., Sharp, P.A., Jacks, T. and Anderson, D.G. (2014) Genome editing with Cas9 in adult mice corrects a disease mutation and phenotype. Nat Biotechnol, 32, 551 -553. Long, C., Amoasii, L., Mireault, A. A., McAnally, J.R., Li, H., Sanchez-Ortiz, E., Bhattacharyya, S., Shelton, J.M., Bassel-Duby, R. and Olson, E.N. (2016) Postnatal genome editing partially restores dystrophin expression in a mouse model of muscular dystrophy. Science, 351 , 400-403. Tabebordbar, M., Zhu, K., Cheng, J.K., Chew, W.L., Widrick, J. J., Yan, W.X., Maesner, C., Wu, E.Y., Xiao, R., Ran, F.A. etal. (2016) In vivo gene editing in dystrophic mouse muscle and muscle stem cells. Science, 351 , 407-41 1 . Amoasii, L., Hildyard, J.C.W., Li, H., Sanchez-Ortiz, E., Mireault, A., Caballero, D., Harron, R., Stathopoulou, T.R., Massey, C., Shelton, J.M. et al. (2018) Gene editing restores dystrophin expression in a canine model of Duchenne muscular dystrophy. Science. Kosicki, M., Tomberg, K. and Bradley, A. (2018) Repair of double-strand breaks induced by CRISPR-Cas9 leads to large deletions and complex rearrangements. Nat Biotechnol, 36, 765-771. Kleinstiver, B.P., Prew, M.S., Tsai, S.Q., Topkar, V.V., Nguyen, N.T., Zheng, Z., Gonzales, A.P., Li, Z., Peterson, R.T., Yeh, J.R. et al. (2015)Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature, 523, 481 -485.23. Zetsche, B., Gootenberg, J.S., Abudayyeh, O.O., Slaymaker, I.M., Makarova, K.S., Essletzbichler, P., Volz, S.E., Joung, J., van der Oost, J., Regev, A. et al. (2015) Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell, 163, 759-771 .24. Burstein, D., Harrington, L.B., Strutt, S.C., Probst, A. J., Anantharaman, K., Thomas, B.C., Doudna, J. A. and Banfield, J.F. (2017) New CRISPR-Cas systems from uncultivated microbes. Nature, 542, 237-241 .25. Liu, J. J., Orlova, N., Oakes, B.L., Ma, E., Spinner, H.B., Baney, K.L.M., Chuck, J., Tan, D., Knott, G.J., Harrington, L.B. et al. (2019) CasX enzymes comprise a distinct family of RNA-guided genome editors. Nature.26. Saito, M., Xu, P., Faure, G., Maguire, S., Kannan, S., Altae-Tran, H., Vo, S., Desimone, A., Macrae, R.K. and Zhang, F. (2023) Fanzor is a eukaryotic programmable RNA-guided endonuclease. Nature, 620, 660- 668.27. Slaymaker, I.M., Gao, L., Zetsche, B., Scott, D.A., Yan, W.X. and Zhang, F. (2016) Rationally engineered Cas9 nucleases with improved specificity. Science, 351 , 84-88.28. Kleinstiver, B.P., Pattanayak, V., Prew, M.S., Tsai, S.Q., Nguyen, N.T., Zheng, Z. and Joung, J.K. (2016) High-fidelity CRISPR-Cas9 nucleases with no detectable genome-wide off-target effects. Nature, 529, 490-495.29. Casini, A., Olivieri, M., Petris, G., Montagna, C., Reginato, G., Maule, G., Lorenzin, F., Prandi, D., Romanel, A., Demichelis, F. et al. (2018) A highly specific SpCas9 variant is identified by in vivo screening in yeast. Nat Biotechnol, 36, 265-271 .30. Gaudelli, N.M., Komor, A.C., Rees, H.A., Packer, M.S., Badran, A.H., Bryson, D.L and Liu, D.R. (2017) Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature, 551 , 464-471.31 . Dominguez, A. A., Lim, W.A. and Qi, L.S. (2016) Beyond editing: repurposing CRISPR-Cas9 for precision genome regulation and interrogation. Nat Rev Mol Cell Biol, 17, 5-15.32. Nishida, K., Arazoe, T., Yachie, N., Banno, S., Kakimoto, M., Tabata, M., Mochizuki, M., Miyabe, A., Araki, M., Hara, K.Y. et al. (2016) Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science, 353.33. Pawluk, A., Amrani, N., Zhang, Y., Garcia, B., Hidalgo-Reyes, Y., Lee, J., Edraki, A., Shah, M., Sontheimer, E.J., Maxwell, K.L. etal. (2016) Naturally Occurring Off-Switches for CRISPR-Cas9. Cell, 167, 1829-1838 e1829.34. Khajanchi, N. and Saha, K. (2022) Controlling CRISPR with small molecule regulation for somatic cell genome editing. Mol Ther, 30, 17-31 .35. Chu, V.T., Weber, T., Wefers, B., Wurst, W., Sander, S., Rajewsky, K. and Kuhn, R. (2015) Increasing the efficiency of homology-directed repair for CRISPR-Cas9-induced precise gene editing in mammalian cells. Nat Biotechnol, 33, 543-548.36. Robert, F., Barbeau, M., Ethier, S., Dostie, J. and Pelletier, J. (2015) Pharmacological inhibition of DNA-PK stimulates Cas9-mediated genome editing. Genome Med, 7, 93.37. Yu, C., Liu, Y., Ma, T., Liu, K., Xu, S., Zhang, Y., Liu, H., La Russa, M., Xie, M., Ding, S. et al. (2015) Small molecules enhance CRISPR genome editing in pluripotent stem cells. Cell Stem Cell, 16, 142-147.38. Xie, K., Minkenberg, B. and Yang, Y. (2015) Boosting CRISPR / Cas9 multiplex editing capability with the endogenous tRNA-processing system. Proc Natl Acad Sci U S A, 112, 3570-3575.39. Fulcher, M.L., Gabriel, S.E., Olsen, J.C., Tatreau, J.R., Gentzsch, M., Livanos, E., Saavedra, M.T., Salmon, P. and Randell, S.H. (2009) Novel human bronchial epithelial cell lines for cystic fibrosis research. Am J Physiol Lung Cell Mol Physiol, 296, L82-91 .40. Liu, P., Begley, M., Michowski, W., Inuzuka, H., Ginzberg, M., Gao, D., Tsou, P., Gan, W., Papa, A., Kim, B.M. et al. (2014) Cell-cycle-regulated activation of Akt kinase by phosphorylation at its carboxyl terminus. Nature, 508, 541 -545.41 .Liu, P., Gan, W., Chin, Y.R., Ogura, K., Guo, J., Zhang, J., Wang, B., Blenis, J., Cantley, L.C., Toker, A. et al. (2015) Ptdlns(3,4,5)P3-Dependent Activation of the mTORC2 Kinase Complex. Cancer Discov, 5, 1194-1209.Liu, P., Gan, W., Guo, C., Xie, A., Gao, D., Guo, J., Zhang, J., Willis, N., Su, A., Asara, J.M. et al. (2015) Akt-mediated phosphorylation of XLF impairs non-homologous end-joining DNA repair. Mol Cell, 57, 648-661 . Guo, J., Chakraborty, A. A., Liu, P., Gan, W., Zheng, X., Inuzuka, H., Wang, B., Zhang, J., Zhang, L., Yuan, M. et al. (2016) pVHL suppresses kinase activity of Akt in a proline-hydroxylation-dependent manner. Science, 353, 929-932. Thakore, P.I., Kwon, J.B., Nelson, C.E., Rouse, D.C., Gemberling, M.P., Oliver, M.L. and Gersbach, C.A. (2018) RNA-guided transcriptional silencing in vivo with S. aureus CRISPR-Cas9 repressors. Nat Common, 9, 1674. Chiarella, A.M., Butler, K.V., Gryder, B.E., Lu, D., Wang, T.A., Yu, X., Pomella, S., Khan, J., Jin, J. and Hathaway, N.A. (2020) Dose-dependent activation of gene expression is achieved using CRISPR and small molecules that recruit endogenous chromatin machinery. Nat Biotechnol, 38, 50-55. Hast, B.E., Cloer, E.W., Goldfarb, D., Li, H., Siesser, P.F., Yan, F., Walter, V., Zheng, N., Hayes, D.N. and Major, M.B. (2014) Cancer-derived mutations in KEAP1 impair NRF2 degradation but not ubiquitination. Cancer Res, 74, 808-817. Zhu, Z., Zhou, X., Du, H., Cloer, E.W., Zhang, J., Mei, L., Wang, Y., Tan,X., Hepperla, A. J., Simon, J.M. et al. (2023) STING Suppresses Mitochondrial VDAC2 to Govern RCC Growth Independent of Innate Immunity. Adv Sci (Weinh), 10, e2203718. Perez- Rivero I, Y., Bai, J., Bandla, C., Garcia-Seisdedos, D., Hewapathirana, S., Kamatchinathan, S., Kundu, D.J., Prakash, A., Frericks-Zipper, A., Eisenacher, M. ef al. (2022) The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res, 50, D543-D552. Ergunay, T., Ayhan, O., Celen, A.B., Georgiadou, P., Pekbilir, E., Abaci,Y.T., Yesildag, D., Rettel, M., Sobhiafshar, U., Ogmen, A. et al. (2022) Sumoylation of Cas9 at lysine 848 regulates protein stability and DNA binding. Life Sci Alliance, 5.Baird, L. and Yamamoto, M. (2020) The Molecular Mechanisms Regulating the KEAP1 -NRF2 Pathway. Mol Cell Biol, 40. Kansanen, E., Kuosmanen, S.M., Leinonen, H. and Levonen, A.L. (2013) The Keap1 -Nrf2 pathway: Mechanisms of activation and dysregulation in cancer. Redox Biol, 1 , 45-49. 0rthwein, A., Noordermeer, S.M., Wilson, M.D., Landry, S., Enchev, R.L, Sherker, A., Munro, M., Pinder, J., Salsman, J., Dellaire, G. etal. (2015) A mechanism for the suppression of homologous recombination in G1 cells. Nature, 528, 422-426. Li, X., Zhang, D., Hannink, M. and Beamer, L.J. (2004) Crystal structure of the Kelch domain of human Keapl . J Biol Chem, 279, 54750-54758. Ran, F.A., Hsu, P.D., Wright, J., Agarwala, V., Scott, D.A. and Zhang, F. (2013) Genome engineering using the CRISPR-Cas9 system. Nat Protoc, 8, 2281 -2308. Akutsu, M., Dikic, I. and Bremm, A. (2016) Ubiquitin chain diversity at a glance. J Cell Sci, 129, 875-880. Kampmann, M. (2018) CRISPRi and CRISPRa Screens in Mammalian Cells for Precision Biology and Medicine. ACS Chem Biol, 13, 406-416. Chavez, A., Tuttle, M., Pruitt, B.W., Ewen-Campen, B., Chari, R., Ter- Ovanesyan, D., Haque, S.J., Cecchi, R.J., Kowal, E.J.K., Buchthal, J. etal. (2016) Comparison of Cas9 activators in multiple species. Nat Methods, 13, 563-567. Butler, K.V., Chiarella, A.M., Jin, J. and Hathaway, N.A. (2018) Targeted Gene Repression Using Novel Bifunctional Molecules to Harness Endogenous Histone Deacetylation Activity. ACS Synth Biol, 7, 38-45.Simhadri, V.L., McGill, J., McMahon, S., Wang, J., Jiang, H. and Sauna, Z.E. (2018) Prevalence of Pre-existing Antibodies to CRISPR-Associated Nuclease Cas9 in the USA Population. Mol Ther Methods Clin Dev, 10, 105-112. Charlesworth, C.T., Deshpande, P.S., Dever, D.P., Camarena, J., Lemgart, V.T., Cromer, M.K., Vakulskas, C.A., Collingwood, M.A., Zhang, L., Bode, N.M. et al. (2019) Identification of preexisting adaptive immunity to Cas9 proteins in humans. Nat Med, 25, 249-254.Ferdosi, S.R., Ewaisha, R., Moghadam, F., Krishna, S., Park, J.G., Ebrahimkhani, M.R., Kiani, S. and Anderson, K.S. (2019) Multifunctional CRISPR-Cas9 with engineered immunosilenced human T cell epitopes. Nat Commun, 10, 1842. Wagner, D.L., Amini, L., Wendering, D.J., Burkhardt, L.M., Akyuz, L., Reinke, P., Volk, H.D. and Schmueck-Henneresse, M. (2019) High prevalence of Streptococcus pyogenes Cas9-reactive T cells within the adult human population. Nat Med, 25, 242-248. Toral, M.A., Charlesworth, C.T., Ng, B., Chemudupati, T., Homma, S., Nakauchi, H., Bassuk, A.G., Porteus, M.H. and Mahajan, V.B. (2022) Investigation of Cas9 antibodies in the human eye. Nat Commun, 13, 1053. Kim, J.E., You, D.J., Lee, C., Ahn, C., Seong, J.Y. and Hwang, J. I. (2010) Suppression of NF-kappaB signaling by KEAP1 regulation of IKKbeta activity through autophagic degradation and inhibition of phosphorylation. Cell Signal, 22, 1645-1654. Kobayashi, A., Kang, M.L, Watai, Y., Tong, K.L, Shibata, T., Uchida, K. and Yamamoto, M. (2006) Oxidative and electrophilic stresses activate Nrf2 through inhibition of ubiquitination activity of Keapl . Mol Cell Biol, 26, 221 -229. Platt, R.J., Chen, S., Zhou, Y., Yim, M.J., Swiech, L., Kempton, H.R., Dahlman, J.E., Parnas, O., Eisenhaure, T.M., Jovanovic, M. et al. (2014) CRISPR-Cas9 knockin mice for genome editing and cancer modeling. Cell, 159, 440-455. Haapaniemi, E., Botla, S., Persson, J., Schmierer, B. and Taipale, J. (2018) CRISPR-Cas9 genome editing induces a p53-mediated DNA damage response. Nat Med. Ihry, R.J., Worringer, K.A., Salick, M.R., Frias, E., Ho, D., Theriault, K., Kommineni, S., Chen, J., Sondey, M., Ye, C. et al. (2018) p53 inhibits CRISPR-Cas9 engineering in human pluripotent stem cells. Nat Med.Sporn, M.B. and Liby, K.T. (2012) NRF2 and cancer: the good, the bad and the importance of context. Nat Rev Cancer, 12, 564-571 . 0hta, T., lijima, K., Miyamoto, M., Nakahara, I., Tanaka, H., Ohtsuji, M., Suzuki, T., Kobayashi, A., Yokota, J., Sakiyama, T. et al. (2008) Loss ofKeapl function activates Nrf2 and provides advantages for lung cancer cell growth. Cancer Res, 68, 1303-1309. Zhou, X.L., Zhu, C.Y., Wu, Z.G., Guo, X. and Zou, W. (2019) The oncoprotein HBXIP competitively binds KEAP1 to activate NRF2 and enhance breast cancer cell growth and metastasis. Oncogene. Xu, C., Huang, M.T., Shen, G., Yuan, X., Lin, W., Khor, T.O., Conney, A.H. and Kong, A.N. (2006) Inhibition of 7,12-dimethylbenz(a)anthracene- induced skin tumorigenesis in C57BL / 6 mice by sulforaphane is mediated by nuclear factor E2-related factor 2. Cancer Res, 66, 8293-8296. Ramos-Gomez, M., Kwak, M.K., Dolan, P.M., Itoh, K., Yamamoto, M., Talalay, P. and Kensler, T.W. (2001 ) Sensitivity to carcinogenesis is increased and chemoprotective efficacy of enzyme inducers is lost in nrf2 transcription factor-deficient mice. Proc Natl Acad Sci U S A, 98, 3410- 3415. Khor, T.O., Huang, M.T., Prawan, A., Liu, Y., Hao, X., Yu, S., Cheung, W.K., Chan, J.Y., Reddy, B.S., Yang, C.S. et al. (2008) Increased susceptibility of Nrf2 knockout mice to colitis-associated colorectal cancer. Cancer Prev Res (Phila), 1 , 187-191 . lida, K., Itoh, K., Kumagai, Y., Oyasu, R., Hattori, K., Kawai, K., Shimazui, T., Akaza, H. and Yamamoto, M. (2004) Nrf2 is essential for the chemopreventive efficacy of oltipraz against urinary bladder carcinogenesis. Cancer Res, 64, 6424-6431. Taguchi, K. and Yamamoto, M. (2017) The KEAP1 -NRF2 System in Cancer. Front Oncol, 7, 85. Wang, Q., Ma, J., Lu, Y., Zhang, S., Huang, J., Chen, J., Bei, J.X., Yang, K., Wu, G., Huang, K. etal. (2017) CDK20 interacts with KEAP1 to activate NRF2 and promotes radiochemoresistance in lung cancer cells. Oncogene, 36, 5321 -5330. Komatsu, M., Kurokawa, H., Waguri, S., Taguchi, K., Kobayashi, A., Ichimura, Y., Sou, Y.S., Ueno, L, Sakamoto, A., Tong, K.l. et al. (2010) The selective autophagy substrate p62 activates the stress responsive transcription factor Nrf2 through inactivation of Keapl . Nat Cell Biol, 12, 213-223.Chen, W., Sun, Z., Wang, X.J., Jiang, T., Huang, Z., Fang, D. and Zhang, D.D. (2009) Direct interaction between Nrf2 and p21 (Cip1 / WAF1 ) upregulates the Nrf2-mediated antioxidant response. Mol Cell, 34, 663-673.Kusebauch, U., Campbell, D.S., Deutsch, E.W., Chu, C.S., Spicer, D.A., Brusniak, M.Y., Slagel, J., Sun, Z., Stevens, J., Grimes, B. et al. (2016)Human SRMAtlas: A Resource of Targeted Assays to Quantify the Complete Human Proteome. Cell, 166, 766-778. Wilburn, D.B., Shannon, A.E., Spicer, V., Richards, A.L., Yeung, D., Swaney, D.L., Krokhin, O.V. and Searle, B.C. (2023) Deep learning from harmonized peptide libraries enables retention time prediction of diverse post translational modifications. bioRxiv.
Claims
CLAIMSWhat is claimed is:1 . A mutant Cas9 protein comprising one or more mutations in a Keapl degron sequence of a Cas9 protein, wherein said mutant Cas9 protein has a half-life greater than the half-life of the corresponding Cas9 protein that does not comprise the one or more mutations in the Keapl degron sequence.
2. The mutant Cas9 protein of claim 1 , wherein the Cas9 protein is a Streptococcus pyogenes Cas9 protein comprising a sequence of SEQ ID NO:1 or a protein having a sequence having at least 90% homology to SEQ ID NO: 1.
3. The mutant Cas9 protein of claim 2, wherein the Keapl degron sequence is the ETGE sequence at amino acids 1068 to 1071 of SEQ ID NO: 1.
4. The mutant Cas9 protein of claim 3, wherein the one or more mutations comprise an alanine substitution at position 1069 or at position 1070 of SEQ ID NO: 1.
5. The mutant Cas9 protein of claim 3, wherein the one or more mutations comprise an alanine substitution at position 1069 and at position 1070 of SEQ ID NO: 1.
6. The mutant Cas9 protein of claim 1 , wherein the Cas9 protein is a Staphylococcus aureus Cas9 of SEQ ID NO: 2 or a protein having at least 90% homology to SEQ ID NO: 2.
7. The mutant Cas9 protein of claim 6, wherein the Keapl degron sequence is the ETGE-like sequence at amino acids 860 to 863 of SEQ ID NO: 2.
8. The mutant Cas9 protein of claim 7, wherein the one or more mutations comprise an alanine substitution at position 860 of SEQ ID NO: 2, a glutamic acid substitution at position 861 of SEQ ID NO: 2, and an alanine substitution at postion 863 of SEQ ID NO: 2.
9. The mutant Cas9 protein of any of claims 1 -8, further comprising one or more mutations outside of the Keapl degron sequence of the Cas9 protein,optionally wherein the one or more mutations outside of the Keapl degron sequence are conservative and / or nonconservative substitutions.
10. The mutant Cas9 protein of claim 9, wherein the one or more mutations outside of the Keapl degron sequence of the Cas9 protein comprise an alanine at position 10 of SEQ ID NO: 1 and an alanine at position 840 of SEQ ID NO: 1.1 1. An isolated nucleic acid sequence comprising a sequence that encodes the mutant Cas9 protein of any one of claims 1 -10.
12. A construct comprising the nucleic acid sequence of claim 11 .
13. The construct of claim 12, further comprising a promoter operably linked to the nucleic acid sequence encoding the mutant Cas9 protein.
14. The construct of claim 12 or 13, further comprising a sequence encoding a marker gene.
15. The construct of any one of claims 12-14, wherein the construct comprises a viral based vector or a non-viral based vector.
16. A composition comprising a construct of any one of claims 12-15 and a pharmaceutically acceptable carrier.
17. A cell line comprising an isolated nucleic acid sequence or construct of any of claims 12-16.
18. A method of enhancing CRISPR efficiency comprising administering one or more guide RNAs (gRNAs) to a cell comprising the mutant Cas9 protein of any one of claims 1 -9 to thereby increase or improve CRISPR efficiency.
19. The method of claim 18, wherein the enhanced CRISPR efficiency increases CRIPSR-mediated knock-out efficiency.
20. The method of claim 18, wherein the enhanced CRISPR efficiency increases CRISPR-mediated knock-in efficiency.21 . The method of any one of claims 18-20, wherein the cell has decreased expression of Keapl .
22. The method of claim 18, wherein the cell comprises a Keapl mutant, and wherein presence of the Keapl mutant leads to decreased degradation of any Cas9 protein or mutant Cas9 in said cell.
23. The method of claim 22, wherein the Keapl mutant is a Keapl protein having the mutation of G333C relative to wild-type Keapl as set forth in SEQ ID NO: 3.
24. A CRISPR-Cas9 system for use in editing a gene, wherein the CRISPR-Cas9 system comprises (i) a guide RNA; and (ii) a mutant Cas9 protein of any one of claims 1 -10.
25. A method of editing a gene, wherein the method comprises introducing a mutant Cas9 of any one of claims 1-10 into a cell or tissue, and introducing one or more guide RNAs into said cell or tissue.
26. A method of suppressing Keapl activity comprising administering a protein or nucleic acid that blocks the Keapl - Kelch domain responsible for Keapl binding to its substrates.
27. The method of claim 26, wherein the protein that blocks the Keapl - Kelch domain comprises an antibody specific to the Kelch domain.
28. The method of claim 26, wherein the nucleic acid that blocks the Keapl - Kelch domain is a nucleic acid that mimics a Kelch domain substrate.
29. The method of claim 28, wherein the Kelch domain substrate is a Keapl degron sequence of a Cas9 protein.
30. The method of any one of claims 26-29, wherein the suppressed Keapl activity is the ability to ubiquitinate the Cas9 protein.31 . The method of any one of claims 26-30, wherein suppression of Keapl activity increases the half-life of the Cas9 protein.
32. A method of promoting cell growth comprising administering an inhibitor of Keapl and NRF2 binding in amounts sufficient to promote cell growth.
33. The method of claim 32, wherein the inhibitor is a mutant Cas9 protein of any one of claims 1 -10.