Compositions and methods for epigenome editing

The CRISPR/Cas9-based gene activation system, combining Cas9 with a histone acetyltransferase, addresses limitations of dCas9 activators by enabling precise and robust gene activation from both promoters and distal enhancers using a single guide RNA, achieving specific epigenetic regulation.

JP7827321B2Active Publication Date: 2026-03-10DUKE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-04-05
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Current methods for manipulating epigenetic traits, such as those using dCas9 activators, are limited by the need for multiple activation domains or combinations of gRNAs to achieve high levels of gene transfer, and lack direct enzymatic function for specific chromatin regulation, hindering the ability to target and test specific epigenetic marks.

Method used

A CRISPR/Cas9-based gene activation system is developed, comprising a fusion protein of Cas9 and a histone acetyltransferase, like the p300 protein, to directly regulate epigenetic structures and activate genes from promoters and distal enhancers using a single guide RNA.

Benefits of technology

The system provides precise and robust activation of target genes, specifically acetylating histone H3 lysine 27, enabling efficient transcriptional activation from both proximal and distal regulatory elements with high specificity and using a single gRNA.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007827321000028
    Figure 0007827321000028
  • Figure 0007827321000029
    Figure 0007827321000029
  • Figure 0007827321000030
    Figure 0007827321000030
Patent Text Reader

Abstract

To provide CRISPR / Cas9-based gene activation systems and methods of using the systems.SOLUTION: Disclosed herein are CRISPR / Cas9-based gene activation systems that include a fusion protein of a Cas9 protein and a protein having histone acetyltransferase activity, and methods of using the systems.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 113,569, filed February 9, 2015, the entire contents of which are incorporated herein by reference.

[0002] STATEMENT OF GOVERNMENT RIGHTS This invention was made with government support under Federal Grant No. 1R01DA036865 awarded by the National Institutes of Health. The government has certain rights in this invention.

[0003] Technical Field The present disclosure is directed to a CRISPR / Cas9-based gene activation system and methods of using said system. [Background technology]

[0004] The Human Genome Project was funded and driven by the premise that sequencing the human genome would reveal the genetic basis of complex diseases with a strong heritable component, including cardiovascular disease, neurodegenerative conditions, and metabolic diseases such as diabetes. It was believed that this information would yield novel drug targets for these widespread diseases. However, thousands of genome-wide association studies (GWAS) have shown that genetic variants associated with these complex diseases do not occur within genes but rather in intergenic regulatory regions that control the levels of specific genes. Similarly, approximately 20% of Mendelian disorders have no detectable coding mutations, suggesting that the causative mutations reside within gene regulatory elements. Importantly, assigning functional roles to these regulatory elements is extremely difficult because they are often located far from their target genes. Furthermore, many genes and regulatory elements are classified as individual positive hits in each GWAS study. Indeed, follow-up projects to the Human Genome Project, such as the NIH-funded Encyclopedia of DNA Elements (ENCODE) and the Roadmap Epigenomics Project, have identified millions of putative regulatory elements throughout the human genome for many human cell types and tissues.

[0005] The first hurdle in functional genomics is developing technologies to directly and precisely manipulate genome function at individual loci. Projects such as ENCODE and the Roadmap Epigenomics Project have identified millions of epigenetic marks across the human genome for many human cell types and tissues. However, studying the function of these marks has been limited primarily to statistical associations with gene expression. Technologies for direct manipulation of these epigenetic traits are essential for translating these association-based discoveries into mechanistic principles of gene regulation. These advances have the potential to benefit human health by enabling gene therapies that alter the epigenetic code of targeted regions of the genome, strategies for regenerative medicine and disease modeling based on lineage-specific epigenetic reprogramming, and the design of epigenome-specific drug screening platforms.

[0006] Manipulation of the epigenome is possible by treating cells with small molecule drugs, such as inhibitors of histone deacetylases or DNA methyltransferases, or by differentiating cells into specific lineages. However, small molecule-based methods globally alter the epigenome and transcriptome and are not suitable for targeting individual loci. Epigenome editing techniques, including fusion of epigenome-modifying enzymes with programmable DNA-binding proteins, such as zinc finger proteins and transcription activator-like effectors (TALEs), are effective in achieving targeted DNA methylation, DNA hydroxymethylation, and histone demethylation, methylation, and deacetylation. Summary of the Invention [Problem to be solved by the invention]

[0007] When fused to an activation domain, such as an oligomer of herpes simplex virus protein 16 (VP16), dCas9 can function as a synthetic transcription regulator. However, limitations remain in the use of dCas9 activators, including the need for multiple activation domains or the need for combinations of gRNAs to achieve high levels of gene transfer through synergy between activation domains. Conventional activator domains used in these engineered transcription factors, such as the VP16 tetramer VP64, function as scaffolds to recruit multiple components of the transcription preinitiation complex and lack direct enzymatic function to specifically regulate chromatin states. This indirect method of epigenetic remodeling does not allow for testing the role of specific epigenetic marks and may not be as powerful as direct programming of epigenetic states. There remains a need for the ability to target direct manipulation of epigenetic traits. [Means for solving the problem]

[0008] The present invention is directed to a fusion protein comprising two heterologous polypeptide domains, wherein a first polypeptide domain comprises a Clustered Regularly Interspaced Short Palindromic Repeat-associated (Cas) protein and a second polypeptide domain comprises a peptide with histone acetyltransferase activity.

[0009] The present invention is directed to a DNA targeting system comprising the fusion protein described above and at least one guide RNA (gRNA).

[0010] The present invention is directed to a method of activating gene expression of a target gene in a cell, comprising contacting the cell with a polynucleotide encoding a DNA targeting system, wherein the DNA targeting system comprises a fusion protein as described above and at least one guide RNA (gRNA). [Brief explanation of the drawings]

[0011] [Figures 1A-1C] This shows that the dCas9p300 Core fusion protein activates transcription of endogenous genes from the proximal promoter region. Figure 1A shows a schematic diagram of the dCas9 fusion proteins dCas9VP64, dCas9FL p300, and dCas9p300 Core. Streptococcus pyogenes dCas9 contains the nuclease-inactivating mutations D10A and H840A. The D1399 catalytic residue within the p300 HAT domain is shown. Figure 1B shows a Western blot demonstrating the expression levels of dCas9 fusion proteins and GAPDH in co-transfected cells (the full blot is shown in Figure 7C). Figure 1C shows the relative mRNA expression of IL1RN, MYOD, and OCT4 by the indicated dCas9 fusion proteins co-transfected with four gRNAs targeting each promoter region, as determined by qRT-PCR (Tukey's test, *P value < 0.05, n = 3 (each independent experiment), error bars: sem). Numbers above the bars indicate the average expression levels. FLAG, epitope tag; NLS, nuclear localization signal; HA, hemagglutinin epitope tag; CH, cysteine-histidine-rich region; Bd, bromodomain; HAT, histone acetyltransferase domain. [Figures 2A-2C]The dCas9p300 Core fusion protein activates endogenous gene transcription from the distal enhancer region. Figure 2A shows the relative MYOD mRNA production in cells cotransfected with dCas9VP64 or dCas9p300 Core and a pool of gRNAs targeting the proximal or distal regulatory region; promoter data from Figure 1C are shown (Tukey's test, *P value < 0.05 compared to mock-transfected cells; Tukey's test †P value < 0.05 between dCas9p300 Core and dCas9VP64, n = 3 independent experiments, error bars: sem). The human MYOD locus is depicted schematically with the corresponding gRNA locations in red. CE, MyoD core enhancer; DRR, MyoD distal regulatory region. Figure 2B shows the relative OCT4 mRNA production in cells co-transfected with dCas9VP64 or dCas9p300 Core and a pool of gRNAs targeting the proximal and distal regulatory regions; promoter data from Figure 1C are shown (Tukey's test, *P value < 0.05 compared to mock-transfected cells; Tukey's test †P value < 0.05 between dCas9p300 Core and dCas9VP64, n = 3 independent experiments; error bars: sem). The human OCT4 locus is depicted schematically with the corresponding gRNA location in red. DE, Oct4 distal enhancer; PE, Oct4 proximal enhancer. Figure 2C shows the human β-globin locus depicted schematically with the appropriately positioned hypersensitive region 2 (HS2) enhancer region and downstream genes (HBE, HBG, HBD, and HBB). The corresponding HS2 gRNA position is shown in red. Relative mRNA production from distal genes in cells co-transfected with four gRNAs targeting the HS2 enhancer and the indicated dCas9 protein. Note the logarithmic y-axis and dashed red line indicating background expression (Tukey's test between conditions for each β-globin gene, †P-value <0.05, n = 3 independent experiments, error bars: sem). ns, not significant difference. [Figure 3A-3B]We demonstrate that transcriptional activation targeted by dCas9p300 Core is specific and robust. Figures 3A-3C show MA plots derived from DEseq2 analysis of genome-wide RNA-seq data from HEK293T cells transiently co-transfected with dCas9VP64 (Figure 3A), dCas9p300 Core (Figure 3B), or dCas9p300 Core(D1399Y) (Figure 3C) and gRNAs targeting four IL1RN promoters, compared with HEK293T cells transiently co-transfected with dCas9 and gRNAs targeting four IL1RN promoters. In each of Figures 3A-3C, mRNAs corresponding to IL1RN isoforms are indicated in blue and circled. The dots labeled in red in Figures 3B and 3C correspond to off-target transcripts that were significantly enriched after multiple hypothesis testing (KDR, (FDR=1.4×10-3); FAM49A, (FDR=0.04); p300, (FDR=1.7×10-4) in Figure 3B). [Figure 3C] We demonstrate that transcriptional activation targeted by dCas9p300 Core is specific and robust. Figures 3A-3C show MA plots derived from DEseq2 analysis of genome-wide RNA-seq data from HEK293T cells transiently co-transfected with dCas9VP64 (Figure 3A), dCas9p300 Core (Figure 3B), or dCas9p300 Core(D1399Y) (Figure 3C) and gRNAs targeting four IL1RN promoters, compared with HEK293T cells transiently co-transfected with dCas9 and gRNAs targeting four IL1RN promoters. In each of Figures 3A-3C, mRNAs corresponding to IL1RN isoforms are indicated in blue and circled. The dots labeled in red in Figures 3B and 3C correspond to off-target transcripts that were significantly enriched after multiple hypothesis testing (KDR, (FDR=1.4×10-3); FAM49A, (FDR=0.04); p300, (FDR=1.7×10-4) in Figure 3B; and p300, (FDR=4.4×10-10) in Figure 3C). [Figure 4A]The dCas9p300 Core fusion protein acetylates chromatin at the targeted enhancer and corresponding downstream gene. Figure 4A shows the region encompassing the human β-globin locus on chromosome 11 (positions 5,304,000–5,268,000; GRCh37 / hg19 assembly). The HS2 gRNA target location is shown in red, and the ChIP-qPCR amplicon region is depicted in black with corresponding green numbers. For comparison, ENCODE / Broad Institute H3K27ac enrichment signals in K562 cells are shown. Enlarged insets for the HS2 enhancer, HBE, and HBG1 / 2 promoter regions are shown below. [Figure 4B-4C] The dCas9p300 Core fusion protein acetylates chromatin at targeted enhancers and corresponding downstream genes. Figures 4B–4D show H3K27ac ChIP-qPCR enrichment (relative to dCas9; red dotted lines) at the HS2 enhancer, HBE promoter, and HBG1 / 2 promoter in cells cotransfected with four gRNAs targeting the HS2 enhancer and the indicated dCas9 fusion proteins. HBG ChIP amplicons 1 and 2 amplify redundant sequences in the HBG1 and HBG2 promoters (represented by ‡). Tukey's test between conditions for each ChIP-qPCR region; *P value < 0.05, n = 3 (independent experiments); error bars: sem). [Figure 4D]The dCas9p300 Core fusion protein acetylates chromatin at targeted enhancers and corresponding downstream genes. Figures 4B–4D show H3K27ac ChIP-qPCR enrichment (relative to dCas9; red dotted lines) at the HS2 enhancer, HBE promoter, and HBG1 / 2 promoter in cells cotransfected with four gRNAs targeting the HS2 enhancer and the indicated dCas9 fusion proteins. HBG ChIP amplicons 1 and 2 amplify redundant sequences in the HBG1 and HBG2 promoters (represented by ‡). Tukey's test between conditions for each ChIP-qPCR region; *P value < 0.05, n = 3 (independent experiments); error bars: sem). [Figure 5A-5B] This shows that the dCas9p300 Core fusion protein, together with a single gRNA, activates transcription of endogenous genes from regulatory regions. Relative levels of IL1RN (Figure 5A), MYOD (Figure 5B), or OCT4 (Figure 5C) mRNA produced in cells co-transfected with dCas9p300 Core or dCas9VP64 and gRNAs targeting the respective promoters (n = 3 independent experiments, error bars: sem). HS2, β-globin locus control region hypersensitive region 2; ns, not significant (Tukey's test). [Figure 5C-5D]These results demonstrate that dCas9p300 Core fusion proteins, together with a single gRNA, activate transcription of endogenous genes from regulatory regions. Relative IL1RN (Figure 5A), MYOD (Figure 5B), or OCT4 (Figure 5C) mRNA levels were produced from cells co-transfected with dCas9p300 Core or dCas9VP64 and gRNAs targeting the respective promoters (n = 3 independent experiments, error bars: sem). Relative MYOD (Figure 5D) or OCT4 (Figure 5E) mRNA levels were produced from cells co-transfected with dCas9p300 Core and the indicated gRNAs targeting the indicated MYOD or OCT4 enhancers (n = 3 independent experiments, error bars: sem). HS2, β-globin locus control region hypersensitive region 2; ns, not significant (Tukey's test). [Figures 5E-5G]We demonstrate that dCas9p300 Core fusion proteins, together with a single gRNA, activate transcription of endogenous genes from regulatory regions. Relative MYOD (Figure 5D) or OCT4 (Figure 5E) mRNA levels were measured in cells cotransfected with dCas9p300 Core and the indicated gRNAs targeting the indicated MYOD or OCT4 enhancers (n = 3 independent experiments; error bars: sem). DRR, MYOD distal regulatory region; CE, MYOD core enhancer; PE, OCT4 proximal enhancer; DE, OCT4 distal enhancer. (Tukey's test between dCas9p300 Core and a single OCT4 DE gRNA compared to mock-transfected cells, *P < 0.05; Tukey's test between dCas9p300 Core and an OCT4 DE gRNA compared to all, †P < 0.05). Relative HBE (Figure 5F) or HBG (Figure 5G) mRNA production in cells co-transfected with dCas9p300 Core and the indicated gRNA targeting the HS2 enhancer (Tukey test between dCas9p300 Core and a single HS2 gRNA compared to mock-transfected cells, *P value < 0.05; Tukey test between dCas9p300 Core and a single HS2 gRNA compared to all, †P value < 0.05, n = 3 independent experiments, error bars: sem). HS2, β-globin locus control region hypersensitive region 2; ns, not significant (Tukey test). [Figures 6A-6C]We demonstrate that p300 Core can target genomic loci using various programmable DNA-binding proteins. Figure 6A shows a schematic diagram of the Neisseria meningitidis (Nm) dCas9 fusion proteins Nm-dCas9VP64 and Nm-dCas9p300 Core. Neisseria meningitidis dCas9 contains the nuclease-inactivating mutations D16A, D587A, H588A, and N611A. Figures 6B-6C show the relative HBE (Figure 6B) or HBG (Figure 6C) mRNA levels in cells cotransfected with the indicated five individual or pooled Nm gRNAs (A-E) targeting the HBE or HBG promoters, along with Nm-dCas9VP64 or Nm-dCas9p300 Core. Tukey's test, *P value < 0.05 (compared to mock-transfected control), n = 3 (each independent experiment), error bars: semNLS, nuclear localization signal; HA, hemagglutinin tag; Bd, bromodomain; CH, cysteine-histidine-rich region; HAT, histone acetyltransferase domain. [Figure 6D-6E] This demonstrates that p300 Core can target genomic loci via various programmable DNA-binding proteins. Figures 6D-6E show the relative HBE (Figure 6D) or HBG (Figure 6E) mRNA levels in cells cotransfected with the five indicated individual or pooled Nm gRNAs (A-E) targeting the HS2 enhancer, along with Nm-dCas9VP64 or Nm-dCas9p300 Core. Tukey's test, *P value < 0.05 (compared to mock-transfected controls), n = 3 (each independent experiment). Error bars: semNLS, nuclear localization signal; HA, hemagglutinin tag; Bd, bromodomain; CH, cysteine-histidine-rich region; HAT, histone acetyltransferase domain. [Figure 6F]This demonstrates that p300 Core can target genomic loci via various programmable DNA-binding proteins. Figure 6F shows a schematic diagram of a TALE with a repeat variable diresidue (repeat domain)-containing domain targeted to IL1RN. Tukey's test, *P value < 0.05 (compared to mock-transfected control), n = 3 (each independent experiment). Error bars: semNLS, nuclear localization signal; HA, hemagglutinin tag; Bd, bromodomain; CH, cysteine-histidine-rich region; HAT, histone acetyltransferase domain. [Figure 6G] This demonstrates that p300 Core can target genomic loci via various programmable DNA-binding proteins. Figure 6G shows the relative IL1RN mRNA levels in cells transfected with individual or pooled (A-D) IL1RN TALEVP64 or IL1RN TALEp300 Core-encoding plasmids. Tukey's test, *P value < 0.05 (compared to mock-transfected controls), n = 3 (each independent experiment). Error bars: semNLS, nuclear localization signal; HA, hemagglutinin tag; Bd, bromodomain; CH, cysteine-histidine-rich region; HAT, histone acetyltransferase domain. [Figures 6H-6I] This demonstrates that p300 Core can target genomic loci via various programmable DNA-binding proteins. Figure 6H shows a schematic diagram of ZF fusion proteins with zinc finger helices 1 to 6 (F1 to F6) targeting the ICAM1 promoter. Figure 6I shows the relative ICAM1 mRNA levels in cells transfected with ICAM1 ZFVP64 or ICAM1 ZFp300 Core. Tukey's test, *P value <0.05 (compared to mock-transfected controls), n = 3 (each independent experiment). Error bars: semNLS, nuclear localization signal; HA, hemagglutinin tag; Bd, bromodomain; CH, cysteine-histidine-rich region; HAT, histone acetyltransferase domain. [Figures 7A-7B]Figure 7A shows a schematic depiction of the WT dCas9p300 Core fusion protein and p300 Core mutant derivatives. The relative positions of the mutated amino acids are indicated as yellow bars within the p300 Core effector domain. Figure 7B shows that dCas9p300 Core variants were transiently co-transfected with four IL1RN promoter gRNAs and screened for hyperactivity (amino acid 1645 / 1646 RR / EE and C1204R mutations) or hypoactivity (denoted by ‡) via mRNA production from the IL1RN locus (upper panel, n = 2 (independent experiments), error bars: sem). Experiments were performed in duplicate, with one well used for RNA isolation and the other for Western blotting to confirm expression (lower panel). Nitrocellulose membranes were cut and incubated with α-FLAG primary antibody (top, Sigma-Aldrich cat.# F7425) or α-GAPDH (bottom, Cell Signaling Technology cat.# 14C10), followed by α-rabbit HRP secondary antibody (Sigma-Aldrich cat.# A6154). [Figure 7C] Figure 7C shows the activity of dCas9p300 Core mutant fusion proteins. Figure 7C shows the whole membrane from the Western blot shown in Figure 1B. Nitrocellulose membranes were cut and incubated with α-FLAG primary antibody (top, Sigma-Aldrich cat.# F7425) or α-GAPDH (bottom, Cell Signaling Technology cat.# 14C10), followed by α-rabbit HRP secondary antibody (Sigma-Aldrich cat.# A6154). After careful rearrangement of the cut pieces, the membranes were developed for the indicated periods. [Figure 8] Figure 1 shows that target gene activation is not affected by overexpression of synthetic dCas9 fusion proteins. [Figure 9A]Figure 9A shows a comparison of Sp.dCas9 and Nm.dCas9 gene transfer from the HS2 enhancer with individual and pooled gRNAs. Figure 9A shows a schematic representation of the human β-globin locus, including Streptococcus pyogenes dCas9 (Sp.dCas9) and Neisseria meningitidis dCas9 (Nm.dCas9) gRNA locations at the HS2 enhancer. Overlaid transcription profiles from nine ENCODE cell lines (GM12878, H1-hESC, HeLa-S3, HepG2, HSMM, HUVEC, K562, NHEK, and NHLF), ranked into eight vertical display ranges, are shown, along with ENCODE p300 binding peaks in K562, A549 (EtOH.02), HeLA-S3, and SKN_SH_RA cell lines. The ENCODE HEK293T DNase hypersensitive region (HEK293T DHS) is shown within the HS2 enhancer inset. [Figure 9B-9E] Figures 9B-9E show a comparison of Sp.dCas9 transfection with individual and pooled gRNAs from the HS2 enhancer. Figures 9B-9E show the relative transcriptional induction of HBE, HBG, HBD, and HBD transcripts from single and pooled Sp.dCas9 gRNAs (A-D) or single and pooled Nm.dCas9 gRNAs (A-E) in response to co-transfection with Sp.dCas9p300 Core or Nm.dCas9p300 Core, respectively. gRNAs are tiled for each dCas9 ortholog corresponding to its position in GRCh37 / hg19. The gray dashed line indicates the background expression level in transiently co-transfected HEK293T cells. Note the common logarithmic scale across Figures 9B-9E. Numbers above the bars in Figures 9B-9E indicate the average expression level (n = at least 3 independent experiments; error bars: s.e.m.). [Figure 10] We show that dCas9VP64 and dCas9p300 Core induce H3K27ac enrichment at chromatin targeted by IL1RN gRNA. [Figures 11A-11C] Figure 11 shows a direct comparison of the VP64 and p300 Core effector domains between TALEs and dCas9 programmable DNA-binding proteins. Figure 11A shows the GRCh37 / hg19 region encompassing the IL1RN transcription start site, shown schematically with the IL1RN TALE binding site and dCas9 IL1RN gRNA target site. Figure 11B shows a direct comparison of IL1RN activation in HEK293T cells upon transfection with individual or pooled (A-D) IL1RN TALEVP64 fusion proteins or cotransfection with individual or pooled (A-D) IL1RN-targeting gRNAs. Figure 11C shows a direct comparison of IL1RN activation in HEK293T cells upon transfection with individual or pooled (A-D) IL1RN TALEp300 Core fusion proteins or cotransfection with dCas9p300 Core and individual or pooled (A-D) IL1RN-targeting gRNAs. Note the common logarithmic scale between Figure 11B and Figure 11C. Numbers above the bars in Figure 11B and Figure 11C indicate mean values. Tukey's test, *P value < 0.05, n = at least 3 (independent experiments). Error bars: sem. [Figures 12A-12B]Figure 12A shows Western blotting performed on cells transiently transfected with individual or pooled IL1RN TALE proteins. Nitrocellulose membranes were cut and probed with α-HA primary antibody (1:1000 dilution in TBST + 5% milk, top, Covance cat.# MMS-101P) or α-GAPDH (bottom, Cell Signaling Technology cat.# 14C10), followed by α-mouse HRP (Santa Cruz, sc-2005) or α-rabbit HRP (Sigma-Aldrich cat.# A6154) secondary antibodies, respectively. Figure 12B shows Western blotting was performed on cells transiently transfected with ICAM1 ZF-effector proteins, and nitrocellulose membranes were cut and probed with α-FLAG primary antibody (top, Sigma-Aldrich cat.# F7425) or α-GAPDH (bottom, Cell Signaling Technology cat.# 14C10), followed by α-rabbit HRP secondary antibody (Sigma-Aldrich cat.# A6154). Red asterisks indicate nonspecific bands. [Figures 13A-13B] Figure 13A shows that dCas9p300 Core and dCas9VP64 do not synergize in transactivation. dCas9p300 Core was co-transfected with the four indicated IL1RN promoter gRNAs at a 1:1 mass ratio to PL-SIN-EF1α-EGFP3 (GFP), dCas9, or dCas9VP64 (n = 2 independent experiments, error bars: sem). Figure 13B shows that dCas9p300 Core was co-transfected with the four indicated MYOD promoter gRNAs at a 1:1 mass ratio to GFP, dCas9, or dCas9VP64 (n = 2 independent experiments, error bars: sem). No significant differences were observed using Tukey's test (not significant). [Figures 14A-14B]The basic chromatin context of the dCas9p300 Core target locus is shown. Figures 14A-14D show the indicated locus along with the relevant Streptococcus pyogenes gRNA used in this study, along with the corresponding genomic location in GRCh37 / hg19. ENCODE HEK293T DNase hypersensitivity enrichment is shown (changes are scaled) along with regions of significant DNase hypersensitivity in HEK293T cells ("DHS"). Additionally, the ENCODE master DNase cluster across 125 cell types is shown. An overlay of ENCODE H3K27ac and H3K4me3 enrichment across seven cell lines (GM12878, H1-hESC, HSMM, HUVEC, K562, NHEK, and NHLF) is also presented, each ranked in a vertical range of 50 to 150. The endogenous p300 binding profile was also shown for each locus and each cell line. [Figure 14C-14D] The basic chromatin context of the dCas9p300 Core target locus is shown. Figures 14A-14D show the indicated locus along with the relevant Streptococcus pyogenes gRNA used in this study, along with the corresponding genomic location in GRCh37 / hg19. ENCODE HEK293T DNase hypersensitivity enrichment is shown (changes are scaled) along with regions of significant DNase hypersensitivity in HEK293T cells ("DHS"). Additionally, the ENCODE master DNase cluster across 125 cell types is shown. An overlay of ENCODE H3K27ac and H3K4me3 enrichment across seven cell lines (GM12878, H1-hESC, HSMM, HUVEC, K562, NHEK, and NHLF) is also presented, each ranked in a vertical range of 50 to 150. The endogenous p300 binding profile was also shown for each locus and each cell line. [Figure 14E] A summary of the information provided in Figures 14A-14D is shown. [Figure 15A] The amino acid sequence of the dCas9 construct is shown. [Figure 15B]The amino acid sequence of the dCas9 construct is shown. [Figure 15C] The amino acid sequence of the dCas9 construct is shown. [Figure 15D] The amino acid sequence of the dCas9 construct is shown. [Figure 15E] The amino acid sequence of the dCas9 construct is shown. [Figure 15F] The amino acid sequence of the dCas9 construct is shown. [Figure 15G] The amino acid sequence of the dCas9 construct is shown. [Figure 15H] The amino acid sequence of the dCas9 construct is shown. [Figure 15I] The amino acid sequence of the dCas9 construct is shown. [Figure 15J] The amino acid sequence of the dCas9 construct is shown. [Figure 16] The amino acid sequence of the ICAM1 zinc finger 10 effector is shown. [Figure 17] gRNA design and screening. [Figure 18] gRNA combinatorial activation is shown. [Figure 19] Pax7 guide screening in 293T. [Figure 20] This shows that gRNA19 is localized to the DHS. [Figure 21] The relative amounts of FGF1A mRNA in 293T with or without dCas9p300 Core are shown. [Figure 22] Expression levels of FGF1B and FGF1C in 293T cells with dCas9p300 Core, dCas9VP64, or dCas9 (alone) are shown. [Figure 23] Expression levels of FGF1A, FGF1B, and FGF1C in 293T cells with dCas9p300 Core, dCas9VP64, or dCas9 (alone) are shown. DETAILED DESCRIPTION OF THE INVENTION

[0012] Disclosed herein is a CRISPR / Cas9-based gene activation system and a method for using the system. This system provides an easily programmable method for facilitating reliable control of epigenome and downstream gene expression. The CRISPR / Cas9-based gene activation system includes a CRISPR / Cas9-based acetyltransferase, which is a fusion protein between the Cas9 protein and a protein with histone acetyltransferase activity (such as the histone acetyltransferase (HAT) catalytic core domain of the human E1A-associated protein p300). The Cas9 protein may not have nuclease activity. An example of a Cas9 protein with abolished nuclease activity is dCas9. Recruitment of acetyltransferase function to genomic target sites by dCas9 and gRNA allows for direct regulation of epigenetic structures, thus providing an effective means of gene activation.

[0013] The disclosed CRISPR / Cas9-based acetyltransferases catalyze the acetylation of histone H3 lysine 27 at their target sites, resulting in robust transcriptional activation of target genes from promoters and proximal and distal enhancers. As disclosed herein, gene activation by these targeted acetyltransferases is highly specific across the genome. CRISPR / Cas9-based acetyltransferases can target any site within the genome and have the unique ability to activate distal regulatory elements. In contrast to conventional dCas9-based activators, CRISPR / Cas9-based acetyltransferases efficiently activate genes from enhancer regions together with individual or single guide RNAs.

[0014] 1.Definition The terms "comprise," "including," "having," "having," "can," "containing," and variations thereof, as used herein, are intended to be open-ended phrases, terms, or words that do not exclude additional acts or structural possibilities. The singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The present disclosure also contemplates other embodiments that "comprise," "consist of," and "consist essentially of" the embodiments or elements provided herein, whether or not explicitly stated.

[0015] For the recitation of numerical ranges herein, each intervening number is specifically contemplated to the same degree of precision. For example, for the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and for the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are specifically contemplated.

[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art. In case of conflict, the present document, including definitions, will control. Preferred methods and materials are described below, although methods and materials similar or equivalent to those described herein can be used in the practice or testing of this specification. All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety. The materials, methods, and examples disclosed herein are illustrative only and are not intended to be limiting.

[0017] "Adeno-associated virus" or "AAV," used interchangeably herein, refers to a small virus belonging to the Dependovirus genus of the Parvoviridae family that infects humans and some other primate species. AAV is not currently known to cause disease, and therefore the virus provokes a very mild immune response.

[0018] As used herein, "chromatin" refers to the organized complex of chromosomal DNA associated with histones.

[0019] As used herein interchangeably, "Cis-regulatory element" or "CRE" refers to a region of non-coding DNA that regulates the transcription of nearby genes. CREs are found near the gene(s) they regulate. CREs generally regulate gene transcription by functioning as binding sites for transcription factors. Examples of CREs include promoters and enhancers.

[0020] "Clustered Regularly Interspaced Short Palindromic Repeats" and "CRISPR," used interchangeably herein, refer to loci containing multiple short direct repeats found in approximately 40% of sequenced bacterial and 90% of sequenced archaeal genomes.

[0021] As used herein, "coding sequence" or "encoding nucleic acid" refers to a nucleic acid (RNA or DNA molecule) comprising a nucleotide sequence that encodes a protein. The coding sequence may further comprise a start and stop signal operably linked to regulatory elements, including a promoter and polyadenylation signal, capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence may be codon-optimized.

[0022] As used herein, "complement" or "complementary" means that a nucleic acid is capable of Watson-Crick (e.g., AT / U and CG) or Hoogsteen base pairing between nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to the property shared between two nucleic acid sequences such that when aligned antiparallel to each other, the nucleotide bases at each position are complementary.

[0023] As used herein, "endogenous gene" refers to a gene that originates within an organism, tissue, or cell. An endogenous gene is native to a cell in its normal genomic and chromatin context and is not heterologous to that cell. Such cellular genes include, for example, animal genes, plant genes, bacterial genes, protozoan genes, fungal genes, mitochondrial genes, and chloroplast genes.

[0024] As used herein, "enhancer" refers to a non-coding DNA sequence containing multiple activator and repressor binding sites. Enhancers can range in length from 200 bp to 1 kb and can be located proximally, i.e., 5' upstream of the promoter or within the first intron of the gene being regulated, or distally, i.e., within the intron or intergenic region of an adjacent gene distant from the locus. Active enhancers contact promoters through DNA looping, depending on the core DNA-binding motif promoter specificity. Four to five enhancers can interact with a promoter. Similarly, enhancers can regulate more than one gene without being limited by linkage, and can "skip" adjacent genes to regulate more distant genes. Transcriptional regulation can involve elements located on a chromosome different from the one on which the promoter resides. The proximal enhancer or promoter of an adjacent gene can serve as a platform for recruiting more distal elements.

[0025] As used herein, "fusion protein" refers to a chimeric protein created through the joining of two or more genes that originally encoded separate proteins. Translation of the fusion gene results in a single polypeptide possessing functional properties from each of the original proteins.

[0026] As used herein, " gene construct " refers to the DNA or RNA molecule that comprises the nucleotide sequence that encodes protein.The coding sequence comprises the start and stop signal that is operably linked to the regulatory elements, including promoter and polyadenylation signal, that can induce expression in the cells of the individual that nucleic acid molecule is administered to.As used herein, the term " expressible form " refers to the gene construct that contains the necessary regulatory elements that are operably linked to the coding sequence that encodes protein, so that when present in the cells of the individual, the coding sequence will be expressed.

[0027] "Histone acetyltransferase" or "HAT" are used interchangeably herein to refer to enzymes that acetylate conserved lysine amino acids on histone proteins by transferring an acetyl group from acetyl-CoA to form ε-N-acetyllysine. DNA wraps around histones, and genes can be turned on or off by transferring an acetyl group to the histone. Generally, histone acetylation leads to transcriptional activation and, as it is associated with euchromatin, increases gene expression. Histone acetyltransferases can also acetylate non-histone proteins, such as nuclear receptors and other transcription factors, to promote gene expression.

[0028] As used herein, "identical" or "identity" in the context of two or more nucleic acid or polypeptide sequences means that the sequences have a specified percentage of residues that are identical over a specified region. This percentage can be calculated by optimally aligning the two sequences, comparing the two sequences over a specified region, determining the number of positions where identical residues exist in both sequences to obtain the number of matching positions, dividing the number of matching positions by the total number of positions in the specified region, and multiplying the result by 100 to obtain the percentage of sequence identity. If the two sequences are of different lengths, or if the alignment results in one or more sticky ends such that only a single sequence is included in a specified region of comparison, the residues of the single sequence are included in the denominator of the calculation, but not in the numerator. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or by using a computer sequence algorithm such as BLAST or BLAST 2.0.

[0029] As used herein, "nucleic acid" or "oligonucleotide" or "polynucleotide" refers to at least two nucleotides covalently linked together. The depiction of a single strand also defines the sequence of the complementary strand. Thus, nucleic acid also encompasses the complementary strand of the depicted single strand. Many variants of nucleic acid can be used for the same purpose as a given nucleic acid. Thus, nucleic acid also encompasses substantially identical nucleic acids and their complements. A single strand provides a probe that can hybridize with a target sequence under stringent hybridization conditions. Thus, nucleic acid also encompasses probes that hybridize under stringent hybridization conditions.

[0030] Nucleic acids can be single-stranded or double-stranded, or can contain portions of both double-stranded and single-stranded sequences. Nucleic acids can be DNA, both genomic and cDNA, RNA, or hybrids, where the nucleic acid can contain combinations of deoxyribonucleotides and ribonucleotides and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine hypoxanthine, isocytosine, and isoguanine. Nucleic acids can be obtained by chemical synthesis methods or by recombinant methods.

[0031] As used herein, "operably linked" means that the expression of a gene is under the control of a promoter spatially linked to the gene. The promoter can be located 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene can be approximately the same as the distance between the promoter and the gene that the promoter controls in the gene from which the promoter is derived. As is known in the art, variations in this distance can be accommodated without loss of promoter function.

[0032] As used interchangeably herein, "p300 protein," "EP300," or "E1A-binding protein p300" refers to the adenovirus E1A-associated cellular p300 transcriptional coactivator protein, encoded by the EP300 gene. p300 is a highly conserved acetyltransferase involved in a wide range of cellular processes. p300 functions as a histone acetyltransferase that regulates transcription through chromatin remodeling and is involved in the processes of cell proliferation and differentiation.

[0033] As used herein, "promoter" refers to a synthetic or naturally occurring molecule capable of conferring, activating, or promoting the expression of a nucleic acid in a cell. A promoter can contain one or more specific transcriptional regulatory sequences to further promote expression and / or alter the spatial and / or temporal expression of the same. A promoter can also contain distal enhancer or repressor elements, which can be located as far as several thousand base pairs from the start site of transcription. Promoters can be obtained from sources including viruses, bacteria, fungi, plants, insects, and animals. A promoter can constitutively or variably regulate the expression of genetic components depending on the cell, tissue, or organ in which expression occurs, the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Representative examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator-promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter, and CMV IE promoter.

[0034] As used herein, a "target enhancer" refers to an enhancer targeted by a gRNA and CRISPR / Cas9-based gene activation system. The target enhancer may be within the target region.

[0035] As used herein, "target gene" refers to any nucleotide sequence that encodes a known or putative gene product. Target genes include regulatory regions such as promoter and enhancer regions, transcription regions including coding regions, and other functional sequence regions.

[0036] As used herein, "target region" refers to the cis- or trans-regulatory region of a target gene for which a guide RNA is designed to recruit a CRISPR / Cas9-based gene activation system to regulate the epigenetic structure and enable activation of gene expression of the target gene.

[0037] As used herein, "target regulatory element" refers to a regulatory element targeted by a gRNA and CRISPR / Cas9-based gene activation system. The target regulatory element may be within the target region.

[0038] As used herein, the term "transcribed region" refers to a region of DNA that is transcribed into a single-stranded RNA molecule known as messenger RNA, which transfers genetic information from a DNA molecule to messenger RNA. During transcription, RNA polymerase reads the template strand in a 3' to 5' direction and synthesizes RNA in a 5' to 3' direction. The mRNA sequence is complementary to the DNA strand.

[0039] "Transcription start site" or "TSS," used interchangeably herein, refers to the first nucleotide of a transcribed DNA sequence at which RNA polymerase begins synthesis of an RNA transcript.

[0040] As used herein, "transgene" refers to a gene or genetic material containing a gene sequence that is isolated from one organism and introduced into another organism. This non-native segment of DNA can retain the ability to produce RNA or protein in transgenic organisms, or can change the normal function of the genetic code of transgenic organisms. The introduction of a transgene can potentially change the phenotype of an organism.

[0041] As used herein, "trans-regulatory element" refers to a region of non-coding DNA that regulates the transcription of a gene distant from the gene it is transcribed from. The trans-regulatory element can be on the same or a different chromosome as the target gene.

[0042] As used herein, "variant" with respect to a nucleic acid means (i) a portion or fragment of a reference nucleotide sequence; (ii) the complement of a reference nucleotide sequence or a portion thereof; (iii) a nucleic acid that is substantially identical to a reference nucleic acid or its complement; or (iv) a nucleic acid that hybridizes to a reference nucleic acid under stringent conditions, its complement, or a sequence substantially identical thereto.

[0043] A "variant" with respect to a peptide or polypeptide refers to a protein that differs in amino acid sequence by amino acid insertion, deletion, or conservative substitution, but retains at least one biological activity. A variant can also refer to a protein having an amino acid sequence substantially identical to a reference protein having an amino acid sequence that retains at least one biological activity. Conservative amino acid substitutions, i.e., replacing an amino acid with another amino acid of similar properties (e.g., hydrophilicity, degree, and distribution of charged regions), are recognized in the art as generally involving minor changes. These minor changes can be determined, to some extent, by considering the hydropathic index of the amino acid, as understood in the art. Kyte et al., J. Mol. Biol. 157:105-132 (1982). The hydropathic index of an amino acid is based on consideration of its hydrophobicity and charge. It is known in the art that amino acids with similar hydropathic indices can be substituted and still retain protein function. In one embodiment, amino acids with hydropathic indices of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that will result in proteins that retain biological function. Consideration of amino acid hydrophilicity in the context of a peptide allows for the calculation of the maximum local average hydrophilicity of the peptide. Substitutions can be made with amino acids having hydrophilicity values ​​within ±2 of each other. Both the hydrophobicity index and hydrophilicity value of an amino acid are influenced by the specific side chain of that amino acid. Consistent with this finding, it will be understood that amino acid substitutions that are compatible with biological function depend on the relative similarity of amino acids, particularly their side chains, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.

[0044] As used herein, "vector" refers to a nucleic acid sequence containing an origin of replication. A vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. A vector can be a DNA or RNA vector. A vector can be a self-replicating extrachromosomal vector, preferably a DNA plasmid. For example, a vector can encode a CRISPR / Cas9-based acetyltransferase having the amino acid sequence of SEQ ID NO: 140, 141, or 149 and / or at least one gRNA nucleotide sequence of any one of SEQ ID NOs: 23-73, 188-223, or 224-254.

[0045] 2. CRISPR / Cas9-based gene activation system Provided herein is a CRISPR / Cas9-based gene activation system for use in activating gene expression of a target gene. The CRISPR / Cas9-based gene activation system comprises a fusion protein of a Cas9 protein, such as dCas9, which does not have nuclease activity, and a histone acetyltransferase or a histone acetyltransferase effector domain. Histone acetylation, performed by histone acetyltransferase (HAT), plays a fundamental role in regulating chromatin dynamics and transcriptional regulation. Histone acetyltransferase proteins release DNA from its heterochromatic state, allowing sustained and robust gene expression by endogenous cellular machinery. The recruitment of acetyltransferase to target sites on the genome by dCas9 can directly regulate epigenetic structure.

[0046] The CRISPR / Cas9-based gene activation system catalyzes the acetylation of histone H3 lysine 27 at its target site, resulting in robust transcriptional activation of target genes from promoters and proximal and distal enhancers. The CRISPR / Cas9-based gene activation system is highly specific and can be directed to target genes using as little as one guide RNA. The CRISPR / Cas9-based gene activation system can activate the expression of a gene or a family of genes by targeting enhancers at distant locations within the genome.

[0047] a) CRISPR system The CRISPR system is a microbial nuclease system involved in defense against invading phages and plasmids, providing a form of adaptive immunity. CRISPR loci in microbial hosts contain a combination of CRISPR-associated (Cas) genes and non-coding RNA elements that can program the specificity of CRISPR-mediated nucleic acid cleavage. Short segments of foreign DNA, called spacers, are integrated between CRISPR repeats in the genome and serve as a "memory" of past exposure. Cas9 forms a complex with the 3' end of a single guide RNA ("sgRNA"). This protein-RNA pair recognizes its genomic target through complementary base pairing between the 5' end of the sgRNA sequence and a predefined 20-bp DNA sequence known as the protospacer. This complex is guided to the homologous locus of pathogen DNA, i.e., the protospacer and protospacer-adjacent motif (PAM) in the pathogen genome, via an encoded region in the CRISPR RNA ("crRNA"). The non-coding CRISPR array is transcribed and cleaved within the direct repeats to produce short crRNAs containing individual spacer sequences, which guide the Cas nuclease to the target site (protospacer). Simply swapping the 20-bp recognition sequence of the expressed chimeric sgRNA allows the Cas9 nuclease to be directed to new targets in the genome. The CRISPR spacers are used to recognize and silence exogenous genetic elements in a manner similar to RNAi in eukaryotes.

[0048] Three classes of CRISPR systems are known: type I, type II, and type III effector systems. Type II effector systems use a single effector enzyme, Cas9, to perform targeted DNA double-strand breaks and cleave dsDNA in four sequential steps. Compared with type I and type III effector systems, which require multiple different effectors acting as a complex, type II effector systems can function in alternative contexts, such as eukaryotic cells. Type II effector systems consist of a long precursor-crRNA (transcribed from a spacer-containing CRISPR locus), Cas9 protein, and tracrRNA (involved in precursor-crRNA processing). The tracrRNA hybridizes to the repeat region separating the spacer of the precursor-crRNA, thus initiating dsRNA cleavage by endogenous RNase III. This cleavage is followed by a second cleavage event within each spacer by Cas9, producing the tracrRNA and mature crRNA that remains associated with Cas9, forming the Cas9:crRNA-tracrRNA complex.

[0049] A genetically engineered form of the Streptococcus pyogenes type II effector system has been shown to function in human cells for genome engineering. In this system, the Cas9 protein is guided to a target site in the genome by a synthetically reconstituted "guide RNA" ("gRNA," also used interchangeably herein with chimeric sgRNA, and generally a crRNA-tracrRNA fusion that obviates the need for RNase III and crRNA processing).

[0050] The Cas9:crRNA-tracrRNA complex unwinds the DNA duplex and searches for a sequence match with the crRNA, resulting in cleavage. Target recognition occurs upon detection of complementarity between the "protospacer" sequence in the target DNA and the remaining spacer sequence in the crRNA. Cas9 mediates cleavage of the target DNA if the correct protospacer adjacent motif (PAM) is also present at the 3' end of the protospacer. For protospacer targeting, the sequence must immediately follow the protospacer adjacent motif (PAM), a short sequence recognized by the Cas9 nuclease required for DNA cleavage. Alternative type II systems have different PAM requirements. The Streptococcus pyogenes (S. pyogenes) CRISPR system can have a PAM sequence for this Cas9 (SpCas9) as 5'-NRG-3' (where R is either A or G), characterizing the specificity of this system in human cells. A unique feature of the CRISPR / Cas9 system is its straightforward ability to simultaneously target multiple different genomic loci by co-expressing a single Cas9 protein with two or more sgRNAs. For example, in genetically engineered systems, the Streptococcus pyogenes type II system naturally prefers the use of the "NGG" sequence (where "N" can be any nucleotide), but also accepts other PAM sequences such as "NAG" (Hsu et al., Nature Biotechnology (2013) doi:10.1038 / nbt.2647). Similarly, the Cas9 from Neisseria meningitidis (NmCas9) typically has a natural PAM sequence of NNNNGATT, but is active across a variety of PAMs, including the highly degenerate NNNNGNNN PAM (Esvelt et al., Nature Methods (2013) doi:10.1038 / nmeth.2681).

[0051] Cas9 CRISPR / Cas9-based gene activation systems can include a Cas9 protein or a Cas9 fusion protein. The Cas9 protein is a nucleic acid-cleaving endonuclease encoded by the CRISPR locus and included in type II CRISPR systems. The Cas9 protein can be derived from any bacterial or archaeal species, such as Streptococcus pyogenes, Streptococcus thermophiles, or Neisseria meningitidis. The Cas9 protein can be mutated to inactivate its nuclease activity. In some embodiments, an inactivated Cas9 protein from Streptococcus pyogenes (iCas9, also known as "dCas9"; SEQ ID NO: 1) can be used. As used herein, "iCas9" and "dCas9" both refer to a Cas9 protein that has the amino acid substitutions D10A and H840A and has inactivated its nuclease activity. In some embodiments, an inactivated Cas9 protein from Neisseria meningitidis can be used, such as NmCas9, which has the amino acid sequence of SEQ ID NO: 10.

[0052] Histone acetyltransferase (HAT) proteins CRISPR / Cas9-based gene activation systems can include histone acetyltransferase proteins, such as p300 protein, CREB-binding protein (CBP; a p300 analog), GCN5, or PCAF, or fragments thereof. p300 proteins regulate the activity of many genes in tissues throughout the body. p300 proteins regulate cell growth and division, promote cell maturation and specialized functions (differentiation), and play a role in preventing the growth of cancerous tumors. p300 proteins can activate transcription by linking transcription factors to the protein complex that carries out transcription in the cell nucleus. p300 proteins also function as histone acetyltransferases, regulating transcription through chromatin remodeling.

[0053] The histone acetylase protein can comprise a human p300 protein or a fragment thereof. The histone acetylase protein can comprise a wild-type human p300 protein or a mutant of the human p300 protein, or a fragment thereof. The histone acetylase protein can comprise the lysine-acetyltransferase core domain of the human p300 protein, i.e., p300 HAT Core (also known as "p300 Core"). In some embodiments, the histone acetylase protein comprises the amino acid sequence of SEQ ID NO: 2 or 3.

[0054] dCas9 p300 Core The CRISPR / Cas9-based gene activation system can include a histone acetylation effector domain. The histone acetylation effector domain can be the histone acetyltransferase (HAT) catalytic core domain of the human E1A-associated protein p300 (also referred to herein as "p300 Core"). In some embodiments, the p300 Core comprises amino acids 1048-1664 of SEQ ID NO:2 (i.e., SEQ ID NO:3). In some embodiments, the CRISPR / Cas9-based gene activation system can include the dCas9 of SEQ ID NO:141.p300 Core Fusion protein or Nm-dCas9 of SEQ ID NO: 149 p300 Core The fusion protein contains p300 Core, which can acetylate lysine 27 on histone H3 (H3K27ac) and provide H3K27ac enrichment.

[0055] dCas9 p300 Core Fusion proteins are a powerful and easily programmable tool for synthetically manipulating acetylation at targeted endogenous loci, resulting in the regulation of genes regulated by proximal and distal enhancers. Fusions of the catalytic core domain of p300 with dCas9 can result in substantially higher transactivation of downstream genes than direct fusions of the full-length p300 protein, despite robust protein expression. p300 Core Fusion proteins can also be used, for example, in the context of an Nm-dCas9 scaffold, particularly in the distal enhancer region (where dCas9 VP64 exhibited little, if any, measurable downstream transcriptional activity), dCas9 VP64 Furthermore, dCas9 may exhibit increased transactivation ability compared to p300 Core dCas9 exhibits precise and robust genome-wide transcriptional specificity. p300 Core may be capable of potent transcriptional activation and simultaneous enrichment of acetylation at promoters targeted by epigenetically modified enhancers.

[0056] dCas9 p300 Corecan activate gene expression through a single gRNA that targets and binds to a promoter and / or characterized enhancer. This technology also provides the ability to artificially transactivate genes distal to putative and known regulatory regions, facilitating transactivation through the application of a single programmable effector and a single target site. These capabilities allow for multiplexing to simultaneously target several promoters and / or enhancers. Mammalian-origin p300 may offer advantages over viral-derived effector domains for in vivo applications by minimizing the potential for immunogenicity.

[0057] gRNA CRISPR / Cas9-based gene activation systems can include at least one gRNA that targets a specific nucleic acid sequence. The gRNA provides the targeting for the CRISPR / Cas9-based gene activation system. The gRNA is a fusion of two non-coding RNAs: crRNA and tracrRNA. The sgRNA can target any desired DNA sequence by replacing the sequence encoding the 20-bp protospacer, which confers targeting specificity through complementary base pairing with the desired DNA target. The gRNA mimics the naturally occurring crRNA:tracrRNA duplex found in type II effector systems. This duplex can contain, for example, a 42-nucleotide crRNA and a 75-nucleotide tracrRNA and acts as a guide for Cas9.

[0058] The gRNA can target and bind to a target region of a target gene. The target region can be a cis-regulatory region or a trans-regulatory region of the target gene. In some embodiments, the target region is a distal or proximal cis-regulatory region of the target gene. The gRNA can target and bind to a cis-regulatory region or a trans-regulatory region of the target gene. In some embodiments, the gRNA can target and bind to an enhancer region, a promoter region, or a transcription region of the target gene. For example, the gRNA can target and bind to at least one of the following target regions: the HS2 enhancer of the human β-globin locus, the distal regulatory region (DRR) of the MYOD gene, the core enhancer (CE) of the MYOD gene, the proximal (PE) enhancer region of the OCT4 gene, or the distal (DE) enhancer region of the OCT4 gene. In some embodiments, the target region can be a viral promoter, such as an HIV promoter.

[0059] The target region can include a target enhancer or a target regulatory element. In some embodiments, the target enhancer or the target regulatory element controls the gene expression of several target genes. In some embodiments, the target enhancer or the target regulatory element controls a cellular phenotype involved in the gene expression of one or more target genes. In some embodiments, the identity of one or more target genes is known. In some embodiments, the identity of one or more target genes is unknown. The CRISPR / Cas9-based gene activation system allows for the determination of the identity of these unknown genes involved in the cellular phenotype. Examples of cellular phenotypes include, but are not limited to, T cell phenotype, cell differentiation such as hematopoietic cell differentiation, oncogenesis, immunoregulation, cellular response to stimuli, cell death, cell growth, drug resistance, or drug sensitivity.

[0060] In some embodiments, at least one gRNA can target and bind to a target enhancer or target regulatory element, thereby activating the expression of one or more genes. For example, 1 to 20 genes, 1 to 15 genes, 1 to 10 genes, 1 to 5 genes, 2 to 20 genes, 2 to 15 genes, 2 to 10 genes, 2 to 5 genes, 5 to 20 genes, 5 to 15 genes, or 5 to 10 genes are activated by at least one gRNA. In some embodiments, at least 1 gene, at least 2 genes, at least 3 genes, at least 4 genes, at least 5 genes, at least 6 genes, at least 7 genes, at least 8 genes, at least 9 genes, at least 10 genes, at least 11 genes, at least 12 genes, at least 13 genes, at least 14 genes, at least 15 genes, or at least 20 genes are activated by at least one gRNA.

[0061] CRISPR / Cas9-based gene activation systems can activate genes both proximal and distal to the transcription start site (TSS). CRISPR / Cas9-based gene activation systems can activate genes at locations at least about 1 base pair to about 100,000 base pairs, at least about 100 base pairs to about 100,000 base pairs, at least about 250 base pairs to about 100,000 base pairs, at least about 500 base pairs to about 100,000 base pairs, at least about 1,000 base pairs to about 100,000 base pairs, at least about 2,000 base pairs to about 100,000 base pairs, at least about 5,000 base pairs to about 100,000 base pairs, or at least about 10,000 base pairs from the TSS. base pairs to about 100,000 base pairs, at least about 20,000 base pairs to about 100,000 base pairs, at least about 50,000 base pairs to about 100,000 base pairs, at least about 75,000 base pairs to about 100,000 base pairs, at least about 1 base pair to about 75,000 base pairs, at least about 100 base pairs to about 75,000 base pairs, at least about 250 base pairs to about 75,000 base pairs, at least about 500 base pairs to about 75,000 base pairs, at least about 1,000 base pairs to about 75,000 base pairs, In all, the amino acid sequence is from about 2,000 base pairs to about 75,000 base pairs, at least about 5,000 base pairs to about 75,000 base pairs, at least about 10,000 base pairs to about 75,000 base pairs, at least about 20,000 base pairs to about 75,000 base pairs, at least about 50,000 base pairs to about 75,000 base pairs, at least about 1 base pair to about 50,000 base pairs, at least about 100 base pairs to about 50,000 base pairs, at least about 250 base pairs to about 50,000 base pairs, at least about 500 base pairs to about 50,000 base pairs base pairs, at least about 1,000 base pairs to about 50,000 base pairs, at least about 2,000 base pairs to about 50,000 base pairs, at least about 5,000 base pairs to about 50,000 base pairs, at least about 10,000 base pairs to about 50,000 base pairs, at least about 20,000 base pairs to about 50,000 base pairs, at least about 1 base pair to about 25,000 base pairs, at least about 100 base pairs to about 25,000 base pairs, at least about 250 base pairs to about 25,000 base pairs, at least about 500 base pairs to about 25,000 base pairs, at least about 1,000 base pairs to about 25,000 base pairs, at least about 2,000 base pairs to about 25,000 base pairs, at least about 5,000 base pairs to about 25,000 base pairs, at least about 10,000 base pairs to about 25,000 base pairs, at least about 20,000 base pairs to about 25,000 base pairs, at least about 1 base pair to about 10,000 base pairs, at least about 100 base pairs to about 10,000 base pairs, at least about 250 base pairs to about 10,000 base pairs, at least about 500 base pairs to about 10,000 base pairs, In any case, the target region can be from about 1,000 base pairs to about 10,000 base pairs, at least about 2,000 base pairs to about 10,000 base pairs, at least about 5,000 base pairs to about 10,000 base pairs, at least about 1 base pair to about 5,000 base pairs, at least about 100 base pairs to about 5,000 base pairs, at least about 250 base pairs to about 5,000 base pairs, at least about 500 base pairs to about 5,000 base pairs, at least about 1,000 base pairs to about 5,000 base pairs, or at least about 2,000 base pairs to about 5,000 base pairs upstream. CRISPR / Cas9-based gene activation systems can target regions that are at least about 1 base pair, at least about 100 base pairs, at least about 500 base pairs, at least about 1,000 base pairs, at least about 1,250 base pairs, at least about 2,000 base pairs, at least about 2,250 base pairs, at least about 2,500 base pairs, at least about 5,000 base pairs, at least about 10,000 base pairs, at least about 11,000 base pairs, at least about 20,000 base pairs, at least about 30,000 base pairs, at least about 46,000 base pairs, at least about 50,000 base pairs, at least about 54,000 base pairs, at least about 75,000 base pairs, or at least about 100,000 base pairs upstream from the TSS.

[0062] CRISPR / Cas9-based gene activation systems can target a region that is at least about 1 base pair to at least about 500 base pairs, at least about 1 base pair to at least about 250 base pairs, at least about 1 base pair to at least about 200 base pairs, at least about 1 base pair to at least about 100 base pairs, at least about 50 base pairs to at least about 500 base pairs, at least about 50 base pairs to at least about 250 base pairs, at least about 50 base pairs to at least about 200 base pairs, at least about 50 base pairs to at least about 100 base pairs, at least about 100 base pairs to at least about 500 base pairs, at least about 100 base pairs to at least about 250 base pairs, or at least about 100 base pairs to at least about 200 base pairs downstream from the TSS. The CRISPR / Cas9 based gene activation system may comprise at least about 1 base pair, at least about 2 base pairs, at least about 3 base pairs, at least about 4 base pairs, at least about 5 base pairs, at least about 10 base pairs, at least about 15 base pairs, at least about 20 base pairs, at least about 25 base pairs, at least about 30 base pairs, at least about 40 base pairs, at least about 50 base pairs, at least about 60 base pairs, at least about 70 base pairs, at least about 80 base pairs, at least about 90 base pairs, at least about 100 base pairs, at least about 150 ...150 base pairs, at least about 150 base pairs, at least about 150 base pairs, at least about 160 base pairs, at least about 160 base pairs, at least about 170 base pairs, at least about 180 base pairs, at least about 180 base pairs, at least about 190 base pairs, at least about 200 base pairs, at least about 210 base pairs, at least about 220 base pairs, at least about 230 base pairs, at least about 24 A region that is at least about 100 base pairs, at least about 110 base pairs, at least about 120, at least about 130, at least about 140 base pairs, at least about 150 base pairs, at least about 160 base pairs, at least about 170 base pairs, at least about 180 base pairs, at least about 190 base pairs, at least about 200 base pairs, at least about 210 base pairs, at least about 220, at least about 230, at least about 240 base pairs, or at least about 250 base pairs downstream can be targeted.

[0063] In some embodiments, the CRISPR / Cas9-based gene activation system can target and bind to a target region that is on the same chromosome as the target gene but that is more than 100,000 base pairs upstream or more than 250 base pairs downstream from the TSS. In some embodiments, the CRISPR / Cas9-based gene activation system can target and bind to a target region that is on a different chromosome than the target gene.

[0064] The CRISPR / Cas9-based gene activation system can use gRNAs of various sequences and lengths. The gRNA can comprise a complementary polynucleotide sequence of the target DNA sequence followed by NGG. The gRNA can comprise a "G" at the 5' end of the complementary polynucleotide sequence. The gRNA can comprise a complementary polynucleotide sequence of at least 10 base pairs, at least 11 base pairs, at least 12 base pairs, at least 13 base pairs, at least 14 base pairs, at least 15 base pairs, at least 16 base pairs, at least 17 base pairs, at least 18 base pairs, at least 19 base pairs, at least 20 base pairs, at least 21 base pairs, at least 22 base pairs, at least 23 base pairs, at least 24 base pairs, at least 25 base pairs, at least 30 base pairs, or at least 35 base pairs of the target DNA sequence followed by NGG. The gRNA can target at least one of the promoter region, enhancer region, or transcription region of the target gene. The gRNA can comprise at least one nucleic acid sequence of SEQ ID NOs: 23-73, 188-223, or 224-254.

[0065] The CRISPR / Cas9-based gene activation system can include at least one gRNA, at least two different gRNAs, at least three different gRNAs, at least four different gRNAs, at least five different gRNAs, at least six different gRNAs, at least seven different gRNAs, at least eight different gRNAs, at least nine different gRNAs, or at least ten different gRNAs. The CRISPR / Cas9-based gene activation system can include at least one gRNA to at least ten different gRNAs, at least one gRNA to at least eight different gRNAs, at least one gRNA to at least four different gRNAs, at least two gRNAs to at least ten different gRNAs, at least two gRNAs to at least eight different gRNAs, at least two different gRNAs to at least four different gRNAs, at least four gRNAs to at least ten different gRNAs, or at least four different gRNAs to at least eight different gRNAs.

[0066] Target gene The CRISPR / Cas9-based gene activation system can be designed to target and activate the expression of any target gene. The target gene can be an endogenous gene, a transgene, or a viral gene in a cell line. In some embodiments, the target region is on a different chromosome from the target gene. In some embodiments, the CRISPR / Cas9-based gene activation system can comprise two or more gRNAs. In some embodiments, the CRISPR / Cas9-based gene activation system can comprise two or more different gRNAs. In some embodiments, the different gRNAs bind to different target regions. For example, different gRNAs can bind to target regions of different target genes, activating the expression of two or more target genes.

[0067] In some embodiments, the CRISPR / Cas9-based gene activation system is capable of activating about 1 to about 10 target genes, about 1 to about 5 target genes, about 1 to about 4 target genes, about 1 to about 3 target genes, about 1 to about 2 target genes, about 2 to about 10 target genes, about 2 to about 5 target genes, about 2 to about 4 target genes, about 2 to about 3 target genes, about 3 to about 10 target genes, about 3 to about 5 target genes, or about 3 to about 4 target genes. In some embodiments, the CRISPR / Cas9-based gene activation system is capable of activating at least 1 target gene, at least 2 target genes, at least 3 target genes, at least 4 target genes, at least 5 target genes, or at least 10 target genes. For example, the hypersensitive region 2 (HS2) enhancer region of the human β-globin locus can be targeted to activate downstream genes (HBE, HBG, HBD, and HBB).

[0068] In some embodiments, the CRISPR / Cas9-based gene activation system increases gene expression by at least about 1 fold, at least about 2 fold, at least about 3 fold, at least about 4 fold, at least about 5 fold, at least about 6 fold, at least about 7 fold, at least about 8 fold, at least about 9 fold, at least about 10 fold, at least 15 fold, at least 20 fold, at least 30 fold, at least 40 fold, at least 50 fold, at least 60 fold, at least 70 fold, at least 80 fold, at least about 100 fold, at least about 150 ...0 fold, at least 300 fold, at least 400 fold, at least 500 fold, at least 600 fold, at least 700 fold, at least 800 fold, at least about 1000 fold, at least about 1500 fold, at least 2000 fold, at least about 1500 fold, at least 3000 fold, at least 4000 fold, at least 5000 fold, at least 6000 fold, at least 1000 fold, at least 1000 fold, at least 1000 In some embodiments, the control level of gene expression of the target gene induces gene expression of the target gene by at least 90 fold, at least 100 fold, at least about 110 fold, at least 120 fold, at least 130 fold, at least 140 fold, at least 150 fold, at least 160 fold, at least 170 fold, at least 180 fold, at least 190 fold, at least 200 fold, at least about 300 fold, at least 400 fold, at least 500 fold, at least 600 fold, at least 700 fold, at least 800 fold, at least 900 fold, or at least 1000 fold. The control level of gene expression of the target gene can be the level of gene expression of the target gene in cells that have not been treated with any CRISPR / Cas9-based gene activation system.

[0069] The target gene can be a mammalian gene. For example, the CRISPR / Cas9-based gene activation system can target mammalian genes such as IL1RN, MYOD1, OCT4, HBE, HBG, HBD, HBB, MYOCD (myocardin), PAX7 (paired box protein Pax-7), FGF1 (fibroblast growth factor-1) genes, for example, FGF1A, FGF1B, and FGF1C. Other target genes include, but are not limited to, Atf3, Axud1, Btg2, c-Fos, c-Jun, Cxcl1, Cxcl2, Edn1, Ereg, Fos, Gadd45b, Ier2, Ier3, Ifrd1, Il1b, Il6, Irf1, Junb, Lif, Nfkbia, Nfkbiz, Ptgs2, Slc25a25, Sqstm1, Tieg, Tnf, Tnfaip3, Zfp36, Birc2, Ccl2, Ccl20, Ccl7, Cebpd, Ch25h, CSF1, Cx3cl1, Cxcl10, Cxcl5, Gch, Icam1, and Ifi47. , Ifngr2, Mmp10, Nfkbie, Npal1, p21, Relb, Ripk2, Rnd1, S1pr3, Stx11, Tgtp, Tlr2, Tmem140, Tnfaip2, Tnfrsf6, Vcam1, 1110004C05Rik (GenBank accession number BC010291), Abca1, AI561871 (GenBank accession number BI143915), AI882074 (GenBank accession number BB730912), Arts1, AW049765 (GenBank accession number BC026642).1), C3, Casp4, Ccl5, Ccl9, Cdsn, Enpp2, Gbp2, H2-D1, H2-K, H2-L, Ifit1, Ii, Il13ra1, Il1rl1, Lcn2, Lhfpl2, LOC677168 (GenBank accession number AK019325), Mmp13, Mmp3, Mt2, Naf1, Ppicap, Prnd, Psmb10, Saa3, Serpina3g, Serpinf1, Sod3, Stat1, Tapbp, U90926 (GenBank accession number NM_020562), Ubd, A2AR (adenosine A2A receptor), B7-H3 (also known as CD276), B7-H4 (also known as VTCN1), BTLA (B and T lymphocyte attenuator) Attenuator; also known as CD272), CTLA-4 (Cytotoxic T-Lymphocyte-Associated protein 4; also known as CD152), IDO (Indoleamine 2,3-dioxygenase), KIR (Killer-cell Immunoglobulin-like Receptor), LAG3 (Lymphocyte Activation Gene-3), PD-1 (Programmed Death 1 (PD-1) Receptor), TIM-3 (T-cell Immunoglobulin Domain and Mucin Domain 3), and VISTA (V-domain Ig suppressor of T cell activation).

[0070] Compositions for gene activation The present invention is directed to a composition for activating the gene expression of a target gene, a target enhancer, or a target regulatory element in a cell or a subject.The composition can comprise a CRISPR / Cas9-based gene activation system as disclosed above.The composition can also comprise a viral delivery system.For example, the viral delivery system can comprise an adeno-associated virus vector or a modified lentivirus vector.

[0071] Methods for introducing nucleic acids into host cells are known in the art, and any known method can be used to introduce nucleic acids (e.g., expression constructs) into cells. Suitable methods include, for example, viral or bacteriophage infection, gene transfer, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated gene transfer, DEAE-dextran-mediated gene transfer, liposome-mediated gene transfer, particle gun technology, calcium phosphate precipitation, direct microinjection, nanoparticle-mediated nucleic acid delivery, and the like. In some embodiments, compositions can be delivered by mRNA delivery and ribonucleoprotein (RNP) complex delivery.

[0072] a) Constructs and Plasmids The composition can include a genetic construct encoding the CRISPR / Cas9-based gene activation system as disclosed herein, as described above. The genetic construct, such as a plasmid or expression vector, can include a nucleic acid encoding the CRISPR / Cas9-based gene activation system, such as a CRISPR / Cas9-based acetyltransferase and / or at least one gRNA. The composition can include a genetic construct encoding a modified AAV vector, as described above, and a nucleic acid sequence encoding the CRISPR / Cas9-based gene activation system as disclosed herein. The genetic construct, such as a plasmid, can include a nucleic acid encoding the CRISPR / Cas9-based gene activation system. The composition can include a genetic construct encoding a modified lentiviral vector, as described above. The genetic construct, such as a plasmid, can include a nucleic acid encoding the CRISPR / Cas9-based acetyltransferase and at least one sgRNA. The genetic construct can exist in a cell as a functional extrachromosomal molecule. The genetic construct can be a linear minichromosome (including the centromere), a telomere, or a plasmid or cosmid.

[0073] Gene constructs can also be part of the genome of recombinant virus vectors, including recombinant lentiviruses, recombinant adenoviruses, and recombinant adenovirus-associated viruses. Gene constructs can be part of the genetic material in attenuated living microorganisms or recombinant microbial vectors that live in cells. Gene constructs can include regulatory elements for gene expression of nucleic acid coding sequences. Regulatory elements can be promoters, enhancers, initiation codons, stop codons, or polyadenylation signals.

[0074] The nucleic acid sequence can constitute a gene construct (which can be a vector). The vector can be capable of expressing a fusion protein, such as a CRISPR / Cas9-based gene activation system, in mammalian cells. The vector can be recombinant. The vector can contain a heterologous nucleic acid encoding a fusion protein, such as a CRISPR / Cas9-based gene activation system. The vector can be a plasmid. The vector can be useful for introducing a nucleic acid encoding a CRISPR / Cas9-based gene activation system into a cell, where the transformed host cell is cultured and maintained under conditions that allow expression of the CRISPR / Cas9-based gene activation system.

[0075] Coding sequences can be optimized for stability and high levels of expression. In some cases, codons are selected to reduce secondary structure formation in RNA, such as those formed due to intramolecular binding.

[0076] The vector can contain a heterologous nucleic acid encoding a CRISPR / Cas9-based gene activation system, and can further contain a start codon that can be upstream of the CRISPR / Cas9-based gene activation system coding sequence and a stop codon that can be downstream of the CRISPR / Cas9-based gene activation system coding sequence. The start and stop codons can be in frame with the CRISPR / Cas9-based gene activation system coding sequence. The vector can also contain a promoter operably linked to the CRISPR / Cas9-based gene activation system coding sequence. The CRISPR / Cas9-based gene activation system can be under light-inducible or chemical-inducible control to enable dynamic control of gene activation in space and time. The promoter operably linked to the CRISPR / Cas9-based gene activation system coding sequence can be a promoter derived from simian virus 40 (SV40), a mouse mammary tumor virus (MMTV) promoter, a human immunodeficiency virus (HIV) promoter, such as the bovine immunodeficiency virus (BIV) long terminal repeat (LTR) promoter, a Moloney virus promoter, an avian leukosis virus (ALV) promoter, a cytomegalovirus (CMV) promoter, such as the CMV immediate early promoter, an Epstein-Barr virus (EBV) promoter, or a Rous sarcoma virus (RSV) promoter. The promoter can also be a promoter derived from a human gene, such as human ubiquitin C (hUbC), human actin, human myosin, human hemoglobin, human muscle creatine, or human metallothionein. The promoter can be a natural or synthetic tissue-specific promoter, such as a muscle- or skin-specific promoter. Examples of such promoters are described in U.S. Patent Application Publication No. 20040175727, the entire contents of which are incorporated herein.

[0077] The vector can also include a polyadenylation signal that can be downstream of the CRISPR / Cas9-based gene activation system.The polyadenylation signal can be an SV40 polyadenylation signal, an LTR polyadenylation signal, a bovine growth hormone (bGH) polyadenylation signal, a human growth hormone (hGH) polyadenylation signal, or a human β-globin polyadenylation signal.The SV40 polyadenylation signal can be the polyadenylation signal from the pCEP4 vector (Invitrogen, San Diego, CA).

[0078] The vector can also include a CRISPR / Cas9-based gene activation system, i.e., an enhancer upstream of the CRISPR / Cas9-based acetyltransferase coding sequence or sgRNA. The enhancer can be essential for DNA expression. The enhancer can be human actin, human myosin, human hemoglobin, human muscle creatine, or a viral enhancer, such as those derived from CMV, HA, RSV, or EBV. Polynucleotide functional enhancers are described in U.S. Patent Nos. 5,593,972, 5,962,428, and WO 94 / 016737, the contents of each of which are incorporated by reference in their entirety. The vector can also include a mammalian origin of replication to maintain the vector extrachromosomally, allowing multiple copies of the vector to be produced in cells. The vector can also include regulatory sequences, which can be sufficiently adapted for gene expression in mammalian or human cells to which the vector is administered. The vector may also include a reporter gene such as green fluorescent protein ("GFP") and / or a selectable marker such as hygromycin ("Hygro").

[0079] The vector can be an expression vector or system for producing the protein by conventional techniques and readily available starting materials (Sambrook et al., Molecular Cloning and Laboratory Manual, Second Ed., Cold Spring Harbor (1989), incorporated by reference in its entirety). In some embodiments, the vector can comprise a nucleic acid sequence encoding a CRISPR / Cas9-based gene activation system, including a nucleic acid sequence encoding a CRISPR / Cas9-based acetyltransferase and at least one gRNA comprising at least one nucleic acid sequence of SEQ ID NOs: 23-73, 188-223, or 224-254.

[0080] combination CRISPR / Cas9-based gene activation system compositions can be combined with orthogonal dCas9s, TALEs, and zinc finger proteins to facilitate the study of independent targeting of specific effector functions to different gene loci. In some embodiments, CRISPR / Cas9-based gene activation system compositions can be multiplexed with a variety of activators, repressors, and epigenetic modifiers to precisely control cellular phenotypes or decipher complex networks of gene regulation.

[0081] How to use The potential applications of the CRISPR / Cas9-based gene activation system are diverse across many fields of science and biotechnology. It can be used to activate gene expression of a target gene or to target enhancers or regulatory elements. It can be used to transdifferentiate cells and / or activate genes in the context of cell and gene therapy, gene reprogramming, and regenerative medicine. It can also be used to reprogram lineage specification. Activation of endogenous genes encoding key regulators of cell fate (rather than forced overexpression of these factors) could potentially provide a more rapid, efficient, stable, or specific method for gene reprogramming and transdifferentiation. The CRISPR / Cas9-based gene activation system may offer a greater diversity of transcriptional activators to complement other means for regulating mammalian gene expression. CRISPR / Cas9-based gene activation systems can be used to compensate for genetic abnormalities, inhibit angiogenesis, inactivate oncogenes, activate silenced tumor suppressors, regenerate tissues, or reprogram genes.

[0082] Methods for activating gene expression The present disclosure provides a mechanism for activating the expression of a target gene based on targeting histone acetyltransferase to a target region via a CRISPR / Cas9-based gene activation system as described above. The CRISPR / Cas9-based gene activation system can activate a silenced gene. The CRISPR / Cas9-based gene activation system target region upstream of the TSS of the target gene substantially induces gene expression of the target gene. The polynucleotide encoding the CRISPR / Cas9-based gene activation system can also be directly introduced into cells.

[0083] The method can include administering to a cell or a subject a CRISPR / Cas9-based gene activation system, a composition of a CRISPR / Cas9-based gene activation system, or one or more polynucleotides or vectors encoding the CRISPR / Cas9-based gene activation system, as described above.The method can include administering to a mammalian cell or a subject a CRISPR / Cas9-based gene activation system, a composition of a CRISPR / Cas9-based gene activation system, or one or more polynucleotides or vectors encoding the CRISPR / Cas9-based gene activation system, as described above.

[0084] Pharmaceutical Composition The CRISPR / Cas9-based gene activation system can be in a pharmaceutical composition. The pharmaceutical composition can contain about 1 ng to about 10 mg of DNA encoding the CRISPR / Cas9-based gene activation system. The pharmaceutical composition according to the present invention is formulated according to the mode of administration to be used. When the pharmaceutical composition is an injectable pharmaceutical composition, it is sterile, pyrogen-free, and particulate-free. It is preferable to use an isotonic formulation. Generally, additives for isotonicity can include sodium chloride, dextrose, mannitol, sorbitol, and lactose. In some cases, an isotonic solution such as phosphate-buffered saline is preferred. Stabilizers include gelatin and albumin. In some embodiments, a vasoconstrictor is added to the formulation.

[0085] Pharmaceutical compositions containing CRISPR / Cas9-based gene activation systems can further comprise a pharmaceutically acceptable excipient. Pharmaceutically acceptable excipients can be functional molecules as vehicles, adjuvants, carriers, or diluents. Pharmaceutically acceptable excipients can be gene transfer promoters (which can include surfactants), such as immunostimulating complexes (ISCOMS), Freund's incomplete adjuvant, LPS analogs (including monophosphoryl lipid A), muramyl peptides, quinone analogs, vesicles (such as squalene and squalene), hyaluronic acid, lipids, liposomes, calcium ions, viral proteins, polyanions, polycations, or nanoparticles, or other known gene transfer promoters.

[0086] The gene transfer promoter is a polyanion, a polycation (including poly-L-glutamic acid (LGS)), or a lipid. The gene transfer promoter is poly-L-glutamic acid, and more preferably, poly-L-glutamic acid is present in a pharmaceutical composition containing a CRISPR / Cas9-based gene activation system at a concentration of less than 6 mg / ml. The gene transfer promoter can also include surfactants such as immune stimulating complexes (ISCOMS), Freund's incomplete adjuvant, LPS analogs (including monophosphoryl lipid A), muramyl peptides, quinone analogs, vesicles such as squalene and squalene, and hyaluronic acid can also be used with the gene construct. In some embodiments, the DNA vector encoding the CRISPR / Cas9-based gene activation system can also include, for example, lipids, liposomes (including lecithin liposomes or other liposomes known in the art), DNA-liposome mixtures (see, e.g., WO 09324640), calcium ions, viral proteins, polyanions, polycations, or nanoparticles, or other known materials. Preferably, is a polyanion, polycation (including poly-L-glutamic acid (LGS)), or lipid.

[0087] Method of delivery Provided herein is a method for delivering pharmaceutical preparations of a CRISPR / Cas9-based gene activation system, providing gene constructs and / or proteins of the CRISPR / Cas9-based gene activation system. Delivery of the CRISPR / Cas9-based gene activation system can be by transfection or electroporation of the CRISPR / Cas9-based gene activation system as one or more nucleic acid molecules that are expressed in cells and delivered to the surface of the cells. CRISPR / Cas9-based gene activation system proteins can be delivered to cells. These nucleic acid molecules can be electroporated using a BioRad Gene Pulser Xcell or Amaxa Nucleofector IIb device or other electroporation device. Several different buffers can be used, including BioRad electroporation solution, Sigma phosphate-buffered saline (product #D8537) (PBS), Invitrogen OptiMEM I (OM), or Amaxa Nucleofector solution V (NV). Transfection can include transfection reagents such as Lipofecamine 2000.

[0088] Vectors encoding CRISPR / Cas9-based gene activation system proteins can be delivered to mammals by DNA injection (also known as DNA vaccination), with or without in vivo electroporation, liposome-mediated delivery, nanoparticle-facilitated delivery, and / or recombinant vectors. Recombinant vectors can be delivered by any type of virus. The virus type can be recombinant lentivirus, recombinant adenovirus, and / or recombinant adeno-associated virus.

[0089] To induce gene expression of a target gene, nucleotides encoding a CRISPR / Cas9-based gene activation system protein can be introduced into cells. For example, one or more nucleotide sequences encoding a CRISPR / Cas9-based gene activation system induced by a target gene can be introduced into mammalian cells. Upon delivery of the CRISPR / Cas9-based gene activation system to a cell, and upon delivery of the vector to a mammalian cell, the transgenic cell will express the CRISPR / Cas9-based gene activation system. The CRISPR / Cas9-based gene activation system can be administered to a mammal to induce or regulate gene expression of a target gene in the mammal. The mammal can be a human, non-human primate, cow, pig, sheep, goat, antelope, bison, water buffalo, bovid, deer, hedgehog, elephant, llama, alpaca, mouse, rat, or chicken, preferably a human, cow, pig, or chicken.

[0090] Route of administration The CRISPR / Cas9-based gene activation system and its composition can be administered to a subject by various routes, including orally, parenterally, sublingually, transdermally, rectally, transmucosally, topically, via inhalation, via buccal administration, intrapleurally, intravenously, intraarterially, intraperitoneally, subcutaneously, intramuscularly, intranasally, intrathecally, and intraarticularly, or a combination thereof. For veterinary use, the composition can be administered in an appropriately acceptable formulation according to standard veterinary practice. A veterinarian can easily determine the most suitable dosage regimen and administration route for a particular animal. The CRISPR / Cas9-based gene activation system and its composition can be administered by conventional syringe, needleless injection device, "microprojectile bombardment guns", or other physical methods, such as electroporation ("EP"), "hydrodynamic methods", or ultrasound. The compositions can be delivered to mammals by several techniques, including DNA injection (also called DNA vaccination) with or without in vivo electroporation, liposome-mediated, nanoparticle-facilitated, recombinant vectors such as recombinant lentiviruses, recombinant adenoviruses, and recombinant adenovirus-associated viruses.

[0091] cell type The CRISPR / Cas9-based gene activation system can be used with any type of cell. In some embodiments, the cell is a bacterial cell, a fungal cell, an archaeal cell, a plant cell, or an animal cell. In some embodiments, the cell is selected from the group consisting of, but not limited to, GM12878, K562, H1 human embryonic stem cells, HeLa-S3, HepG2, HUVEC, SK-N-SH, IMR90, A549, MCF7, HMEC or LHCM, CD14+, CD20+, primary heart or liver cells, differentiated H1 cells, 8988T, Adult_CD4_naive, Adult_CD4_Th0, Adult_CD4_Th1, AG04449, AG04450, AG09309, AG09319, AG10803, AoAF, AoSMC, BC_Adipose_UHN00001, BC_Adrenal_Gland_H12803N, BC_Bladder_01-11002, BC_Brain_H11058N, BC_Breast_02-03015, BC_Colon_ 01-11002, BC_Colon_H12817N, BC_Esophagus_01-11002, BC_Esophagus_H12817N, BC_Jejunum_H12817N, BC_Kidney_01-11002, BC_Kidney _H12817N, BC_Left_Ventricle_N41, BC_Leukocyte_UHN00204, BC_Liver_01-11002, BC_Lung_01-11002, BC_Lung_H12817N, BC_Pan creas_H12817N, BC_Penis_H12817N, BC_Pericardium_H12529N, BC_Placenta_UHN00189, BC_Prostate_Gland_H12817N, BC_Rectum _N29, BC_Skeletal_Muscle_01-11002, BC_Skeletal_Muscle_H12817N, BC_Skin_01-11002, BC_Small_Intestine_01-11002, BC_Sp leen_H12817N, BC_Stomach_01-11002, BC_Stomach_H12817N, BC_Testis_N30, BC_Uterus_BN0765, BE2_C, BG02ES, BG02ES-EBD, BJ,bone_marrow_HS27a、bone_marrow_HS5、bone_marrow_MSC、Breast_OC、Caco-2、CD20+_RO01778、CD20+_RO01794、CD34+_Mobilized、CD4+_Naive_Wb11970640、CD4+_Naive_Wb78495824、Cerebellum_OC、Cerebrum_frontal_OC、Chorion、CLL、CMK、Colo829、Colon_BC、Colon_OC、Cord_CD4_naive、Cord_CD4_Th0、Cord_CD4_Th1、Decidua、Dnd41、ECC-1、Endometrium_OC、Esophagus_BC、Fibrobl、Fibrobl_GM03348、FibroP、FibroP_AG08395、FibroP_AG08396、FibroP_AG20443、Frontal_cortex_OC、GC_B_cell、Gliobla、GM04503、GM04504、GM06990、GM08714、GM10248、GM10266、GM10847、GM12801、GM12812、GM12813、GM12864、GM12865、GM12866、GM12867、GM12868、GM12869、GM12870、GM12871、GM12872、GM12873、GM12874、GM12875、GM12878-XiMat、GM12891、GM12892、GM13976、GM13977、GM15510、GM18505、GM18507、GM18526、GM18951、GM19099、GM19193、GM19238、GM19239、GM19240、GM20000、H0287、H1-neurons、H7-hESC、H9ES、H9ES-AFP-、H9ES-AFP+、H9ES-CM、H9ES-E、H9ES-EB、H9ES-EBD、HAc、HAEpiC、HA-h、HAL、HAoAF、HAoAF_6090101.11、HAoAF_6111301.9、HAoEC、HAoEC_7071706.1、HAoEC_8061102.1、HA-sp、HBMEC、HBVP、HBVSMC、HCF、HCFaa、HCH、HCH_0011308.2P、HCH_8100808.2、HCM、HConF、HCPEpiC、HCT-116、Heart_OC、Heart_STL003、HEEpiC、HEK293、HEK293T、HEK293-T-REx、Hepatocytes、HFDPC、HFDPC_0100503.2、HFDPC_0102703.3、HFF,HFF-Myc、HFL11W、HFL24W、HGF、HHSEC、HIPEpiC、HL-60、HMEpC、HMEpC_6022801.3、HMF、hMNC-CB、hMNC-CB_8072802.6、hMNC-CB_9111701.6、hMNC-PB、hMNC-PB_0022330.9、hMNC-PB_0082430.9、hMSC-AT、hMSC-AT_0102604.12、hMSC-AT_9061601.12、hMSC-BM、hMSC-BM_0050602.11、hMSC-BM_0051105.11、hMSC-UC、hMSC-UC_0052501.7、hMSC-UC_0081101.7、HMVEC-dAd、HMVEC-dBl-Ad、HMVEC-dBl-Neo、HMVEC-dLy-Ad、HMVEC-dLy-Neo、HMVEC-dNeo、HMVEC-LBl、HMVEC-LLy、HNPCEpiC、HOB、HOB_0090202.1、HOB_0091301、HPAEC、HPAEpiC、HPAF、HPC-PL、HPC-PL_0032601.13、HPC-PL_0101504.13、HPDE6-E6E7、HPdLF、HPF、HPIEpC、HPIEpC_9012801.2、HPIEpC_9041503.2、HRCEpiC、HRE、HRGEC、HRPEpiC、HSaVEC、HSaVEC_0022202.16、HSaVEC_9100101.15、HSMM、HSMM_emb、HSMM_FSHD、HSMMtube、HSMMtube_emb、HSMMtube_FSHD、HT-1080、HTR8svn、Huh-7、Huh-7.5、HVMF、HVMF_6091203.3、HVMF_6100401.3、HWP、HWP_0092205、HWP_8120201.5、iPS、iPS_CWRU1、iPS_hFib2_iPS4、iPS_hFib2_iPS5、iPS_NIHi11、iPS_NIHi7、Ishikawa、Jurkat、Kidney_BC、Kidney_OC, LHCN-M2, LHSR, Liver_OC, Liver_STL004, Liver_STL011, LNCaP, Loucy, Lung_BC, Lung_OC, Lymphoblastoid_cell_line, M059J, MCF10A-Er-Src, MCF-7, MDA-MB-231, Medullo, Medullo_D341, Mel_2183, Melano, Monocytes-CD14+, Monocytes-CD14+_RO01746, Monocytes-CD14+_RO01826, MRT_A204, MRT_G401, MRT_TTC549, Myometr, Naive_B_cell, NB4, NH-A, NHBE, NHBE_RA, NHDF, NHDF_0060801.3, NHDF_7071701.2, NHDF-Ad, NHDF-neo, NHEK, NHEM.f_M2, NHEM.f_M2_5071302.2, NHEM.f_M2_6022001, NHEM_M2, NHEM_M2_7011001.2, NHEM_M2_7012303, NHLF, NT2-D1, Olf_neurosphere, Osteobl, ovcar-3, PANC-1, Pancreas_OC, PanIsletD, PanIslets, PBDE, PBDEFetal, PBMC, PFSK-1, pHTE, Pons_OC, PrEC, ProgFib, Prostate, Prostate_OC, Psoas_muscle_OC, Raji, RCC_7860, RPMI-7951, RPTEC, RWPE1, SAEC, SH-SY5Y, Skeletal_Muscle_BC, SkMC, SKMC, SkMC_8121902.17, SkMC_9011302, SK-N-MC, SK-N-SH_RA, Small_intestine_OC, Spleen_OC, Stellate, Stomach_BC, T_cells_CD4+, T-47D, T98G, TBEC, Th1, Th1_Wb33676984, Th1_Wb54553204, Th17, Th2, Th2_Wb33676984, Th2_Wb54553204, Treg_Wb78495824, Treg_Wb83319432, U2OS, U87, UCH-1, Urothelia, WERI-Rb-1, and WI-38, includingThe cell line may be an ENCODE cell line.

[0092] kit Provided herein is a kit that can be used to activate gene expression of a target gene. The kit includes a composition for activating gene expression, as described above, and instructions for using the composition. The instructions included in the kit can be affixed to packaging materials or included as a package insert. The instructions are typically, but not limited to, written or printed materials. Any medium capable of storing such instructions and communicating them to an end user is contemplated by the present disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic disks, tapes, cartridges, chips), optical media (e.g., CD ROMs), and the like. As used herein, the term "instructions" can include the address of an internet site that provides the instructions.

[0093] The composition for activating gene expression can include a modified AAV vector and a nucleotide sequence encoding a CRISPR / Cas9-based gene activation system as described above. The CRISPR / Cas9-based gene activation system can include a CRISPR / Cas9-based acetyltransferase as described above that specifically binds to and targets a cis- or trans-regulatory region of a target gene. The kit can include a CRISPR / Cas9-based acetyltransferase as described above that specifically binds to and targets a specific regulatory region of a target gene. [Example]

[0094] The foregoing can be better understood by reference to the following examples, which are given for purposes of illustration and not to limit the scope of the invention.

[0095] Example 1 Methods and Materials—Activators Cell Lines and Transfection. HEK293T cells were obtained from the American Tissue Collection Center (ATCC, Manassas, VA) through the Duke University Cell Culture Facility. Cells were cultured in Dulbecco's modified Eagle's medium (DMEM) supplemented with 10% FBS and 1% penicillin / streptomycin and maintained at 37°C and 5% CO2. Transfection was performed in 24-well plates using 375 ng of each dCas9 expression vector and 125 ng of equimolar pooled or individual gRNA expression vectors (mixed with Lipofectamine 2000 (Life Technologies, cat. #11668019)) according to the manufacturer's instructions. For ChIP-qPCR experiments, HEK293T cells were transfected in 15 cm dishes with Lipofectamine 2000 and 30 μg of each dCas9 expression vector and 10 μg of equimolar pooled gRNA expression vectors according to the manufacturer's instructions.

[0096] Plasmid construct: pcDNA-dCas9 VP64 (dCas9 VP64 dCas9 was inserted via the AscI / PacI restriction enzyme sites (Addgene, Plasmid #47107) (Perez-Pinera, P. et al., Nature Methods 10:973-976 (2013)). VP64 An HA epitope tag was added to dCas9 (without the effector) by removing the VP64 effector domain from pcDNA-dCas9 and using isothermal assembly (Gibson et al. Nat. Methods 6:343-345 (2009)) and including a set of annealed oligos containing the appropriate sequences as per the manufacturer's instructions (NEB cat. #2611). FLp300 (dCas9 FLp300) was used to amplify the full-length p300 from pcDNA3.1-p300 (Addgene, Plasmid #23252) (Chen et al. EMBO J. 21:6539-6548 (2002)) in two separate fragments, and these fragments were then used to construct dCas9 via isothermal assembly. VP64 It was created by cloning into the dCas9 backbone. FLp300 A substitution within the full-length p300 protein (L553M) located outside the HAT Core region in the precursor pcDNA3.1-p300 was identified during sequence verification. p300 Core (dCas9 p300 Core ) was first used to amplify amino acids 1048 to 1664 of human p300 from cDNA, and then the resulting amplicon was inserted into pCR-Blunt (pCR-Blunt p300 Core The AscI site, HA-epitope tag, and PmeI site were inserted into pCR-Blunt p300 Core pCR-Blunt (pCR-Blunt p300 Core+HA ) (Life Technologies cat. #K2700). The HA-tagged p300 Core was cloned into pCR-Blunt p300 Core+HA from dCas9 VP64 The backbone was cloned via the common AscI / PmeI restriction enzyme sites. pcDNA-dCas9 p300 Core(D1399Y) (dCas9 p300 Core(D1399Y) ) using a primer set containing a specific nucleic acid mutation p300 Core p300 Core was generated by amplification from dCas9 using the common AscI / PmeI restriction enzyme sites, followed by linkage PCR to generate overlapping fragments. p300Core The dCas9 constructs were cloned into the backbone. All PCR amplifications were performed using Q5 high-fidelity DNA polymerase (NEB cat. #M0491). The protein sequences of all dCas9 constructs are shown in Figures 15A-15J.

[0097] The IL1RN, MYOD, and OCT4 promoter gRNA protospacers have been previously described (Perez-Pinera, P. et al., Nature Methods 10:973-976 (2013); Hu, J. et al., Nucleic Acids Res 42:4375-4390 (2014)). VP64 (Nm-dCas9 VP64 ) was obtained from Addgene (plasmid #48676). p300 Core dCas9 using primers p300 Core Amplify HA-tagged p300 Core from AleI / AgeI-digested Nm-dCas9 using isothermal construct (NEB cat.# 2611). VP64 The IL1RN TALE was generated by facilitating subcloning into the IL1RN backbone. p300 Core TALE is dCas9 p300 Core The HA-tagged p300 Core domain derived from the IL1RN TALE was synthesized using the previously published (Perez-Pinera, P. et al., Nature Methods 10:973-976 (2013)). VP64 The constructs were generated by subcloning through the common AscI / PmeI restriction enzyme sites. The IL1RN TALE target sites are shown in Table 1.

[0098] [Table 1]

[0099] ICAM1 ZFVP64 and ICAM1 ZF p300 Core pMX-CD54-31Opt-VP64 54 The ICAM1 ZF derived from dCas9 was synthesized using an isothermal construct (NEB cat.# 2611). VP64 and dCas9 p300 Core ICAM1 ZF constructs were constructed by subcloning into each backbone. The protein sequences of the ICAM1 ZF constructs are shown in Figure 16. In all experiments, transfection efficiency was consistently greater than 90% when analyzed by cotransfection of PL-SIN-EF1α-EGFP (Addgene plasmid #21320) with an empty gRNA vector. All Streptococcus pyogenes gRNAs were annealed and cloned into pZdonor-pSPgRNA (Addgene plasmid #47108) using NEB BbsI and T4 ligase (Cat. #s R0539 and M0202) for expression with minor modifications (Cong, L. et al., Science 339:819-823 (2013)). Nm-dCas9 gRNA oligos were rationally designed using published PAM requirements (Esvelt, K. et al., Nature Methods 10:1116-1121 (2013)) and then cloned into pZDonor-Nm-Cas9-gRNA-hU6 (Addgene, Plasmid #61366) via the BbsI site. The plasmid is available through Addgene (Table 2).

[0100] [Table 2]

[0101] All gRNA protospacer targets are listed in Tables 3 and 4.

[0102] [Table 3] TIFF0007827321000004.tif104170

[0103] [Table 4]

[0104] Western blotting. 20 μg of protein was loaded for SDS-PAGE and transferred to a nitrocellulose membrane for Western blotting. Primary antibodies (α-FLAG; Sigma-Aldrich cat. #F7425 and α-GAPDH; Cell Signaling Technology cat. #14C10) were used at a 1:1000 dilution in TBST + 5% milk. Secondary α-rabbit HRP (Sigma-Aldrich cat. #A6154) was used at a 1:5000 dilution in TBST + 5% milk. The membrane was exposed to light after the addition of ECL (Bio-Rad cat. #170-5060).

[0105] Quantitative reverse transcription PCR. RNA was isolated from transfected cells using the RNeasy Plus mini kit (Qiagen, cat. #74136), and 500 ng of purified RNA was used as a template for cDNA synthesis (Life Technologies, cat. #11754). Real-time PCR was performed using PerfeCTa SYBR Green FastMix (Quanta Biosciences, cat. #95072) and a CFX96 Real-Time PCR Detection System (Bio-Rad) equipped with a C1000 thermal cycler. Baseline values ​​were subtracted using the baseline subtraction curve fit analysis mode, and thresholds were calculated automatically using Bio-Rad CFX Manager software version 2.1. Results are expressed as fold changes over control mock-transfected cells (no DNA) after normalization to GAPDH expression using the ΔΔCt method (Schmittgen et al., Nat. Protoc. 3:1101–1108 (2008)). All qPCR primers and conditions are listed in Table 5.

[0106] [Table 5] TIFF0007827321000007.tif198170 TIFF0007827321000008.tif119170

[0107] RNA-seq. RNA-seq was performed using three replicates per experimental condition. RNA was isolated from transfected cells using the RNeasy Plus mini kit (Qiagen, cat. #74136), and 1 μg of purified mRNA was used as a template for cDNA synthesis and library construction using the PrepX RNA-Seq Library Kit (Wafergen Biosystems, cat. #400039). Libraries were prepared using an Apollo 324 liquid handling platform according to the manufacturer's instructions. Indexed libraries were checked for quality and size distribution using a Tapestation 2200 (Agilent) and quantified by qPCR using the KAPA Library Quantification Kit (KAPA Biosystems; KK4835), followed by multiplex pooling and sequencing at the Duke University Genome Sequencing Shared Resource facility. Libraries were pooled, and 50-bp single-end reads were sequenced on a Hiseq 2500 (Illumina), demultiplexed, and aligned to the HG19 transcriptome using Bowtie 2 (Langmead et al. Nat. Methods 9:357-359 (2012)). Transcript abundance was calculated using the SAMtools suite (Li et al. Bioinformatics 25:2078:2079 (2009)), and differential expression was determined in R using the DESeq2 analysis package. Multiplex hypothesis correction was performed using the method of Benjamini and Hochberg with an FDR of <5%. RNA-seq data have been deposited in NCBI's Gene Expression Omnibus and are accessible under GEO Series accession number GSE66742.

[0108] ChIP-qPCR. HEK293T cells were co-transfected with the four HS2 enhancer gRNA constructs and the indicated dCas9 fusion expression vectors in 15 cm plates, with biological triplicates for each condition tested. Cells were cross-linked with 1% formaldehyde (final concentration; Sigma F8775-25ML) for 10 minutes at room temperature, and then the reaction was stopped by adding glycine (to a final concentration of 125 mM). For H3K27ac ChIP enrichment, approximately 2.5e7 cells were used from each plate. Chromatin was sheared to a median fragment size of 250 bp using a Bioruptor XL (Diagenode). H3K27ac enrichment was performed by incubating with 5 μg of Abcam ab4729 and 200 μl of sheep anti-rabbit IgG magnetic beads (Life Technologies 11203D) for 16 hours at 4°C. Crosslinks were reversed by overnight incubation with sodium dodecyl sulfate at 65°C, and DNA was purified using MinElute DNA purification columns (Qiagen). 10 ng of DNA was used for subsequent qPCR reactions using a CFX96 Real-Time PCR Detection System equipped with a C1000 Thermal Cycler (Bio-Rad). Baseline values ​​were subtracted using the baseline subtraction curve fit analysis mode, and thresholds were automatically calculated using Bio-Rad CFX Manager software version 2.1. Results are expressed as fold changes over cells cotransfected with dCas9 and the four HS2 gRNAs after normalization to β-actin enrichment using the ΔΔCt method (Schmittgen et al., Nat. Protoc. 3:1101-1108 (2008)). All ChIP-qPCR primers and conditions are listed in Table 5.

[0109] Example 2 dCas9 fusion with the p300 HAT domain activates target genes. The full-length p300 protein was synthesized using dCas9 (dCas9 FLp300dCas9 was fused to the IL1RN gene (Figure 1A-1B) and analyzed for its ability to transactivate human HEK293T cells by transient co-transfection with four gRNAs targeting the endogenous promoters of IL1RN, MYOD1 (MYOD), and POU5F1 / OCT4 (OCT4) (Figure 1C). A combination of four gRNAs targeting each promoter was used. FLp300 The VP64 acidic activation domain (dCas9) was expressed sufficiently. VP64 ) elicited moderate activation above background compared to the base dCas9 activator fused to the p300 HAT core domain (Figures 1A-1C). The full-length p300 protein is a promiscuous acetyltransferase that interacts primarily through its termini with numerous endogenous proteins. To attenuate these interactions, we isolated a proximal region (amino acids 1048-1664) of full-length p300 (2414 aa) that is exclusively required for intrinsic HAT activity, known as the p300 HAT core domain (p300 Core). When fused to the C-terminus of dCas9 (dCas9 Core), p300 Core , Figure 1A-1B), the p300 Core domain induced high levels of transcription from promoters targeted by endogenous gRNAs (Figure 1C). When targeting the IL1RN and MYOD promoters, dCas9 p300 Core The fusion is dCas9 VP64The dCas9-effector fusion proteins exhibited significantly higher levels of transactivation than the dCas9-effector fusion proteins (P values ​​0.01924 and 0.0324, respectively; Figure 1). These dCas9-effector fusion proteins were expressed at similar levels (Figure 1B, Figure 7), indicating that the observed differences were due to differences in transactivation ability. Furthermore, when the effector fusions were transfected without gRNA, no changes in target gene expression were observed (Figure 8). For Figure 8, dCas9 fusion proteins were transiently cotransfected with an empty gRNA vector backbone, and mRNA expression of IL1RN, MYOD, and OCT4 was analyzed as described in the main text. The dashed red line indicates the background expression level from no-DNA transfected cells. n = 2 (independent experiments). Error bars: s.e.m. No significant activation was observed for any of the target genes analyzed.

[0110] p300 Core acetyltransferase activity is expressed by dCas9 p300 Core To confirm that the fusions were responsible for gene transactivation, a panel of dCas9 p300 Core HAT domain mutant fusion proteins were screened (Figure 7). p300 Core (dCas9 p300 Core(D1399Y) A single inactivating amino acid substitution within the HAT core domain of dCas9 (WT residue D1399 in full-length p300) (Figure 1A) abolished the transactivation ability of the fusion protein (Figure 1C), indicating that intact p300 Core acetyltransferase activity is not present in the dCas9 fusion protein. p300 Core We have shown that it is required for ATP-mediated transactivation.

[0111] Example 3 dCas9 p300 Core activates genes from proximal and distal enhancers Since p300 plays a role in and localizes to endogenous enhancers, dCas9 p300Core The distal regulatory region (DRR) and core enhancer (CE) of the human MYOD locus can be efficiently induced by dCas9 with appropriately targeted gRNAs. VP64 or dCas9 p300 Core Targeting was achieved by co-transfection with either dCas9 or dCas9 (Figure 2A). VP64 did not show any induction when targeting the MYOD DRR or CE region. In contrast, dCas9 p300 Core induced significant transcription when targeting any of the MYOD regulatory elements with the corresponding gRNAs (P values ​​0.0115 and 0.0009 for the CE and DRR regions, respectively). The upstream proximal (PE) and distal (DE) enhancer regions of the human OCT4 gene were also significantly enhanced by the six gRNAs and dCas9. VP64 or dCas9 p300 Core Targeting was achieved by co-introduction with either dCas9 or dCas9 (Figure 2B). p300 Core induced significant transcription from these regions (P value ≤ 0.0001 and P value ≤ 0.003 for DE and PE, respectively), whereas dCas9 VP64 failed to activate OCT4 above background levels when targeting either the PE or DE region.

[0112] The well-characterized mammalian β-globin locus control region (LCR) regulates the transcription of downstream hemoglobin genes: hemoglobin ε1 (HBE, ~11 kb), hemoglobin γ1 and γ2 (HBG, ~30 kb), hemoglobin δ (HBD, ~46 kb), and hemoglobin β (HBB, ~54 kb) (Figure 2C). DNase hypersensitive regions within the β-globin LCR serve as docking sites for transcriptional and chromatin modifiers, including p300, that regulate distal target gene expression. We generated four gRNAs targeting the DNase hypersensitive region 2 within the LCR enhancer region (HS2 enhancer). These four HS2-targeting gRNAs are designated dCas9 and dCas9. VP64 , dCas9 p300 Core , or dCas9 p300 Core(D1399Y) The resulting mRNA production from HBE, HBG, HBD, and HBB was analyzed (Figure 2C). VP64 , and dCas9 p300 Core(D1399Y) failed to transactivate any downstream genes when targeting the HS2 enhancer. p300 Core Targeting of dCas9 to the HS2 enhancer resulted in significant expression of downstream HBE, HBG, and HBD genes (P values ​​≤ 0.0001, 0.0056, and 0.0003 for HBE, HBG, and HBD, respectively). p300 Core (vs. mock-transfected cells). Taken together, HBD and HBE appear to be relatively insensitive to synthetic p300 Core-mediated activation from the HS2 enhancer; a finding consistent with the general lower rates of transcription from these two genes across several cell lines (Figures 9A-9E).

[0113] Nevertheless, except for the most distal HBB gene, dCas9 p300 Coreshows the ability to activate transcription from downstream genes when targeting all characterized enhancer regions analyzed, i.e., dCas9 VP64 Taken together, these results suggest that dCas9 p300 Core We demonstrate that β-actin is a potent, programmable transcription factor that can be used to regulate gene expression from a variety of promoter-proximal and promoter-distal locations.

[0114] Example 4 dCas9 p300 Core Gene activation by is highly specific. Recent reports have shown that dCas9, in combination with some gRNAs, can have a wide range of off-target binding events in mammalian cells, potentially leading to off-target changes in gene expression. p300 Core To evaluate the transcription specificity of the fusion proteins, we tested four IL1RN-targeting gRNAs, dCas9, and dCas9. VP64 , dCas9 p300 Core , or dCas9 p300 Core(D1399Y) We performed transcriptome profiling by RNA-seq in cells co-transfected with either dCas9 or dCas9. VP64 , dCas9 p300 Core , or dCas9 p300 Core(D1399Y) The results were compared between either dCas9 or dCas9 (Figure 3). VP64 and dCas9 p300 Core Both upregulated all four IL1RN isoforms, but dCas9 p300 Core Only the effect of dCas9 reached genome-wide significance (Figures 3A-3B, Table 6; dCas9 VP64 For P value 1.0×10 -3~5.3×10 -4 ;dCas9 p300 Core For the P value 1.8×10 -17 ~1.5×10 -19 ).

[0115] [Table 6] TIFF0007827321000010.tif158170

[0116] In contrast, dCas9 p300 Core(D1399Y) did not significantly induce any IL1RN expression (Figure 3C; P values ​​> 0.5 for all four IL1RN isoforms). Comparative analysis with dCas9 revealed a false positive rate (FDR) < 5%: KDR (FDR = 1.4 × 10 -3 and FAM49A (FDR = 0.04), only two transcripts were significantly induced above background, limiting dCas9 p300 Core Off-target gene transfer was evident (Figure 3B, Table 6). p300 Core and dCas9 p300 Core(D1399Y) Increased expression of p300 mRNA was observed in cells transfected with dCas9. This finding is likely explained by RNA-seq read mapping to mRNA from the transiently transfected p300 core fusion domain. p300 Core The fusions exhibited high genome-wide targeted transcriptional specificity and robust gene transfer of all four targeted IL1RN isoforms.

[0117] Example 5 dCas9 p300 Core acetylates H3K27 at enhancers and promoters The activity of regulatory elements correlates with covalent histone modifications, such as acetylation and methylation. Among these histone modifications, acetylation of lysine 27 on histone H3 (H3K27ac) is one of the most widely documented indicators of enhancer activity. H3K27 acetylation is catalyzed by p300 and correlates with endogenous p300 binding profiles. Therefore, H3K27ac enrichment can be used to assess the relative dCas9 activity at genomic target sites. p300 Core was used as a measure of dCas9-mediated acetylation. p300 Core To quantify H3K27 acetylation by HS2 enhancers, we used four HS2 enhancer-targeting gRNAs, dCas9, and dCas9. VP64 , dCas9 p300 Core , or dCas9 p300 Core(D1399Y) Chromatin immunoprecipitation using anti-H3K27ac antibody followed by quantitative PCR (ChIP-qPCR) was performed in HEK293T cells co-transfected with either dCas9 or dCas9 (Figure 4). Three amplicons were analyzed at or around target sites within the HS2 enhancer or promoter regions of the HBE and HBG genes (Figure 4A). Notably, in the human K562 erythroid cell line, which has high levels of globin gene expression, H3K27ac is enriched in each of these regions (Figure 4A). dCas9 at the HS2 enhancer target locus VP64 (P value 0.0056 for ChIP region 1 and P value 0.0029 for ChIP region 3) and dCas9 p300 Core Significant H3K27ac enrichment was observed in both the co-transfected and dCas9 treated samples (P value 0.0013 for ChIP region 1 and P value 0.0069 for ChIP region 3) compared to dCas9 treatment (Figure 4B).

[0118] dCas9 VP64 or dCas9 p300 CoreA similar trend of H3K27ac enrichment was observed when targeting the IL1RN promoter using dCas9 (Figure 10). Figure 10 shows the IL1RN locus on GRCh37 / hg19 along with the IL1RN gRNA target site. Additionally, an overlay of ENCODE H3K27ac enrichment from seven cell lines (GM12878, H1-hESC, HSMM, HUVEC, K562, NHEK, and NHLF) is shown with the vertical range set to 50. Tiled IL1RN ChIP qPCR amplicons (1-13) are also shown at the corresponding locations on GRCh37 / hg19. dCas9 co-transfected with four IL1RN-targeting gRNAs and normalized to dCas9 co-transfected with four IL1RN-targeting gRNAs. VP64 and dCas9 p300 Core H3K27ac enrichment for each gene is shown for each ChIP qPCR locus analyzed. For each reaction, 5 ng of ChIP-prepared DNA was used (n=3 independent experiments, error bars: sem).

[0119] dCas9 VP64 and dCas9 p300 Core In contrast to these increases in H3K27ac at target sites by both dCas9 and dCas9, the strong enrichment of H3K27ac at the HS2-regulated HBE and HBG promoters was observed. p300 Core This was observed only in the dCas9 treatment (Figure 4C-D). p300 Core We demonstrate that dCas9 specifically catalyzes H3K27ac enrichment at the gene locus targeted by the gRNA and at the distal promoter targeted by the enhancer. p300 Core Acetylation established by dCas9 p300 Core By dCas9 VP64 dCas9, as shown by distal activation of the HBE, HBG, and HBD genes from the HS2 enhancer independent of dCas9 VP64These results suggest that enhancer activity can be catalyzed by a mechanism distinct from the direct recruitment of pre-initiation complex components by ATP (Figure 2C, Figures 9A-9E).

[0120] Example 6 dCas9 p300 Core activates genes together with a single gRNA Robust transactivation using dCas9-effector fusion proteins currently relies on the application of multiple gRNAs, multiple effector domains, or both. Transcriptional activation could be simplified using a single gRNA in conjunction with a single dCas9-effector fusion. This would also facilitate multiplexing of additional target genes and the incorporation of additional functionality into the system. dCas9 with a single gRNA p300 Core The transactivation potential of dCas9 with four pooled gRNAs targeting the IL1RN, MYOD, and OCT4 promoters was examined. p300 core The transactivation potential of dCas9 was compared with that of dCas9 (Figures 5A-5B). p300 Core Substantial activation was observed when co-transfecting dCas9 with a single gRNA for each promoter tested. For the IL1RN and MYOD promoters, there was no significant difference between the pooled gRNAs and the best individual gRNA (Figures 5A-5B; IL1RN gRNA "C", P value 0.78; MYOD gRNA "D", P value 0.26). The four gRNAs were shown to enhance the activity of dCas9. p300 Core The most potent single gRNA (gRNA “D”) was the only gRNA that produced an additive effect when pooled with dCas9. VP64 dCas9 induced gene expression at levels statistically comparable to that observed when co-transfected with an equimolar pool of all four promoter-gRNAs (P value 0.73; Figure 5C). p300 Core Compared with dCas9 VP64 The level of gene activation with a single gRNA was substantially lower.p300 Core In contrast to dCas9 VP64 In all cases, the combination with gRNA showed synergistic effects (Figures 5A-5C).

[0121] dCas9 with a single gRNA in the enhancer region and in the promoter region p300 Core Based on the transactivation ability of dCas9 p300 Core However, we hypothesize that it may also be possible to transactivate enhancers via a single targeting gRNA. The MYOD (DRR and CE), OCT4 (PE and DE), and HS2 enhancer regions were tested together with equimolar pools or single gRNAs (Figures 5D-5G). For both MYOD enhancer regions, dCas9 p300 Core Co-transfection of dCas9 with a gRNA targeting a single enhancer p300 Core The dCas9 gRNAs were sufficient to activate gene expression to a level similar to that of cells co-transfected with a pool of four PE-targeted gRNAs (Figure 5D). Similarly, OCT4 gene expression was also enhanced by dCas9 gRNAs co-transfected with a pool of six PE-targeted gRNAs. p300 Core dCas9 with a single gRNA to a similar level p300 Core dCas9 was activated via localization from the PE (Figure 5E), OCT4 from the DE (Figure 5E), and HBE and HBG genes from the HS2 enhancer (Figure 5F-5G). p300 Core Although the mediated induction showed increased expression with pooled gRNAs compared to single gRNAs, there was still activation of target gene expression above the control for some single gRNAs at these enhancers (Figures 5E-5G).

[0122] Example 7 The p300 HAT domain can be transplanted into other DNA-binding proteins. The dCas9 / gRNA system from Streptococcus pyogenes has been widely adopted due to its robust, versatile, and easily programmable nature. However, several other programmable DNA-binding proteins, including orthogonal dCas9 systems, TALEs, and zinc finger proteins from other species, are also under development for various applications and may be preferred for specific applications. To determine whether the p300 Core HAT domain can be transplanted into these other systems, we generated fusions with dCas9 from Neisseria meningitidis (Nm-dCas9), the IL1RN promoter targeting four different TALEs, and the ICAM1 targeting zinc finger protein (Figure 6). Nm-dCas9 p300 Core Co-transfection of Nm-dCas9 with five Nm-gRNAs targeting the HBE or HBG promoter resulted in significant gene transduction compared to mock-transfected controls (P values ​​0.038 and 0.0141 for HBE and HBG, respectively) (Figures 6B-6C). Co-transfection of five Nm-gRNAs targeting the HS2 enhancer resulted in significant gene transduction. p300 Core dCas9 also significantly activated the distal HBE and HBG globin genes compared to mock-transfected controls (p = 0.0192 and p = 0.0393, respectively) (Figures 6D-6E). p300 Core Similarly, Nm-dCas9 p300 Core also activated gene expression from the promoter and HS2 enhancer via a single gRNA. VP64 The four TALEs targeting the IL1RN promoter, together with single or multiple gRNAs, exhibited negligible ability to transactivate HBE or HBG, regardless of their localization to the promoter region or to the HS2 enhancer (Figures 6B-6E). p300 Core Fusion protein (IL1RN TALE p300 Core) combinations, the introduction of expression plasmids into the four corresponding TALE VP64 Fusion (IL1RN TALE VP64 ) activated downstream gene expression, albeit to a lesser extent than the single p300 Core effector (Figures 6F-6G). However, a single p300 Core effector, when fused to the IL1RN TALE, was much more potent than a single VP64 domain. Interestingly, dCas9 induced by either single binding site p300 Core dCas9 resulted in comparable IL1RN expression to single or pooled IL1RN TALE effectors, and direct comparisons suggested that dCas9 may be a more robust activator than TALEs when fused to the larger p300 Core fusion domain (Figures 11A-11C). The p300 Core effector domain did not exhibit synergy with additional gRNAs or TALEs (see Figures 5, 6, 9, and 11), nor in combination with VP64 (see Figures 13A-13B). p300 Core The basic chromatin context of the target locus is shown in Figures 14A-14E.

[0123] ZF targeting the ICAM1 promoter p300 Core Fusion (ICAM1 ZF p300 Core ) compared to the control, and ZF VP64 (ICAM1 ZF VP64 ) activated its target genes at levels similar to those of p300 Core (Figures 6H-6I). The versatility of p300 Core fusions with multiple targeting domains is evidence that this is a robust approach for targeted acetylation and gene regulation. Although various p300 Core fusion proteins were well expressed as determined by Western blot (Figures 12A-12B), differences in p300 Core activity between different fusion proteins may be due to binding affinity or protein folding.

[0124] Example 8 Myocardin Thirty-six gRNAs were designed, spanning the −2000 bp to +250 bp (coordinates relative to the TSS) region of the MYOCD gene (Table 7).

[0125] [Table 7] TIFF0007827321000012.tif47170

[0126] These gRNAs were cloned into the spCas9 gRNA expression vector containing the hU6 promoter and BbsI restriction enzyme site. These gRNAs were then transfected into HEK293T cells to express dCas9 p300 Core The resulting mRNA production for myocardin was analyzed in samples collected 3 days after transfection (Figure 17). Cr32, Cr13, Cr30, Cr28, Cr31, and Cr34 were transiently co-transfected with dCas9. p300 Core The combinations of were analyzed (Table 8; Figure 18).

[0127] [Table 8]

[0128] Example 9 Pax7 gRNAs spanning the region surrounding the PAX7 gene were designed (Table 9). These gRNAs were cloned into the spCas9 gRNA expression vector, which contains an hU6 promoter and a BbsI restriction enzyme site. These gRNAs were transfected into HEK293T cells to express dCas9 p300 Core or dCas9 VP64 The resulting mRNA production for Pax7 was analyzed in samples collected 3 days after transfection (Figure 19). Further experiments used gRNA19 ("g19"), which was shown to localize to DNase hypersensitive regions (DHS) (Figure 20).

[0129] [Table 9]

[0130] Example 10 FGF1 gRNAs were designed for the FGF1A, FGF1B, and FGF1C genes (Tables 10 and 11). 25 nM of these gRNAs were transfected into HEK293T cells using dCas9 p300 Core or dCas9 VP64 The resulting mRNA production for FGF1 expression was determined (Figures 21-23). ​​In Figure 23, the number of stable cell lines transduced with lentiviral vectors was 2, except for FGF1A, where n=1.

[0131] [Table 10]

[0132] [Table 11] TIFF0007827321000017.tif79170

[0133] It is understood that the foregoing detailed description and accompanying examples are merely illustrative and should not be taken as limitations on the scope of the invention, which is defined solely by the appended claims and equivalents thereof.

[0134] Various changes and modifications to the disclosed embodiments will be apparent to those skilled in the art. Such changes and modifications, including but not limited to, with respect to the chemical structure, substituents, derivatives, intermediates, compounds, compositions, formulations, or methods of use of the invention, can be made without departing from the spirit and scope thereof.

[0135] For reasons of completeness, the various aspects of the invention are presented in the following numbered clauses:

[0136] Item 1. A fusion protein comprising two heterologous polypeptide domains, wherein a first polypeptide domain comprises a Clustered Regularly Interspaced Short Palindromic Repeat-associated (Cas) protein and a second polypeptide domain comprises a peptide having histone acetyltransferase activity.

[0137] Section 2. A fusion protein according to section 1, which activates transcription of a target gene.

[0138] Section 3. The fusion protein of section 1 or 2, wherein the Cas protein comprises Cas9.

[0139] Section 4. The fusion protein of section 3, wherein the Cas9 comprises at least one amino acid mutation that knocks out the nuclease activity of the Cas9.

[0140] Section 5. The fusion protein of section 4, wherein the Cas protein comprises SEQ ID NO:1 or SEQ ID NO:10.

[0141] Item 6. The fusion protein of any one of items 1 to 5, wherein the second polypeptide domain comprises a histone acetylase effector domain.

[0142] Item 7. The fusion protein of item 6, wherein the histone acetylase effector domain is a p300 histone acetylase effector domain.

[0143] Paragraph 8. The fusion protein of any one of paragraphs 1 to 7, wherein the second polypeptide domain comprises SEQ ID NO:2 or SEQ ID NO:3.

[0144] Item 9. The fusion protein of any one of items 1 to 8, wherein the first polypeptide domain comprises SEQ ID NO: 1 or SEQ ID NO: 10, and the second polypeptide domain comprises SEQ ID NO: 2 or SEQ ID NO: 3.

[0145] Paragraph 10. The fusion protein of any one of paragraphs 1 to 9, wherein the first polypeptide domain comprises SEQ ID NO: 1 and the second polypeptide domain comprises SEQ ID NO: 3, or the first polypeptide domain comprises SEQ ID NO: 10 and the second polypeptide domain comprises SEQ ID NO: 3.

[0146] Item 11. The fusion protein of any one of items 1 to 10, further comprising a linker connecting the first polypeptide domain to the second polypeptide domain.

[0147] Item 12. The fusion protein of any one of items 1 to 11, comprising the amino acid sequence of SEQ ID NO: 140, 141, or 149.

[0148] Item 13. A DNA targeting system comprising the fusion protein of any one of items 1 to 12 and at least one guide RNA (gRNA).

[0149] Item 14. The DNA targeting system of item 13, wherein at least one gRNA comprises a 12-22 base pair complementary polynucleotide sequence of the target DNA sequence followed by a protospacer adjacent motif.

[0150] Item 15. The DNA targeting system of item 13 or 14, wherein at least one gRNA targets a target region, and the target region comprises a target enhancer, a target regulatory element, a cis-regulatory region of a target gene, or a trans-regulatory region of a target gene.

[0151] Item 16. The DNA targeting system of item 15, wherein the target region is a distal or proximal cis-regulatory region of the target gene.

[0152] Item 17. The DNA targeting system of items 15 or 16, wherein the target region is an enhancer or promoter region of the target gene.

[0153] Item 18. The DNA targeting system of any one of items 15 to 17, wherein the target gene is an endogenous gene or a transgene.

[0154] Paragraph 19. The DNA targeting system of paragraph 15, wherein the target region comprises a target enhancer or a target regulatory element.

[0155] Paragraph 20. The DNA targeting system of paragraph 19, wherein the targeted enhancer or targeted regulatory element controls gene expression of two or more target genes.

[0156] Paragraph 21. The DNA targeting system of any one of paragraphs 15 to 20, wherein the DNA targeting system comprises 1 to 10 different gRNAs.

[0157] Paragraph 22. The DNA targeting system of any one of paragraphs 15 to 21, wherein the DNA targeting system comprises one gRNA.

[0158] Paragraph 23. A DNA targeting system according to any one of paragraphs 15 to 22, wherein the target region is located on the same chromosome as the target gene.

[0159] Item 24. The DNA targeting system of item 23, wherein the target region is located from about 1 base pair to about 100,000 base pairs upstream of the transcription start site of the target gene.

[0160] Item 25. The DNA targeting system of item 24, wherein the target region is located between 1000 base pairs and about 50,000 base pairs upstream of the transcription start site of the target gene.

[0161] Item 26. A DNA targeting system according to any one of items 15 to 22, wherein the target region is located on a different chromosome than the target gene.

[0162] Paragraph 27. A DNA targeting system described in any one of paragraphs 15 to 28, wherein different gRNAs bind to different target regions.

[0163] Item 28. The DNA targeting system of item 27, wherein different gRNAs bind to target regions of different target genes.

[0164] Item 29. The DNA targeting system of item 27, wherein expression of two or more target genes is activated.

[0165] Paragraph 30. The DNA targeting system of any one of paragraphs 15 to 29, wherein the target gene is selected from the group consisting of IL1RN, MYOD1, OCT4, HBE, HBG, HBD, HBB, MYOCD, PAX7, FGF1A, FGF1B, and FGF1C.

[0166] Paragraph 31. The DNA targeting system of paragraph 30, wherein the target region is at least one of the HS2 enhancer of the human β-globin locus, the distal regulatory region (DRR) of the MYOD gene, the core enhancer (CE) of the MYOD gene, the proximal (PE) enhancer region of the OCT4 gene, or the distal (DE) enhancer region of the OCT4 gene.

[0167] Clause 32. The DNA targeting system of any one of clauses 13 to 31, wherein the gRNA comprises at least one of SEQ ID NOs: 23-73, 188-223, or 224-254.

[0168] Paragraph 33. An isolated polynucleotide encoding the fusion protein of any one of paragraphs 1 to 12 or the DNA targeting system of any one of paragraphs 13 to 32.

[0169] Item 34. A vector comprising the isolated polynucleotide of Item 33.

[0170] Item 35. A cell comprising the isolated polynucleotide of Item 33 or the vector of Item 34.

[0171] Paragraph 36. A kit comprising a fusion protein according to any one of paragraphs 1 to 12, a DNA targeting system according to paragraphs 13 to 32, an isolated polynucleotide according to paragraph 33, a vector according to paragraph 34, or a cell according to paragraph 35.

[0172] Clause 37. A method of activating gene expression of a target gene in a cell, comprising contacting the cell with a fusion protein described in any one of clauses 1-12, a DNA targeting system described in clauses 13-32, an isolated polynucleotide described in clause 33, or a vector described in clause 34.

[0173] Clause 38. A method for activating gene expression of a target gene in a cell, comprising contacting the cell with a polynucleotide encoding a DNA targeting system, wherein the DNA targeting system comprises a fusion protein described in any one of clauses 1 to 12 and at least one guide RNA (gRNA).

[0174] Item 39. The method of item 38, wherein at least one gRNA comprises a 12-22 base pair complementary polynucleotide sequence of the target DNA sequence followed by a protospacer adjacent motif.

[0175] Paragraph 40. The method of paragraph 38 or 39, wherein at least one gRNA targets a target region, and the target region is a cis-regulatory region or a trans-regulatory region of a target gene.

[0176] Item 41. The method of item 40, wherein the target region is a distal or proximal cis-regulatory region of the target gene.

[0177] Item 42. The method of items 40 or 41, wherein the target region is an enhancer or promoter region of the target gene.

[0178] Item 43. The method according to items 40 to 42, wherein the target gene is an endogenous gene or a transgene.

[0179] Item 44. The method of item 43, wherein the DNA targeting system comprises 1 to 10 different gRNAs.

[0180] Item 45. The method of item 43, wherein the DNA targeting system comprises one gRNA.

[0181] Item 46. The method of items 40 to 45, wherein the target region is located on the same chromosome as the target gene.

[0182] Item 47. The method of item 46, wherein the target region is located from about 1 base pair to about 100,000 base pairs upstream of the transcription start site of the target gene.

[0183] 48. The method of claim 46, wherein the target region is located between 1000 and about 50,000 base pairs upstream of the transcription start site of the target gene.

[0184] Item 49. The method according to items 40 to 45, wherein the target region is located on a different chromosome than the target gene.

[0185] Item 50. The method according to items 40 to 45, wherein different gRNAs bind to different target regions.

[0186] Item 51. The method of item 50, wherein different gRNAs bind to target regions of different target genes.

[0187] Item 52. The method of item 51, wherein expression of two or more target genes is activated.

[0188] Item 53. The method of items 40 to 52, wherein the target gene is selected from the group consisting of IL1RN, MYOD1, OCT4, HBE, HBG, HBD, HBB, MYOCD, PAX7, FGF1A, FGF1B, and FGF1C.

[0189] Item 54. The method of item 53, wherein the target region is at least one of the HS2 enhancer of the human β-globin locus, the distal regulatory region (DRR) of the MYOD gene, the core enhancer (CE) of the MYOD gene, the proximal (PE) enhancer region of the OCT4 gene, or the distal (DE) enhancer region of the OCT4 gene.

[0190] Clause 55. The method of clauses 37-54, wherein the gRNA comprises at least one of SEQ ID NOs: 23-73, 188-223, or 224-254.

[0191] Paragraph 56. The method of any one of paragraphs 37 to 55, wherein the DNA targeting system is delivered to the cell virally or non-virally.

[0192] Item 57. The method of any one of items 37 to 56, wherein the cell is a mammalian cell.

[0193] Appendix - Sequence

[0194] Streptococcus pyogenes Cas 9 (D10A, with H849A) (SEQ ID NO: 1) TIFF0007827321000018.tif112160

[0195] Human p300 (with L553M mutation) (SEQ ID NO: 2) TIFF0007827321000019.tif196160

[0196] p300 Core effector (aa 1048 to 1664 of SEQ ID NO: 2) (SEQ ID NO: 3) TIFF0007827321000020.tif57160

[0197] p300 Core effector (aa 1048-1664 of SEQ ID NO: 2 with D1399Y mutation) (SEQ ID NO: 4) TIFF0007827321000021.tif54160

[0198] p300 Core effector (aa 1048-1664 of SEQ ID NO: 2 with 1645 / 1646 RR / EE mutations) (SEQ ID NO: 5) TIFF0007827321000022.tif55160

[0199] p300 Core effector (aa 1048-1664 of SEQ ID NO: 2 with C1204R mutation) (SEQ ID NO: 6) TIFF0007827321000023.tif54160

[0200] p300 Core effector (aa 1048-1664 of SEQ ID NO: 2 with Y1467F mutation) (SEQ ID NO: 7) TIFF0007827321000024.tif55160

[0201] p300 Core effector (aa 1048-1664 of SEQ ID NO:2 with 1396 / 1397 SY / WW mutations) (SEQ ID NO:8) TIFF0007827321000025.tif57160

[0202] p300 Core effector (aa 1048-1664 of SEQ ID NO: 2 with H1415A, E1423A, Y1424A, L1428S, Y1430A, and H1434A mutations) (SEQ ID NO: 9) TIFF0007827321000026.tif55160

[0203] Neisseria meningitidis Cas9 (with D16A, D587A, H588A, and N611A mutations) (SEQ ID NO: 10) TIFF0007827321000027.tif93160

[0204] 3X "Flag" epitope (SEQ ID NO: 11) DYKDHDGDYKDHDIDYKDDDDK

[0205] Nuclear localization sequence (SEQ ID NO: 12) PKKKRKVG

[0206] HA epitope (SEQ ID NO: 13) YPYDVPDYAS

[0207] VP64 effector (SEQ ID NO: 14) DALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDML

Claims

1. a fusion protein comprising two heterologous polypeptide domains, wherein a first polypeptide domain comprises a clustered regularly interspaced short palindromic repeats-associated (Cas) protein and a second polypeptide domain comprises a p300 histone acetyltransferase effector domain, and wherein the fusion protein comprises the amino acid sequence of SEQ ID NO: 141; At least one guide RNA (gRNA) encoded by a polynucleotide comprising a sequence selected from SEQ ID NOs: 38, 45, 59-61, 152-153, 158, 160-161, 164-165, 167-168, 170, 172, and 175-185. and a DNA targeting system comprising:

2. a fusion protein comprising two heterologous polypeptide domains, wherein a first polypeptide domain comprises the polypeptide sequence of SEQ ID NO: 1 and a second polypeptide domain comprises the polypeptide sequence of SEQ ID NO: 3; At least one guide RNA (gRNA) encoded by a polynucleotide comprising a sequence selected from SEQ ID NOs: 38, 45, 59-61, 152-153, 158, 160-161, 164-165, 167-168, 170, 172, and 175-185. and a DNA targeting system comprising:

3. The DNA targeting system of claim 2 , further comprising a linker connecting the first polypeptide domain to the second polypeptide domain.

4. A DNA targeting system described in any one of claims 1 to 3, wherein the fusion protein activates transcription of a target gene.

5. 5. The DNA targeting system of claim 1, wherein the at least one gRNA targets a target region, and the target region comprises an enhancer region of a target gene, a promoter region of a target gene, a regulatory element of a target gene, a cis-regulatory region of a target gene, or a trans-regulatory region of a target gene.

6. The DNA targeting system of claim 5 , wherein the target region is a distal or proximal cis-regulatory region of a target gene.

7. An isolated polynucleotide encoding the DNA targeting system of any one of claims 1 to 6.

8. A vector comprising the isolated polynucleotide of claim 7.

9. A composition for activating gene expression of a target gene in a cell, comprising the DNA targeting system of any one of claims 1 to 6, a polynucleotide encoding the DNA targeting system of any one of claims 1 to 6, or the vector of claim 8.

10. The composition of claim 9 , wherein expression of two or more target genes is activated.

Citation Information

Patent Citations

  • RNA-guided targeting of genetic and epigenomic regulatory proteins to specific genomic loci

    WO2014152432A2

  • RNA-guided gene editing and gene regulation

    WO2014197748A2