Compositions and methods for genome editing

By designing fusion molecular constructs containing DNMT3A, DNMT3L, dCas9, and KRAB, the delivery and safety challenges of the CRISPR/Cas9 system in vivo were addressed, achieving efficient and specific targeted regulation of gene expression, reducing off-target modifications, and minimizing toxicity and immune responses.

CN119095972BActive Publication Date: 2025-11-28EPIGENIC THERAPEUTICS PTE LTD
View PDF 63 Cites 0 Cited by

Patent Information

Application Number
CN202380023772.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-04
Filing Date
2023-03-03
Publication Date
2025-11-28
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

Existing CRISPR/Cas9 systems face challenges in in vivo applications, including limitations in viral vector delivery, safety, toxicity, immunogenicity, and off-target effects, making it difficult to effectively target and silence endogenous genes.

Method used

A fusion molecular construct containing DNMT3A, DNMT3L, dCas9, and KRAB was designed. Through the combination of polynucleotide sequences and peptides, sgRNA is recruited to genomic loci to achieve targeted regulation and silencing of gene expression, reducing off-target modifications.

Benefits of technology

It achieves efficient and specific targeted regulation of gene expression in vivo, reduces the rate of off-target modifications, decreases the expression of gene products, and has low toxicity and immune response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present disclosure provides CRISPR / Cas9-based fusion molecules and guide RNAs for in vivo targeted reduction or elimination of VEGFA gene products. The present disclosure also relates to formulations thereof, methods of production, and methods of use.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to the fields of molecular biology, immunology, and medicine. More specifically, it concerns CRISPR / Cas9-based fusion molecules for targeted reduction or elimination of gene products in vivo and methods of use thereof. BACKGROUND

[0002] Engineered DNA-binding proteins that can be tailored to target any gene in mammalian cells have enabled rapid advances in biomedical research and are promising platforms for gene therapy. The RNA-guided CRISPR-Cas9 system has emerged as a promising platform for programmable targeted gene regulation. Fusion of catalytically inactive "dead" Cas9 (dCas9) with a Kruppel-associated box (KRAB) domain generates a synthetic repressor that can highly specifically and efficiently regulate or silence target genes in cell culture experiments.

[0003] However, sustained regulation and silencing of endogenous genes using synthetic dCas9-KRAB fusion proteins presents challenges for use in in vivo therapies. The synthetic repressor exceeds the packaging size limitations of viral vector delivery methods. Safety, toxicity, immunogenicity, and off-target effects are other challenges that limit the use of synthetic repressors in vivo. There is a need in the art for alternative methods for producing genetically engineered synthetic gene repressors, and in vivo delivery of synthetic gene repressors for use as therapeutic agents. The present disclosure addresses unmet needs in the art. SUMMARY

[0004] The present disclosure provides a construct of Formula I: 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3' (I), wherein: one of A and B is a polynucleotide encoding DNMT3A or a portion thereof; the other of A and B is a polynucleotide encoding DNMT3L or a portion thereof; CasN is a polynucleotide encoding an N-terminal portion of dCas9; CasC is a polynucleotide encoding a C-terminal portion of dCas9; E is 5'-(A m5 -B m6 ) n3 -K r -D q -3' or 5' -K r -D q -(A m5 -B m6 ) n3-3'; K is a polynucleotide encoding KRAB or a portion thereof; D is a polynucleotide encoding a gene expression modulator; T comprises a polynucleotide encoding (i) a polypeptide sequence capable of binding to an epitope of an antibody or antigen binding fragment thereof or (ii) a polynucleotide sequence capable of binding to a nucleic acid structural element; ml, m2, m3, m4, m5, and m6 are each independently an integer selected from 0 to 3; nl, n2, and n3 are each independently an integer selected from 0 to 2; p is an integer selected from 0 to 20; q is an integer selected from 0 to 5; r is an integer selected from 0 to 5; and wherein when p is 0, at least one of ml, m2, m3, and m4 is not 0, at least one of nl and n2 is not 0, and at least one of q and r is not 0.

[0005] In some embodiments, the construct comprises Formula II: 5'-CasN-(A m3 -B m4 ) n2 -CasC-K r -D q -3' (II), wherein: n2 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; q is an integer selected from 0 to 5; and at least one of m3 and m4 is not 0. In some embodiments, the construct comprises Formula IIa: 5'-CasN-(A-B)-CasC-K r -D q -3' (IIa).

[0006] In some embodiments, the construct comprises Formula III: 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3' (III), wherein p is an integer selected from 1 to 20. In some embodiments, the construct comprises Formula IIIa: 5'-CasN-CasC-T p -3' (IIIa).

[0007] In some embodiments, the construct comprises Formula IIIb: 5'-(A m1 -B m2 )-CasN-CasC-T p -E-3' (IIIb), wherein at least one of ml and m2 is not 0. In some embodiments, the construct comprises Formula IIIb-1: 5'-(A-B)-CasN-CasC-T p -K r -D q-3’ (IIIb-1), wherein: r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5.

[0008] In some embodiments, the construct comprises Formula IIIc: 5'-CasN-CasC-T p -3’ (IIIc-1), wherein: n3 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5; and at least one of m5 and m6 is not 0. In some embodiments, the construct comprises Formula IIIc-1: 5'-CasN-CasC-Tp-(A m5 -B m6 )n3-K r -D q -3’ (IIIc-1). In some embodiments, the construct comprises Formula IIIc-2: 5'-CasN-CasC-T p -K r -D q -(A m5 -B m6 ) n3 -3’ (IIIc-2).

[0009] In some embodiments, T comprises (a) a polynucleotide encoding an antibody or antigen binding fragment thereof capable of binding to an epitope, and further comprises (b) a polynucleotide encoding a self-cleaving peptide at the 3’ end of the polynucleotide in (a). In some embodiments, the 3’ end of the nucleotide encoding the self-cleaving peptide further comprises (c) a polynucleotide encoding an antibody or antigen binding fragment thereof capable of binding to the epitope.

[0010] In some embodiments, T comprises (a) a polynucleotide encoding a polypeptide sequence capable of binding to a nucleotide sequence element, and further comprises a polynucleotide encoding a self-cleaving peptide at the 3’ end of the polynucleotide in (a).

[0011] In some embodiments, the self-cleaving peptide is selected from the group consisting of a T2A, P2A, E2A, and F2A self-cleaving peptide. In some embodiments, the self-cleaving peptide is T2A. In some embodiments, the T2A comprises the amino acid sequence of SEQ ID NO: 99.

[0012] In some embodiments, the DNMT3A comprises the amino acid sequence of SEQ ID NO: 69. In some embodiments, the polynucleotide encoding the DNMT3A comprises the nucleic acid sequence of SEQ ID NO: 83.

[0013] In some embodiments, the DNMT3L comprises Homo sapiens DNMT3L, Mus musculus DNMT3L, Mus caroli DNMT3L, Rattus norvegicus DNMT3L, or Gallus gallus DNMT3L. Mus caroli Mus Pahari) DNMT3L, Tinnevelly mouse Rattus norvegicus ) DNMT3L, Brown rat Rattus Rattus ) DNMT3L, Black rat Arvicanthis niloticus ) DNMT3L, Nile grass rat Grammomys surdaster ) DNMT3L, Brush-tailed rat Mastomys coucha ) DNMT3L or White-nosed coati Staphylococcus aureus ) DNTM3L. In some embodiments, the DNMT3L comprises an amino acid sequence of any one of SEQ ID NOs: 74-82. In some embodiments, the polynucleotide encoding the DNMT3L comprises a nucleic acid sequence of any one of SEQ ID NOs: 84-92.

[0014] In some embodiments, the KRAB comprises an amino acid sequence of SEQ ID NO: 51, 53, or 230-241. In some embodiments, the polynucleotide encoding the KRAB comprises a nucleic acid sequence of SEQ ID NO: 52, 54, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, or 228.

[0015] In some embodiments, the gene expression modulator comprises Kruppel-associated box (KRAB), enhancer of zeste homolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), Friend of GATA protein 1 (FOG1), histone deacetylase (HDAC3), and / or DOT1L. In some embodiments, the gene expression modulator comprises an amino acid sequence of SEQ ID NO: 51, 53, 55, 57, 59, 61, 63, 65, or 67. In some embodiments, the polynucleotide encoding the gene expression modulator comprises a nucleic acid sequence of SEQ ID NO: 52, 54, 56, 58, 60, 62, 64, 66, or 68.

[0016] In some embodiments, the dCas9 comprises Staphylococcus aureus Streptococcus pyogenes dCas9, Streptococcus pyogenes Campylobacter jejuni dCas9, Campylobacter jejuni Corynebacterium diphtheria Eubacterium ventriosum dCas9, Corynebacterium diphtheriae Streptococcus pasteurianus dCas9, Eubacterium dolichum Lactobacillus farciminis dCas9, Streptococcus pasteurianus Sphaerochaeta globus dCas9, Lactobacillus farciminus Azospirillum) dCas9, Chaetomium globosum Gluconacetobacter diazotrophicus ) dCas9, Azospirillum sp. Neisseria cinerea ) dCas9, Gluconacetobacter diazotrophicus (e.g., strain B510) Roseburia intestinalis ) dCas9, Neisseria cinerea Parvibaculum lavamentivorans ) dCas9, Roseburia intestinalis Nitratifractor salsuginis ) dCas9, Algoriphagus cleanseris Campylobacter lari ) dCas9, Halonitrnosus alimentus Streptococcus thermophilus Figure 1A ) dCas9, Campylobacter lari (e.g., strain DSM 16511) Figure 1B ) dCas9, Streptococcus thermophilus (e.g., strain CF89-12) Figure 1C ) dCas9, Lactococcus lactis (e.g., strain LMD-9). In some embodiments, the dCas9 comprises the amino acid sequence of any one of SEQ ID NOs: 106-122. In some embodiments, the polynucleotide encoding the dCas9 comprises the nucleic acid sequence of SEQ ID NO: 26.

[0017] In some embodiments, the dCas9 comprises the amino acid sequence of SEQ ID NO: 1. In some embodiments, the N-terminal portion of the dCas9 comprises the amino acid sequence of SEQ ID NO: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, or 24. In some embodiments, the C-terminal portion of the dCas9 comprises the amino acid sequence of SEQ ID NO: 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, or 25.

[0018] In some embodiments, the polynucleotide encoding the dCas9 comprises the nucleic acid sequence of SEQ ID NO: 26. In some embodiments, the polynucleotide encoding the N-terminal portion of the dCas9 comprises the amino acid sequence of SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, or 49. In some embodiments, the polynucleotide encoding the C-terminal portion of the dCas9 comprises the nucleic acid sequence of SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50.

[0019] In some embodiments, the polypeptide capable of binding an antibody or antigen binding fragment thereof is selected from the group consisting of GCN4, T2A, 10xGCN4, and 10xGFP-11. In some embodiments, the GCN4 comprises the amino acid sequence of SEQ ID NO: 97. In some embodiments, the polynucleotide encoding GCN4 comprises the nucleic acid sequence of SEQ ID NO: 101.

[0020] In some embodiments, the antibody or antigen binding fragment is a single domain antibody, scFv, Fab, VH, VHH, or antibody mimetic. In some embodiments, the antigen binding fragment is a scFv.

[0021] In some embodiments, the polypeptide sequence capable of binding a nucleic acid structural element is selected from the group consisting of MS2 bacteriophage coat protein (MCP), PP7, and PCP. In some embodiments, the polypeptide sequence capable of binding a nucleic acid structural element is MCP. In some embodiments, the MCP comprises the amino acid sequence of SEQ ID NO: 100. In some embodiments, the polynucleotide encoding MCP comprises the nucleic acid sequence of SEQ ID NO: 105.

[0022] In some embodiments, the nucleic acid structural element is an RNA hairpin motif. In some embodiments, the RNA hairpin motif is selected from the group consisting of MS2 and PP7. In some embodiments, the RNA hairpin motif is an MS2 RNA hairpin motif. In some embodiments, the MS2 RNA hairpin motif comprises the nucleic acid sequence of SEQ ID NO: 104.

[0023] In some embodiments, the construct comprises the nucleic acid sequence of SEQ ID NO: 134-162.

[0024] The present disclosure provides a polypeptide expressed by any one of the constructs described herein. The present disclosure provides a vector comprising any one of the constructs described herein.

[0025] In some embodiments, the vector further comprises a polynucleotide encoding a single guide RNA (sgRNA). In some embodiments, the polynucleotide encoding the sgRNA further comprises 2-20 copies of the nucleic acid structural element.

[0026] The present disclosure provides a cell comprising any one of the constructs described herein. The present disclosure provides a cell comprising any one of the polypeptides described herein. The present disclosure provides a cell comprising any one of the vectors described herein.

[0027] In some embodiments, the cell further comprises at least one sgRNA. In some embodiments, the at least one sgRNA further comprises 2-20 copies of the nucleic acid structural element.

[0028] The present disclosure provides a composition comprising any one of the constructs described herein. The present disclosure provides a composition comprising any one of the polypeptides described herein. The present disclosure provides a composition comprising any one of the vectors described herein.

[0029] In some embodiments, the composition further comprises at least one sgRNA. In some embodiments, the sgRNA further comprises 2-20 copies of a nucleic acid structural element. In some embodiments, the composition further comprises a pharmaceutically acceptable carrier.

[0030] The present disclosure provides a method of altering expression of a gene product in a population of cells and minimizing off-target modifications, comprising the step of introducing into the population of cells: i) any one of the constructs described herein or one or more polypeptides expressed by the construct; and ii) at least one sgRNA, wherein the KRAB and / or gene expression modulator provides a modification of at least one nucleotide in the vicinity of the gene and / or within a gene regulatory element, thereby altering expression of the gene product.

[0031] In some embodiments, (a) comprises a 5’-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3’ portion, 5’-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3’ portion, 5’-(A m1 -B m2 )-CasN-CasC-T p -3’ portion, 5’-(A-B)-CasN-CasC-T p -3’ portion, 5’-CasN-CasC-T p -3’ portion, 5’-CasN-CasC-T p -3’ portion, 5’-CasN-CasC-T pThe polypeptide of i) and the sgRNA of ii) in the -3' portion are recruited to genomic loci, and (b) multiple copies of the 5'-E-3' portion of formula I, the 5'-E-3' portion of formula III, the 5'-E-3' portion of formula IIIb, the 5'-Kr-Dq-3' portion of formula IIIb-1, the 5'-E-3' portion of formula IIIc, and the 5'-(A) of formula IIIc-1. m5 -B m6 ) n3 -K r -D q -3' part or 5'-K of formula IIIc-2 r -D q -(A m5 -B m6 ) n3 The polypeptide of part i) of the -3' portion is recruited to the genomic locus by binding of an antibody or its antigen-binding fragment to the epitope, thereby recruiting the polypeptide expressed by the construct to the genomic locus and altering the expression of the gene product in the cell population.

[0032] In some embodiments, the method further includes introducing into the cell: iii) 5'-(A) containing E. m5 -B m6 ) n3 -K r -D q -3' or 5'-K r -D q -(A m5 -B m6 ) n3 -3' second construct or polypeptide expressed by the second construct, containing 5'-CasN-CasC-T of formula IIIa p The polypeptide and sgRNA of the -3' portion i) are recruited to the genomic locus, and the multi-copy polypeptide of iii) is recruited to the genomic locus by binding of an antibody or its antigen-binding fragment to the epitope, thereby recruiting the polypeptide expressed by the constructs of i) and iii) to the genomic locus and altering the expression of gene products in the cell population.

[0033] In some implementations, (a) includes 5'-(A) of formula I. m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3' part, 5'-(A) of Formula III m1 -B m2 ) n1- CasN-(A m3 - B m4 ) n2 - CasC-T p - 3’ portion, the 5’-(A m1 - B m2 )- CasN-CasC-T p - 3’ portion, the 5’-(A-B)-CasN-CasC-T of Formula IIIb-1 p - 3’ portion, the 5’-CasN-CasC-T of Formula IIIc p - 3’ portion, the 5’-CasN-CasC-T of Formula IIIc-1 p - 3’ portion, or the 5’-CasN-CasC-T of Formula IIIc-2 p polypeptide of i) and the sgRNA of ii) are recruited to a genomic locus, and (b) multiple copies of a construct comprising a 5’-E-3’ portion of Formula I, a 5’-E-3’ portion of Formula III, a 5’-E-3’ portion of Formula IIIb, a 5’-Kr-Dq-3’ portion of Formula IIIb-1, a 5’-E-3’ portion of Formula IIIc, a 5’-(A m5 - B m6 ) n3 - K r - D q - 3’ portion, or the 5’-K r - D q - (A m5 - B m6 ) n3 polypeptide of i) is recruited to a genomic locus by binding of the polypeptide sequence capable of binding to a nucleic acid structural element to the nucleic acid structural element, thereby recruiting the polypeptide expressed by the construct to the genomic locus and altering expression of a gene product in the cell population.

[0034] In some embodiments, the method further comprises introducing into the cell: iii) a second construct comprising a 5’-(A m5 - B m6 ) n3 - K r - D q - 3’ or 5’-K r - D q - (A m5 - B m6 ) n3 second construct or a polypeptide expressed by the second construct, wherein the 5’-CasN-CasC-T of Formula IIIa p- the polypeptide of i) of the 3' portion and the sgRNA are recruited to the genomic locus, wherein the multiple copies of the polypeptide of iii) are recruited to the genomic locus by binding of the polypeptide sequence capable of binding to a nucleic acid structural element to the nucleic acid structural element, thereby recruiting the polypeptide expressed from the construct to the genomic locus and altering expression of the gene product in the population of cells.

[0035] The present disclosure provides an in vivo method of reducing or eliminating expression of a gene product in a subject comprising the step of introducing into a cell of the subject: i) any one of the constructs described herein or one or more polypeptides expressed from the construct; and ii) at least one sgRNA, wherein the KRAB and / or gene expression modulator provides a modification of at least one nucleotide in the vicinity of the gene and / or within a gene regulatory element, thereby altering expression of the gene product in the subject.

[0036] The present disclosure provides a method for treating or alleviating a symptom of a gene product-associated disorder in a subject comprising the step of introducing into a cell of the subject: i) any one of the constructs described herein or one or more polypeptides expressed from the construct; and ii) at least one sgRNA, wherein the KRAB and / or gene expression modulator provides a modification of at least one nucleotide in the vicinity of the gene and / or within a gene regulatory element, thereby altering expression of the gene product and treating or alleviating a symptom of a gene product-associated disorder in the subject.

[0037] In some embodiments, the expression of the gene product is reduced by 50-100% in the plurality of modified cells compared to a population of wild-type cells. In some embodiments, the ratio of on-site modification of the gene product to off-site modification of the gene product is about 10: 1.

[0038] In some embodiments, the modification of the at least one nucleotide is DNA methylation or histone modification. In some embodiments, the modification of the at least one nucleotide is DNA methylation.

[0039] In some embodiments, the gene regulatory element is a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, or a locus control region.

[0040] In some embodiments, the modification of at least one nucleotide near a gene and / or within a gene regulatory element is located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the transcription start site of the gene.

[0041] In some embodiments, the modification of at least one nucleotide near a gene and / or within a gene regulatory element is located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp downstream of the transcription start site of the gene.

[0042] In some embodiments, the construct is a deoxyribonucleic acid (DNA). In some embodiments, the construct is a messenger ribonucleic acid (mRNA). In some embodiments, the construct is formulated in a liposome or a lipid nanoparticle. In some embodiments, the construct and the sgRNA are formulated in a liposome or a lipid nanoparticle. In some embodiments, the construct and the sgRNA are formulated in the same liposome or lipid nanoparticle. In some embodiments, the construct and the sgRNA are formulated in different liposomes or lipid nanoparticles.

[0043] In some embodiments, the liposome or the lipid nanoparticle comprises an ionizable lipid (20-70%, molar ratio), a PEGylated lipid (0-30%, molar ratio), a supporting lipid (30-50%, molar ratio), and cholesterol (10-50%, molar ratio). In some embodiments, the ionizable lipid is selected from the group consisting of a pH-responsive ionizable lipid, a thermally-responsive ionizable lipid, and a light-responsive ionizable lipid.

[0044] In some embodiments, the construct is formulated in an AAV vector. In some embodiments, the construct and the sgRNA are formulated in an AAV vector. In some embodiments, the construct and the sgRNA are formulated in the same AAV vector. In some embodiments, the construct and the sgRNA are formulated in different AAV vectors.

[0045] In some embodiments, the construct or the one or more polypeptides expressed from the construct is delivered to the cell by local injection, systemic infusion, or a combination thereof.

[0046] In some embodiments, the gene product is selected from the group consisting of: a VEGFA gene product, a PCSK9 gene product, an ANGPTL3 gene product, a PTBP1 gene product, a TTR gene product, a Ube3a-ATS gene product, a Ptp1b gene product, an APOC3 gene product, a hsd17b13 gene product, a bcl11a gene product, and a TGF-beta gene product.

[0047] In some embodiments, the subject is a human. In some embodiments, the disease or disorder is selected from the group consisting of: familial hypercholesterolemia (FH), nonalcoholic steatohepatitis (NASH), Parkinson's disease, liver fibrosis (HF), age-related macular degeneration (AMD), Angelman syndrome (AS), type II diabetes, beta-thalassemia, and hepatocellular carcinoma. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figures 1A-1B is a schematic showing DNA interaction of constructs comprising DNMT3A, DNMT3L, KRAB, and dCas9 that is either intact or cleaved. Figure 2A A schematic of the construct is shown. Figure 2B A schematic of the construct is shown. Figures 2C-2E The editing efficiency of the constructs shown to suppress expression of two target genes with gRNAs targeting CD81 or CD151, respectively, where: Dnmt-dCas9-Krab represents Dnmt3A-Dnmt3L-dCas9-Krab; 768 INS represents dCas9N-2-Dnmt3A-Dnmt3L-dCas9C-2-Krab; 776 INS represents dCas9N-3-Dnmt3A-Dnmt3L-dCas9C-3-Krab; 1009 INS represents dCas9N-4-Dnmt3A-Dnmt3L-dCas9C-4-Krab; 1048-1063 INS represents dCas9N-5-Dnmt3A-Dnmt3L-dCas9C-5-Krab; 1072 INS represents dCas9N-6-Dnmt3A-Dnmt3L-dCas9C-6-Krab; 1246 INS represents dCas9N-7-Dnmt3A-Dnmt3L-dCas9C-7-Krab; 1248 INS represents dCas9N-8-Dnmt3A-Dnmt3L-dCas9C-8-Krab; 1260 INS represents dCas9N-10-Dnmt3A-Dnmt3L-dCas9C-10-Krab; Control represents Dnmt3A-Dnmt3L-dCas9-Krab with Nt gRNA.

[0049] Figures 2A-2B is a schematic showing that DNMT3A, DNMT3L, and KRAB are recruited to dCas9 by scFv binding to 10xGCN4. Figure 3A schematics of constructs are shown. Figure 3B schematics of constructs are shown. Figure 4A The editing efficiency of constructs shown to suppress the expression of two target genes with gRNAs targeting CD81 or CD151, respectively, wherein: Suntag means Dnmt3A-Dnmt3L-dCas9-10xGCN4-T2A-scFv-KRAB; Dnmt-dCas9+scFv-Krab also means Dnmt3A-Dnmt3L-dCas9-10xGCN4-T2A-scFv-KRAB; dCas9+scFv-Dnmt-Krab means dCas9-10xGCN4-T2A-scFv-Dnmt3A-Dnmt3L-KRAB; dCas9+scFv-Dnmt+scFv-Krab means dCas9-10xGCN4-T2A-scFv-Dnmt3A-Dnmt3L-scFv-KRAB; Control means Dnmt3A-Dnmt3L-dCas9-Krab in Figure 1 with Nt gRNA.

[0050] Figure 4B is a schematic showing a construct comprising dCas9 and a guide RNA containing MS2, which binds to DNMT3A, DNMT3L, and KRAB via MCP. Figures 4C-4D schematics of constructs comprising DNMT3A, DNMT3L, MCP, dCas9, and KRAB are shown.

[0051] Figures 4A-4B is a schematic showing the interaction between a complex comprising dCas9, DNMT3A, and DNMT3L, and a gene expression regulator selected from the group consisting of Kruppel-associated box (KRAB), enhancer of zeste homolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), friend of GATA protein 1 (FOG1), histone deacetylase (HDAC3), and / or DOT1L, and their DNA target sites. Figures 4E-4FA schematic showing a construct comprising dCas9, DNMT3A and DNMT3L, and a gene expression modulator selected from Kruppel-associated box (KRAB), enhancer of zeste homolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), GATA protein friend of GATA 1 (FOG1), histone deacetylase (HDAC3), and / or DOT1L. Figures 4A-4B A schematic showing a construct comprising dCas9, DNMT3A and DNMT3L, and a gene expression modulator selected from Kruppel-associated box (KRAB), enhancer of zeste homolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), GATA protein friend of GATA 1 (FOG1), histone deacetylase (HDAC3), and / or DOT1L. Figure 5A The editing efficiency of the construct shown to inhibit target gene expression, wherein the gene expression modulator is selected from the types listed on the right side of the figure. Figure 5B The editing efficiency of the construct shown to inhibit target gene expression, wherein the gene expression modulator is selected from the types listed on the right side of the figure. Figures 5C-5D The editing efficiency of the construct shown to inhibit target gene expression, wherein the KRAB modulator is selected from the sources listed on the right side of the figure. Control refers to Dnmt3A-Dnmt3L-dCas9-Krab in Figure 1 with Nt gRNA.

[0052] Figures 5A-5B A schematic showing the interaction between complexes comprising dCas9, DNMT3A and DNMT3L from different species. Figure 6A A schematic showing a construct. Figure 6B The editing efficiency of the construct shown to inhibit target gene expression, wherein the KRAB modulator is selected from the sources listed on the right side of the figure. Control refers to Dnmt3A-Dnmt3L-dCas9-Krab in Figure 1 with Nt gRNA. Figure 1B The editing efficiency of the construct shown to inhibit target gene expression, wherein the KRAB modulator is selected from the sources listed on the right side of the figure. Control refers to Dnmt3A-Dnmt3L-dCas9-Krab in Figure 1 with Nt gRNA.

[0053] Figure 2B Three examples of constructs of the application are shown. Figure 3B The editing efficiency of the three constructs shown to inhibit target gene expression. DETAILED DESCRIPTION

[0054] The present disclosure overcomes problems associated with current technology by providing constructs comprising DNMT3A, DNMT3L, dCas9 and KRAB for targeted modification of gene product expression. For example, targeted modification of gene products in cells for in vivo gene therapy.

[0055] TERMINOLOGY

[0056] As used herein, the term "coding sequence" or "coding nucleic acid" means a nucleic acid (RNA or DNA molecule) comprising a nucleotide sequence that encodes a protein. The coding sequence can also comprise initiation and termination signals operably linked to regulatory elements, including a promoter and polyadenylation signal capable of directing expression in the cells of an individual or mammal to which the nucleic acid is administered. The coding sequence can be codon optimized.

[0057] The term "complement" or "complementary" as used herein with respect to nucleic acids can mean Watson-Crick (e.g., A-T / U and C-G) or Hoogsteen base pairing between the nucleotides or nucleotide analogs of a nucleic acid molecule. "Complementarity" refers to the property of two nucleic acid sequences shared, which, when placed in close proximity to each other in an antiparallel arrangement, the nucleotide bases at each position will be complementary.

[0058] The terms "correcting," "genome editing," and "restoring" refer to altering a mutant gene that encodes a mutant protein, a truncated protein, or no protein at all, to obtain full-length functional or partial full-length functional protein expression. Correcting or restoring a mutant gene can include using a repair mechanism, such as homology directed repair (HDR), to replace a region of the gene with a mutation or to replace the entire mutant gene with a copy of the gene that does not have the mutation. Correcting or restoring a mutant gene can also include repairing a frameshift mutation that results in a premature stop codon, an aberrant splice acceptor site, or an aberrant splice donor site by creating a double-stranded break in the gene and then repairing it using non-homologous end joining (NHEJ). NHEJ can add or delete at least one base pair during repair, which can restore the correct reading frame and eliminate the premature stop codon. Correcting or restoring a mutant gene can also include disrupting an aberrant splice acceptor site or splice donor sequence. Correcting or restoring a mutant gene can also include deleting an unnecessary gene segment by the simultaneous action of two nucleases on the same DNA strand so that the correct reading frame is restored by removing the DNA between the two nuclease target sites and repairing the DNA break by NHEJ.

[0059] As used herein, the terms "donor DNA," "donor template," and "repair template" refer to a double-stranded DNA fragment or molecule that comprises at least a portion of a gene of interest. The donor DNA can encode a fully functional protein or a partially functional protein.

[0060] As used herein, the terms "frameshift" or "frameshift mutation" are used interchangeably to refer to a type of genetic mutation in which the addition or deletion of one or more nucleotides results in a shift in the reading frame of codons in the mRNA. The shift in reading frame can result in a change in the amino acid sequence of the protein as it is translated, such as a missense mutation or a premature stop codon.

[0061] As used herein, the terms "functional" and "fully functional" describe a protein that has biological activity. A "functional gene" refers to a gene that is transcribed into mRNA that is translated into a functional protein.

[0062] As used herein, the term "fusion protein" refers to a chimeric protein produced by the direct or indirect covalent or non-covalent joining of two or more genes that originally encoded separate proteins. In some embodiments, translation of the fusion gene results in a single polypeptide with functional properties derived from each of the original proteins.

[0063] As used herein, the term "genetic construct" refers to a DNA or RNA molecule comprising a nucleotide sequence encoding a protein. The coding sequence includes initiation and termination signals operably linked to regulatory elements including a promoter and polyadenylation signal capable of directing expression in a cell.

[0064] The terms "homology directed repair" or "HDR" used interchangeably herein refer to a mechanism by which cells repair double-stranded DNA damage when a homologous DNA fragment is present in the nucleus, primarily during the G2 and S phases of the cell cycle. HDR uses a donor DNA template to direct repair and can be used to make specific sequence changes to the genome, including targeted addition of entire genes. If the donor template is provided with a site-specific nuclease, such as with a CRISPR / Cas9-based system, then the cellular machinery will repair the break by homologous recombination, which is enhanced several orders of magnitude in the presence of DNA cleavage. When a homologous DNA fragment is not present, non-homologous end joining can occur instead.

[0065] The term "genome editing" as used herein refers to altering a gene. Genome editing can include correcting or restoring a mutated gene. Genome editing can include knocking out a gene, such as a mutated gene or a normal gene. Genome editing can treat a disease by altering a gene of interest.

[0066] The terms "identical" or "identity" in the context of two or more nucleic acid or polypeptide sequences, refer to the sequences having a specified percentage of residues that are the same. The percentage of sequence identity can be readily determined by optimally aligning the two sequences and counting the number of positions at which the identical residue occurs in both sequences. The percentage of sequence identity is then determined by comparing the number of identical positions to the total number of positions in the specified comparison region. If the two sequences are of different lengths or if the alignment produces one or more staggered ends, the residues of the single sequence that are not included in the comparison are not included in the denominator of the calculation. When comparing DNA and RNA, thymine (T) and uracil (U) can be considered equivalent. Identity can be performed manually or by using a computer algorithm, such as BLAST or BLAST 2.0. Identity of related peptides can be readily calculated by known methods. Such methods include, but are not limited to, those described in Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part 1, Griffin, A. M. and Griffin, H. G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M. Stockton Press, New York, 1991; and Carillo et al., SIAM J. Applied Math. 48, 1073 (1988), which are incorporated by reference herein in their entireties.

[0067] As used herein, the terms "mutant gene" or "mutated gene" are used interchangeably herein to refer to a gene that has undergone a detectable mutation. A mutant gene has undergone an alteration, such as a loss, gain, or exchange of genetic material, which affects the normal transmission and expression of the gene. As used herein, a "disrupted gene" refers to a mutant gene having a mutation that results in a premature stop codon. The disrupted gene product is truncated relative to the full-length non-disrupted gene product.

[0068] As used herein, the term "modulator of epigenetic modification" refers to an agent that targets gene expression through epigenetic modification (e.g., through histone acetylation or methylation, or DNA methylation at regulatory elements of a target gene, such as a promoter, enhancer, or transcriptional start site). Chromatin remodeling and DNA methylation are two major mechanisms that regulate gene transcription. Specific epigenetic marks (e.g., DNA methylation) structurally or biochemically direct gene transcription or gene silencing / repression. For example, DNA methylation of regions that modulate transcriptional activity alters gene expression without changing the underlying DNA sequence. Transcriptional modulation using epigenetic modifications (e.g., DNA methylation) allows for targeted modulation of gene expression without affecting the expression of other gene products.

[0069] As used herein, the term "non-homologous end joining (NHEJ) pathway" refers to a pathway that repairs double-stranded breaks in DNA by directly joining the broken ends without the need for a homologous template. Template-independent re-joining of DNA ends by NHEJ is a random, error-prone repair process that introduces random microinsertions and microdeletions (indels) at the DNA breakpoint. This method can be used to intentionally disrupt, delete, or alter the reading frame of a target gene sequence. NHEJ often uses short homologous DNA sequences, known as microhomologies, to direct repair. These microhomologies often occur in single-stranded overhangs at the ends of double-stranded breaks. When the overhangs are perfectly compatible, NHEJ often repairs the break exactly, however, imprecise repair that results in loss of nucleotides can also occur, but is much more common when the overhangs are incompatible.

[0070] As used herein, the term "normal gene" refers to a gene that has not undergone an alteration, such as a loss, gain, or exchange of genetic material. A normal gene undergoes normal gene transmission and gene expression.

[0071] As used herein, the term "nuclease-mediated NHEJ" refers to NHEJ initiated after a nuclease, such as cas9, cuts double-stranded DNA.

[0072] As used herein, the term "nucleic acid" or "oligonucleotide" or "polynucleotide" refers to at least two nucleotides covalently linked together. The description of a single strand also defines the sequence of the complementary strand. Thus, a nucleic acid also includes the complementary strand of a described single strand. Numerous variants of a nucleic acid can be used for the same purpose as a given nucleic acid. Thus, a nucleic acid also includes substantially identical nucleic acids and their complements. A single strand provides a probe that can hybridize to a target sequence under stringent hybridization conditions. Thus, a nucleic acid also includes probes that hybridize under stringent hybridization conditions. A nucleic acid can be single-stranded or double-stranded, or can contain portions of both double-stranded and single-stranded sequences. A nucleic acid can be DNA (genomic DNA and cDNA), RNA, or a hybrid, wherein the nucleic acid can contain combinations of deoxyribonucleotides and ribonucleotides, and combinations of bases including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. A nucleic acid can be obtained by chemical synthesis methods or recombinant methods.

[0073] As used herein, the term "operably linked" means that the expression of a gene is under the control of a promoter spatially connected thereto. The promoter can be located 5' (upstream) or 3' (downstream) of the gene it controls. The distance between the promoter and the gene can be approximately the same as the distance between the promoter and the gene it originates from that it controls. Variations in this distance can be accommodated without loss of promoter function, as is known in the art.

[0074] The term "partially functional" as used herein describes a protein encoded by a mutant gene that has a biological activity lower than that of a functional protein, but higher than that of a non-functional protein. In one embodiment, a partially functional protein exhibits a biological activity that is lower than 95%, 90%, 85%, 80%, 75%, 70%, 65%, 60%, 55%, 50%, 45%, 40%, 35%, or 30% of the biological activity of the corresponding functional protein.

[0075] The terms "premature stop codon" or "out-of-frame stop codon" used interchangeably herein refer to a nonsense mutation in a DNA sequence that results in a stop codon at a position that does not normally exist in the wild-type gene. A premature stop codon can result in a protein being truncated or shortened compared to the full-length form of the protein.

[0076] The term "promoter" as used herein refers to a synthetic or naturally occurring molecule capable of conferring, activating or enhancing expression of a nucleic acid in a cell. A promoter can comprise one or more specific transcriptional regulatory sequences to further enhance expression of the nucleic acid and / or to alter spatial expression and / or temporal expression of the nucleic acid. A promoter can also comprise distal enhancer or repressor elements, which can be located up to several kilobases from the transcriptional start site. Promoters can be derived from sources including viruses, bacteria, fungi, plants, insects and animals. A promoter can constitutively regulate expression of a genetic component, or differentially regulate expression of a genetic component relative to the cell, tissue or organ in which expression occurs, or relative to the stage of development at which expression occurs, or in response to an external stimulus such as a physiological stress, pathogen, metal ion or an inducing agent. Representative examples of promoters include the bacteriophage T7 promoter, the bacteriophage T3 promoter, the SP6 promoter, the lac operator-promoter, the tac promoter, the SV40 late promoter, the SV40 early promoter, the RSV-LTR promoter and the CMV IE promoter.

[0077] The term "target gene" as used herein refers to any nucleotide sequence that encodes a known or putative gene product. A target gene can be a mutated gene associated with a genetic disease or disorder.

[0078] The term "target region" as used herein refers to a region of a target gene to which a site-specific nuclease is designed to bind.

[0079] The term "transgene" as used herein refers to a gene or genetic material that contains a sequence of genes that has been isolated from one organism and introduced into a different organism. Alternatively, the term "transgene" also refers to a gene or genetic material that is chemically synthesized and introduced into an organism. Such non-native segments of DNA can retain the ability to produce RNA or proteins in the transgenic organism, or it can alter the normal functioning of the genetic code of the transgenic organism. Introduction of a transgene has the potential to alter the phenotype of the organism.

[0080] The term "variant" as used herein when used in reference to a nucleic acid means (i) a portion or fragment of a reference nucleotide sequence; (ii) a complement of a reference nucleotide sequence or portion thereof; (iii) a nucleic acid that is substantially identical to the reference nucleic acid or complement thereof; or (iv) a nucleic acid that hybridizes under stringent conditions to the reference nucleic acid, complement thereof, or to a sequence substantially identical thereto. A "variant" refers to a peptide or polypeptide that differs in amino acid sequence by insertion, deletion or conservative substitution, but which retains at least one biological activity.

[0081] A variant can also mean a protein having an amino acid sequence that is substantially identical to a reference protein having an amino acid sequence that retains at least one biological activity. Conservative substitution of amino acids, i.e., replacing an amino acid with a different amino acid of similar properties (e.g., hydrophilicity, degree and distribution of charged regions) is recognized in the art as typically involving a minor change. As understood in the art, these minor changes can be identified by considering the hydropathic index of amino acids. Kyte et al., J. Mol. Biol. 157: 105-132 (1982), incorporated by reference in its entirety herein. The hydropathic index of an amino acid is based on a consideration of its hydrophobicity and charge. It is known in the art that amino acids of similar hydropathic index can be substituted without a significant loss of protein function. In one aspect, an amino acid having a hydropathic index of ±2 is substituted. The hydrophilicity of amino acids can also be used to reveal substitutions that will result in proteins that retain biological function. A consideration of the hydrophilicity of amino acids in the context of a peptide permits the calculation of the greatest local average hydrophilicity of that peptide. Substitutions can be made with amino acids having hydropathic values within ±2 of each other. Both the hydrophobicity index and the hydrophilicity values of amino acids are influenced by the particular side chain of that amino acid. Consistent with this observation, amino acid substitutions that are compatible with biological function are understood to depend on the relative similarity of the amino acids, as revealed by the hydrophobicity, hydrophilicity, charge, size, and other properties.

[0082] As used herein, the term "vector" refers to a nucleic acid sequence that contains an origin of replication. The vector can be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. The vector can be a DNA or RNA vector. The vector can be a self-replicating extrachromosomal vector, such as a DNA plasmid.

[0083] As used herein, the terms "gene transfer," "gene delivery," and "gene transduction" refer to a method or system for reliably inserting a particular nucleotide sequence (e.g., DNA or RNA), fusion protein, polypeptide, etc. into a target cell.

[0084] As used herein, the terms "adeno-associated virus (AAV) vector," "AAV gene therapy vector," and "gene therapy vector" refer to a vector having functional or partially functional ITR sequences and a transgene. As used herein, the term "ITR" refers to inverted terminal repeat sequences (ITRs). The ITR sequences can be derived from an adeno-associated virus serotype, including but not limited to AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, and AAV-6. However, the ITRs need not be wild-type nucleotide sequences, and can be altered (e.g., by insertion, deletion, or substitution of nucleotides), so long as the sequence retains the function of providing functional rescue, replication, and packaging. An AAV vector can have one or more AAV wild-type genes (preferably rep and / or cap genes) that are all or partially deleted but retain functional flanking ITR sequences. Functional ITR sequences function, for example, to rescue, replicate, and package AAV virions or particles. Thus, an "AAV vector" is defined herein to include at least those sequences required for insertion of a transgene into a subject's cell. Optionally included are those cis-sequences necessary for viral replication and packaging (e.g., functional ITRs).

[0085] As used herein, the term "gene therapy" refers to a method of treating a patient in which a polypeptide or nucleic acid sequence is transferred into the patient's cells, thereby modulating the activity and / or expression of a particular gene. In certain embodiments, the expression of a gene is inhibited. In certain embodiments, the expression of a gene is enhanced. In certain embodiments, the temporal or spatial pattern of gene expression is modulated.

[0086] A "transgene" can comprise a transgenic sequence or a native or wild-type DNA sequence. A transgene can become part of the genome of a primate subject. A transgene sequence can be partially or entirely species heterologous, i.e., the transgene sequence or portion thereof can be from a different species than the cell into which it is introduced.

[0087] As used herein, the term "stably maintained" refers to a characteristic of a transgenic subject (e.g., a human or non-human primate) that maintains at least one of its transgenic elements (i.e., the desired element) over multiple generations of cells. For example, the term is intended to include many cell division cycles of the originally transfected cell. The term "stably transfected" or "stably trans fected" refers to the introduction of foreign DNA into a cell's genome and integration. The term "stable transfectant" refers to a cell that has stably integrated foreign DNA into its genomic DNA.

[0088] As used herein, the terms "transgene encodes," "nucleic acid molecule encodes," "DNA sequence encodes," and "DNA encodes" refer to the order or sequence of deoxyribonucleotides along a deoxyribonucleic acid strand. For example, the order of these deoxyribonucleotides can dictate the order of amino acids along a polypeptide (protein) strand. Thus, a DNA sequence can encode an amino acid sequence.

[0089] As used herein, the term "wild type" (wt) refers to a gene or gene product that has the characteristics of that gene or gene product as it is found in a natural source. A wild type gene is the most commonly observed gene in a population and is arbitrarily designed the "normal" or "wild type" form of the gene. In contrast, the term "modified" or "mutant" refers to a gene or gene product that exhibits a modification (i.e., an altered characteristic) in sequence and / or functional properties when compared to a wild type gene or gene product. Note that naturally occurring mutants can be isolated that are identified by acquiring an altered characteristic when compared to a wild type gene or gene product.

[0090] As used herein, the term "transfection" refers to the uptake of foreign nucleic acid (e.g., DNA or RNA) by a cell. A cell is "transfected" when exogenous nucleic acid (DNA or RNA) has been introduced into the cell membrane. Numerous transfection techniques are well known in the art (see, e.g., Graham et al., Virol., 52:456 (1973); Sambrook et al., Molecular Cloning, a Laboratory Manual, Cold Spring Harbor Laboratories, New York (1989); Davis et al., Basic Methods in Molecular Biology, Elsevier, (1986); and Chu et al., Gene 13:197 (1981), incorporated by reference in their entireties herein). Such techniques can be used to introduce one or more exogenous DNA moieties, such as gene transfer vectors and other nucleic acid molecules, into suitable recipient cells.

[0091] As used herein, the terms "stably transfected" and "stably transfected" refer to the introduction and integration of foreign DNA into the genome of a transfected cell. The term "stable transfectant" refers to a cell that has stably integrated foreign DNA into the genomic DNA.

[0092] As used herein, the term "transient transfection" or "transient transfected" refers to the introduction of foreign DNA into a cell, where the foreign DNA fails to integrate into the genome of the transfected cell and is maintained as an episome. During this time, the foreign DNA is subject to the regulatory controls of the endogenous genes in the governing chromosome. The term "transient transfectant" refers to a cell that has taken up foreign DNA but failed to integrate the DNA. As used herein, the term "transduction" denotes the delivery of a DNA molecule to a recipient cell, either in vivo or in vitro, by a replication-defective viral vector, such as by a recombinant AAV virion.

[0093] As used herein, the term "recipient cell" refers to a cell that has been or is capable of being transfected or transduced with a nucleic acid construct or vector carrying a selected nucleotide sequence of interest. The term includes the progeny of the parent cell, whether or not the progeny is identical to the original parent cell in morphology or in genomic makeup, so long as the selected nucleotide sequence is present. A recipient cell can be a cell of a subject to which a gene therapy particle and / or a gene therapy vector has been administered.

[0094] As used herein, the term "recombinant DNA molecule" refers to a DNA molecule composed of segments of DNA joined together by means of molecular biological techniques.

[0095] As used herein, the term "regulatory element" refers to a genetic element that controls the expression of a nucleic acid sequence. For example, a promoter is a regulatory element that promotes the initiation of transcription of an operably linked coding region. Other regulatory elements are splicing signals, polyadenylation signals, termination signals, and the like.

[0096] The term DNA "control sequences" collectively refers to regulatory elements such as promoter sequences, polyadenylation signals, transcription termination sequences, upstream regulatory domains, origins of replication, internal ribosome entry sites ("IRES"), enhancers, and the like, which collectively provide for replication, transcription, and translation of a coding sequence in a recipient cell. Not all of these control sequences need to be present.

[0097] Transcription control signals in eukaryotes generally include "promoter" and "enhancer" elements. Promoters and enhancers are composed of short arrays of DNA sequences that specifically interact with cellular proteins involved in transcription (Maniatis et al., Science 236:1237 (1987), incorporated herein by reference). Promoter and enhancer elements have been isolated from genes in a variety of eukaryotic sources, including yeast, insect and mammalian cells, as well as viruses (similar control sequences, i.e., promoters, also exist in prokaryotes). The choice of particular promoters and enhancers depends on the recipient cell type. Some eukaryotic promoters and enhancers have a broad host range, while other promoters and enhancers are functional in a limited subset of cell types (for reviews, see, e.g., Voss et al., Trends Biochem. Sci., 11 :287 (1986); and Maniatis et al., supra, incorporated herein by reference in their entireties). For example, the SV40 early gene enhancer is very active in a wide variety of cell types from many mammalian species and has been used to express proteins in a broad range of mammalian cells (Dijkema et al., EMBO J. 4:761 (1985), incorporated herein by reference in its entirety). Promoter and enhancer elements derived from the human elongation factor 1-alpha gene (Uetsuki et al., J. Biol. Chem., 264:5791 (1989); Kim et al., Gene 91 :217 (1990); and Mizushima and Nagata, Nucl. Acids. Res., 18:5322 (1990)), the long terminal repeat of the Rous Sarcoma Virus (Gorman et al., Proc. Natl. Acad. Sci. U.S.A. 79:6777 (1982)), and the human cytomegalovirus (Boshart et al., Cell 41 :521 (1985)), incorporated herein by reference in their entireties, can also be used to express proteins in different mammalian cell types. Promoters and enhancers can occur naturally, either alone or together. For example, the retroviral long terminal repeat contains both promoter and enhancer elements. In general, promoters and enhancers can act independently of the gene being transcribed or translated. Thus, the enhancer and promoter used can be "endogenous," "exogenous," or "heterologous" with respect to the gene with which they are operably linked. An "endogenous" enhancer / promoter is one that is naturally associated with a given gene in the genome. An "exogenous" or "heterologous" enhancer or promoter is one that is placed in juxtaposition to a gene by genetic manipulation, i.e., molecular biology techniques, such that transcription of that gene is directed by the linked enhancer / promoter.

[0098] As used herein, the term "tissue-specific" refers to a regulatory element or control sequence, such as a promoter, enhancer, etc., in which expression of a nucleic acid sequence is significantly higher in one or more specific cell types or tissues.

[0099] The presence of a "splicing signal" on an expression vector generally results in higher levels of recombinant transcript expression. Splicing signals mediate the removal of introns from primary RNA transcripts, consisting of splice donor and acceptor sites (Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nded., Cold Spring Harbor Laboratory Press, New York (1989), pp. 16.7-16.8, incorporated herein by reference in its entirety). A commonly used splice donor and acceptor site is the splice junction from the 16S RNA of SV40.

[0100] A transcription termination signal is usually present downstream of the polyadenylation signal and is several hundred nucleotides in length. As used herein, the term "poly A site" or "poly A sequence" means a DNA sequence that directs termination and polyadenylation of a nascent RNA transcript. Efficient polyadenylation of a recombinant transcript is desirable because transcripts lacking a poly A tail are unstable and are rapidly degraded. The poly A signal used in an expression vector can be "heterologous" or "endogenous." An endogenous poly A signal is a poly A signal that naturally occurs at the 3' end of the coding region of a given gene in the genome. A heterologous poly A signal is a poly A signal that has been isolated from one gene and operably linked to the 3' end of another gene. A commonly used heterologous poly A signal is the SV40 poly A signal. The SV40 poly A signal is contained on a 237 bp BamHI / BcII restriction fragment and directs termination and polyadenylation (Sambrook et al., supra, at 16.6-16.7, incorporated herein by reference in its entirety).

[0101] As used herein, the terms "subject" and "patient" are used interchangeably herein to refer to humans and non-human animals. The term "non-human animals" of the present disclosure includes all vertebrates, e.g., mammals and non-mammals, such as non-human primates, sheep, dogs, cats, horses, cows, chickens, amphibians, reptiles, etc.

[0102] As defined herein, a "therapeutically effective amount" or "therapeutically effective dose" is an amount or dosage of a fusion protein, polypeptide, nucleic acid, lipid nanoparticle, liposome, AAV particle(s), or virion(s) that is capable of producing a sufficient amount of a desired protein to modulate protein activity in a desired manner (thereby providing a means for clinical intervention). In some embodiments, a therapeutically effective amount or dose of a transfected fusion protein, polypeptide, nucleic acid, AAV particle(s), or virion(s) as described herein is sufficient to inhibit a gene targeted by the fusion protein / gene therapy construct.

[0103] As used herein, the term "treat," e.g., a disorder, means a subject (e.g., a human) having, at risk of having, and / or experiencing symptoms of a disorder, in embodiments, will have less severe symptoms and / or will recover faster when administered a fusion molecule or a nucleic acid encoding a fusion molecule and / or a gRNA or a nucleic acid encoding a gRNA (e.g., as described herein) compared to not being administered a fusion molecule or a nucleic acid encoding a fusion molecule and / or a gRNA or a nucleic acid encoding a gRNA. DETAILED DESCRIPTION

[0105] Construct

[0106] In one aspect, provided herein is a construct comprising a polynucleotide encoding DNMT3A or a portion thereof, a polynucleotide encoding DNMT3L or a portion thereof, a polynucleotide encoding dCas9, a polynucleotide encoding KRAB or a portion thereof, a polynucleotide encoding a gene expression modulator, and a polynucleotide encoding (i) an epitope capable of binding to an antibody or antigen binding fragment thereof or (ii) a polypeptide sequence capable of binding to a nucleic acid structural element. In some embodiments, the construct provided herein has the structure of Formula I: 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3'.

[0107] In Formula I, one of A and B is a polynucleotide encoding DNMT3A or a portion thereof; and the other of A and B is a polynucleotide encoding DNMT3L or a portion thereof. CasN and CasC are polynucleotides encoding an N-terminal portion of dCas9 and a C-terminal portion of dCas9, respectively. E is 5'-(A m5 -B m6 ) n3 -K r -D q -3' or 5'-K r -D q -(Am5 -B m6 ) n3 -3'. K is a polynucleotide encoding KRAB or a portion thereof. D is a polynucleotide encoding a gene expression regulator. The gene expression regulator can be, for example, Kruppel-associated box (KRAB), enhancer of zete homeolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), GATA friend of GATA 1 (FOG1), histone deacetylase (HDAC3), and / or DOT1L. T comprises a polynucleotide encoding (i) an epitope capable of binding to an antibody or antigen-binding fragment thereof (e.g., a single domain antibody, scFv, Fab, VH, VHH, or antibody mimetic) or (ii) a polypeptide sequence capable of binding to a nucleic acid structural element. The polypeptide capable of binding to an antibody or antigen-binding fragment thereof can be, for example, GCN4. The polypeptide sequence capable of binding to a nucleic acid structural element can be, for example, MS2 bacteriophage coat protein (MCP). Polypeptides encoding DNMT3A, DNMT3L, CasN, CasC, KRAB, and gene expression regulators are described in more detail below.

[0108] In Formula I, each of ml, m2, m3, m4, m5, and m6 is independently an integer selected from 0 to 3; each of nl, n2, and n3 is independently an integer selected from 0 to 2; p is an integer selected from 0 to 20; q is an integer selected from 0 to 5; r is an integer selected from 0 to 5; and when p is 0, at least one of ml, m2, m3, and m4 is not 0, at least one of nl and n2 is not 0, and at least one of q and r is not 0.

[0109] In some embodiments, the constructs provided herein have the structure of Formula II: 5'-CasN-(A m3 -B m4 ) n2 -CasC-K r -D q -3', wherein: n2 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; q is an integer selected from 0 to 5; and at least one of m3 and m4 is not 0. In some embodiments, the construct comprises the structure of Formula Ila: 5'-CasN-(A-B)-CasC-K r -D q -3'. Exemplary constructs having the structure of Formula Ila are shown in Figure 2B the second, third, and fourth schematic drawings.

[0110] In some embodiments, the constructs provided herein have the structure of Formula III: 5'-(A m1 -B m2 )n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3’, wherein p is an integer selected from 1 to 20. An exemplary construct having the structure of Formula III is shown in Figure 3B and Figure 2B In some embodiments, the construct has the structure of Formula Ilia: 5’-CasN-CasC-T p -3’. In the case of the construct of Formula Ilia, DNMT3A, DNMT3L, and KRAB are expressed on separate constructs.

[0111] In some embodiments, the constructs provided herein have the structure of Formula Illb: 5’-(A m1 -B m2 )-CasN-CasC-T p -E-3’, wherein at least one of ml and m2 is other than 0. An exemplary construct having the structure of Formula Illb is shown in the second schematic in Figure 3B and the second schematic in dCas9 is shown in bold underlined In some embodiments, the construct has the structure of Formula Illb-l: 5’-(A-B)-CasN-CasC-T p -K r -D q -3’, wherein: r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5.

[0112] In some embodiments, the constructs provided herein have the structure of Formula IIIc: 5’-CasN-CasC-T p -E-3’, wherein: n3 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5; and at least one of m5 and m6 is other than 0. An exemplary construct having the structure of Formula IIIc is shown in the third and fourth schematics in DNMT3A is shown in italic and the third and fourth schematics in DNMT3L is shown underlined In some embodiments, the construct has the structure of Formula IIIc-l: 5’-CasN-CasC-T p -(A m5 -B m6 ) n3 -K r -D q -3’. In some embodiments, the construct has the structure of Formula IIIc-2: 5’-CasN-CasC-T p -K r -D q -(Am5 - B m6 ) n3 -3’.

[0113] The sequences of the constructs described herein are shown in Table 1, wherein KRAB is shown in bold italic , (2) DNMT3A-DNMT3L-dCas9-KRAB (ZIM3) RTLVTFEDVTVNFTQGEWQRLNPEQRNLYRDVMLENYSNLVSVGQGETTKPDVILRLEQGKEPWL (SEQ ID NO: 204) , RTLLTFRDVAIEFSLEEWQCLDTAQRNLYRKVMFENYRNLVFLGIAVSKPHLITCLEQGKEPWN (SEQ ID NO: 205) , RTLVTFEDVSMDFSQEEWELLEPAQKNLYREVMLENYRNVVSLEALKNQCTDVGIKEGPLSPAQT (SEQ ID NO: 207) , EZH2, G9A, LSD1, HP1, FOG1, DHAC3, DOT1L, 10xGCN4, MCP and ScFV are shown in bold.

[0114] Table 1: Exemplary construct sequences

[0115]

[0116]

[0117]

[0118]

[0119]

[0120]

[0121]

[0122]

[0123]

[0124]

[0125]

[0126]

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229] wherein the nucleic acid sequence of KRAB (KOX1) is:

[0230] CGGACACTGGTGACCTTCAAGGATGTATTTGTGGACTTCACCAGGGAGGAGTGGAAGCTGCTGGACACTGCTCAGC AGATCGTGTACAGAAATGTGATGCTGGAGAACTATAAGA ACCTGGTTTCCTTGGGTTATCAGCTTACTAAGCCAGATGTGATCCTCCGGTTGGAGAAGGGAGAAGAGCCC (SEQ ID NO: 52).

[0231] RTLVTFDDVAVTFTKEEWGQLDLAQRTLYQEVMLENCGLLVSLGCPVPKAELICHLEHGQEPWT (SEQ ID NO: 209)

[0232] RTLVTFEDVTVNFTQGEWQRLNPEQRNLYRDVMLENYSNLVSVGQGETTKPDVILRLEQGKEPWL (SEQ ID NO: 204) RTLLTFRDVAIEFSLEEWQCLDTAQRNLYRKVMFENYRNLVFLGIAVSKPHLITCLEQGKEPWN (SEQ ID NO: 205) ,

[0233] wherein the nucleic acid sequence of KRAB(ZIM3) is: CGGACACTGGTGACCTTCGAGGACGTGACCGTGAACTTCACCCAAGGCGAGTGGCAGAGACTGAACCCCGAGCAGAGAAACCTGTACCGGGACGTGATGCTGGAAAACTACAGCAACCTGGTGTCCGTCGGCCAGGGCGAGACAACAAAGCCTGACGTGATCCTGCGGCTGGAACAGGGCAAAGAACCTTGGCTG (SEQ ID NO: 54).

[0234] (3) DNMT3A-DNMT3L-dCas9-KRAB(ZNF680)

[0235] RTLVTFEDVSMDFSQEEWELLEPAQKNLYREVMLENYRNVVSLEALKNQCTDVGIKEGPLSPAQT (SEQ ID NO: 207) RTLVTFDDVAVTFTKEEWGQLDLAQRTLYQEVMLENCGLLVSLGCPVPKAELICHLEHGQEPWT (SEQ ID NO: 209) ,

[0236] wherein the nucleic acid sequence of KRAB (ZNF680) is: CGGACACTGCTGACCTTCAGAGATGTGGCCATCGAGTTCAGCCTGGAAGAGTGGCAGTGCCTGGACACAGCCCAGCGGAACCTGTACAGAAAAGTGATGTTCGAGAACTACCGGAACCTCGTGTTCCTGGGAATCGCCGTGTCTAAGCCCCACCTGATCACCTGTCTCGAGCAAGGGAAAGAACCCTGGAAC (SEQ ID NO: 206).

[0237] (4) DNMT3A-DNMT3L-dCas9-KRAB (ZNF554)

[0238] ​ ​ ,

[0239] wherein the nucleic acid sequence of KRAB (ZNF554) is: CGGACACTGGTGACATTTGAGGATGTCTCCATGGACTTCAGCCAAGAGGAATGGGAGCTGCTGGAACCCGCTCAGAAGAACCTGTATAGGGAAGTCATGCTCGAAAATTACCGGAATGTGGTGTCCCTGGAAGCCCTGAAGAACCAGTGTACCGACGTGGGCATCAAAGAGGGCCCACTGTCTCCAGCTCAGACC (SEQ ID NO: 208).

[0240] (5) DNMT3A-DNMT3L-dCas9-KRAB (ZNF264)

[0241] ​ ​ ,

[0242] wherein the nucleic acid sequence of KRAB (ZNF264) is: CGGACACTGGTGACTTTCGACGATGTGGCCGTGACCTTTACCAAAGAAGAGTGGGGCCAGCTGGATCTGGCCCAGAGAACACTGTACCAAGAAGTTATGCTGGAGAACTGCGGCCTGCTGGTGTCTCTGGGATGTCCTGTGCCTAAGGCCGAGCTGATCTGCCACCTGGAACACGGACAAGAGCCTTGGACC (SEQ ID NO: 210).

[0243] (6) DNMT3A-DNMT3L-dCas9-KRAB (ZNF582)

[0244] RTLELFRDVAIVFSQEEWQWLAPAQRDLYRDVMLETYSNLVSLGLAVSKPDVISFLEQGKEPWM (SEQ ID NO: 211) ,

[0245] wherein the nucleic acid sequence of KRAB (ZNF582) is: CGGACACTGGAGCTGTTCCGCGACGTGGCCATTGTGTTCTCTCAAGAGGAATGGCAGTGGCTGGCCCCTGCTCAGAGAGATCTCTACCGCGACGTTATGCTGGAGACATACTCCAATCTCGTGTCCCTCGGCCTGGCCGTGTCCAAACCTGATGTGATCAGCTTCCTCGAGCAGGGAAAAGAACCTTGGATG (SEQ ID NO: 212).

[0246] (7) DNMT3A-DNMT3L-dCas9-KRAB (ZNF324)

[0247] RTLMAFEDVAVYFSQEEWGLLDTAQRALYRRVMLDNFALVASLGLSTSRPRVVIQLERGEEPWV (SEQ ID NO: 213) ,

[0248] wherein the nucleic acid sequence of KRAB (ZNF 324) is: CGGACACTGATGGCCTTTGAGGACGTCGCCGTGTATTTTTCCCAAGAAGAATGGGGACTGCTCGACACTGCCCAGAGAGCCCTGTATCGCAGAGTCATGCTGGACAACTTCGCCCTGGTGGCCTCCCTGGGCCTGTCTACAAGCAGACCCAGAGTGGTCATTCAGCTGGAAAGAGGCGAGGAACCCTGGGTC (SEQ ID NO: 214).

[0249] (8) DNMT3A-DNMT3L-dCas9-KRAB (ZNF669)

[0250] RTLVAFEDVAVNFTQEEWALLDSSQKNLYREVMQETCRNLASVGSQWKDQNIEDHFEKPGKDIRN (SEQ ID NO: 215) ,

[0251] wherein the nucleic acid sequence of KRAB (ZNF669) is: CGGACACTGGTCGCCTTTGAAGATGTGGCTGTGAATTTCACCCAAGAAGAGTGGGCTCTGCTGGACAGCAGCCAGAAGAATCTCTACCGCGAAGTGATGCAAGAGACATGCCGGAACCTGGCCTCTGTGGGCTCTCAGTGGAAGGACCAGAACATCGAGGACCACTTCGAGAAGCCCGGCAAGGACATCAGAAAC (SEQ ID NO: 216).

[0252] (9) DNMT3A-DNMT3L-dCas9-KRAB (ZNF354A)

[0253] RTLLTFEDVAVLFTRDEWRKLAPSQRNLYRDVMLENYRNLVSLGLPFTKPKVISLLQQGEDPWE (SEQ ID NO: 217) ,

[0254] wherein the nucleic acid sequence of KRAB (ZNF354A) is: CGGACACTGCTGACATTCGAAGATGTCGCAGTGCTGTTCACCCGGGACGAGTGGAGGAAACTGGCTCCCAGCCAGCGCAATCTGTATAGAGATGTGATGCTGGAGAATTACAGAAATCTCGTCAGCCTGGGGCTGCCCTTCACCAAGCCTAAAGTGATCAGCCTGCTGCAACAAGGCGAGGATCCTTGGGAA (SEQ ID NO: 218).

[0255] (10) DNMT3A-DNMT3L-dCas9-KRAB (ZNF82)

[0256] RTLVMFSDVSIDFSPEEWEYLDLEQKDLYRDVMLENYSNLVSLGCFISKPDVISSLEQGKEPWK (SEQ ID NO: 219) ,

[0257] wherein the nucleic acid sequence of KRAB (ZNF82) is: CGGACACTGGTTATGTTCTCCGACGTGTCCATCGACTTTAGCCCCGAAGAGTGGGAGTATCTGGACCTGGAACAGAAGGATCTGTATCGCGACGTGATGCTCGAGAATTACTCTAACCTGGTGAGCCTGGGCTGCTTCATCAGCAAGCCCGATGTGATCTCTAGCCTGGAGCAAGGCAAAGAGCCATGGAAA (SEQ ID NO: 220).

[0258] (11) DNMT3A-DNMT3L-dCas9-KRAB (ZNF595)

[0259] RTLVTFRDVAIEFSPEEWKCLDPAQQNLYRDVMLENYRNLVSLGFVISNPDLVTCLEQIKEPCN (SEQ ID NO: 221) ,

[0260] wherein the nucleic acid sequence of KRAB (ZNF595) is: CGGACACTGGTCACCTTCCGGGACGTCGCAATCGAATTCAGCCCCGAGGAATGGAAATGTCTGGACCCCGCACAGCAAAACCTCTACAGAGATGTCATGCTCGAAAACTATAGGAACCTGGTCTCCCTGGGCTTCGTGATCAGCAACCCTGATCTCGTGACCTGCCTCGAACAGATCAAAGAGCCCTGCAAC (SEQ ID NO: 222).

[0261] (12) DNMT3A-DNMT3L-dCas9-KRAB (ZNF419)

[0262] RTLVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFM (SEQ ID NO: 223) ,

[0263] wherein the nucleic acid sequence of KRAB (ZNF 419) is: CGGACACTGGTTACCTTCGAAGATGTTGCCGTGTACTTCAGCCAAGAAGAGTGGCGGCTGCTGGATGACGCCCAGAGACTGCTGTATCGGAATGTTATGCTCGAAAACTTCACCCTGCTGGCTTCCCTGGGACTCGCCAGCTCTAAGACCCACGAGATTACCCAGCTGGAATCCTGGGAAGAACCCTTCATG (SEQ ID NO: 224).

[0264] (13) DNMT3A-DNMT3L-dCas9-KRAB (ZNF566)

[0265] RTLVMFSDVSVDFSQEEWECLNDDQRDLYRDVMLENYSNLVSMGHSISKPNVISYLEQGKEPWL (SEQ ID NO: 225) ,

[0266] wherein the nucleic acid sequence of KRAB (ZNF566) is: CGGACACTGGTCATGTTCAGCGACGTGTCCGTGGACTTTAGCCAAGAGGAATGGGAATGCCTGAACGACGACCAGCGGGACCTCTATAGGGATGTCATGCTCGAGAACTACAGCAATCTCGTTTCCATGGGCCACAGCATCTCCAAGCCAAACGTCATCAGCTATCTCGAACAAGGCAAAGAGCCCTGGCTG (SEQ ID NO: 226).

[0267] (14) DNMT3A-DNMT3L-dCas9-KRAB(ZIM2)

[0268] RTLVTFEDVLVDFSPEELSSLSAAQRNLYREVMLENYRNLVSLGHQFSKPDIISRLEEEESYAM (SEQ ID NO: 227) ,

[0269] wherein the nucleic acid sequence of KRAB(ZIM2) is: CGGACACTGGTCACCTTCGAGGATGTGCTGGTGGATTTCAGCCCTGAGGAACTGAGCAGCCTGTCCGCCGCACAGAGAAATCTCTATCGGGAAGTGATGCTGGAGAACTATCGGAATCTGGTGTCTCTGGGCCACCAGTTCAGCAAGCCTGACATCATCAGCAGACTGGAAGAAGAGGAATCCTACGCCATG (SEQ ID NO: 228).

[0270] The construct can be DNA or RNA. In some embodiments, the construct is mRNA. In some embodiments, the construct is double stranded DNA. In some embodiments, the construct is double stranded RNA. In some embodiments, the construct is single stranded DNA. In some embodiments, the construct is single stranded RNA.

[0271] CRISPR-Cas system

[0272] The present disclosure provides CRISPR-Cas9 based engineered systems for genome editing and treatment of genetic diseases. The CRISPR-Cas9 based engineered systems can be designed to target any gene, including genes involved in angiogenesis, such as VEGFA. The present disclosure provides CRISPR-Cas systems comprising genetically engineered Cas proteins and / or guide RNAs with desired specificity and activity (e.g., reducing or eliminating expression of VEGFA gene products). The CRISPR-Cas9 based systems can include a Cas9 protein, a mutated Cas9 protein, or a Cas9 fusion protein (e.g., DNMT3A-DNMT3L (3A3L)-dCas9-KRAB fusion molecule) and at least one sgRNA (e.g., VEGFA sgRNA). For example, the Cas9 fusion protein can include a domain that differs from the endogenous activity of Cas9 (e.g., DNMT3A, DNMT3L, or KRAB).

[0273] The Cas9 protein can be cleaved into an N-terminal portion (CasN) and a C-terminal portion (CasC).

[0274] Generally, a Cas protein (interchangeably used herein with CRISPR protein, CRISPR enzyme, CRISPR-Cas protein, CRISPR-Cas enzyme, Cas, CRISPR effector, or Cas effector protein) and / or a guide sequence are components of a CRISPR-Cas system. A CRISPR-Cas system or CRISPR system collectively refers to transcripts and other elements involved in the expression or directing the activity of CRISPR-associated (“Cas”) genes, including sequences encoding Cas genes, a tracr (trans-activating CRISPR) sequence (e.g., a tracrRNA or an active partial tracrRNA), a tracr-mate sequence (including a “direct repeat sequence” and a partial direct repeat sequence processed from a tracrRNA in the context of an endogenous CRISPR system), a guide sequence (also referred to as a “spacer” in the context of an endogenous CRISPR system), or “one or more RNAs” as the term is used herein (e.g., one or more RNAs that direct a Cas, such as a CRISPR RNA and a trans-activating (tracr) RNA or a single guide RNA (a.k.a. sgRNA; a chimeric RNA)), or other sequences and transcripts from a CRISPR locus.

[0275] Generally, a CRISPR system is characterized by elements that facilitate the formation of a CRISPR complex at the site of a target sequence (also referred to as a protospacer in the context of an endogenous CRISPR system). In the engineered systems of the present disclosure, a direct repeat sequence can include a naturally occurring sequence or a non-naturally occurring sequence. The direct repeat sequences of the present disclosure are not limited to naturally occurring lengths and sequences. In addition, the direct repeat sequences of the present disclosure can include insertions of nucleotides such as aptamers or sequences that bind to adaptor proteins (for association with functional domains). In certain embodiments, a direct repeat sequence comprising an insertion such as is approximately the first half of a short direct repeat (DR) on one end and approximately the second half of a short direct repeat (DR) on the other end

[0276] In the context of forming a CRISPR complex, a “target sequence” or “target polynucleotide” refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between the target sequence and the guide sequence facilitates the formation of a CRISPR complex. A target sequence can comprise any polynucleotide, such as a DNA or RNA polynucleotide. In some embodiments, a target sequence is located in a nucleus or cytoplasm of a cell.

[0277] Generally, a guide sequence (or spacer sequence) can be any polynucleotide sequence that has sufficient complementarity to a target polynucleotide sequence to hybridize to the target sequence and direct sequence-specific binding of a CRISPR complex to the target sequence. In some embodiments, the degree of complementarity between a guide sequence and its corresponding target sequence, when optimally aligned using a suitable alignment algorithm, is about or more than about 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more.

[0278] In certain embodiments, cleavage efficiency can be modulated by introducing mismatches (e.g., 1 or more mismatches, such as 1 or 2 mismatches between the spacer sequence and the target sequence, including the position of the mismatches along the spacer / target). For example, the closer the doublet of mismatches is to the center (i.e., not 3’ or 5’), the greater the impact on cleavage efficiency. Thus, by choosing the position of the mismatch along the spacer, cleavage efficiency can be modulated. For example, if less than 100% target cleavage is desired (e.g., in a population of cells), 1 or more, such as preferably 2, mismatches between the spacer and target sequence can be introduced in the spacer sequence. The closer the position of the mismatch to the center of the spacer, the lower the percentage of cleavage.

[0279] The CRISPR-Cas system or components thereof can be used to introduce one or more mutations in a target locus or nucleic acid sequence. The one or more mutations can include the introduction, deletion, or substitution of one or more nucleotides produced by the one or more guide RNAs or one or more sgRNAs at each target sequence of the one or more cells. The mutations can include the introduction, deletion, or substitution of 1-75 nucleotides produced by the one or more guide RNAs at each target sequence of the one or more cells.

[0280] Generally, in the case of an endogenous CRISPR-Cas system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage in or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from) the target sequence, although this can depend on, for example, secondary structure, particularly in the case of an RNA target. In some cases, in the case of an endogenous CRISPR system, formation of a CRISPR complex (comprising a guide sequence hybridized to a target sequence and complexed with one or more Cas proteins) results in cleavage of one or both strands (if applicable) in or near (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from) the target sequence.

[0281] In some embodiments, a guide RNA (capable of guiding a Cas to a target locus) can comprise (1) a guide sequence capable of hybridizing to a target locus (polynucleotide target locus, such as an RNA target locus) in a eukaryotic cell; (2) a direct repeat (DR) sequence present in a single RNA (i.e., sgRNA (arranged in 5' to 3' direction) or crRNA).

[0282] For general information concerning CRISPR-Cas systems, components thereof, and delivery of such components, including methods, materials, delivery vehicles, vectors, particles, AAVs, and their preparation and use, including amounts and formulations, all of which are useful in the practice of the present disclosure, reference is made to: U.S. Patent Nos. 8,999,641; 8,993,233; 8,945,839; 8,932,814; 8,906,616; 8,895,308; 8,889,418; 8,889,356; 8,871,445; 8,865,406; 8,795,965; 8,771,945; and 8,697,359; U.S. Patent Publication Nos. US 2014-0310830; US 2014-0287938 Al; US 2014-0273234 Al; US 2014-0273232 Al; US 2014-0273231 Al; US 2014-0256046 Al; US 2014-0248702 Al; US 2014-0242700 Al; US 2014-0242699 Al; US 2014-0242664 Al; US 2014-0234972 Al; US 2014-0227787 Al; US 2014-0189896 Al; US 2014-0186958; US 2014-0186919 Al; US 2014-0186843 Al; US 2014-0179770 Al; and US 2014-0179006 Al; US 2014-0170753; European Patent Nos. EP 2784162 Bl and EP 2771468 Bl; European Patent Application Nos. EP 2771468; EP 2764103; and EP 2784162;and PCT Patent Publications WO 2021 / 183807 Al (PCT / US2021 / 021973), WO 2014 / 093661 (PCT / US2013 / 074743), WO 2014 / 093694 (PCT / US2013 / 074790), WO 2014 / 093595 (PCT / US2013 / 074611), WO 2014 / 093718 (PCT / US2013 / 074825), WO 2014 / 093709 (PCT / US2013 / 074812), WO 2014 / 093622 (PCT / US2013 / 074667), WO 2014 / 093635 (PCT / US2013 / 074691), WO 2014 / 093655 (PCT / US2013 / 074736), WO 2014 / 093712 (PCT / US2013 / 074819), WO 2014 / 093701 (PCT / US2013 / 074800), WO 2014 / 018423 (PCT / US2013 / 051418), WO 2014 / 204723 (PCT / US2014 / 041790), WO 2014 / 204724 (PCT / US2014 / 041800), WO 2014 / 204725 (PCT / US2014 / 041803), WO 2014 / 204726 (PCT / US2014 / 041804), WO 2014 / 204727 (PCT / US2014 / 041806), WO 2014 / 204728 (PCT / US2014 / 041808), WO 2014 / 204729 (PCT / US2014 / 041809), each of which is incorporated by reference herein in its entirety.

[0283] Cas protein

[0284] A Cas protein (e.g., an engineered Cas protein) can have substantially the same (e.g., 80% to 100%, 90% to 100%, 95% to 100%, 98% to 100%, 99% to 100%, 99.9% to 100%, or about 100%) nuclease activity as a wild-type counterpart Cas protein. In certain cases, an engineered Cas protein has higher (e.g., at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90%) nuclease activity than a wild-type counterpart Cas protein.

[0285] Alternatively or additionally, a Cas protein (e.g., an engineered Cas protein) can have a specificity that is at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, or at least 90% higher than that of a wild-type counterpart Cas protein. In particular examples, a Cas protein (e.g., an engineered Cas protein) has a specificity that is at least 30% higher than that of a wild-type counterpart Cas protein. As used herein, the term “specificity” of a Cas can correspond to the number or percentage of on-target polynucleotide cleavage events relative to the number or percentage of all polynucleotide cleavage events (including on-target and off-target events). The activity and specificity of a Cas protein are consistent with those described in Hsu PD, et al. DNA targeting specificity of RNA-guided Cas9 nucleases. Nat Biotechnol. 2013 Sep; 31(9):827-832; and Slaymaker IM, et al. Rationally engineered Cas9 nucleases with improved specificity. Science. 2016 Jan 1; 351(6268): 84-88, which also describe examples of methods for detecting the activity and specificity of a Cas protein, and are incorporated by reference in their entirety herein, and are detailed elsewhere herein.

[0286] In some embodiments, a Cas protein (e.g., its RuvC domain) can slide one base upstream (relative to the PAM) and create a staggered cut, which can be filled in and result in replication of a single base (i.e., a +1 insertion). Examples of +1 insertion positions are described in Zuo, Z. and Liu, J. (2016) Evidence for Cas9-catalyzed DNA cleavage creates a staggered end from Molecular Dynamics Simulations. Scientific Reports 6, 37584. In some embodiments, an engineered Cas protein has a different +1 insertion frequency than a wild-type counterpart Cas protein. For example, the +1 insertion frequency when a guanine is present at the -2 position relative to the PAM is higher than the +1 insertion frequency when a thymidine, cytidine, or adenine is present at the -2 position relative to the PAM. In some cases, +1 insertion is dependent on host mechanisms in human cells. In some examples, a Cas protein can produce a staggered cut. The staggered cut can be a 1-bp or 1-nucleotide 5’ overhang. The staggered cut can be a 1-bp or 1-nucleotide 3’ overhang.

[0287] The nucleic acid molecule encoding the Cas can be codon-optimized. In this case, examples of codon-optimized sequences are sequences optimized for expression in a eukaryote, e.g., human (i.e., optimized for expression in a human), or another eukaryote, animal, or mammal discussed herein; see, e.g., the SaCas9 human codon-optimized sequence in WO 2014 / 093622 (PCT / US2013 / 074667). While this is preferred, it will be appreciated that other examples are possible, and codon optimization for host species other than human, or for specific organs, are known. In some embodiments, the enzyme-encoding sequence encoding the Cas is codon-optimized for expression in a particular cell, such as a eukaryotic cell. The eukaryotic cell can be a cell of or derived from a particular organism, such as a mammal, including but not limited to a human or a non-human eukaryote or animal or mammal described herein, e.g., a mouse, rat, rabbit, dog, livestock, or non-human mammal or primate. In some embodiments, methods for altering the human germ line genetic identity and / or methods for altering the genetic identity of an animal that can cause the human or animal to suffer without any substantial medical benefit to the human or animal and also to the animals resulting from such methods can be excluded. Generally, codon optimization refers to the process of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with a codon that is more frequently or most frequently used in the genes of that host cell, while maintaining the native amino acid sequence. Different species exhibit a particular bias for particular codons for a particular amino acid. Codon bias (differences in codon usage between organisms) is generally correlated with the efficiency of translation of messenger RNA (mRNA), which is in turn believed to be a function of the nature of the codon being translated and the availability of the particular transfer RNA (tRNA) molecule, among other things. The predominance of a selected tRNA in a cell is generally a reflection of the most frequently used codons in peptide synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables are readily available, e.g., in the “Codon Usage Database” available at www.kazusa.orjp / codon / , and these tables can be modified in a variety of ways.See Nakamura, Y. et al. "Codon usage tabulated from the international DNA sequence databases: status for the year 2000" Nucl. Acids Res. 28:292 (2000). Computer algorithms for codon optimization of particular sequences for expression in particular host cells are also available, such as Gene Forge (Aptagen; Jacobus, PA). In some embodiments, one or more codons (e.g., 1, 2, 3, 4, 5, 10, 15, 20, 25, 50 or more, or all codons) in a sequence encoding a Cas have the most commonly used codon for the particular amino acid.

[0288] In some embodiments, a Cas protein can have nucleic acid cleavage activity. A Cas protein can have RNA binding and DNA cleavage functions. In some embodiments, a Cas can direct cleavage of one or two nucleic acid strands at or near the location of a target sequence, such as within the target sequence and / or within the complement of the target sequence or on a sequence related to the target sequence (e.g., within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence). In some embodiments, a Cas protein can direct more than one cleavage (e.g., 1, 2, 3, 4, 5 or more cleavages) of one or two strands within the target sequence and / or within the complement of the target sequence or on a sequence related to the target sequence and / or within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500 or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the cleavage can be blunt, i.e., produce a blunt end. In some embodiments, the cleavage can be staggered, i.e., produce a sticky end.

[0289] In some embodiments, the vector encodes a nucleic acid-targeting Cas protein that can be mutated relative to the corresponding wild-type enzyme such that the mutated nucleic acid-targeting Cas protein lacks the ability to cleave one or both strands of a target polynucleotide containing a target sequence, e.g., an alteration or mutation in the HNH domain results in a mutated Cas that is substantially devoid of all DNA cleavage activity, e.g., the DNA cleavage activity of the mutated enzyme is no more than about 25%, 10%, 5%, 1%, 0.1%, 0.01%, or less of the nucleic acid cleavage activity of the non-mutated form of the enzyme; one example can be when the nucleic acid cleavage activity of the mutated form is zero or negligible when compared to the non-mutated form. As used herein, the term "derived" with respect to an enzyme means that the derived enzyme is largely based on the wild-type enzyme in the sense of having a high degree of sequence homology to the wild-type enzyme, but which has been mutated (modified) in some way known in the art or described herein.

[0290] Generally, in the case of an endogenous nucleic acid-targeting system, formation of a nucleic acid-targeting complex (comprising a guide RNA or crRNA that hybridizes to a target sequence and is complexed with one or more nucleic acid-targeting effector proteins) results in cleavage of one or more DNA strands in or near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs). As used herein, the term "one or more sequences associated with the target locus of interest" refers to sequences near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from the target sequence, where the target sequence is contained within the target locus of interest).

[0291] It will be appreciated that the effector protein is based on or derived from an enzyme, and thus in some embodiments, the term "effector protein" certainly includes "enzyme." However, it will also be appreciated that the effector protein can have DNA or RNA binding activity, but not necessarily cleavage or nicking activity, including dead Cas protein functionality, as required in some embodiments.

[0292] In some embodiments, Cas proteins can form components of an inducible system. The inducible nature of this system will allow for temporal and spatial control of gene editing or gene expression using a form of energy. Forms of energy can include, but are not limited to, electromagnetic radiation, sonic energy, chemical energy, and thermal energy. Examples of inducible systems include tetracycline inducible promoters (Tet-On or Tet-Off), small molecule two-hybrid transcriptional activation systems (FKBP, ABA, etc.), or light inducible systems (phytochromes, LOV domains, or cryptochromes). In one embodiment, a CRISPR effector protein can be part of a light inducible transcriptional effector (LITE) to direct changes in transcriptional activity in a sequence specific manner. Components of light can include a CRISPR effector protein, a light responsive cytochrome heterodimer (e.g., from Arabidopsis thaliana), and a transcriptional activation / repression domain. Other examples of inducible DNA binding proteins and methods of use thereof are provided in US 61 / 736465 and US 61 / 721,283 and WO 2014018423 A2, which are hereby incorporated by reference in their entirety.

[0293] In some embodiments, a mutated Cas can have one or more mutations that result in reduced off-target effects, e.g., improved CRISPR enzymes for effecting modification of a target locus but reducing or eliminating activity toward off-targets (such as when complexed with a guide RNA), and improved CRISPR enzymes for enhancing CRISPR enzyme activity (such as when complexed with a guide RNA). It will be appreciated that mutated enzymes as described below can be used in any of the methods of the present disclosure described elsewhere herein. Any of the methods, products, compositions, and uses described elsewhere herein are equally applicable to mutated CRISPR enzymes, as described in further detail below.

[0294] Methods and mutations that can be used in various combinations to enhance or reduce on-target activity and / or specificity, or to increase or decrease on-target binding and / or specificity relative to off-target binding, can be used to compensate for or enhance mutations or modifications made to facilitate other effects. Such mutations or modifications made to facilitate other effects include mutations or modifications to Cas and / or mutations or modifications made to guide RNAs. The methods and mutations of the present disclosure are used to modulate Cas nuclease activity and / or binding to chemically modified guide RNAs.

[0295] In certain embodiments, the catalytic activity of the Cas proteins of the present disclosure is altered or improved. It will be appreciated that a mutated Cas has altered or improved catalytic activity if the catalytic activity is different from that of the corresponding wild-type Cas protein (e.g., a non-mutated Cas protein). Catalytic activity can be determined by methods known in the art. For example, but not by way of limitation, catalytic activity can be determined in vitro or in vivo by determining the percent indel (e.g., after a given time, or at a given dose). In certain embodiments, the catalytic activity is enhanced. In certain embodiments, the catalytic activity is enhanced by at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 100%. In certain embodiments, the catalytic activity is decreased. In certain embodiments, the catalytic activity is decreased by at least 5%, preferably at least 10%, more preferably at least 20%, such as at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or (substantially) 100%. One or more mutations herein can inactivate the catalytic activity, which can significantly reduce all catalytic activity, reduce activity to below detectable levels, or reduce to no measurable catalytic activity.

[0296] One or more features of an engineered Cas protein can differ from the corresponding wild-type Cas protein. Examples of such features include catalytic activity, gRNA binding, specificity of the Cas protein (e.g., specificity to edit a defined target), stability of the Cas protein, off-target binding, target binding, protease activity, nickase activity, PFS recognition. In some examples, an engineered Cas protein can comprise one or more mutations of the corresponding wild-type Cas protein. In some embodiments, the catalytic activity of an engineered Cas protein is enhanced compared to the corresponding wild-type Cas protein. In some embodiments, the catalytic activity of an engineered Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the gRNA binding of an engineered Cas protein is enhanced compared to the corresponding wild-type Cas protein. In some embodiments, the gRNA binding of an engineered Cas protein is attenuated compared to the corresponding wild-type Cas protein. In some embodiments, the specificity of the Cas protein is enhanced compared to the corresponding wild-type Cas protein. In some embodiments, the specificity of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the stability of the Cas protein is enhanced compared to the corresponding wild-type Cas protein. In some embodiments, the stability of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the engineered Cas protein further comprises one or more mutations that inactivate the catalytic activity. In some embodiments, the off-target binding of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the off-target binding of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the target binding of the Cas protein is increased compared to the corresponding wild-type Cas protein. In some embodiments, the target binding of the Cas protein is decreased compared to the corresponding wild-type Cas protein. In some embodiments, the engineered Cas protein has higher protease activity or polynucleotide binding capacity compared to the corresponding wild-type Cas protein. In some embodiments, the PFS recognition is altered compared to the corresponding wild-type Cas protein.

[0297] Examples of Cas proteins

[0298] Examples of Cas proteins include Class 1 (e.g., Type I, Type III, and Type IV) and Class 2 (e.g., Type II, Type V, and Type VI) Cas proteins, e.g., Cas9, Cas12 (e.g., Cas12a, Cas12b, Cas12c, Cas12d), Cas13 (e.g., Cas13a, Cas13b, Cas13c, Cas13d), CasX, CasY, Cas14, variants (e.g., mutant forms, truncated forms) thereof, homologs thereof, and orthologs thereof. The terms “ortholog” and “homolog” are well known in the art. By way of further guidance, a “homolog” of a protein as used herein is a protein of the same species that performs the same or similar function as the protein of which it is a homolog. The homologous protein can be, but is not necessarily, structurally related, or only partially structurally related. As used herein, an “ortholog” of a protein is a protein of a different species that performs the same or similar function as the protein of which it is an ortholog. The orthologous protein can be, but is not necessarily, structurally related, or only partially structurally related.

[0299] Cas proteins of Class 2

[0300] In some embodiments, the Cas protein is a Class 2 Cas protein, i.e., a Cas protein of a Class 2 CRISPR-Cas system. The Class 2 CRISPR-Cas system can be a subtype, e.g., Type II-A, Type II-B, Type II-C, Type V-A, Type V-B, Type V-C, or Type V-U. In some embodiments, the Cas protein is Cas9, Cas12a, Cas12b, Cas12c, or Cas12d. In some embodiments, the Cas9 can be SpCas9, SaCas9, StCas9, and other Cas9 orthologs. The Cas12 can be Cas12a, Cas12b, and Cas12c, including FnCas12a or homologs or orthologs thereof. Definitions and exemplary members of CRISPR-Cas systems include those described in Kira S. Makarova and Eugene V. Koonin, Annotation and Classification of CRISPR-Cas systems, Methods Mol Biol. 2015; 1311: 47-75; and Sergey Shmakov et al., Diversity and evolution of class 2 CRISPR-Cas systems, Nat Rev Microbial. 2017 Mar; 15(3): 169-182.

[0301] Cas protein linkers

[0302] In some examples, the Cas protein comprises at least one RuvC domain and at least one HNH domain. The Cas protein can also comprise a first and second linker domain connecting the RuvC domain and the HNH domain. The first linker (LI) and the second linker (L2) connecting the HNH and RuvC domains in Cas9 are described in Nishimasu, H, et al. “Crystal structure of Cas9 in complex with guide RNA and target RNA” Cell 156 (Feb. 27, 2014): 935-949 and Ribeiro, L. et al. (2018) “Protein engineering strategies to expand CRISPR-Cas9 applications” International Journal of Genomics Volume 2018, Article ID 1652567 (doi.org / 10.1155 / 2018 / 1652567). Figure 1 of Ribeiro (specifically incorporated by reference herein) shows the overall organization, structure, and function of Cas9. In particular, Figure 1A A schematic showing the domain organization of SpCas9, indicating the genetic architecture of the HNH and RuvC domains including linkers LI (spanning amino acids 765-780) and L2 (spanning amino acids 906-918) as described herein.

[0303] Similarly, the domain organization of Staphylococcus aureus Cas9 (SaCas9) can be utilized when referring to the first and second linker domains. In one aspect, the linker 1 domain region spans residues 481-519 and will connect the RuvC-II domain with the HNH domain in SaCas9. In some embodiments, the linker 2 region spans residues 629-649 and connects the RuvC-III domain and the HNH domain of SaCas9. Thus, the first and / or second linker domains can be mutated in Cas9 orthologs and reference can be made to the amino acid residues corresponding to the amino acids of wild-type SaCas9. See, Nishimasu, Cell. 2015 Aug 27; 162(5): 1113-1126; doi: 10.1016 / j.cell.2015.08.007, incorporated by reference herein. In particular, Figure 1, S1-S3 of Nishimasu detail the domain organization of Cas9 proteins and are specifically incorporated by reference for their teachings.

[0304] The first and second linkers can comprise about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, or more amino acids. The first and second linkers can correspond to wild-type linkers. In some aspects, the first and second linkers can comprise one or more mutations in the first and / or second linkers. In one aspect, the first and / or second linkers comprise one or more mutations that increase specificity of the Cas9 protein.

[0305] In some embodiments, the linkers LI and L2 connecting the HNH and RuvC domains of Cas9 contain wild-type amino acid sequences. In some embodiments, the linkers connecting the HNH and RuvC domains contain mutations in one or more amino acids. In exemplary embodiments, the first linker (LI) comprises a mutation corresponding to amino acid T769I of SpCas9, and / or the second linker (L2) comprises a mutation corresponding to amino acid G915M of SpCas9. In exemplary embodiments, one or more linker mutations, such as T769I and G915M, confer increased specificity of the Cas9 protein.

[0306] In one embodiment, one or more mutations in the first and second linkers can be combined with one or more mutations in other portions of the Cas9 protein to further increase specificity and / or to maintain substantially equivalent activity to the wild-type Cas9 protein, as described herein. In one embodiment, mutations in the linkers and / or additional mutations in the Cas protein can be identified using the methods detailed herein that enhance / increase specificity and substantially retain wild-type activity of the wild-type Cas9.

[0307] Cas proteins of Class 2, Type II (e.g., Cas9)

[0308] In some embodiments, the Cas protein can be a Cas protein of a class 2 type II CRISPR-Cas system (type II Cas protein). In some embodiments, the Cas protein can be a class 2 type II Cas protein, e.g., Cas9. In some embodiments, a CRISPR / Cas9-based system can include a Cas9 protein or fragment thereof, a Cas9 fusion protein, a nucleic acid encoding a Cas9 protein or fragment thereof, or a nucleic acid encoding a Cas9 fusion protein. “Cas9 (CRISPR-associated protein 9)” refers to a polypeptide or fragment thereof having at least about 85% amino acid identity to NCBI Accession No. NP_269215 and having RNA binding activity, DNA binding activity, and / or DNA cleavage activity (e.g., endonuclease or nickase activity). “Cas9 function” can be determined by any of a variety of assays, including but not limited to a fluorescence polarization-based nucleic acid binding assay, a fluorescence polarization-based strand invasion assay, a transcription assay, an EGFP disruption assay, a DNA cleavage assay, and / or a Surveyor assay, e.g., as described herein. “Cas9 nucleic acid molecule” refers to a polynucleotide encoding a Cas 9 polypeptide or fragment thereof. An exemplary Cas9 nucleic acid molecule sequence is provided in Genomic Sequence No. NC_002737. In some embodiments, an inhibitor of Cas9 (e.g., naturally occurring Cas9 in Streptococcus pyogenes (SpCas9) or Staphylococcus aureus (SaCas9) or a variant thereof is disclosed herein. Cas9 recognizes foreign DNA using a protospacer adjacent motif (PAM) sequence and a guide RNA (gRNA) for base pairing to the target DNA. The relative ease with which Cas9 induces targeted strand breaks at any genomic locus enables efficient genome editing in a variety of cell types and organisms. Cas9 derivatives can also be used as transcriptional activators / repressors.

[0309] In some cases, the CRISPR-Cas protein is Cas9 or a variant thereof. In some examples, the Cas9 can be a wild-type Cas9, including any naturally occurring bacterial Cas9. Cas9 orthologs generally share the general organization of 3-4 RuvC domains and one HNH domain. The 5' most RuvC domain cleaves the non-complementary strand, and the HNH domain cleaves the complementary strand. All designations are with reference to the guide sequence. By homology comparison of Cas9 of interest with other Cas9 orthologs (from S. pyogenes type II CRISPR locus, S. thermophilus CRISPR locus 1, S. thermophilus CRISPR locus 3, and F. novicida type II CRISPR locus), catalytic residues in the 5' RuvC domain were identified, and the conserved Asp residue (D10) was mutated to alanine to convert Cas9 to a complementary strand nickase. Thus, the Cas enzyme can be a wild-type Cas9, including any naturally occurring bacterial Cas9. The CRISPR, Cas, or Cas9 enzyme can be a codon-optimized or modified version, including any chimeric, mutant, homolog, or ortholog. In another aspect of the disclosure, the Cas9 enzyme can comprise one or more mutations and can be used as a universal DNA binding protein, fused or not fused to functional domains.

[0310] The mutations can be artificially introduced mutations or gain-of-function mutations or loss-of-function mutations. In some embodiments, the transcriptional activation domain can be VP64. In some embodiments, the transcriptional repressor domain can be KRAB or SID4X. Other aspects of the disclosure relate to mutated Cas9 enzymes fused to domains including, but not limited to, nucleases, transcriptional activators, repressors, recombinases, transposases, histone remodelers, demethylases, DNA methyltransferases, cryptochromes, light-inducible / controllable domains, or chemical-inducible / controllable domains. The disclosure can relate to sgRNAs or tracrRNAs or guide sequences or chimeric guide sequences that allow for enhanced performance of these RNAs in cells. Such Type II CRISPR enzymes can be any Cas enzyme. In some cases, the Cas9 enzyme is from or derived from SpCas9 or SaCas9. As used herein, the term "derived" with respect to an enzyme means that the derived enzyme is largely based on the wild-type enzyme in the sense of having a high degree of sequence homology to the wild-type enzyme, but that it has been mutated (modified) in some way known in the art or described herein. In examples, the mutations can include one or more mutations in the first linker domain, the second linker domain, and / or other portions of the protein. High degree of sequence homology can include at least 80%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more relative to the wild-type enzyme.

[0311] The Cas enzyme can be identified as Cas9 in that this can refer to the general class of enzymes that share homology with the largest nuclease having multiple nuclease domains from Type II CRISPR systems. In some cases, the Cas9 enzyme is from or derived from SpCas9 (Streptococcus pyogenes Cas9) or saCas9 (Staphylococcus aureus Cas9). "StCas9" refers to the wild-type Cas9 from Streptococcus thermophilus (UniProt ID: G3ECR1). Similarly, "SpCas9" refers to the wild-type Cas9 from Streptococcus pyogenes (UniProt ID: Q99ZW2). As used herein, the term "derived" with respect to an enzyme means that the derived enzyme is largely based on the wild-type enzyme in the sense of having a high degree of sequence homology to the wild-type enzyme, but that it has been mutated (modified) in some way known in the art or described herein. It should be understood that the terms Cas and CRISPR enzyme are generally used interchangeably herein unless otherwise specified. As noted above, many of the residue numbers used herein refer to the Cas9 enzyme from the Type II CRISPR locus in Streptococcus pyogenes.

[0312] In particular embodiments, the effector protein is a Cas9 effector protein from or derived from an organism of the genus Streptococcus Streptococcus ), Campylobacter Campylobacter , Nitratifractor , Staphylococcus Staphylococcus , Parvibaculum , Ralstonia Roseburia , Neisseria Neisseria , Gluconacetobacter Gluconacetobacter , Azospirillum Azospirillum , Caviomonas Sphaerochaeta , Lactobacillus Lactobacillus , Eubacterium Eubacterium , Corynebacterium Corynebacte , Carnobacterium Carnobacterium , Rhodobacter Rhodobacter , Listeria Listeria , Porphyridium Paludibacter , Clostridium Clostridium , Lachnospiraceae Lachnospiraceae , Clostridium Clostridiaridium , Leptotrichia Leptotrichia , Francisella Francisella , Legionella Legionella , Lipsonia Alicyclobacillus , Methylophaga Methanomethyophilus , Porphyromonas Porphyromonas , Prevotella Prevotella , Bacteroides Bacteroidetes , Petrimonas Helcococcus , Leptospira Letospira , Desulfovibrio Desulfovibrio , Desulfococcus Desulfonatronum , Hyphomicrobiaceae Opitutaceae , M. bulgaricus Tuberibacillus , Bacillus, Brevibacillus Brevibacilus , Methylobacterium Methylobacterium , or Acidaminococcus Acidaminococcus , Streptococcus, Campylobacter, Nitratifractor Staphylococcus, Parvibaculum Ralstonia, Neisseria, Gluconacetobacter, Azospirillum, Caviomonas, Lactobacillus Lactobacillus , Eubacterium, Corynebacterium, Sutterella Sutterella , Legionella, Treponema Treponema , Filifactor Filifactor), Eubacteria, Streptococcus Streptococcus Lactobacillus () Lactobacillus ), Mycoplasma genus ( Mycoplasma Bacteroides, Freundii Flaviivola Flavobacterium ( Flavobacterium ), *Trichophyton* spp., *Azotobacter* spp., *Gluconobacter* spp., *Neisseria* spp., *Rhodotorula* spp., *Corynebacterium* spp. ( Parvibaculum ), Staphylococcus spp., nitrate-lysing bacteria spp. Nitratifractor ), Mycoplasma or Campylobacter.

[0313] In some implementations, the Cas9 protein is derived from or selected from Streptococcus mutans (Streptococcus mutans). S. mutans ), agalactococcus ( S. agalactiae Streptococcus equi ( S. equisimilis ), Streptococcus sanguinis ( S. sanguinis Streptococcus pneumoniae () S. pneumonia Campylobacter jejuni ( C. jejuni ), Escherichia coli ( C. coli ); brine nitrate lysis bacteria ( N salsuginis ), ( N. tergarcus Staphylococcus aureus ( S. auricularis Staphylococcus aureus ( S. carnosus ); Neisseria meningitidis ( N meningitides ), Neisseria gonorrhoeae ( N gonorrhoeae Listeria monocytogenes ( ) L. monocytogenes Listeria monocytogenes ( ) L. ivanovii Clostridium botulinum ( ); Clostridium botulinum ( C. botulinum Clostridium difficile ( C. difficile Clostridium tetani C. tetani ) or Clostridium solvae ( C. sordellii ), Tulafrancsis ( Francisella tularensis1 ), a new culprit subspecies of Francisella tularensis ( Francisella tularensis subsp. novicida ), Alberprevotella ( Prevotella albensis ), Styloides MC2017 1 ( Lachnospiraceae bacterium MC2017 1 ), Vibrio butyricum ( Butyrivibrio proteoclasticus ), Heterophytes bacteria GW2011 GWA2_33_10 ( Peregrinibacteria bacterium GW2011 GWA2_33_10 ), Bacteroidetes GW2011 GWC2_44_17 ( Parcubacteria bacterium GW2011 GWC2_44_17 ), a certain SCADC of the genus Smithella ( Smithella sp. SCADC ), a certain species of BV3L6 of the genus *Coccus* (amino acid) Acidaminococcus sp. BV3L6), Lachnospiraceae bacterium MA2020 Lachnospiraceae bacterium MA2020 ), Methanomassiliicoccus alvum sp. nov. Candidatus Methanoplasma termitum ), Eubacterium eligens Eubacterium eligens ), Moraxella 237 Moraxella bovoculi 237 ), Leptospira borgpetersenii Leptospira inadai ), Lachnospiraceae bacterium ND2006 Lachnospiraceae bacterium Nd2006 ), Porphyromonas cangingivalis 3 Porphyromonas crevioricanis 3 ), Prevotella disiens Prevotella disiens ), and Porphyromonas macacae Porphyromonas macacae In some embodiments, the Cas9 effector protein from an organism is from or derived from a Streptococcus pyogenes, Staphylococcus aureus, or Streptococcus thermophilus Cas9.

[0314] In more preferred embodiments, the Cas9 protein is derived from a bacterial species selected from the group consisting of Streptococcus pyogenes, Staphylococcus aureus, or Streptococcus thermophilus Cas9. In certain embodiments, the Cas9 is derived from a bacterial species selected from the group consisting of Francisella tularensis 1, Prevotella albensis, Lachnospiraceae bacterium MC20171, Psychrobacillus lookhanii, GW Taxon 275 bacterium GW2011 GWA2_33_10, GW Superkingdom Bacteria bacterium GW2011 GWC2_44_17, a species of the genus Smithella SCADC, a species of the genus Acidaminococcus BV3L6, Lachnospiraceae bacterium MA2020, Methanomassiliicoccus alvum sp. nov., Eubacterium eligens, Moraxella 237, Leptospira borgpetersenii, Lachnospiraceae bacterium ND2006, Porphyromonas cangingivalis 3, Prevotella disiens, and Porphyromonas macacae. In certain embodiments, the Cas9 protein is derived from a bacterial species selected from the group consisting of a species of the genus Acidaminococcus BV3L6, Lachnospiraceae bacterium MA2020. In certain embodiments, the effector protein is derived from a subspecies of Francisella tularensis 1, including but not limited to Francisella tularensis novicida subsp. novicida.

[0315] Cas9 enzymes include, but are not limited to, Streptococcus pyogenes serotype Ml (UniProt ID: Q99ZW2), Staphylococcus aureus Cas9 (UniProt ID: J7RUA5), Eubacterium dolichum Cas9 (UniProt ID: A5Z395), Azospirillum sp. (strain B510) Cas9 (UniProt ID: D3NT09), Acidovorax tempericus (strain ATCC 49037) Cas9 (UnitProt ID: A9HKP2), Neisseria cinerea Cas9 (UniProt ID: D0W2Z9), Roseburia intestinalis Cas9 (UniProt ID: C7G697), Allobaculum cetoedamentale (strain DS-1) Cas9 (UniProt ID: A7HP89), Halonitrosus nitratogenes (strain DSM 16511) Cas9 (UniProt ID: E6WZS9), Campylobacter laridis Cas9 (UniProt ID: G1UFN3).

[0316] The enzymatic action of Cas9 from Streptococcus pyogenes or any closely related Cas9 produces a double-stranded break at a target site sequence that hybridizes to a 20 nucleotides of a guide sequence and has a protospacer adjacent motif (PAM) sequence following the 20 nucleotides of the target sequence (examples include NGG / NRG or PAMs that can be determined as described herein). CRISPR activity by Cas9 for site-specific DNA recognition and cleavage is defined by a guide sequence, a tracr sequence that partially hybridizes to the guide sequence, and a PAM sequence. Further aspects of the CRISPR system are described in Karginov and Hannon, The CRISPR system: small RNA-guided defense in bacteria and archaea, Mole Cell 2010, January 15; 37(1): 7. The type II CRISPR locus from Streptococcus pyogenes SF370, which contains a cluster of four genes, Cas9, Casl, Cas2, and Csnl, as well as two non-coding RNA elements, a tracrRNA and a characteristic array of repeat sequences (direct repeats) interspaced by short segments of non-repetitive sequence (spacers, ~30 bp each). In this system, a targeted DNA double-strand break (DSB) is generated in four sequential steps. First, two non-coding RNAs, a pre-crRNA array and a tracrRNA, are transcribed from the CRISPR locus. Second, the tracrRNA hybridizes to the direct repeats of the pre-crRNA, which is then processed into mature crRNAs containing a single spacer sequence. Third, the mature crRNA:tracrRNA complex directs Cas9 to a DNA target composed of a protospacer sequence and a corresponding PAM through a heteroduplex between the spacer of the crRNA and the protospacer sequence DNA. Finally, Cas9 mediates cleavage of the target DNA upstream of the PAM to generate a DSB within the protospacer sequence. A pre-crRNA array consisting of a single spacer flanked by two direct repeats (DRs) is also included in the term "tracr-mate sequence." In certain embodiments, Cas9 can be constitutively present or inducibly present or conditionally present or administered or delivered. Cas9 optimization can be used to enhance function or develop new functions. One can generate chimeric Cas9 proteins, Cas9 can be used as a general DNA binding protein. The structural information provided for Cas9 can be used to further engineer and optimize the CRISPR-Cas system, and this can also be extrapolated to interrogate structure-function relationships in other CRISPR enzyme systems, particularly other type II CRISPR enzymes or Cas9 orthologs.Crystal structure information (described in U.S. provisional applications 61 / 915,251, filed December 12, 2013, 61 / 930,214, filed January 22, 2014, 61 / 980,012, filed April 15, 2014; and Nishimasu et al., "Crystal Structure of Cas9 in Complex with Guide RNA and Target DNA," Cell 156(5):935-949, DOI: http: / / dx.doi.org / 10.1016 / j.cell.2014.02.001 (2014), each and all of which are incorporated herein by reference) provides structural information for truncation and production of modular or multi-part CRISPR enzymes that can be integrated into inducible CRISPR Cas systems. In particular, structural information is provided for Streptococcus pyogenes Cas9 (SpCas9), which can be extrapolated to other Cas9 orthologs or other Type II CRISPR enzymes. Cas9 genes are found in several different bacterial genomes, often located in the same locus as casl, cas2, and cas4 genes and CRISPR cassettes. In addition, Cas9 proteins comprise an easily identifiable C-terminal region that is homologous to transposon ORF-B and includes an active RuvC-like nuclease, an arginine-rich region.

[0317] dCas9

[0318] A Cas9 protein can be mutated such that nuclease activity is inactivated. An inactivated Cas9 protein from S. pyogenes without endonuclease activity (iCas9, also called "dCas9") has recently been targeted by gRNAs to genes in bacteria, yeast, and human cells to silence gene expression by steric hindrance. As used herein, a "dCas molecule" can refer to a dCas protein or a fragment thereof. As used herein, a "dCas9 molecule" can refer to a dCas9 protein or a fragment thereof. As used herein, the terms "iCas" and "dCas" are used interchangeably to refer to a CRISPR-associated protein without catalytic activity. In one embodiment, a dCas molecule comprises one or more mutations in the DNA cleavage domain. In one embodiment, a dCas molecule comprises one or more mutations in the RuvC or HNH domain. In one embodiment, a dCas molecule comprises one or more mutations in both the RuvC and HNH domains. In one embodiment, a dCas molecule is a fragment of a wild-type Cas molecule. In one embodiment, a dCas molecule comprises a functional domain from a wild-type Cas molecule, wherein the functional domain is selected from a Reel domain, a bridge helix domain, or a PAM interaction domain. In one embodiment, a dCas molecule has at least 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% less nuclease activity compared to the nuclease activity of the corresponding wild-type Cas molecule. An exemplary amino acid sequence of a wild-type dCas9 protein is set forth in SEQ ID NO: 1. An exemplary nucleic acid sequence encoding a wild-type dCas9 is set forth in SEQ ID NO: 26.

[0319] A suitable dCas molecule can be derived from a wild-type Cas molecule. The Cas molecule can be from a Type I, Type II, or Type III CRISPR-Cas system. In one embodiment, a suitable dCas molecule can be derived from a Casl, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, or CaslO molecule. In one embodiment, the dCas molecule is derived from a Cas9 molecule. For example, a dCas9 molecule can be obtained by introducing point mutations (e.g., substitutions, deletions, or additions) in the DNA cleavage domain, e.g., nuclease domain, e.g., RuvC and / or HNH domain, of a Cas9 molecule. See, e.g., Jinek et al. (2012) 337:816-21, incorporated by reference in its entirety. For example, introducing two point mutations in the RuvC and HNH domains reduces Cas9 nuclease activity while retaining Cas9 sgRNA and DNA binding activity. In one embodiment, the two point mutations within the RuvC and HNH active sites are D10A and H840A mutations of a S. pyogenes Cas9 molecule. Alternatively, D10 and H840 of a S. pyogenes Cas9 molecule can be deleted to eliminate Cas9 nuclease activity while retaining its sgRNA and DNA binding activity. In one embodiment, the two point mutations within the RuvC and HNH active sites are D10A and N580A mutations of a S. pyogenes Cas9 molecule.

[0320] In one embodiment, the dCas molecule is a S. aureus dCas9 molecule comprising a mutation at D10 and / or N580 (numbered according to SEQ ID NO: 1). In one embodiment, the dCas molecule is a S. aureus dCas9 molecule comprising a D10A and / or N580A mutation (numbered according to SEQ ID NO: 1).

[0321] In one embodiment, the dCas9 molecule is a S. aureus dCas9 molecule comprising the amino acid sequence of SEQ ID NO: 1, a sequence having substantial identity (e.g., at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 1, or a sequence having 1, 2, 3, 4, 5 or more changes (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 1 or any fragment thereof.

[0322] Similar mutations can also be applied to any other naturally occurring Cas9 (e.g., Cas9 from other species) or engineered Cas9 molecule. In certain embodiments, the dCas9 comprises a S. pyogenes dCas9 molecule, a S. aureus dCas9 molecule, a C. jejuni dCas9 molecule, a C. diphtheriae dCas9 molecule, a E. truebdontium dCas9 molecule, a S. parasanguinis dCas9 molecule, a L. farciminus dCas9 molecule, a C. globoformis dCas9 molecule, a Azospirillum (e.g., strain B510) dCas9 molecule, a Acetobacter diazotrophicus dCas9 molecule, a N. cinerea dCas9 molecule, a R. intestinalis dCas9 molecule, a C. hygienicola dCas9 molecule, a Haliangium sp. (e.g., strain DSM 16511) dCas9 molecule, a C. gullinarum (e.g., strain CF89-12) dCas9 molecule, a S. thermophilus (e.g., strain LMD-9) dCas9 molecule, or a fragment thereof.

[0323] In certain embodiments, the present application provides a vector comprising a nucleotide encoding a S. pyogenes dCas9 molecule, a S. aureus dCas9 molecule, a C. jejuni dCas9 molecule, a C. diphtheriae dCas9 molecule, a E. truebdontium dCas9 molecule, a S. parasanguinis dCas9 molecule, a L. farciminus dCas9 molecule, a C. globoformis dCas9 molecule, a Azospirillum (e.g., strain B510) dCas9 molecule, a Acetobacter diazotrophicus dCas9 molecule, a N. cinerea dCas9 molecule, a R. intestinalis dCas9 molecule, a C. hygienicola dCas9 molecule, a Haliangium sp. (e.g., strain DSM 16511) dCas9 molecule, a C. gullinarum (e.g., strain CF89-12) dCas9 molecule, a S. thermophilus (e.g., strain LMD-9) dCas9 molecule, or a fragment thereof.

[0324] In some embodiments, the dCas9 protein comprises a sequence set forth in any one of SEQ ID NOs: 106-122.

[0325] The dCas9 molecule in the constructs provided herein can be continuous or cleaved into an N-terminal portion (dCas9N) and a C-terminal portion (dCas9C). In some embodiments, the N-terminal and C-terminal portions are separated by a DNMT3A and / or a DNMT3L. Accordingly, provided herein are constructs comprising, in order from N-terminal to C-terminal: a dCas9N, a DNMT3A, a DNMT3L, and a dCas9C.

[0326] It will be understood by one of skill in the art that any dCas9 protein can be cleaved at different sites in the protein sequence, so long as the fusion protein can refold into a functional dCas9 molecule. Exemplary sequences for dCas9-N and dCas9-C sequences are listed in Table 2. The N-terminal and C-terminal sequences are preferably paired according to their numbering, e.g., dCas9N-1 is used with dCas9C-1, dCas9N-2 is used with dCas9C-2, and so on. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 2 and a dCas9C comprising the sequence set forth in SEQ ID NO: 3. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 4 and a dCas9C comprising the sequence set forth in SEQ ID NO: 5. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 6 and a dCas9C comprising the sequence set forth in SEQ ID NO: 7. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 8 and a dCas9C comprising the sequence set forth in SEQ ID NO: 9. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 10 and a dCas9C comprising the sequence set forth in SEQ ID NO: 11. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 12 and a dCas9C comprising the sequence set forth in SEQ ID NO: 13. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 14 and a dCas9C comprising the sequence set forth in SEQ ID NO: 15. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 16 and a dCas9C comprising the sequence set forth in SEQ ID NO: 17. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 18 and a dCas9C comprising the sequence set forth in SEQ ID NO: 19. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 20 and a dCas9C comprising the sequence set forth in SEQ ID NO: 21. In some embodiments, the constructs provided herein comprise a dCas9N comprising the sequence set forth in SEQ ID NO: 22 and a dCas9C comprising the sequence set forth in SEQ ID NO: 23.In some embodiments, the constructs provided herein comprise a dCas9N comprising a sequence set forth in SEQ ID NO: 24 and a dCas9C comprising a sequence set forth in SEQ ID NO: 25.

[0327] In some embodiments, the dCas9N is encoded by a polynucleotide comprising the nucleic acid sequence in SEQ ID NO: 27, 29, 31, 33, 35, 37, 39, 41, 43, 45, 47, or 49. In some embodiments, the dCas9C is encoded by a polynucleotide comprising the nucleic acid sequence set forth in SEQ ID NO: 28, 30, 32, 34, 36, 38, 40, 42, 44, 46, 48, or 50.

[0328] Table 2: Exemplary dCas9 and cleaved dCas9 sequences

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335]

[0336]

[0337]

[0338] Cas9 fusion proteins

[0339] A CRISPR / Cas9-based system can include a fusion molecule (e.g., DNMT3A-DNMT3L (3A3L)-dCas9-KRAB). In some embodiments, the fusion molecule includes a dCas9, a KRAB, a DNMT3A, and a DNMT3L,

[0340] In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to DNMT3A or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to DNMT3L or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to DNMT3L and DNMT3L or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to a DNMT3A-DNMT3L fusion peptide.

[0341] DNMT3A:

[0342] MNHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQVDRYIASEVCEDSITVGMVRHQGKIMYVGDVRSVTQKHIQEWGPFDLVIGGSPCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPFFWLFENVVAMGVSDKRDISRFLESNPVMIDAKEVSAAHRARYFWGNLPGMNRPLASTVNDKLELQECLEHGRIAKFSKVRTITTRSNSIKQGKDQHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRSWSVPVIRHLFAPLKEYFACV (SEQ ID NO: 69)

[0343] DNMT3L:

[0344] MGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGTLKYVEDVTNVVRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFHRILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQDVRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVRSRSKLDAPKVDLLVKNCLLPLREYFKYFSQNSLPL (SEQ ID NO: 70)

[0345] DNMT3A-DNMT3L fusion peptide:

[0346] MNHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQVDRYIASEVCEDSITVGMVRHQGKIMYVGDVRSVTQKHIQEWGPFDLVIGGSPCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPFFWLFENVVAMGVSDKRDISRFLESNPVMIDAKEVSAAHRARYFWGNLPGMNRPLASTVNDKLELQECLEHGRIAKFSKVRTITTRSNSIKQGKDQHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRSWSVPVIRHLFAPLKEYFACVSSGNSNANSRGPSFSSGLVPLSLRGSHMGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGTLKYVEDVTNVVRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFHRILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQDVRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVRSRSKLDAPKVDLLVKNCLLPLREYFKYFSQNSLPL (SEQ ID NO: 71)

[0347] In some embodiments, the DNMT3A is a human DNMTA. In some embodiments, the DNMT3A is a non-human DNMT3A, such as a rodent DNMT3A. In some embodiments, the DNMT3A is selected from Mus musculus DNMT3L, Mus castanlus DNMT3L, Mus sipreeus DNMT3L, Rattus norvegicus DNMT3L, Rattus rattus DNMT3L, Arvicanthis niloticus DNMT3L, Peromyscus eremicus DNMT3L, or Peromyscus maniculatus DNTM3L.

[0348] In some embodiments, the DNMT3L is human DNMT3L (SEQ ID NO: 74). In some embodiments, the DNMT3L is a non-human DNMT3L, such as a rodent DNMT3L. In some embodiments, the DNMT3L is selected from Mus musculus DNMT3L (SEQ ID NO: 75), Mus castaneus DNMT3L (SEQ ID NO: 76), Mus reductus DNMT3L (SEQ ID NO: 77), Rattus norvegicus DNMT3L (SEQ ID NO: 78), Rattus rattus DNMT3L (SEQ ID NO: 79), Arvicanthis niloticus DNMT3L (SEQ ID NO: 80), Peromyscus eremicus DNMT3L (SEQ ID NO: 81), and Peromyscus maniculatus DNMT3L (SEQ ID NO: 82).

[0349] In some embodiments, the DNMT3L is encoded by a polynucleotide comprising the nucleic acid sequence set forth in any one of SEQ ID NOs: 83-92.

[0350] In one embodiment, the Cas9 fusion protein further comprises a nuclear localization sequence (NLS), such as n NLSs fused to the N-terminus and / or C-terminus of Cas9.

[0351] Nuclear localization sequences are known in the art. In one embodiment, the NLS comprises the amino acid sequence of SEQ ID NO: 72, 73, 125, or 126, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO: 72, 73, 125, or 126, or a sequence having 1, 2, 3, 4, 5 or more alterations (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 72 or 73, or any fragment thereof.

[0352] SEQ ID NO: 72 (exemplary nuclear localization sequence): APKKKRKVGIHGVPAA

[0353] SEQ ID NO: 73 (exemplary nuclear localization sequence): KRPAATKKAGQAKKKK

[0354] SEQ ID NO: 125 (BPNLS): KRTADGSEFESPKKKRKV

[0355] SEQ ID NO: 126 (SV40 NLS): PKKKRKV

[0356] The construct can also comprise a gene expression modulator. Different gene expression modulators are known in the art, see, e.g., Thakore et al., Nat Methods. 2016; 13: 127-37, incorporated by reference herein in its entirety. Non-limiting examples of gene expression modulators include Kruppel-associated box (KRAB), enhancer of zeste homolog 2 (EZH2), G9A, lysine-specific histone demethylase 1A (LSD1), heterochromatin protein 1 (HP1), friend of GATA protein 1 (FOG1), histone deacetylase (HDAC3), and DOT1L, each of which will be described in more detail below.

[0357] In some embodiments, the modulator of epigenetic modification can have DNA methylase activity. For example, the modulator of epigenetic modification can have methylase activity that involves transferring a methyl group to DNA, RNA, a protein, a small molecule, a cytosine, or an adenine.

[0358] In some embodiments, the CRISPR / Cas9-based system can comprise a dCas9 molecule and a gene expression modulator, or a nucleic acid encoding a dCas9 molecule and a gene expression modulator. In one embodiment, the dCas9 molecule and the gene expression modulator are covalently linked. In one embodiment, the gene expression modulator is directly covalently fused to the dCas9 molecule. In one embodiment, the gene expression modulator is indirectly covalently fused to the dCas9 molecule (e.g., through a non-modulator or linker, or through a second modulator). In one embodiment, the gene expression modulator is located at the N-terminus and / or C-terminus of the dCas9 molecule. In one embodiment, the dCas9 molecule and the gene expression modulator are non-covalently linked. Exemplary sequences include, but are not limited to, those listed in Table 3. In some embodiments, the linker between the dCas9 and the at least one gene expression modulator comprises an amino acid sequence corresponding to a linker listed in Table 3.

[0359] Table 3: Exemplary linker sequences

[0360]

[0361] In one embodiment, the dCas9 molecule is fused to a first tag, e.g., a first peptide tag. In one embodiment, the gene expression modulator is fused to a second tag, e.g., a second peptide tag. In one embodiment, the first and second tags, e.g., the first and second peptide tags, non-covalently interact with each other, thereby bringing the dCas9 molecule and the gene expression modulator in close proximity.

[0362] In one embodiment, the CRISPR / Cas9-based system includes a fusion molecule or a nucleic acid encoding a fusion molecule. In one embodiment, the fusion molecule includes a sequence including a dCas9 fused to a gene expression modulator. In one embodiment, the dCas9 molecule includes a Streptococcus pyogenes dCas9 molecule, a Staphylococcus aureus dCas9 molecule, a Campylobacter jejuni dCas9 molecule, a Corynebacterium diphtheriae dCas9 molecule, a Eubacterium dolichum dCas9 molecule, a Streptococcus parasanguinis dCas9 molecule, a Lactobacillus farciminus dCas9 molecule, a Chaetomium globosum dCas9 molecule, a Azospirillum (e.g., strain B510) dCas9 molecule, a Gluconacetobacter diazotrophicus dCas9 molecule, a Neisseria cinerea dCas9 molecule, a Roseburia intestinalis dCas9 molecule, a Parvibacter pectinophilus dCas9 molecule, a Halorhodinacetae (e.g., strain DSM 16511) dCas9 molecule, a Campylobacter gromini (e.g., strain CF89-12) dCas9 molecule, a Streptococcus thermophilus (e.g., strain LMD-9) dCas9 molecule, or a fragment thereof.

[0363] Gene expression modulators

[0364] In some embodiments, the constructs provided herein include one or more gene expression modulators. The gene expression modulator can be a repressor of gene expression or an activator of gene expression. The repressor can be any known repressor of gene expression, for example, a repressor selected from the group consisting of a Kruppel-associated box (KRAB) domain, a mSin3 interaction domain (SID), a MAX-interacting protein 1 (MXI1), a chromo shadow domain, an EAR repression domain (SRDX), a eukaryotic release factor 1 (ERF1), a eukaryotic release factor 3 (ERF3), a tetracycline repressor, a lad repressor, a CATHARABELOS box binding factors 1 and 2, a Drosophila Groucho, a tripartite motif 28 (TRTM28), a nuclear receptor co-repressor 1, a nuclear receptor co-repressor 2, or a fragment or fusion thereof. The activator can be any known activator of gene expression, for example, a VP16 activation domain, a VP64 activation domain, a p65 activation domain, an Epstein-Barr virus R transactivation Rta molecule, or a fragment thereof. Activators that can be used with dCas9 molecules are known in the art. See Chavez et al., Nat Methods. (2016) 13:563-67, incorporated by reference in its entirety.

[0365] In particular embodiments, the methods and compositions disclosed herein include fusion molecules comprising a dCas9 molecule fused to a gene expression modulator. In some embodiments, the gene expression modulator includes a regulator of epigenetic modification. In one embodiment, the fusion molecule modulates expression of a target gene by making an epigenetic modification on a regulatory element (e.g., a promoter, enhancer, or transcription start site) of the target gene, e.g., by histone acetylation or methylation, or DNA methylation. The modulator can be any known regulator of epigenetic modification, e.g., a histone acetyltransferase (e.g., a p300 catalytic domain), a histone deacetylase, a histone methyltransferase (e.g., SUV39H1 or G9a (EHMT2)), a histone demethylase (e.g., LSD1), a DNA methyltransferase (e.g., DNMT3a or DNMT3a-DNMT3L), a DNA demethylase (e.g., a TET1 catalytic domain or TDG), or a fragment thereof.

[0366] Kruppel-associated box (KRAB)

[0367] The KRAB domain is a type of transcriptional repression domain that is found in the N-terminal portion of many zinc-finger-based transcription factors. When tethered to target DNA by a DNA-binding domain, the KRAB domain functions as a transcriptional repressor. The KRAB domain is rich in charged amino acids and can be divided into subdomains A and B. The KRAB A and B subdomains can be separated by a variable spacer segment, and many KRAB proteins contain only the A subdomain. The sequence of 45 amino acids in the KRAB A subdomain has been shown to be important for transcriptional repression. The B subdomain does not repress transcription by itself, but enhances the repression exerted by the KRAB A subdomain. The KRAB domain recruits the co-repressor KAP1 (KRAB- associated protein-1, also known as transcription intermediary factor 1 beta, KAP1-A interacting protein, and tripartite motif protein 28) and heterochromatin protein 1 (Hpl), as well as other chromatin-modulating proteins, leading to transcriptional repression through heterochromatin formation. In one embodiment, the methods and compositions disclosed herein include fusion molecules comprising a dCas9 molecule fused to a KRAB domain or fragment thereof. In one embodiment, the KRAB domain or fragment thereof is fused to the N-terminus of the dCas9 molecule. In one embodiment, the KRAB domain or fragment thereof is fused to the C-terminus of the dCas9 molecule. In one embodiment, the KRAB domain or fragment thereof is fused to both the N-terminus and the C-terminus of the dCas9 molecule. In one embodiment, the fusion molecule comprises a KRAB domain comprising the sequence of SEQ ID NO: 51, 53, or 230-241, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO: 51, 53, or 230-241, or a sequence having 1, 2, 3, 4, or 5 or more alterations (e.g., amino acid substitutions, insertions, or deletions) relative to SEQ ID NO: 51, 53, or 230-241, or any fragment thereof.

[0368] Exemplary KRAB domain sequences:

[0369] RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEP (SEQ ID NO: 51).

[0370] In one embodiment, the fusion molecule is a DNMT3A-DNMT3L (3A3L)-dCas9-KRAB fusion molecule comprising, from N-terminus to C-terminus: a DNMT3A-DNMT3L fusion peptide (3A3L), a dCas9 peptide, and a KRAB peptide domain fused directly or indirectly (e.g., via a linker).

[0371] In one embodiment, the fusion molecule is a DNMT3A-DNMT3L (3A3L)-dCas9-KRAB fusion molecule comprising, from N-terminus to C-terminus: a DNMT3A-DNMT3L fusion peptide (3A3L), a dCas9 peptide, and a KRAB peptide domain fused directly or indirectly (e.g., via a linker).

[0372] In one embodiment, the fusion molecule comprises the amino acid sequence of SEQ ID NO: 96, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 97, or a sequence having 1, 2, 3, 4, 5 or more changes (e.g., substitutions, insertions or deletions) relative to SEQ ID NO: 96 or any fragment thereof.

[0373] DNMT3A-DNMT3L (3A3L)-dCas9-KRAB:

[0374] MNHDQEFDPPKVYPPVPAEKRKPIRVLSLFDGIATGLLVLKDLGIQVDRYIASEVCEDSITVGMVRHQ GKIMYVGDVRSVTQKHIQEWGPFDLVIGGSPCNDLSIVNPARKGLYEGTGRLFFEFYRLLHDARPKEGDDRPFFWL FENVVAMGVSDKRDISRFLESNPVMIDAKEVSAAHRARYFWGNLPGMNRPLASTVNDKLELQECLEHGRIAKFSKV RTITTRSNSIKQGKDQHFPVFMNEKEDILWCTEMERVFGFPVHYTDVSNMSRLARQRLLGRSWSVPVIRHLFAPLK EYFACV SSGNSNANSRGPSFSSGLVPLSLRGSHMGPMEIYKTVSAWKRQPVRVLSLFRNIDKVLKSLGFLESGSGSGGGTLKYVEDVTNVVRRDVEKWGPFDLVYGSTQPLGSSCDRCPGWYMFQFHRILQYALPRQESQRPFFWIFMDNLLLTEDDQETTTRFLQTEAVTLQDVRGRDYQNAMRVWSNIPGLKSKHAPLTPKEEEYLQAQVRSRSKLDAPKVDLLVKNCLLPLREYFKYFSQNSLPL GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPT STEEGTSTEPSEGSAPGTSTEPSE PKKKRKV MDKKYSIGLAIGTNSVGWAVITDEYKVPSKKFKVLGNTDRHSIKK NLIGALLFDSGETAEATRLKRTARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGN IVDEVAYHEKYPTIYHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFE ENPINASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKDTYDD DLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRVNTEITKAPLSASMIKRYDEHHQDLTLLKALVRQQLPEKYKE IFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLVKLNREDLLRKQRTFDNGSIPHQIHLGELHAILRR QEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEETITPWNFEEVVDKGASAQSFIERMTNFDKN LPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKPAFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDS VEISGVEDRFNASLGTYHDLLKIIKDKDFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKR RRYTGWGRLSRKLINGIRDKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAG SPAIKKGILQTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENTQ LQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDAIVPQSFLKDDSIDNKVLTRSDKNRGKSDNVPSEEVVKKMKN YWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQILDSRMNTKYDENDKLIREVKVIT LKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKKYPKLESEFVYGDYKVYDVRKMIAKSEQEIGKAT AKYFFYSNIMNFFKTEITLANGEIRKRPLIETNGETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKES ILPKRNSDKLIARKKDWDPKKYGGFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAK GYKEVKKDLIIKLPKYSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHYLDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPAAFKYFDTTIDRKRYT STKEVLDATLIHQSITGLYETRIDLSQLGGD AYPYDVPDYASLGSGS PKKKRKV ED PKKKRKV DG SGSETPGTSES ATPES RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEP (SEQ ID NO: 96)

[0375] In some embodiments, the KRAB domain is a zinc finger imprint 3 (ZIM3) KRAB domain. The ZIM3 KRAB domain is a potent repressor of gene expression. In some embodiments, the ZIM3 KRAB domain comprises the sequence of SEQ ID NO: 53, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 53, or a sequence having 1, 2, 3, 4, 5 or more alterations (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 53 or any fragment thereof. In some embodiments, the ZIM3 KRAB domain is encoded by the sequence set forth in SEQ ID NO: 54.

[0376] Enhancer of Zeste Homolog 2 (EZH2)

[0377] In one embodiment, the construct provided herein comprises a Enhancer of Zeste Homolog 2 (EZH2) domain. EZH2 is a histone methyltransferase that modulates several aspects of cell cycle progression. It removes methyl groups from monomethylated and dimethylated lysine 4 and / or lysine 9 of histone H3 (H3K4mel / 2 and H3K9mel / 2),

[0378] In some embodiments, the EZH2 domain comprises the sequence of SEQ ID NO: 55, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 55, or a sequence having 1, 2, 3, 4, 5 or more alterations (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 55 or any fragment thereof. In some embodiments, the EZH2 domain is encoded by the sequence set forth in SEQ ID NO: 56.

[0379] G9a (Euchromatic histone-lysine N-methyltransferase 2, EHMT2) ,

[0380] In some embodiments, the constructs provided herein comprise a G9A or EhMT2 domain. EHMT2 is a methyltransferase that methylates lysine residues of histone H3. In some embodiments, the EHMT2 domain is fused to the C-terminal domain of a dCas9 protein. In some embodiments, the EHMT2 domain is fused to the C-terminus of a KRAB domain.

[0381] In some embodiments, the EHMT2 domain comprises the sequence of SEQ ID NO: 57, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 57, or a sequence having 1, 2, 3, 4, 5 or more changes (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 57 or any fragment thereof. In some embodiments, the EZH2 domain is encoded by the sequence set forth in SEQ ID NO: 58.

[0382] Lysine-specific histone demethylase 1A (LSD1) ,

[0383] In some embodiments, the constructs provided herein comprise a LSD1 domain. LSD1 is a histone demethylase that removes methyl groups from mono- and dimethylated lysine 4 and / or lysine 9 on histone H3 (H3K4mel / 2 and H3K9mel / 2). In some embodiments, the LSD1 domain is fused to the C-terminus of a dCas9 protein. In some embodiments, the LSD1 domain is fused to the C-terminus of a KRAB domain.

[0384] In some embodiments, the LSD1 domain comprises the sequence of SEQ ID NO: 59, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 59, or a sequence having 1, 2, 3, 4, 5 or more changes (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 59 or any fragment thereof. In some embodiments, the LSD1 domain is encoded by the sequence set forth in SEQ ID NO: 60.

[0385] Heterochromatin protein 1 (HP1)

[0386] In some embodiments, the constructs provided herein comprise a HP1 domain. HP1 promotes the formation of heterochromatin structures. In some embodiments, the HP1 domain is fused to the C-terminus of a dCas9 protein. In some embodiments, the HP1 domain is fused to the C-terminus of a KRAB domain.

[0387] In some embodiments, the HP1 domain comprises the sequence of SEQ ID NO: 61, a sequence that is substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more sequence identity) to SEQ ID NO: 61, or a sequence that has 1, 2, 3, 4, 5 or more changes (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 61 or any fragment thereof. In some embodiments, the LSD1 domain is encoded by the sequence set forth in SEQ ID NO: 62.

[0388] Friend of GATA protein 1 (FOG1)

[0389] In some embodiments, the constructs provided herein comprise Friend of GATA 1 (FOG1). FOG1 is a co-factor of the GATA1 transcription factor that regulates cell differentiation. In some embodiments, the FOG1 domain is fused to the C-terminus of a dCas9 protein. In some embodiments, the FOG1 domain is fused to the C-terminus of a KRAB domain.

[0390] In some embodiments, the FOG1 domain comprises the sequence of SEQ ID NO: 63, a sequence that is substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identity) to SEQ ID NO: 63, or a sequence that has 1, 2, 3, 4, 5 or more changes (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 63 or any fragment thereof. In some embodiments, the FOG1 domain is encoded by the sequence set forth in SEQ ID NO: 64.

[0391] Histone deacetylase (HDAC3)

[0392] In some embodiments, the constructs provided herein comprise HDAC3. HDAC3 promotes deacetylation of lysine residues on the N-terminal portion of core histones (H2A, H2B, H3 and H4). In some embodiments, the HDAC3 domain is fused to the C-terminus of a dCas9 protein. In some embodiments, the HDAC3 domain is fused to the C-terminus of a KRAB domain.

[0393] In some embodiments, the HDAC3 domain comprises the sequence of SEQ ID NO: 65, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO: 65, or a sequence having 1, 2, 3, 4, 5 or more alterations (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 65 or any fragment thereof. In some embodiments, the HDAC3 domain is encoded by the sequence set forth in SEQ ID NO: 66.

[0394] DOT1-like histone lysine methyltransferase (DOT1L)

[0395] In some embodiments, the constructs provided herein comprise DOT1L. DOT1L is a histone methyltransferase that methylates lysine-79 of histone H3. In some embodiments, the DOT1L domain is fused to the C-terminus of a dCas9 protein. In some embodiments, the DOT1L domain is fused to the C-terminus of a KRAB domain.

[0396] In some embodiments, the DOT1L domain comprises the sequence of SEQ ID NO: 67, a sequence substantially identical (e.g., at least 80%, 85%, 90%, 92%, 95%, 97%, 98%, 99% or more identical) to SEQ ID NO: 67, or a sequence having 1, 2, 3, 4, 5 or more alterations (e.g., amino acid substitutions, insertions or deletions) relative to SEQ ID NO: 67 or any fragment thereof. In some embodiments, the DOT1L domain is encoded by the sequence set forth in SEQ ID NO: 68.

[0397] histone modification activity

[0398] In some embodiments, the modulators of epigenetic modifications can have histone modification activity. Histone modification activity can include, but is not limited to, histone deacetylase, histone acetyltransferase, histone demethylase, or histone methyltransferase activity.

[0399] In some embodiments, the modulators of epigenetic modifications can have histone acetyltransferase activity. The histone acetyltransferase can be a p300 or a CREB binding protein (CBP) protein or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to an acetyltransferase p300 or a fragment thereof (e.g., the catalytic core of p300). In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to a CREB binding protein (CBP) protein or a fragment thereof.

[0400] In some embodiments, the modulator of epigenetic modification can have histone demethylase activity. For example, the modulator of epigenetic modification can include an enzyme that removes methyl (CH3-) groups from nucleic acids or proteins, such as histones. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to Lys-specific histone demethylase 1 (LSD1) or a fragment thereof.

[0401] In some embodiments, the modulator of epigenetic modification can have histone methyltransferase activity. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to SUV39H1 or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to G9a (EHMT2) or a fragment thereof.

[0402] DNA demethylase activity

[0403] In some embodiments, the modulator of epigenetic modification can have DNA demethylase activity. For example, the modulator of epigenetic modification can convert methyl groups to hydroxymethylcytosine, a mechanism of DNA demethylation. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to 10-11 translocation methylcytosine dioxygenase 1 (TET1) or a fragment thereof. In some embodiments, the methods and compositions disclosed herein include a fusion molecule comprising a dCas9 molecule fused to thymine DNA glycosylase (TDG) or a fragment thereof.

[0404] gRNA

[0405] As used herein, the term "guide sequence" in the context of a CRISPR-Cas system includes any polynucleotide sequence that has sufficient complementarity to a target nucleic acid sequence to hybridize to the target nucleic acid sequence and direct sequence-specific binding of a nucleic acid-targeting complex to the target nucleic acid sequence. A guide sequence can form a duplex with a target sequence. The duplex can be a DNA duplex, an RNA duplex, or an RNA / DNA duplex. The terms "guide molecule," "guide RNA," and "single guide RNA" are used interchangeably herein to refer to an RNA-based molecule that is capable of forming a complex with a CRISPR-Cas protein and comprises a guide sequence that has sufficient complementarity to a target nucleic acid sequence to hybridize to the target nucleic acid sequence and direct sequence-specific binding of the complex to the target nucleic acid sequence. As described herein, a guide molecule or guide RNA specifically includes an RNA-based molecule having one or more chemical modifications (e.g., by chemically linking two ribonucleotides or by replacing one or more ribonucleotides with one or more deoxyribonucleotides).

[0406] A guide molecule or guide RNA for a CRISPR-Cas protein can include a tracr-mate sequence (including a "forward repeat sequence" in the context of an endogenous CRISPR system) and a guide sequence (also referred to as a "spacer" in the context of an endogenous CRISPR system). In some embodiments, the CRISPR-Cas systems or complexes described herein do not comprise a tracr sequence and / or are not dependent on the presence of a tracr sequence. In certain embodiments, a guide molecule can comprise, consist essentially of, or consist of a forward repeat sequence fused or linked to a guide sequence or spacer sequence.

[0407] Generally, a CRISPR-Cas system is characterized by elements that facilitate formation of a CRISPR complex at the site of a target sequence. A "target sequence" in the context of formation of a CRISPR complex refers to a sequence to which a guide sequence is designed to have complementarity, where hybridization between the target DNA sequence and the guide sequence facilitates formation of a CRISPR complex.

[0408] In certain embodiments, a guide sequence or spacer of a guide molecule is 15 to 50 nucleotides in length. In certain embodiments, a spacer of a guide RNA is at least 15 nucleotides in length. In certain embodiments, a spacer is 15 to 17 nucleotides in length, 17 to 20 nucleotides in length, 20 to 24 nucleotides in length, 23 to 25 nucleotides in length, 24 to 27 nucleotides in length, 27 to 30 nucleotides in length, 30 to 35 nucleotides in length, or greater than 35 nucleotides in length.

[0409] In some embodiments, the guide sequence is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in length.

[0410] In some embodiments, the sequence of the guide molecule (forward repeat sequence and / or spacer) is selected to reduce the extent of secondary structure within the guide molecule. In some embodiments, about or less than about 75%, 50%, 40%, 30%, 25%, 20%, 15%, 10%, 5%, 1%, or less of the nucleotides of the nucleic acid targeting guide RNA participate in self-complementary base pairing when optimally folded. Optimal folding can be determined by any suitable polynucleotide folding algorithm. Some programs are based on calculating the minimum Gibbs free energy. An example of one such algorithm is mFold, as described by Zuker and Stiegler (Nucleic Acids Res. 9 (1981), 133-148). Another example folding algorithm is the online web server RNAfold developed by the Institute for Theoretical Chemistry of the University of Vienna, which uses a centroid structure prediction algorithm (see, e.g., A.R. Gruber et al., 2008, Cell 106(1): 23-24; and PA Carr and GM Church, 2009, Nature Biotechnology 27(12): 1151-62).

[0411] As described above, the CRISPR / Cas9 system utilizes gRNA to provide targeting for the CRISPR / Cas9-based system. gRNA is a fusion of two non-coding RNAs (crRNA and tracrRNA). By exchanging the sequence encoding a 20 bp protospacer, sgRNA can target any desired DNA sequence, which imparts targeting specificity through complementary base pairing with the desired DNA target. gRNA mimics the naturally occurring crRNA:tracRNA duplex involved in type II effector systems. This duplex can include, for example, a 42 nucleotide crRNA and a 75 nucleotide tracrRNA as a guide for Cas9 to cleave a target nucleic acid.

[0412] The term "target region," "target sequence," or "protospacer sequence" used interchangeably herein refers to a region of a target gene targeted by the CRISPR / Cas9-based system. The CRISPR / Cas9-based system can include at least one gRNA, where the gRNA targets different DNA sequences. The target DNA sequences can be overlapping. The target sequence or protospacer sequence is followed by a PAM sequence located at the 3' end of the protospacer sequence. Different type II systems have different PAM requirements. For example, the S. pyogenes type II system uses an "NGG" sequence, where "N" can be any nucleotide.

[0413] In some embodiments, the number of gRNAs administered to the cell can be at least 1 gRNA, at least 2 different gRNAs, at least 3 different gRNAs, at least 4 different gRNAs, at least 5 different gRNAs, at least 6 different gRNAs, at least 7 different gRNAs, at least 8 different gRNAs, at least 9 different gRNAs, at least 10 different gRNAs, at least 11 different gRNAs, at least 12 different gRNAs, at least 13 different gRNAs, at least 14 different gRNAs, at least 15 different gRNAs, at least 16 different gRNAs, at least 17 different gRNAs, at least 18 different gRNAs, at least 19 different gRNAs, at least 20 different gRNAs, at least 25 different gRNAs, at least 30 different gRNAs, at least 35 different gRNAs, at least 40 different gRNAs, at least 45 different gRNAs, or at least 50 different gRNAs.

[0414] In some embodiments, the number of gRNAs administered to the cell can be at least 1 gRNA to at least 50 different gRNAs, at least 1 gRNA to at least 45 different gRNAs, at least 1 gRNA to at least 40 different gRNAs, at least 1 gRNA to at least 35 different gRNAs, at least 1 gRNA to at least 30 different gRNAs, at least 1 gRNA to at least 25 different gRNAs, at least 1 gRNA to at least 20 different gRNAs, at least 1 gRNA to at least 16 different gRNAs, at least 1 gRNA to at least 12 different gRNAs, at least 1 gRNA to at least 8 different gRNAs, at least 1 gRNA to at least 4 different gRNAs, at least 4 gRNAs to at least 50 different gRNAs, at least 4 different gRNAs to at least 45 different gRNAs, at least 4 different gRNAs to at least 40 different gRNAs, at least 4 different gRNAs to at least 35 different gRNAs, at least 4 different gRNAs to at least 30 different gRNAs, at least 4 different gRNAs to at least 25 different gRNAs, at least 4 different gRNAs to at least 20 different gRNAs, at least 4 different gRNAs to at least 16 different gRNAs, at least 4 different gRNAs to at least 12 different gRNAs, at least 4 different gRNAs to at least 8 different gRNAs, at least 8 different gRNAs to at least 50 different gRNAs, at least 8 different gRNAs to at least 45 different gRNAs, at least 8 different gRNAs to at least 40 different gRNAs, at least 8 different gRNAs to at least 35 different gRNAs, 8 different gRNAs to at least 30 different gRNAs, at least 8 different gRNAs to at least 25 different gRNAs, 8 different gRNAs to at least 20 different gRNAs, at least 8 different gRNAs to at least 16 different gRNAs, or 8 different gRNAs to at least 12 different gRNAs.

[0415] In some embodiments, a gRNA is selected to increase or decrease transcription of a target gene. In some embodiments, a gRNA targets a region upstream of the transcription start site (TSS) of a target gene, e.g., VEGFA, e.g., between 0-1000 bp upstream of the target gene transcription start site. In some embodiments, a gRNA targets a region between 0 and 50 bp, between 0 and 100 bp, between 0 and 150 bp, between 0 and 200 bp, between 0 and 250 bp, between 0 and 300 bp, between 0 and 350 bp, between 0 and 400 bp, between 0 and 450 bp, between 0 and 500 bp, between 0 and 550 bp, between 0 and 600 bp, between 0 and 650 bp, between 0 and 700 bp, between 0 and 750 bp, between 0 and 800 bp, between 0 and 850 bp, between 0 and 900 bp, between 0 and 950 bp, or between 0 and 1000 bp upstream of the target gene transcription start site. In some embodiments, a gRNA targets a region within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the target gene transcription start site. In one embodiment, a gRNA targets a region 0 to 300 bp upstream of the TSS of a target gene.

[0416] In some embodiments, the gRNA targets a region downstream of the target gene transcription start site (e.g., between 0 and 1000 bp downstream of the target gene transcription start site). In some embodiments, the gRNA targets a region between 0 and 50 bp, between 0 and 100 bp, between 0 and 150 bp, between 0 and 200 bp, between 0 and 250 bp, between 0 and 300 bp, between 0 and 350 bp, between 0 and 400 bp, between 0 and 450 bp, between 0 and 500 bp, between 0 and 550 bp, between 0 and 600 bp, between 0 and 650 bp, between 0 and 700 bp, between 0 and 750 bp, between 0 and 800 bp, between 0 and 850 bp, between 0 and 900 bp, between 0 and 950 bp, or between 0 and 1000 bp downstream of the target gene transcription start site. In some embodiments, the gRNA targets a region within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp downstream of the target gene transcription start site. In one embodiment, the gRNA targets a region between 0 and 300 bp downstream of the target gene TSS.

[0417] The present disclosure provides sgRNA sequences that target the human VEGFA target gene and sgRNA sequences that target the mouse VEGFA target gene. The sequence of human VEGFA is provided in NCBI Reference Sequence: NG_008732.1. The sequence of mouse VEGFA is provided in NCBI Reference Sequence: NC_000083.7. The sequence of rhesus monkey VEGFA is provided in NCBI Reference Sequence: NC_041757.1. The present disclosure provides sgRNA sequences that also target the CD151, CD81, and PCSK9 target genes. Exemplary sgRNAs include, but are not limited to, those listed in Table 4.

[0418] Table 4. Exemplary sgRNAs

[0419]

[0420] In one embodiment, the gRNA targets a promoter region of the target gene. In one embodiment, the gRNA targets an enhancer region of the target gene. The gRNA can be divided into a target binding region, a Cas9 binding region, and a transcriptional termination region. The target binding region hybridizes to the target region in the target gene. Methods of designing such target binding regions are known in the art, see, e.g., Doench et al., Nat Biotechnol. (2014) 32: 1262-7; and Doench et al., Nat Biotechnol. (2016) 34: 184-91, incorporated by reference in their entireties herein. Design tools are available from target Finder by the Feng Zhang lab, Target Finder (E-CRISP) by the Michael Boutros lab, RGEN Tools (Cas-OF Finder), CasFinder, and CRISPR Optimal Target Finder, among others. In certain embodiments, the target binding region can be between about 15 and about 50 nucleotides in length (about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or about 50 nucleotides in length). In certain embodiments, the target binding region can be between about 19 and about 21 nucleotides in length. In one embodiment, the target binding region is 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.

[0421] In one embodiment, the target binding region is complementary, e.g., fully complementary, to the target region in the target gene. In one embodiment, the target binding region is substantially complementary to the target region in the target gene. In one embodiment, the target binding region comprises no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that are not complementary to the target region in the target gene.

[0422] In one embodiment, the target binding region is engineered to improve stability or to increase half-life, for example, by incorporating non-natural nucleotides or modified nucleotides in the target binding region, by removing or modifying RNA destabilizing sequence elements, by adding RNA stabilizing sequence elements, or by enhancing the stability of the Cas9 / gRNA complex. In one embodiment, the target binding region is engineered to enhance its transcription. In one embodiment, the target binding region is engineered to reduce secondary structure formation. In one embodiment, the Cas9 binding region of the gRNA is modified to enhance transcription of the gRNA. In one embodiment, the Cas9 binding region of the gRNA is modified to improve the stability or assembly of the Cas9 / gRNA complex.

[0423] Delivery system

[0424] The present disclosure also provides delivery systems for introducing components of the systems and compositions herein into a cell, tissue, organ, or organism. The delivery system can include one or more delivery vehicles and / or cargos.

[0425] cargo

[0426] The delivery system can include one or more cargos. The cargo can comprise one or more components of the systems and compositions herein. The cargo can comprise one or more of: i) a plasmid encoding one or more Cas proteins; ii) a plasmid encoding one or more guide RNAs, iii) an mRNA of one or more Cas proteins; iv) one or more guide RNAs; v) one or more Cas proteins; vi) any combination thereof. In some examples, the cargo can comprise a plasmid encoding one or more Cas proteins and one or more (e.g., a plurality of) guide RNAs. In some embodiments, the cargo can comprise an mRNA encoding one or more Cas proteins and one or more guide RNAs.

[0427] In some examples, the cargo can comprise one or more Cas proteins and one or more guide RNAs, for example, in the form of a ribonucleoprotein complex (RNP). The ribonucleoprotein complex can be delivered by the methods and systems herein. In some cases, the ribonucleoprotein can be delivered by a polypeptide-based shuttle agent. In one example, the ribonucleoprotein can be delivered using a synthetic peptide comprising an endosome leakage domain (ELD) operably linked to a cell-penetrating domain (CPD), a histidine-rich domain, and a CPD, for example, as described in WO2016161516.

[0428] physical delivery

[0429] In some embodiments, cargo can be introduced into a cell by physical delivery methods. Examples of physical methods include microinjection, electroporation, and hydrodynamic delivery.

[0430] microinjection

[0431] Direct microinjection of cargo into a cell can achieve high efficiency, e.g., greater than 90% or about 100%. In some embodiments, microinjection can be performed using a microscope and a needle (e.g., 0.5-5.0 pm in diameter) to pierce the cell membrane and deliver cargo directly to a target site within the cell. Microinjection can be used for in vitro and ex vivo delivery.

[0432] Plasmids containing coding sequences for Cas proteins and / or guide RNAs, mRNAs, and / or guide RNAs can be microinjected. In some cases, microinjection can be used to i) deliver DNA directly to the nucleus, and / or ii) deliver mRNA (e.g., in vitro transcribed) to the nucleus or cytoplasm. In certain embodiments, microinjection can be used to deliver sgRNA directly to the nucleus and Cas-encoding mRNA to the cytoplasm, e.g., to facilitate translation of Cas and shuttling of Cas to the nucleus.

[0433] Microinjection can be used to generate genetically modified animals. For example, gene editing cargo can be injected into a zygote to allow efficient germline modification. This approach can generate normal embryos and full-term mouse pups containing one or more desired modifications. Microinjection can also be used to transiently up- or down-regulate (e.g., using CRISPRa and CRISPRi) specific genes within the genome of a cell.

[0434] electroporation

[0435] In some embodiments, cargo and / or delivery vehicles can be delivered by electroporation. Electroporation can use pulsed high-voltage electric current to temporarily open nanometer-sized pores within the cell membrane of a cell suspended in a buffer, allowing components with a hydrodynamic diameter of tens of nanometers to flow into the cell. In some cases, electroporation can be used for a variety of cell types and efficiently transfer cargo into cells. Electroporation can be used for in vitro and ex vivo delivery.

[0436] Electroporation can also be used to deliver cargo into the nucleus of a mammalian cell, by applying specific voltages and reagents, such as by nucleofection. Such methods include those described in Wu Y, et al. (2015). Cell Res 25:67-79; Ye L, et al. (2014). Proc Natl Acad Sci USA 111:9591-6; Choi PS, Meyerson M. (2014). Nat Commun 5:3728; Wang J, Quake SR. (2014). Proc Natl Acad Sci 111:13157-62. Electroporation can also be used to deliver cargo in vivo, for example using the methods described in Zuckermann M, et al. (2015). Nat Commun 6:7391.

[0437] hydrodynamic delivery

[0438] Fluid dynamic delivery can also be used to deliver cargo, for example for in vivo delivery. In some examples, fluid dynamic delivery can be performed by rapidly pushing a large volume (8-10% body weight) solution containing gene editing cargo into the bloodstream of a subject (e.g., an animal or human) via the tail vein (e.g., for mice). Because blood is incompressible, a large bolus of liquid can cause an increase in fluid dynamic pressure, which temporarily enhances the permeability of endothelial and parenchymal cells, allowing cargo that would normally not be able to cross the cell membrane to enter the cell. This method can be used to deliver naked DNA plasmids and proteins. The delivered cargo can be enriched in the liver, kidney, lung, muscle, and / or heart.

[0439] transfection

[0440] Cargo, such as a nucleic acid, can be introduced into a cell by a transfection method that introduces the nucleic acid into the cell. Examples of transfection methods include calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendrimer transfection, heat shock transfection, magnetic transfection, lipofection, impalefection, optical transfection, nucleic acid uptake enhanced by proprietary reagents.

[0441] delivery vehicle

[0442] Delivery systems can include one or more delivery vehicles. Delivery vehicles can carry cargo into a cell, tissue, organ, or organism (e.g., an animal or plant). Cargo can be packaged, transported, or otherwise associated with a delivery vehicle. Delivery vehicles can be selected based on the type of cargo to be delivered, and / or the delivery is in vitro and / or in vivo. Examples of delivery vehicles include vectors, viruses, non-viral vehicles, and other delivery reagents described herein.

[0443] Delivery vehicles according to the present disclosure can have a maximum dimension (e.g., diameter) of less than 100 micrometers (pm). In some embodiments, the delivery vehicle has a maximum dimension of less than 10 pm. In some embodiments, the delivery vehicle can have a maximum dimension of less than 2000 nanometers (nm). In some embodiments, the delivery vehicle can have a maximum dimension of less than 1000 nanometers (nm). In some embodiments, the delivery vehicle can have a maximum dimension (e.g., diameter) of less than 900 nm, less than 800 nm, less than 700 nm, less than 600 nm, less than 500 nm, less than 400 nm, less than 300 nm, less than 200 nm, less than 150 nm, or less than 100 nm, less than 50 nm. In some embodiments, the delivery vehicle can have a maximum dimension of between 25 nm and 200 nm.

[0444] In some embodiments, the delivery vehicle can be or comprise a particle. For example, the delivery vehicle can be or comprise a nanoparticle (e.g., a particle having a maximum dimension (e.g., diameter) of no more than 1000 nm). Particles can be provided in different forms, e.g., as solid particles (e.g., metals such as silver, gold, iron, titanium, etc., non-metals, lipid-based solids, polymers), suspensions of particles, or combinations thereof. Metal, dielectric, and semiconductor particles, as well as hybrid structures (e.g., core-shell particles) can be made.

[0445] vector

[0446] The system, composition, and / or delivery system can comprise one or more vectors. The present disclosure also includes vector systems. A vector system can comprise one or more vectors. In some embodiments, a vector refers to a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. Vectors include single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, no free ends (e.g., circular), comprising DNA, RNA, or both; and other polynucleotide species known in the art. A vector can be a plasmid, e.g., a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Certain vectors can be capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors). Some vectors (e.g., non-episomal mammalian vectors) can be integrated into the genome of a host cell upon introduction into the host cell and thereby are replicated along with the host genome. In certain instances, a vector can be an expression vector, e.g., capable of directing the expression of genes to which they are operably linked. In some cases, an expression vector can be used for expression in eukaryotic cells. Expression vectors commonly used in recombinant DNA technologies are often in the form of plasmids.

[0447] Examples of vectors include pGEX, pMAL, pRIT5, E. coli expression vectors (e.g., pTrc, pET lid), yeast expression vectors (e.g., pYepSecl, pMFa, pJRY88, pYES2, and picZ), baculovirus expression vectors (e.g., for expression in insect cells such as SF9 cells) (e.g., the pAc series and the pVL series), mammalian expression vectors (e.g., pCDM8 and pMT2PC).

[0448] A vector can comprise i) one or more Cas-encoding sequences, and / or ii) a single or at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 12, at least 14, at least 16, at least 32, at least 48, at least 50 guide RNA-encoding sequences. In a single vector, there can be one promoter per RNA-encoding sequence. Alternatively or additionally, in a single vector, there can be a promoter that controls (e.g., drives transcription and / or expression of) multiple RNA-encoding sequences.

[0449] regulatory element

[0450] A vector can comprise one or more regulatory elements. A regulatory element can be operably linked to a coding sequence for a Cas protein, an accessory protein, a guide RNA (e.g., a single guide RNA, a crRNA, and / or a tracrRNA), or a combination thereof. The term “operably linked” is intended to mean that a nucleotide sequence of interest is linked to one or more regulatory elements in a manner that allows for expression of the nucleotide sequence (e.g., in an in vitro transcription / translation system, or in a host cell when the vector is introduced into the host cell). In certain examples, a vector can comprise a first regulatory element operably linked to a nucleotide sequence encoding a Cas protein, and a second regulatory element operably linked to a nucleotide sequence encoding a guide RNA.

[0451] Examples of regulatory elements include promoters, enhancers, internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and poly-U sequences). Such regulatory elements are described, for example, in Goeddel, GENE EXPRESSION TECHNOLOGY: METHODS IN ENZYMOLOGY 185, Academic Press, San Diego, Calif (1990). Regulatory elements include elements that direct constitutive expression of a nucleotide sequence in many types of host cells and elements that direct expression of a nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Tissue-specific promoters can direct expression primarily in a desired tissue of interest, such as muscle, neuronal, bone, skin, blood, a particular organ (e.g., liver, pancreas), or a particular cell type (e.g., lymphocytes). Regulatory elements can also direct expression in a time-dependent manner, such as in a cell cycle-dependent or developmental stage-dependent manner, which can or can not also be tissue- or cell type-specific.

[0452] Examples of promoters include one or more pol III promoters (e.g., 1, 2, 3, 4, 5, or more pol III promoters), one or more pol II promoters (e.g., 1, 2, 3, 4, 5, or more pol II promoters), one or more pol I promoters (e.g., 1, 2, 3, 4, 5, or more pol I promoters), or a combination thereof. Examples of pol III promoters include, but are not limited to, U6 and HI promoters. Examples of pol II promoters include, but are not limited to, the retroviral Rous sarcoma virus (RSV) LTR promoter (optionally with the RSV enhancer), the cytomegalovirus (CMV) promoter (optionally with the CMV enhancer), the SV40 promoter, the dihydrofolate reductase promoter, the actin promoter, the phosphoglycerol kinase (PGK) promoter, and the EF1a promoter.

[0453] viral vector

[0454] Cargo can be delivered by a virus. In some embodiments, a viral vector is used. Viral vectors can include DNA or RNA sequences of viral origin used for packaging into a virus (e.g., retrovirus, replication-defective retrovirus, adenovirus, replication-defective adenovirus, and adeno-associated virus). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Viruses and viral vectors can be used for delivery in vitro, ex vivo, and / or in vivo.

[0455] adeno-associated virus (AAV)

[0456] The systems and compositions herein can be delivered by an adeno-associated virus (AAV). AAV vectors can be used for such delivery. AAV belongs to the Dependovirus genus and Parvoviridae family, and is a single-stranded DNA virus. In some embodiments, AAV can provide a persistent source of the provided DNA, as AAV-delivered genomic material can persist indefinitely in cells, for example, either as extrachromosomal DNA or, with certain modifications, be directly integrated into host DNA. In some embodiments, AAV does not cause or is not associated with any disease in humans. The virus itself is capable of efficiently infecting cells while eliciting little or no innate or adaptive immune response or associated toxicity.

[0457] Examples of AAVs that can be used herein include AAV-1, AAV-2, AAV-3, AAV-4, AAV-5, AAV-6, AAV-8, and AAV-9. The type of AAV can be selected based on the cell to be targeted; for example, AAV serotypes 1, 2, 5, or hybrid capsid AAVl, AAV2, AAV5, or any combination thereof can be selected to target brain or neuronal cells; and AAV4 can be selected to target cardiac tissue. AAV8 is used for delivery to the liver. AAV-2-based vectors were initially proposed for delivery of CFTR to the CF airway, other serotypes such as AAV-1, AAV-5, AAV-6, and AAV-9 have shown improved gene transfer efficiency in various lung epithelial models. Examples of cell types targeted by AAV are described in Grimm, D. et al., J. Virol. 82: 5887-5911 (2008)) and WO 2021 / 183807 Al, which are incorporated by reference herein in their entirety.

[0458] CRISPR-Cas AAV particles can be produced in HEK 293 T cells. Once particles with a particular tropism have been produced, they are used to infect target cell lines in a manner very similar to that of natural viral particles. This can allow the CRISPR-Cas components to persist in the infected cell type, making this form of delivery particularly suitable for situations where long-term expression is desired. Examples of dosages and formulations of AAV that can be used include those described in U.S. Patent Nos. 8,454,972; 8,404,658.

[0459] Various strategies can be used to deliver the systems and compositions herein with AAV. In some examples, the coding sequences for Cas and gRNA can be packaged directly on one DNA plasmid vector and delivered by one AAV particle. In some examples, AAV can be used to deliver gRNA into cells that have been previously engineered to express Cas. In some examples, the coding sequences for Cas and gRNA can be made into two separate AAV particles for co-transfection of target cells. In some examples, markers, tags, and other sequences can be packaged in the same AAV particle as the coding sequences for Cas and / or gRNA.

[0460] lentivirus

[0461] The systems and compositions herein can be delivered by lentivirus. Lentiviral vectors can be used for such delivery. Lentiviruses are complex retroviruses with the ability to infect and express their genes in both mitotic and post-mitotic cells.

[0462] Examples of lentiviruses include human immunodeficiency virus (HIV), which can be used to target a wide range of cell types using envelope glycoproteins from other viruses; minimal non-primate lentiviral vectors based on equine infectious anemia virus (EIAV), which can be used for ocular treatments. In certain embodiments, self-inactivating lentiviral vectors with siRNAs targeting the HIV tat / rev common exon, nucleolar-localizing TAR decoys, and anti-CCR5 specific hammerhead ribozymes (see, e.g., DiGiusto et al. (2010) Sci Transl Med 2:36ra43) can be used and / or adapted for the nucleic acid targeting systems herein.

[0463] Lentiviruses can be pseudotyped with other viral proteins, such as the G protein of vesicular stomatitis virus. In this process, the cellular tropism of the lentivirus can be altered to be broad or narrow as desired. In some cases, to increase safety, second and third generation lentiviral systems can split the essential genes onto three plasmids, which can reduce the likelihood of accidental reconstitution of live virus particles inside a cell.

[0464] In some examples, by virtue of their integration capacity, lentiviruses can be used to create libraries of cells containing various genetic modifications, e.g., for screening and / or studying genes and signaling pathways.

[0465] adenovirus

[0466] The systems and compositions herein can be delivered by adenovirus. Adenovirus vectors can be used for such delivery. Adenoviruses include nonenveloped viruses with an icosahedral nucleocapsid containing a double-stranded DNA genome. Adenoviruses can infect both dividing and non-dividing cells. In some embodiments, adenoviruses do not integrate into the genome of the host cell, which can be useful to limit off-target effects of the CRISPR-Cas system in gene editing applications.

[0467] non-viral vehicle

[0468] Delivery vehicles can include non-viral vehicles. In general, methods and vehicles capable of delivering nucleic acids and / or proteins can be used to deliver the systems compositions herein. Examples of non-viral vehicles include lipid nanoparticles, cell-penetrating peptides (CPPs), DNA nanoclews, gold nanoparticles, streptolysin O, multifunctional envelope-type nanodevices (MENDs), lipid-coated mesoporous silica particles, and other inorganic nanoparticles.

[0469] lipid particle

[0470] Delivery vehicles can comprise lipid particles, such as lipid nanoparticles (LNP) and liposomes.

[0471] lipid nanoparticle (LNP)

[0472] LNP can encapsulate nucleic acids within cationic lipid particles (e.g., liposomes) and can be relatively easily delivered to cells. In some examples, lipid nanoparticles do not contain any viral components, which helps to minimize safety and immunogenicity issues. Lipid particles can be used for in vitro, ex vivo, and in vivo delivery. Lipid particles can be used for cell populations of various scales.

[0473] In some examples, LNP can be used to deliver DNA molecules (e.g., those containing coding sequences for Cas and / or gRNAs) and / or RNA molecules (e.g., mRNA for Cas, gRNAs). In certain cases, LNP can be used to deliver RNP complexes of Cas / gRNAs.

[0474] In some embodiments, the LNP is used to deliver mRNA and gRNA (e.g., mRNA fusion molecules comprising DNMT3A-DNMT3L (3A-3L)-dCas9-KRAB and at least one sgRNA targeting VEGFA).

[0475] The components of the LNP can comprise the cationic lipid 1,2-dioleoyl-3-dimethylammonium- propane (DLinDAP), 1,2-dioleoyloxy-3-N,N-dimethylammoniumopropane (DLinDMA), 1,2-dioleoyloxy ketone-N,N-dimethyl-3-aminopropane (DLinK-DMA), 1,2-dioleoyl-4-(2-dimethylaminoethyl)-[1,3]- dioxolane (DLinKC2-DMA), 3-o-[2-(methoxypolyethylene glycol 2000) succinoyl]-1,2- dimyristyloxypropyl-3-amine (PEG-S-DMG), R-3-[(ro-methoxy-poly(ethylene glycol) 2000) carbamoyl]-1,2-dimyristyloxypropyl-3-amine (PEGC-DOMG), and any combination thereof. The preparation and encapsulation of the LNP can be adapted from Conway et al., Molecular Therapy, Volume 27, Issue 4, Pages 866-877, April 2019 and Rosin et al., Molecular Therapy, Volume 19, Issue 12, Pages 1286-2200, December 2011.

[0476] In some embodiments, the LNP can comprise an ionizable lipid. In some embodiments, the ionizable lipid includes, but is not limited to, pH-responsive ionizable lipids, thermally-responsive ionizable lipids, and light-responsive ionizable lipids. In some embodiments, the ionizable lipid includes cationic and anionic lipids that ionize under certain conditions, such as, but not limited to, pH, temperature, or light. In some embodiments, the molar ratio of ionizable lipids of the LNP is from 20% to about 70% (e.g., from about 20% to about 70%, from about 20% to about 65%, from about 20% to about 60%, from about 20% to about 55%, from about 20% to about 50%, from about 20% to about 45%, from about 20% to about 40%, from about 20% to about 35%, from about 20% to about 30%, from about 20% to about 25%, from about 30% to about 70%, from about 30% to about 65%, from about 30% to about 60%, from about 30% to about 55%, from about 30% to about 50%, from about 30% to about 45%, from about 30% to about 40%, from about 30% to about 35%, from about 40% to about 70%, from about 40% to about 65%, from about 40% to about 60%, from about 40% to about 55%, from about 40% to about 50%, from about 40% to about 45%, from about 50% to about 70%, from about 50% to about 65%, from about 50% to about 60%, from about 50% to about 55%, from about 60% to about 70%, or from about 60% to about 65%).

[0477] In some embodiments, the LNP can comprise a PEGylated lipid. In some embodiments, the molar ratio of PEGylated lipids of the LNP is from 0% to about 30% (e.g., from about 0% to about 30%, from about 0% to about 25%, from about 0% to about 20%, from about 0% to about 15%, from about 0% to about 10%, from about 10% to about 30%, from about 10% to about 25%, from about 10% to about 20%, from about 10% to about 15%, from about 20% to about 30%, or from about 20% to about 25%).

[0478] In some embodiments, the LNP can comprise a stabilizer lipid. In some embodiments, the molar ratio of stabilizer lipids of the LNP is from 30% to about 50% (e.g., from about 30% to about 50%, from about 30% to about 45%, from about 30% to about 40%, from about 30% to about 35%, from about 40% to about 50%, or from about 40% to about 45%).

[0479] In some embodiments, the LNP can comprise cholesterol. In some embodiments, the molar ratio of cholesterol to the LNP is 10% to about 50% (e.g., about 10% to about 50%, about 10% to about 45%, about 10% to about 40%, about 10% to about 35%, about 10% to about 30%, about 10% to about 25%, about 10% to about 20%, about 10% to about 15%, about 20% to about 50%, about 20% to about 45%, about 20% to about 40%, about 20% to about 35%, about 20% to about 30%, about 20% to about 25%, about 30% to about 50%, about 30% to about 45%, about 30% to about 40%, about 30% to about 35%, about 40% to about 50%, or about 40% to about 45%).

[0480] In some embodiments, the LNP can comprise a mixture of an ionizable lipid (20-70%, molar ratio), a PEGylated lipid (0-30%, molar ratio), a supporting lipid (30-50%, molar ratio), and cholesterol (10-50%, molar ratio).

[0481] liposome

[0482] In some embodiments, the lipid particle can be a liposome. Liposomes are spherical vesicular structures composed of a single or multiple lipid bilayers surrounding an internal aqueous compartment and an outer, relatively impermeable, lipophilic phospholipid bilayer. In some embodiments, liposomes are biocompatible, non-toxic, can deliver hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their load across biological membranes and the blood-brain barrier (BBB).

[0483] Liposomes can be made from several different types of lipids, e.g., phospholipids. Liposomes can comprise natural phospholipids and lipids such as 1,2-distearoyl-sn-glycero-3-phosphatidylcholine (DSPC), sphingomyelin, egg phosphatidylcholine class, monosialic ganglioside, or any combination thereof.

[0484] To alter the structure and properties of the liposome, several other additives can be added to the liposome. For example, the liposome can further comprise cholesterol, sphingomyelin, and / or 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE), e.g., to enhance stability and / or prevent leakage of the internal cargo of the liposome.

[0485] stable nucleic acid-lipid particle (SNALP)

[0486] In some embodiments, the lipid particle can be a stable nucleic acid-lipid particle (SNALP). SNALPs can comprise an ionizable lipid (DLinDMA) (e.g., cationic at low pH), a neutral helper lipid, cholesterol, a diffusible polyethylene glycol (PEG)-lipid, or any combination thereof. In some examples, SNALPs can comprise synthetic cholesterol, dipalmitoylphosphatidylcholine, 3-N-[(w-methoxypolyethylene glycol)2000)carbamoyl]-l,2-dimyrestyl oxypropylamine, and the cationic 1,2-dilinoleyl-3-N,N dimethylaminopropane. In some examples, SNALPs can comprise synthetic cholesterol, 1,2-distearoyl-sn-glycero-3-phosphocholine, PEG-cDMA, and 1,2-dilinoleyl-3-(N;N-dimethyl)aminopropane (DLinDMA).

[0487] other lipids

[0488] The lipid particle can also comprise one or more other types of lipids, for example, cationic lipids such as the amino lipid 2,2-dilinoleyl-4-dimethylaminoethyl-[l,3]- dioxolane (DLin-KC2-DMA), DLin-KC2-DMA4, Cl2-200, and colipids distearoylphosphatidylcholine, cholesterol, and PEG-DMG.

[0489] lipoplexes and / or polyplexes

[0490] In some embodiments, the delivery vehicle comprises a lipoplex and / or a polyplex. Lipoplexes can bind to negatively charged cell membranes and induce endocytosis into cells. Examples of lipoplexes can be complexes comprising one or more lipids and non-lipid components. Examples of lipoplexes and polyplexes include FuGENE-6 reagent, non-liposomal solutions containing lipids and other components, zwitterionic amino lipids (ZAL), Ca2+ (e.g., forming DNA / Ca2+ microplexes), polyethylenimine (PEI) (e.g., branched PEI), and poly(L-lysine) (PLL).

[0491] cell-penetrating peptide

[0492] In some embodiments, the delivery vehicle comprises a cell penetrating peptide (CPP). CPPs are short peptides that facilitate cellular uptake of various molecular cargos, e.g., from nano-sized particles to small chemical molecules and large DNA fragments.

[0493] CPPs can have different sizes, amino acid sequences, and charges. In some examples, CPPs can transport across the plasma membrane and facilitate delivery of various molecular cargoes to the cytoplasm or organelles. CPPs can be introduced into cells through different mechanisms (e.g., direct membrane penetration, endocytosis-mediated entry, and transfer through the formation of a transient structure).

[0494] CPPs can have an amino acid composition comprising a relatively rich content of positively charged amino acids such as lysine or arginine, or have a sequence comprising an alternating pattern of polar / charged amino acids with nonpolar, hydrophobic amino acids. These two types of structures are referred to as polycationic or amphipathic, respectively. A third class of CPPs are hydrophobic peptides, comprising only nonpolar residues, with a low net charge or with hydrophobic amino acid groups that are critical for cellular uptake. Another type of CPP is the trans-activating transcriptional activator (Tat) from human immunodeficiency virus I (HIV-I). Examples of CPPs include Penetratin, Tat (48-60), Transportan, and (R-AhX-R4) (Ahx refers to aminohexanoyl). Examples of CPPs and related applications also include those described in U.S. Patent 8,372,951.

[0495] CPPs can be readily used for in vitro and ex vivo work, and often require extensive optimization for each cargo and cell type. In some examples, a CPP can be directly covalently linked to a Cas protein, which is then complexed with a gRNA and delivered to a cell. In some examples, CPP-Cas and CPP-gRNA can be delivered to a plurality of cells separately. CPPs can also be used to deliver RNPs.

[0496] DNA nanowire

[0497] In some embodiments, the delivery vehicle comprises a DNA nanoball. A DNA nanoball refers to a globular structure of DNA (e.g., having the shape of a ball of yarn). Nanoballs can be synthesized by rolling circle amplification with palindromic sequences, which facilitate self-assembly of the structure. The ball can then be loaded with a payload. Examples of DNA nanoballs are described in Sun W et al., J Am Chem Soc. 2014 Oct 22; 136(42): 14722-5 and Sun W et al., Angew Chem Int Ed Engl. 2015 Oct 5; 54(41): 12029-33. A DNA nanoball can have a palindromic sequence that is complementary to a portion of a gRNA in a Cas:gRNA ribonucleoprotein complex. The DNA nanoball can be coated (e.g., with PEI) to induce endosome escape.

[0498] gold nanoparticle

[0499] In some embodiments, the delivery vehicle comprises gold nanoparticles (also referred to as AuNPs or colloidal gold). Gold nanoparticles can form a complex with cargo Cas:gRNA RNP. Gold nanoparticles can be coated (e.g., in silicates and endosome-disrupting polymer PAsp(DET)). Examples of gold nanoparticles include AuraSense Therapeutics' Spherical Nucleic Acid (SNA™) constructs, as well as those described in MMout R, et al. (2017). ACS Nano 11 :2452-8; Lee K, et al. (2017). Nat Biomed Eng 1 :889-901.

[0500] iTOP

[0501] In some embodiments, the delivery vehicle comprises iTOP. iTOP refers to the combination of small molecules that drive efficient intracellular delivery of native proteins independent of any transduction peptide. iTOP can be used for induced transduction by osmocytosis and propanebetaine, using NaCl-mediated hypertonicity with a transduction compound (propanebetaine) to trigger the extracellular macromolecule-to- intracellular macropinocytosis. Examples of iTOP methods and reagents include those described in D'Astolfo DS, Pagliero RJ, Pras A, et al. (2015). Cell 161 :674-690.

[0502] polymer-based particle

[0503] In some embodiments, the delivery vehicle can comprise a polymer-based particle (e.g., nanoparticle). In some embodiments, the polymer-based particle can mimic the membrane fusion viral mechanism. The polymer-based particle can be a synthetic copy of the influenza viral mechanism and forms a transfection complex with various types of nucleic acids (siRNA, miRNA, plasmid DNA or shRNA, mRNA) taken up by cells through the endocytic pathway, a process that involves the formation of acidic compartments. The low pH of the late endosome acts as a chemical switch, making the particle surface hydrophobic and facilitating transmembrane. Once in the cytosol, the particle releases its payload for cellular activity. This active endosomal escape technique is safe and maximizes transfection efficiency because it uses the natural uptake pathway. In some embodiments, the polymer-based particle can comprise alkylated and carboxyalkylated branched polyethyleneimine. In some examples, the polymer-based particle is VIROMER, such as VIROMER RNAi, VIROMER RED, VIROMER mRNA, VIROMER CRISPR. Exemplary methods of delivering the systems and compositions herein include those described in Bawage SS, et al., Synthetic mRNA expressed Casl3a mitigates RNA virus infections, www.biorxiv.org / content / 10 / 1 / 370460v1.full doi: doi.org / 10.1101 / 370460, Viromer® RED, a powerful tool for transfection of keratinocytes. doi: 10.13140 / RG.2.2.16993.61281, Viromer® Transfection - Factbook 2018: technology, product overview, users' data., doi: 10.13140 / RG.2.2.23912.16642.

[0504] streptolysin O (SLO)

[0505] The delivery vehicle can be streptolysin O (SLO). SLO is a toxin produced by group A Streptococcus that acts by creating pores in mammalian cell membranes. SLO can act in a reversible manner, which allows for delivery of proteins (e.g., up to 100 kDa) to the cytosol of cells without compromising overall viability. Examples of SLO include those described in Sierig G, et al. (2003). Infect Immun 71 :446-55; Walev I, et al. (2001). Proc Natl Acad Sci U S A 98:3185-90; Teng Kw, et al. (2017). Elife 6:e25460.

[0506] multifunctional envelope-type nanodevices (MEND)

[0507] The delivery vehicle can include a multifunctional enveloped nanodevice (MEND). MENDs can comprise condensed plasmid DNA, a PLL core, and a lipid membrane shell. MENDs can also comprise a cell-penetrating peptide (e.g., stearyl octaarginine). The cell-penetrating peptide can be present in the lipid shell. The lipid envelope can be modified with one or more functional components, such as one or more of the following: polyethylene glycol (e.g., to increase vascular circulation time), a ligand that targets a particular tissue / cell, an additional cell-penetrating peptide (e.g., for greater cell delivery), a lipid that enhances endosomal escape, and a nuclear delivery tag. In some examples, the MEND can be a four-layer MEND (T-MEND), which can target the nucleus and mitochondria. In certain examples, the MEND can be a PEG-peptide-DOPE-conjugated MEND (PPD-MEND), which can target bladder cancer cells. Examples of MENDs include those described in Kogure K, et al. (2004). J Control Release 98:317-23; Nakamura T, et al. (2012). Ace Chem Res 45:1113-21.

[0508] lipid-coated mesoporous silica particles

[0509] The delivery vehicle can comprise a lipid-coated mesoporous silica particle. The lipid-coated mesoporous silica particle can comprise a mesoporous silica nanoparticle core and a lipid membrane shell. The silica core can have a large internal surface area, resulting in a high cargo loading capacity. In some embodiments, the pore size, pore chemical properties, and overall particle size can be modified in order to load different types of cargo. The lipid coating of the particle can also be modified to maximize cargo loading, increase circulation time, and provide precise targeting and cargo release. Examples of lipid-coated mesoporous silica particles include those described in Du X, et al. (2014). Biomaterials 35:5580-90; Durfee Pn, et al. (2016). ACS Nano 10:8325-45.

[0510] inorganic nanoparticle

[0511] The delivery vehicle can comprise an inorganic nanoparticle. Examples of inorganic nanoparticles include carbon nanotubes (CNTs) (e.g., as described in Bates K and Kostarelos K. (2013). Adv Drug Deliv Rev 65:2023-33), bare mesoporous silica nanoparticles (MSNPs) (e.g., as described in Luo GF, et al. (2014). Sci Rep 4:6064), and dense silica nanoparticles (SiNPs) (as described in Luo D and Saltzman WM. (2000). Nat Biotechnol 18:893-5).

[0512] Methods of use

[0513] The compositions and systems herein can be used for a variety of applications, including modifying non-animal organisms such as plants and fungi, and modifying animals, treating and diagnosing diseases in plants, animals, and humans. In general, the compositions and systems can be introduced into a cell, tissue, organ, or organism, where they alter the expression and / or activity of one or more genes.

[0514] cells and organisms

[0515] The present disclosure provides cells, tissues, organisms comprising engineered Cas proteins, CRISPR-Cas systems, constructs, polynucleotides encoding one or more components of a CRISPR-Cas system, and / or vectors comprising the polynucleotides. The present disclosure also provides nucleotide sequences encoding effector proteins that are codon optimized for expression in a eukaryote or eukaryotic cell in any of the methods or compositions described herein. In embodiments of the present disclosure, the codon optimized effector proteins are any of the Cas proteins discussed herein and are codon optimized for operability in a eukaryotic cell or organism, such as the cells or organisms mentioned elsewhere herein, for example, but not limited to, a yeast cell or a mammalian cell or organism, including a mouse cell, a rat cell, and a human cell or a non-human eukaryote, such as a plant.

[0516] In certain embodiments, the modification of the target locus can result in: a eukaryotic cell comprising altered expression of at least one gene product; a eukaryotic cell comprising altered expression of at least one gene product, wherein expression of the at least one gene product is increased; the eukaryotic cell comprising altered expression of at least one gene product, wherein expression of the at least one gene product is decreased; or a eukaryotic cell comprising an edited genome.

[0517] In certain embodiments, the eukaryotic cell can be a mammalian cell or a human cell.

[0518] In further embodiments, the non-naturally occurring or engineered compositions, vector systems, or delivery systems described in the present specification can be used for: site-specific gene knockout; site-specific genome editing; RNA sequence-specific interference; or multiplexed genome engineering.

[0519] Also provided are gene products from the cells, cell lines, or organisms described herein. In certain embodiments, the amount of the expressed gene product can be greater than or less than the amount of the gene product from a cell without altered expression or an edited genome. In certain embodiments, the gene product can be altered as compared to the gene product from a cell without altered expression or an edited genome.

[0520] Also provided herein are compositions comprising the cells provided herein. In some embodiments, provided herein are pharmaceutical compositions comprising the cells provided herein and a pharmaceutically acceptable carrier.

[0521] methods of modifying gene expression

[0522] In another aspect, provided herein are methods of altering gene expression in a cell or altering gene expression in a subject in vivo. In particular, the methods provided herein can be used to alter expression of a gene product in a cell while minimizing off-target modifications.

[0523] In some embodiments, provided herein are methods of altering expression of a gene product in a population of cells and minimizing off-target modifications in the population of cells, or altering expression of a gene product in a subject in vivo and minimizing off-target modifications, the method comprising the step of introducing into the population of cells or into cells of the subject: (i) a construct described herein or one or more polypeptides expressed from the construct; and (ii) at least one sgRNA, wherein the KRAB and / or gene expression modulator provides for modification of at least one nucleotide in the vicinity of the gene and / or within a gene regulatory element, thereby altering expression of the gene product.

[0524] In some embodiments, the construct of step (i) has the structure of a construct of Formula I: 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3' (I), wherein one of A and B is a polynucleotide encoding DNMT3A or a portion thereof, the other of A and B is a polynucleotide encoding DNMT3L or a portion thereof; CasN is a polynucleotide encoding the N-terminal portion of dCas9; CasC is a polynucleotide encoding the C-terminal portion of dCas9; E is 5'-(A m5 -B m6 ) n3 -K r -D q -3' or 5'-K r -D q -(A m5 -B m6 ) n3 -3'; K is a polynucleotide encoding KRAB or a portion thereof; D is a polynucleotide encoding a gene expression modulator; T comprises a polynucleotide encoding (i) an epitope capable of binding to an antibody or antigen binding fragment thereof or (ii) a polypeptide sequence capable of binding to a nucleic acid structural element; each of ml, m2, m3, m4, m5, and m6 is independently an integer selected from 0 to 3; each of nl, n2, and n3 is independently an integer selected from 0 to 2; p is an integer selected from 0 to 20; q is an integer selected from 0 to 5; r is an integer selected from 0 to 5; and wherein when p is 0, at least one of ml, m2, m3, and m4 is other than 0, at least one of nl and n2 is other than 0, and at least one of q and r is other than 0.

[0525] In some embodiments, the construct has the structure of Formula II: 5'-CasN-(A m3 -B m4 ) n2 -CasC-K r -D q -3' (II), wherein n2 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; q is an integer selected from 0 to 5; and at least one of m3 and m4 is not 0.

[0526] In some embodiments, the construct has the structure of Formula IIa: 5'-CasN-(A-B)-CasC-K r -D q -3' (IIa). In some embodiments, the construct has the structure of Formula III: 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3' (III), wherein p is an integer selected from 1 to 20.

[0527] In some embodiments, the construct has the structure of Formula IIIa: 5'-CasN-CasC-T p -3' (IIIa). In some embodiments, the construct has the structure of Formula IIIb: 5'-(A m1 -B m2 )-CasN-CasC-T p -E-3' (IIIb), wherein at least one of m1 and m2 is not 0. In some embodiments, the construct has the structure of Formula IIIb-1: 5'-(A-B)-CasN-CasC-T p -K r -D q -3' (IIIb-1), wherein: r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5.

[0528] In some embodiments, the construct has the structure of Formula IIIc: 5'-CasN-CasC-T p -E-3' (IIIc), wherein: n3 is an integer selected from 1 and 2; r is an integer selected from 1 to 5; and q is an integer selected from 0 to 5; and at least one of m5 and m6 is not 0. In some embodiments, the construct has the structure of Formula IIIc-1: 5'-CasN-CasC-T p -(Am5 -B m6 ) n3 -K r -D q -3’ (IIIc-1). In some embodiments, the construct has the structure of Formula IIIc-2: 5’-CasN-CasC-T p -K r -D q -(A m5 -B m6 ) n3 -3’ (IIIc-2).

[0529] Without wishing to be bound by theory, it is hypothesized that when T comprises a polynucleotide encoding an epitope capable of binding to an antibody or antigen binding fragment thereof, the 5’-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3’ portion, the 5’-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3’ portion, the 5’-(A m1 -B m2 )-CasN-CasC-T p -3’ portion, the 5’-(A-B)-CasN-CasC-T p -3’ portion, the 5’-CasN-CasC-T p -3’ portion, the 5’-CasN-CasC-T p -3’ portion, or the 5’-CasN-CasC-T p -3’ portion of (i), the sgRNA of (ii), and multiple copies of a polynucleotide comprising the 5’-E-3’ portion of Formula I, the 5’-E-3’ portion of Formula III, the 5’-E-3’ portion of Formula IIIb, the 5’-K r -D q -3’ portion, the 5’-E-3’ portion of Formula IIIc, the 5’-(A m5 -B m6 ) n3 -K r -D q-3' portion or 5'-K of Formula IIIc-2 r -D q -(A m5 -B m6 ) n3 -3' portion of (i) is recruited to the genomic locus by binding of the antibody or antigen binding fragment thereof to the epitope, thereby recruiting the polypeptide expressed from the construct to the genomic locus and altering expression of the gene product in the population of cells.

[0530] In some embodiments, the method further comprises (iii) introducing into the cell a second construct comprising 5'-(A m5 -B m6 ) n3 -K r -D q -3' or 5'-K of Formula IIIc-2 r -D q -(A m5 -B m6 ) n3 -3' or the polypeptide expressed from the second construct. Without wishing to be bound by theory, it is believed that the construct comprising 5'-(A p -3' portion of (i), the sgRNA, and the multiple copies of the polypeptide of (iii) are recruited to the genomic locus by binding of the antibody or antigen binding fragment thereof to the epitope, thereby recruiting the polypeptides expressed from the construct of (i) and the construct of (iii) to the genomic locus and altering expression of the gene product in the population of cells.

[0531] Without wishing to be bound by theory, it is believed that when T comprises a polypeptide sequence capable of binding to a nucleic acid structural element, the construct comprising 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3' portion, 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -3' portion, 5'-(A m1 -B m2 )-CasN-CasC-T p -3' portion, 5'-(A-B)-CasN-CasC-T of Formula IIIb-1 p- 3' portion, 5'-CasN-CasC-T of Formula IIIc p - 3' portion, 5'-CasN-CasC-T of Formula IIIc-1 p - 3' portion, or 5'-CasN-CasC-T of Formula IIIc-2 p - polypeptide of (i), sgRNA of (ii), and multiple copies of (iii) a polypeptide comprising a 5'-E-3' portion of Formula I, a 5'-E-3' portion of Formula III, a 5'-E-3' portion of Formula IIIb, a 5'-K r - D q - 3' portion, 5'-E-3' portion of Formula IIIc, 5'- (A m5 - B m6 ) n3 - K r - D q - 3' portion, or 5'-K of Formula IIIc-2 r - D q - (A m5 - B m6 ) n3 - polypeptide of (i), sgRNA of (ii), and multiple copies of (iii) a polypeptide comprising a 5'-E-3' portion of Formula I, a 5'-E-3' portion of Formula III, a 5'-E-3' portion of Formula IIIb, a 5'-K

[0532] In some embodiments, the method further comprises (iii) introducing into the cell a second construct comprising a 5'- (A m5 - B m6 ) n3 - K r - D q - 3' or 5'-K r - D q - (A m5 - B m6 ) n3 - 3' or a polypeptide expressed by the second construct. Without wishing to be bound by theory, it is believed that the inclusion of a 5'-CasN-CasC-T of Formula IIIa p - polypeptide of (i), sgRNA of (ii), and multiple copies of (iii) a polypeptide comprising a 5'-E-3' portion of Formula I, a 5'-E-3' portion of Formula III, a 5'-E-3' portion of Formula IIIb, a 5'-K

[0533] The methods provided herein can be used to alter gene expression in cultured cells, isolated cells, or cells in vivo (e.g., cells in a subject). In vivo methods of altering gene expression can be used to treat a disease, as further described below.

[0534] In some embodiments, the methods provided herein reduce expression of a target gene. Reduction in target gene expression can be measured, for example, as compared to gene expression of a target gene product in a cell prior to exposure to a construct provided herein, or as compared to an unmodified cell. In some embodiments, the methods provided herein reduce expression of a target gene in a plurality of modified cells by 10-20%, 20-30%, 30-40%, 40-50%, 50-60%, 60-70%, 70-80%, 80-90%, 90-100%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or 100%, as compared to a population of wild-type cells. In some embodiments, the ratio of on-target modification to off-target modification of a gene product is about 10: 1, 9: 1, 8: 1, 7: 1, 6: 1, 5: 1, 4: 1, 3: 1, or 2: 1. In some embodiments, the ratio of on-target modification to off-target modification is no more than 10: 1.

[0535] The modification of at least one nucleotide introduced by the KRAB and / or gene expression modulator can be, for example, DNA methylation or histone modification. The modification can be located in any region of the target gene (including, for example, the coding sequence of the gene or a regulatory element) in which the modification achieves the desired effect (e.g., reduction in gene expression). In some embodiments, the modification occurs in a core promoter, a proximal promoter, a distal enhancer, a silencer, an insulator element, a boundary element, or a locus control region.

[0536] The modification can occur in a nucleotide located upstream or downstream of the transcription start site of the target gene. In some embodiments, the modification occurs at a nucleotide located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp upstream of the transcription start site of the gene. In some embodiments, the modification occurs at a nucleotide located within about 100 bp, about 200 bp, about 300 bp, about 400 bp, about 500 bp, about 600 bp, about 700 bp, about 800 bp, about 900 bp, about 1000 bp, about 1100 bp, about 1200 bp, about 1300 bp, about 1400 bp, or about 1500 bp downstream of the transcription start site of the gene.

[0537] exemplary therapies

[0538] In another aspect, provided herein are methods for treating or alleviating a symptom of a gene product-associated disorder in a subject, comprising the step of introducing into a cell of the subject: (i) a construct provided herein or one or more polypeptides expressed from the construct; and (ii) at least one sgRNA, wherein the KRAB and / or gene expression modulator provides a modification of at least one nucleotide in the vicinity of the gene and / or within a gene regulatory element, thereby altering expression of the gene product and treating or alleviating a symptom of the gene product-associated disorder in the subject.

[0539] The present disclosure provides the use of CRISPR-Cas systems for the treatment of a variety of diseases and disorders. In some embodiments, the disclosure described herein relates to methods of treatment in which cells are edited ex vivo to modulate at least one gene by CRISPR or a base editor, followed by administration of the edited cells to a patient in need. In some embodiments, the editing comprises knocking in, knocking out, or knocking down expression of at least one target gene in the cell. In particular embodiments, the editing inserts an exogenous gene, minigene, or sequence, which can comprise one or more exons in trans or native or synthetic, into a safe harbor locus of the target gene’s genomic location, in which new genes or genetic elements can be introduced without disrupting expression or regulation of neighboring genes, or corrected by inserting or deleting one or more mutations in a DNA sequence encoding a regulatory element of the target gene. In some embodiments, the editing comprises introducing one or more point mutations in a nucleic acid (e.g., genomic DNA) of the target cell.

[0540] In some embodiments, the treatment is for a disease / condition of an organ, including liver disease, eye disease, muscle disease, heart disease, blood disease, brain disease, kidney disease, or can include treatment for autoimmune diseases, central nervous system diseases, cancer and other proliferative diseases, neurodegenerative disorders, inflammatory diseases, metabolic disorders, musculoskeletal disorders, and the like. In some embodiments, the disease is a liver disease. In some embodiments, the disease is an eye disease. In some embodiments, the disease is a CNS disease. In some embodiments, the disease is cancer. In some embodiments, the disease is selected from the group consisting of familial hypercholesterolemia (FH), nonalcoholic steatohepatitis (NASH), Parkinson’s disease, liver fibrosis (HF), age-related macular degeneration (AMD), Angelman syndrome (AS), type II diabetes, beta-thalassemia, and hepatocellular carcinoma.

[0541] In some embodiments, the modulation is affected by a modification in the target gene VEGFA, PCSK9, ANGPTL3, PTBP1, TTR, Ube3a-ATS, Ptp1b, APOC3, hsd17b13, bcl11a, or TGF. VEGFA plays an important role in angiogenesis and is overexpressed in several cancers. PCSK9 plays a key role in cholesterol management. ANGPTL3 plays a role in the metabolism of triglycerides. PTBP1 encodes a nuclear protein of the nucleus and is involved in the regulation of transcription. TTR encodes the transport protein transthyretin. Ube3a-ATS encodes a ubiquitin ligase, and mutations in Ube3a-ATS are associated with Angelman syndrome. Ptp1b encodes protein tyrosine phosphatase 1B, which regulates insulin signaling. APOC3 encodes apolipoprotein C3, which plays a key role in lipid transport and metabolism. hsd17b13 encodes liver-specific hydroxysteroid 17B-dehydrogenase 13. BCL11A is a transcriptional repressor. TGF regulates several cellular pathways, importantly those associated with proliferation and differentiation.

[0542] The terms “subject” and “patient” are used interchangeably herein. In some embodiments, the subject is a human. Examples

[0543] Design and construction of plasmids comprising the constructs of the disclosure

[0544] Epigenetic modification editors with HA epitope, P2A and BFP (e.g., including Dnmt3a CD, Dnmt3l CD, dSpCas9, KRAB and / or other histone modifiers) were optimized into nucleic acid sequences suitable for mammalian expression and synthesis by Genscript, and then cloned into pLV-CAG vector with CAG promoter and WPRE, which expresses the complete epigenetic modification editor and self-cleaved BFP.

[0545] In optimizing different functional elements and the order of different functional elements, different functional elements were optimized into nucleic acid sequences suitable for mammalian expression and synthesis by Genscript. First, the vector except the element to be replaced was amplified by PCR, then the element to be replaced was amplified from the sequence synthesized by the company while introducing the homologous arm sequence, and finally the different elements were recombined into the vector by NEBuilder reagent to construct the final expression plasmid.

[0546] Example 1: Construct for epigenetic modification of CD151 gene and CD81 gene

[0547] Different groups of constructs targeting CD151 gene and / or CD81 were generated according to the above method, the structures of which are described in FIG. IB 、 FIG. 2B 、 FIG. 4B and FIG. 5B . The constructs shown in FIG. 2E were generated as follows: (1) Dnmt-dCas9+scFv-Krab, generated by fusing Dnmt3A-Dnmt3L-dCas9-10xGCN4 (SEQ ID NO: 151 and 179), T2A (SEQ ID NO: 99 and 103) and scFv-Krab (SEQ ID NO: 152 and 180); (2) dCas9+scFv-Dnmt-Krab, generated by fusing dCas9-10xGCN4 (SEQ ID NO: 153 and 181), T2A (SEQ ID NO: 99 and 103) and scFv-Dnmt3A-Dnmt3L-Krab (SEQ ID NO: 154 and 182); (3) dCas9+scFv-Dnmt+scFv-Krab, generated by fusing dCas9-10xGCN4 (SEQ ID NO: 153 and 181), T2A (SEQ ID NO: 99 and 103), scFv-Dnmt3A-Dnmt3L (SEQ ID NO: 200 and 201) and scFv-Krab (SEQ ID NO: 202 and 203).

[0548] The repression of CD151 and / or CD81 by the constructs using different targeting sgRNAs (Table 4) was measured by FACS. The results are shown in FIG. 1C , FIGS. 2C-2E , FIGS. 4C-4F and FIGS. 5C-5D . The editors with different effectors or conformations varied greatly in their ability to suppress gene expression, strongly suggesting that various modifications must be made in optimizing epigenetic editors.

[0549] Example 2: EPICAS constructs for VEGFA knockdown

[0550] Three versions of EPICAS specific for VEGFA (EPICAS-V1, EPICAS-V2 and EPICAS-V3) were generated. The structures of these constructs are shown in FIG. 6A EPICAS-V1 contains a KRAB domain, while V2 and V3 contain a ZIM3-KRAB domain, and the orientation of DNMT3L and DNMT3A is reversed in V2 (3A-3L, rather than 3L-3A in V1 and V3).

[0551] The editing efficiency of the constructs using two different sgRNAs targeting human VEGFA (Table 4) on VEGFA was measured by qPCR. The results are shown in FIG. 6B The reduction in mRNA expression by sgRNA1 was comparable between the various constructs.

[0552] Example 3: Constructs for in vivo epigenetic modification of mouse PCSK9

[0553] To test editing efficiency in vivo experiments, the following four forms of constructs were generated: (1) DNMT3A-DNMT3L-dCas9-KRAB (SEQ ID NO: 138 and 166); (2) DNMT3A-DNMT3L-dCas9-EZH2 (SEQ ID NO: 139 and 167); (3) DNMT3A-DNMT3L-dCas9-G9A (SEQ ID NO: 140 and 168); (4) DNMT3A-DNMT3L-dCas9-HDAC3 (SEQ ID NO: 144 and 172); (5) dCas9N-7-Dnmt3A-Dnmt3L-dCas9C-7-Krab (this construct was generated by fusing dCas9N-7 (SEQ ID NO: 14), Dnmt3A-Dnmt3L (SEQ ID NO: 71), dCas9C-7 (SEQ ID NO: 15), and Krab (SEQ ID NO: 51)).

[0554] DNA sequences corresponding to each form were transcribed in vitro to obtain mRNA, which was then mixed with two gRNAs targeting PCSK9 (SEQ ID NO: 198-199) at a ratio of 2:1:1 to prepare LNP (molar ratio of LNP elements: MC3: cholesterol: DSPC: DMG-PEG2000 = 50:38.5:10:1.5) which was injected into wild-type C57 mice by tail vein injection. The injection dose was 4.5 mg (RNA mass) per kilogram of body weight per mouse. Blood samples were collected from mice on days 7, 14, and 21 after injection, and the expression of PCSK9 protein in blood was determined by Elisa, comparing mice that were not injected with LNP with mice that were injected with mRNA + control gRNA (sequence GAAGAGCCTGAGGCTCTTCT).

[0555] The results showed that the protein expression of PCSK9 in the serum of the experimental group mice was significantly lower than that of the control group at all three time points, with a decrease of more than 50%.

Claims

1. A construct of Formula III: 5' - (A-B) -CasN-CasC-T-K - 3' (III) wherein: one of A and B is a polynucleotide encoding DNMT3A; the other of A and B is a polynucleotide encoding DNMT3L; CasN is a polynucleotide encoding an N-terminal portion of dCas9; CasC is a polynucleotide encoding a C-terminal portion of dCas9; K is a polynucleotide encoding KRAB; D is a polynucleotide encoding a gene expression modulator; T comprises (a) a polynucleotide encoding an epitope capable of binding to an antibody or antigen-binding fragment thereof, and further comprises (b) a polynucleotide encoding a self-cleaving peptide 3' to the polynucleotide in (a), and further comprises (c) a polynucleotide encoding an antibody or antigen-binding fragment thereof capable of binding to the epitope 3' to the nucleotide encoding the self-cleaving peptide in (b); each of ml, m2, m3, m4, m5, and m6 is independently 1; n2 is 0; nl is 1 and n3 is 0, or nl is 0 and n3 is 1; p is 1; q is 0; and r is 1. 5'-(A m1 -B m2 ) n1 -CasN-(A m3 -B m4 ) n2 -CasC-T p -E-3' (III), 2. The construct of claim 1, having Formula IIIb-1: 5' - (A-B) -CasN-CasC-T-K - 3' (IIIb-1).

3. The construct of claim 1, having Formula IIIc-1: 5' -CasN-CasC-T-(A-B)-K - 3' (IIIc-1).

4. The construct of claim 1, wherein the self-cleaving peptide is T2A.

5. The construct of claim 1, wherein the DNMT3A is an amino acid sequence of SEQ ID NO:

69. E is 5'-(A m5 -B m6 ) n3 -K r -D q -3' 6. The construct of claim 5, wherein the polynucleotide encoding DNMT3A is a nucleic acid sequence of SEQ ID NO:

83.

7. The construct of claim 1, wherein the DNMT3L is an amino acid sequence of any one of SEQ ID NOs: 74-82.

8. The construct of claim 7, wherein the polynucleotide encoding DNMT3L is a nucleic acid sequence of any one of SEQ ID NOs: 84-92.

9. The construct of claim 1, wherein the KRAB is an amino acid sequence of SEQ ID NO: 51, 53, or 230-241.

10. The construct of claim 9, wherein the polynucleotide encoding KRAB is a nucleic acid sequence of SEQ ID NO: 52, 54, 206, 208, 210, 212, 214, 216, 218, 220, 222, 224, 226, or 228.

11. The construct of claim 1, wherein the dCas9 comprises a Staphylococcus aureus dCas9.

12. The construct of claim 11, wherein the dCas9 is an amino acid sequence of SEQ ID NO:

1. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 13. The construct of claim 11, wherein the polynucleotide encoding the dCas9 is the nucleic acid sequence of SEQ ID NO:

26.

14. The construct of claim 1, wherein the epitope capable of binding to an antibody or antigen binding fragment thereof is selected from the group consisting of GCN4, 10xGCN4.

15. The construct of claim 14, wherein the GCN4 is the amino acid sequence of SEQ ID NO:

97.

16. The construct of claim 15, wherein the polynucleotide encoding the GCN4 is the nucleic acid sequence of SEQ ID NO:

101.

17. The construct of claim 16, wherein the antibody or antigen binding fragment is a single domain antibody, scFv, Fab, VH, VHH, or antibody mimetic.

18. The construct of claim 17, wherein the antigen binding fragment is a scFv.

19. The construct of claim 1, which is the nucleic acid sequence of any one of SEQ ID NOs: 162-189.

20. A polypeptide expressed from the construct of any one of claims 1-19.

21. A vector comprising the construct of any one of claims 1-19.

22. The vector of claim 21, further comprising a polynucleotide encoding a single guide RNA (sgRNA).

23. A cell comprising the construct of any one of claims 1-19.

24. A cell comprising the polypeptide of claim 20.

25. A cell comprising the vector of any one of claims 21-22.

26. The cell of any one of claims 23-25, further comprising at least one sgRNA.

27. A composition comprising the construct of any one of claims 1-19.

28. A composition comprising the polypeptide of claim 20.

29. A composition comprising the vector of any one of claims 21-22.

30. The composition of any one of claims 27-29, further comprising at least one sgRNA.

31. The composition of claim 30, further comprising a pharmaceutically acceptable carrier.

Citation Information

Patent Citations

  • Crispr-CAS systems and methods for altering expression of gene products

    EP2764103A2

  • Engineering of systems, methods and optimized guide compositions for sequence manipulation

    EP2771468A1

  • Engineering of systems, methods and optimized guide compositions for sequence manipulation

    EP2771468B1

  • Engineering of systems, methods and optimized guide compositions for sequence manipulation

    EP2784162A1

  • Engineering of systems, methods and optimized guide compositions for sequence manipulation

    EP2784162B1