Compositions and methods for epigenetic regulation of f8 expression
Epigenetic editors with DNA-binding domains are used to specifically activate F8 gene expression, addressing the inefficiencies and risks of traditional genetic engineering methods and achieving stable, increased F8 protein production.
Patent Information
- Application Number
- PCT/US2024/055689
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Current genetic engineering strategies for increasing Factor VIII (F8) expression in humans are risky and inefficient, often leading to chromosomal translocations, insertions, deletions, and off-target mutations.
Development of epigenetic editors that use fusion proteins with DNA-binding domains, such as dead Cas9, ZFP, or TALE domains, to specifically activate F8 gene expression without making DNA breaks, thereby avoiding the risks associated with traditional genetic engineering.
The epigenetic editors achieve stable and increased expression of functional F8 protein, reducing the risk of chromosomal alterations and ensuring durable, inheritable activation of the F8 gene.
Smart Images

Figure IMGF000123_0001 
Figure IMGF000124_0001 
Figure IMGF000130_0001
Abstract
Description
[0001] COMPOSITIONS AND METHODS FOR EPIGENETIC REGULATION OF F8 EXPRESSION
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit under 35 .U.S.C. § 119(e) of U.S. Provisional Application no. 63 / 598,931, filed November 14, 2023, entitled “COMPOSITIONS AND METHODS FOR EPIGENETIC REGULATION OF F8 EXPRESSION,” the disclosure of which is hereby incorporated by reference in its entirety.
[0004] REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
[0005] The contents of the electronic sequence listing (C169870048WO00-SEQ-AXW.xml; Size: 1,213,317 bytes; and Date of Creation: November 13, 2024) is herein incorporated by reference in its entirety.
[0006] BACKGROUND
[0007] Factor VIII (FVIII), also known as anti-hemophilic factor (AHF), is a protein essential in clotting. In humans, factor VIII is encoded by the F8 gene. Defects in the expression of F8 gene can result in hemophilia A. A subset of hemophilia A cases are caused by insufficient production from a healthy F8 allele. Methods to activate expression of functional F8 are needed in the field. However, traditional genetic engineering strategies typically rely on permanent manipulation of cells at the genomic level, which is associated with certain risks, including, for example, chromosomal translocations, undesired insertions and deletions of nucleotides at the targeted site, and off-target mutations. There remains a need for efficient and safe methods of genetically engineering increased expression of healthy factor VIII protein.
[0008] SUMMARY
[0009] The present disclosure provides systems and compositions for epigenetic modification (“epigenetic editors” or “epigenetic editing systems” herein), and methods of using the same to generate epigenetic modification at F8, including in host cells and organisms.
[0010] In some aspects, the present disclosure provides a system for activating transcription of a human F8 gene in a human cell, comprising a) one or more fusion proteins that collectively comprise a transcriptional activating domain, or an effector domain that catalyzes the removal of DNA methylation, each domain being linked to a DNA-binding domain that binds to a target region in the human F8 promoter or downstream promoter sequence, or b) one or more nucleic acid molecules encoding the one or more fusion proteins, wherein the system does not generate a DNA break in the F8 gene.
[0011] In some embodiments, the DNA- binding domain comprises a dead CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain. In some embodiments, the DNA-binding domain comprises a dCas9 domain and the system further comprises (i) one or more guide RNAs comprising a sequence that binds to any one of the sequences in Table 1 or Table 2, or (ii) nucleic acid molecules coding for the one or more guide RNAs.
[0012] In some aspects, the present disclosure provides a human cell comprising a system of the present disclosure, or progeny of said cell. In some embodiments, the human cell is modified by a system of the present disclosure, or is the progeny of said cell modified by a system of the present disclosure.
[0013] In some aspects, the present disclosure provides a pharmaceutical composition comprising a system of the present disclosure. In some embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable excipient. In some embodiments, the pharmaceutical composition comprises human cells disclosed herein and, optionally, further comprises a pharmaceutically acceptable excipient.
[0014] In some aspects, the present disclosure provides a method of treating a subject in need thereof, comprising administering a system of the present disclosure, human cells of the present disclosure, or a pharmaceutical composition of the present disclosure to the subject. In some embodiments, the subject has or is suspected of having hemophilia.
[0015] In some aspects, the present disclosure provides an epigenetically modified human cell comprising an F8 genomic locus, wherein the epigenetically modified human cell is of a cell type that, in its naturally-occurring (wildtype) state, is characterized by expression of F8, wherein the F8 genomic locus comprises two alleles, and wherein both alleles of the F8 genomic locus comprise a naturally occurring genomic sequence associated with expression of functional F8, and wherein the epigenetically modified human cell is characterized by active or increased expression (as compared to a wildtype cell of the same cell type) of F8.
[0016] DETAILED DESCRIPTION
[0017] The present disclosure provides epigenetic editors for activating expression of the human F8 gene. By altering expression of F8, the editors herein may be used to address indications such as hemophilia A.
[0018] Some aspects of the present disclosure are based on the recognition F8 expression can be activated or increased in cells. Some aspects of the present disclosure are based on the recognition that cells modified according to the present disclosure to express F8. or to express F8 at a substantially higher level as compared to wildtype cells of the same cell type. Some aspects of the present disclosure are based on the recognition that cells modified according to the present disclosure to express F8, or to express F8 at a substantially higher level, may have enhanced functionality.
[0019] In some embodiments, the present disclosure provides compositions, strategies, methods, and modalities for the generation of cells with increased F8 expression. In some embodiments, the present disclosure provides compositions, strategies, methods, and modalities for the generation of cells with increased FSexpression via epigenetic activation of this gene. The presently provided compositions, strategies, methods, and modalities for epigenetic editing of F8, e.g., epigenetic editors targeting F8 and their use, have several advantages compared to other genome engineering methods, including reversibility, decreased risk of chromosomal translocation, and durable, inheritable activation.
[0020] Unless otherwise stated, “F8” refers herein to a human F8 gene. The human F8 gene is well known to those of skill in the art. An exemplary human F8 gene sequence can be found at Ensembl Accession No. ENSG00000185010. The targeted regions of F8 include the area around the transcription start site (TSS) chrX: 155026487-155027450, based on the GRCh38 Genome Reference Consortium Human Build 38, and the downstream promoter region, chrX:155021611-155022971 based on the GRCh38 Genome Reference Consortium Human Build 38. Targeted sequences are below.
[0021] Factor 8 (F8): promoter (TSS) sequence:
[0022] ATTAAATACATAAATAAAAtgtgtcttcctccactaaaagagtaggacacatcgcagaagcc gataaatatttgtt caatgTTTACATGAACAAACACGGTCACTCAGCCGTCTTTATTTTCAG TGAGTCATAGGTAAAATATTCTTCACCAGGCGAGGATTGAGTCGCAGTCGCGTCAGGAACTC GCCTCCAGGTACTGAGCGAGCGAGCCCCGACGGCCGCGGTGGGTGGAGGGGGGAAAGtgggg gagagtggggcgggagtagagggagagcgagggagggagCCTCACCCCCGGGACGCTTTCTG CGCACGCGCATCTCGGTGCATCTGTTGGGCGGGACCGCGGGGCCTGTGACACCGCACGCTGA GCTCTGTGATGTAGCCGCTTGCGGAGACTGCAAGCAGCCGCGGCGCGCCCGGCCCTCCCTCT TCCGCTGCCGCCGTGGGAATGGAAACATCTGCCCCACGTGCCGGAAGCCAAGTGGTGGCGAC AACTGCGCGCCACTCCGCGGCCTACCGCGCAGATCCTCTACGTGTGTCCTCGCGAGACAAGC TCACCGAAATGGCCGCGTCCAGTCAAGGTAAGGGCGGGCGCGCGCGCGGCCGGCGCCGCGGG GAGCCGCCTTCCTGGGCTCCGGAAGGGGAAGGGCTCCGCGCCCGCCAGTCGCGACCGAGCTC CCCGGGGAGTCCAAGCGGCGCGGGGACGGGGAGGGCGCTCGGGGATCGCCCCTGGCTCAGCC GCTCCCAGGGTCCGGCCAAGGTCACGCGCGTGTCTGGGGCCTAACCACCCTGGGGTGTCGTC TCCAGTGGAGCTCACCCAGCCTTCCGGCGCCGCCCCCGGCGCAGAAGCCCGAAGACATTCTG TCCCCGTTCTCGGAAGCGGTCTGGCTTCTTTGTCCCTTTGGACGTTCATCTCAAATGCCTCG TGGCGTCATTCTGTGTACAAGGCACCTCCCCAAC ( SEQ ID NO : 650 ) Factor 8 (F8): downstream promoter sequence caataaaatcaagaaaactcctggaaatattatttttaaaagagaggaagagaacatataca acataaggaatgaaaaagaaaatataactacagacacagaacagtttaacaaaatcacagga gactcttacataacttgataccaatacattatcaaatctagggacagtagattgctttctat gaaagtatagaccaaaatataccaatatttactgaagaagaggtagaaaaatctgaatCAAC T G AT G AC C AT AG AC AAT AT T G AT AAAG TTATTTAATACTCTGCC T AAAT G C AC C AG AAC C C C CAAAACACGAAAAGCTTCTGAACTCAAGGCTCATATTATCCTTACACCCAAAGTAGGAAAAG ATATTATAAAAACACACATTGCATGTACACACAAGAGAACCAGACATGAAACTTGTGTGTGA TATTGATTGGATTTCATTGCCTTGATTCTCTCTATGGAAGCTTGGGCAGGCAGAGGGAAGTT TTGTCTTTGTCCCTTCACCCTCCCTCTACTCCAGTCTCTAGATTCTGGAGGAGCAGCATATG AC G AAG AG AAG C AAT G G AC AAAAC AG AAC AAAAT AT C AAC AAAT AG GTCAGACCTAT G AAAA GAATTGGCAGTTAGCATTTCTGCTAATAACACAGATGTGTGCACACCTTACCCAGAAATGTT TCTTTGGGGCCCAGGTAGCATCACAACCATCCTAACCCGATGTCTGCACCTTCCTCCAAGCA GACTTACATCCCCACAATCCTGGCCCCGATCAGACCCTACAGGACATGCCTTTACCTTGCGT CCACAGGCAGCTCACCGAGATCACTTTGCATATAGTCCCATGACAGTTCCACTGCACCCAGG TAGTATCTTCTGGTGGCACTAAAGCAGAATCGCAAAAGGCACAGAAAGAAGCAGGTGGAGAG CTCTATTTGCATGACTTATTGCTACAAATGTTCAACTGGAGAAGCAAAAGGTTAATTCTTCT CTAAAATATCTTTAGCTCCCAGGAGGGGAAAAAAGTAAAATTTCTGATTTAATGAAAAGTCC CAATTTCTTTGCAGAGCATTTTAAGGAACTTTACCCACTGGATGTGCTCAGCACTAAGCAGT AACCGATAGGATTGCTTCCTTTTTATCAGTGGGAAGCAGCCACAGGAAGAACTGAAGTAGCA AAAGGGAGGCTGCTAAACTTAACCCATTCACTTCTGAGAGCAGTAGGGGCATAAGTCTGCTT TAATCAGAACCTTTAAGGAAAGGAGGAGAGGGATGGAGAAGTCAAAGTGAACAGGAGCTTTA ATTTGTGCTCCTATTCCTAGCCTACTCCAACCCTCTTCCTGAGTGGCAGCAGCAAGAGA
[0023] ( SEQ ID NO : 651_
[0024] In some embodiments, an epigenetic editor as described herein may comprise one or more fusion proteins, wherein each fusion protein comprises a DNA-binding domain linked to one or more effector domains for epigenetic modification. In certain embodiments, where the DNA-binding domain is a polynucleotide guided DNA-binding domain, the epigenetic editor may further comprise one or more guide polynucleotides. DNA-binding domains, effector domains, and guide polynucleotides of an epigenetic editor as described herein may be selected, e.g., from those described below, in any functional combination.
[0025] The epigenetic editors described herein may be expressed in a host cell transiently, or may be integrated in a genome of the host cell; such cells and their progeny are also contemplated by the present disclosure. Both transiently expressed and integrated epigenetic editors or components thereof can effect stable epigenetic modifications. For example, after introducing to a host cell an epigenetic editor described herein, the target gene in the host cell may be stably or permanently increased or activated. In some embodiments, expression of the target gene is activated or increased for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or for the entire lifetime of the cell or the subject carrying the cell, as compared to the level of expression in the absence of the epigenetic editor. The epigenetic modification may be inherited by the progeny of the host cells into which the epigenetic editor was introduced.
[0026] The present disclosure provides epigenetic editors for regulating expression of gene targets in a cell or organism. In particular, epigenetic editor compositions and methods for activating expression of the F8 gene are provided.
[0027] I. DNA-Binding Domains
[0028] An epigenetic editor described herein may comprise one or more DNA-binding domains that direct the effector domain(s) of the epigenetic editor to target sequences within or close to a target gene locus. A DNA-binding domain as described herein may be, e.g., a polynucleotide guided DNA-binding domain, a zinc finger protein (ZFP) domain, a transcription activator like effector (TALE) domain, a meganuclease DNA-binding domain, and the like. Examples of DNA-binding domains can be found in U.S. Patent 11,162,114, which is incorporated by refence herein in its entirety.
[0029] In some embodiments, a DNA-binding domain described herein is encoded by its native coding sequence. In other embodiments, the DNA-binding domain is encoded by a nucleotide sequence that has been codon-optimized for optimal expression in human cells.
[0030] A. Polynucleotide Guided DNA-Binding Domains
[0031] In some embodiments, a DNA-binding domain herein may be a protein domain directed by a guide nucleic acid sequence (e.g., a guide RNA sequence) to a target site in a target gene locus. In certain embodiments, the protein domain may be derived from a CRISPR-associated nuclease, such as a Class I or II CRIS PR-associated nuclease. In some embodiments, the protein domain may be derived from a Cas nuclease such as a Type II, Type IIA, Type IIB, Type IIC, Type V, or Type VI Cas nuclease. In certain embodiments, the protein domain may be derived from a Class II Cas nuclease selected from Casl, Cas IB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas 10, Cas 14a, Cas 14b, Cas 14c, CasX, CasY, CasPhi, C2c4, C2c8, C2c9, C2cl0, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, CsxlS, Csfl, Csf2, CsO, Csf4, and homologues and modified versions thereof. “Derived from” is used to mean that the protein domain comprises the full polypeptide sequence of the parent protein, or comprises a variant thereof (e.g., with amino acid residue deletions, insertions, and / or substitutions). The variant retains the desired function of the parent protein (e.g., the ability to form a complex with the guide nucleic acid sequence and the target DNA). In some embodiments, the CRISPR-associated protein domain may be a Cas9 domain described herein. Cas9 may, for example, refer to a polypeptide with at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to a wildtype Cas9 polypeptide described herein. In some embodiments, said wildtype polypeptide is Cas9 from Streptococcus pyogenes (NCBI Ref. No. NC_002737.2 (SEQ ID NO: 1)) and / or UniProt Ref. No. Q99ZW2 (SEQ ID NO: 2). In some embodiments, said wildtype polypeptide is Cas9 from Staphylococcus aureus (SEQ ID NOs: 3 and 4). In some embodiments, said wildtype polypeptide is Cas9 from Neisseria meningitidis (SEQ ID NOs: 5 and 6). In some embodiments, the CRISPR-associated protein domain is a Cpfl domain or protein, or a polypeptide with at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to a wildtype Cpfl polypeptide described herein (e.g., Cpfl from Francisella novicida (UniProt Ref. No. U2UMQ6 or SEQ ID NO: 7). In certain embodiments, the CRISPR-associated protein domain may be a modified form of the wildtype protein comprising one or more amino acid residue changes such as a deletion, an insertion, or a substitution; a fusion or chimera; or any combination thereof.
[0032] Cas9 sequences and structures of variant Cas9 orthologs have been described for various organisms. Exemplary organisms from which a Cas9 domain herein can be derived include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Listeria innocua, Lactobacillus gasseri, Francisella novicida, Wolinella succinogenes, Sutterella wadsworthensis, Gamma proteobacterium, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Fibrobacter succinogene, Rhodospirillum rubrum, Nocardiopsis dassonvillei, Streptomyces pristinaespiralis, Streptomyces viridochromo genes, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Synechococcus sp. , Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionium, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena varied? ills, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Streptococcus pasteurianus, Neisseria cinerea, Campylobacter lari, Parvibaculum lavamentivorans, Corynebacterium diphtheria, and Acaryochloris marina. Cas9 sequences also include those from the organisms and loci disclosed in Chylinski et al., RNA Biol. (2013) 10(5):726-37.
[0033] In some embodiments, the Cas9 domain is from Streptococcus pyogenes (SpCas9). In some embodiments, the Cas9 domain is from Staphylococcus aureus (SaCas9). In some embodiments, the Cas9 domain is from Neisseria meningitidis (Nme2Cas9).
[0034] Other Cas domains are also contemplated for use in the epigenetic editors herein. These include, for example, those from CasX (Casl2E) (e.g., SEQ ID NO: 8), CasY (Casl2d) (e.g., SEQ ID NO: 9), Cascp (CasPhi) (e.g., SEQ ID NO: 10), Casl2fl (Casl4a) (e.g., SEQ ID NO: 11), Casl2f2 (Casl4b) (e.g., SEQ ID NO: 12), Casl2f3 (Casl4c) (e.g., SEQ ID NO: 13), and C2c8 (e.g., SEQ ID NO: 14).
[0035] For epigenetic editing, the nuclease-derived protein domain (e.g., a Cas9 or Cpfl domain) may have reduced or no nuclease activity through mutations such that the protein domain does not cleave DNA or has reduced DNA-cleaving activity while retaining the ability to complex with the guide nucleic acid sequence (e.g., guide RNA) and the target DNA. For example, the nuclease activity may be reduced by at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% compared to the wildtype domain. In some embodiments, a CRIS PR-associated protein domain described herein is catalytically inactive (“dead”). Examples of such domains include, for example, dCas9 (“dead” Cas9), dCpfl, ddCpfl, dCasPhi, ddCasl2a, dLbCpfl, and dFnCpfl. A dCas9 protein domain, for example, may comprise one, two, or more mutations as compared to wildtype Cas9 that abrogate its nuclease activity. The DNA cleavage domain of Cas9 is known to include two subdomains: the HNH nuclease subdomain and the RuvCl subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, whereas the RuvCl subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, the mutations D10A (in RuvCl) and H840A (in HNH) completely inactivate the nuclease activity of SpCas9. SaCas9, similarly, may be inactivated by the mutations D10A and N580A. In some embodiments, the dCas9 comprises at least one mutation in the HNH subdomain and / or the RuvCl subdomain that reduces or abrogates nuclease activity. In some embodiments, the dCas9 only comprises a RuvCl subdomain, or only comprises an HNH subdomain. It is to be understood that any mutation that inactivates the RuvCl and / or the HNH domain may be included in a dCas9 herein, e.g., insertion, deletion, or single or multiple amino acid substitution in the RuvCl domain and / or the HNH domain.
[0036] In some embodiments, a dCas9 protein herein comprises a mutation at position(s) corresponding to position DIO (e.g., D10A), H840 (e.g., H840A), or both, of a wildtype SpCas9 sequence as numbered in the sequence provided at UniProt Accession No. Q99ZW2 (SEQ ID NO: 2). In particular embodiments, the dCas9 comprises the amino acid sequence of dSpCas9 (D10A and H840A) (SEQ ID NO: 15).
[0037] In some embodiments, a dCas9 protein as described herein comprises a mutation at position(s) corresponding to position DIO (e.g., D10A), N580 (e.g., N580A), or both, of a wildtype SaCas9 sequence (e.g., SEQ ID NO: 3). In particular embodiments, the dCas9 comprises the amino acid sequence of dSaCas9 (D10A and N580A) (SEQ ID NO.: 16). Additional suitable mutations that inactivate Cas9 will be apparent to those of skill in the art based on this disclosure and knowledge in the field and are within the scope of this disclosure. Such mutations may include, but are not limited to, D839A, N863A, and / or K603R in SpCas9. The present disclosure contemplates any mutations that reduce or abrogate the nuclease activity of any Cas9 described herein (e.g., mutations corresponding to any of the Cas9 mutations described herein).
[0038] A dCpfl protein domain may comprise one, two, or more mutations as compared to wildtype Cpfl that reduce or abrogate its nuclease activity. The Cpfl protein has a RuvC- like endonuclease domain that is similar to the RuvC domain of Cas9, but does not have an HNH endonuclease domain, and the N-terminal of Cpfl does not have the alpha-helical recognition lobe of Cas9. In some embodiments, the dCpfl comprises one or more mutations corresponding to position D917A, E1006A, or D1255A as numbered in the sequence of the Francisella novicida Cpfl protein (FnCpfl; SEQ ID NO: 7). In certain embodiments, the dCpfl protein comprises mutations corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A, or corresponding mutation(s) in any of the Cpfl amino acid sequences described herein. In some embodiments, the dCpfl comprises a D917A mutation. In particular embodiments, the dCpfl comprises the amino acid sequence of dFnCpfl (SEQ ID NO: 17).
[0039] Further nuclease inactive CRISPR-associated protein domains contemplated herein include those from, for example, dNmeCas9 (e.g., SEQ ID NO: 18), dCjCas9 (e.g., SEQ ID NO: 19), dStlCas9 (e.g., SEQ ID NO: 20), dSt3Cas9 (e.g., SEQ ID NO: 21), dLbCpfl (e.g., SEQ ID NO: 22), dAsCpfl (e.g., SEQ ID NO: 23), denAsCpfl (e.g., SEQ ID NO: 24), dHFAsCpfl (e.g., SEQ ID NO: 25), dRVRAsCpfl (e.g., SEQ ID NO: 26), dRRAsCpfl (e.g., SEQ ID NO: 27), dCasX (e.g., SEQ ID NO: 28), and dCasPhi (e.g., SEQ ID NO: 29).
[0040] In some embodiments, a Cas9 domain described herein may be a high fidelity Cas9 domain, e.g., comprising one or more mutations that decrease electrostatic interactions between the Cas9 domain and the sugar-phosphate backbone of DNA to confer increased target binding specificity. In certain embodiments, the high fidelity Cas9 domain may be nuclease inactive as described herein.
[0041] A CRIS PR-associated protein domain described herein may recognize a protospacer adjacent motif (PAM) sequence in a target gene. A “PAM” sequence is typically a 2 to 6 bp DNA sequence immediately following the sequence targeted by the CRISPR-associated protein domain. The PAM sequence is required for CRISPR protein binding and cleavage but is not part of the target sequence. The CRISPR-associated protein domain may either recognize a naturally occurring or canonical PAM sequence or may have altered PAM specificity. CRISPR-associated protein domains that bind to non-canonical PAM sequences have been described in the art. For example, Cas9 domains that bind non-canonical PAM sequences have been described in Kleinstiver et al., Nature (2015) 523(7561):481-5 and Kleinstiver et al., Nat Biotechnol. (2015) 33:1293-8. Such Cas9 domains may include, for example, those from “VRER” SpCas9, “EQR” SpCas9, “VQR” SpCas9, “SpG Cas9,” “SpRYCas9,” and “KKH” SaCas9. Nuclease inactive versions of these Cas9 domains are also contemplated, such as nuclease inactive VRER SpCas9 (e.g., SEQ ID NO: 30), nuclease inactive EQR SpCas9 (e.g., SEQ ID NO: 31), nuclease inactive VQR SpCas9 (e.g., SEQ ID NO: 32), nuclease inactive SpG Cas9 (e.g., SEQ ID NO: 33), nuclease inactive SpRY Cas9 (e.g., SEQ ID NO: 34), and nuclease inactive KKH SaCas9 (e.g., SEQ ID NO: 35). Another example is the Cas9 of Francisella novicida engineered to recognize 5’-YG-3’ (where “Y” is a pyrimidine).
[0042] Additional suitable CRISPR-associated proteins, orthologs, and variants, including nuclease inactive variants and sequences, will be apparent to those of skill in the art based on this disclosure.
[0043] Guide RNAs that can be used in conjunction with the CRISPR-associated protein domains herein are further described in Section II below.
[0044] B. Zinc Finger Protein Domains
[0045] In some embodiments, the DNA-binding domain of an epigenetic editor described herein comprises a zinc finger protein (ZFP) domain (or “ZF domain” as used herein). ZFPs are proteins having at least one zinc finger, and bind to DNA in a sequence-specific manner. A “zinc finger” (ZF) or “zinc finger motif’ (ZF motif) refers to a polypeptide domain comprising a beta-beta-alpha (PPa)-protein fold stabilized by a zinc ion. A ZF binds from two to four base pairs of nucleotides, typically three or four base pairs (contiguous or noncontiguous). Each ZF typically comprises approximately 30 amino acids. ZFP domains may contain multiple ZFs that make tandem contacts with their target nucleic acid sequence. A tandem array of ZFs may be engineered to generate artificial ZFPs that bind desired nucleic acid targets. ZFPs may be rationally designed by using databases comprising triplet (or quadruplet) nucleotide sequences and individual ZF amino acid sequences, in which each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of ZFs that bind the particular triplet or quadruplet sequence. See, e.g., U.S. Patents 6,453,242, 6,534,261, and 8,772,453.
[0046] ZFPs are widespread in eukaryotic cells, and may belong to, e.g., C2H2 class, CCHC class, PHD class, or RING class. An exemplary motif characterizing one class of these proteins (C2H2 class) is -Cys-(X)2-4-Cys-(X)i2-His-(X)3-5-His- (SEQ ID NO: 764), where X is any independently chosen amino acid. In some embodiments, a ZFP domain herein may comprise a ZF array comprising sequential C2H2-ZFs each contacting three or more sequential nucleotides.
[0047] A ZFP domain of an epigenetic editor described herein may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more ZFs. The ZFP domain may include an array of two-finger or three- finger units, e.g., 3, 4, 5, 6, 7, 8, 9 or 10 or more units, wherein each unit binds a subsite in the target sequence. In some embodiments, a ZFP domain comprising at least three ZFs recognizes a target DNA sequence of 9 or 10 nucleotides. In some embodiments, a ZFP domain comprising at least four ZFs recognizes a target DNA sequence of 12 to 14 nucleotides. In some embodiments, a ZFP domain comprising at least six ZFs (F1-F6) recognizes a target DNA sequence of 18 to 21 nucleotides.
[0048] In some embodiments, ZFs in a ZFP domain described herein are connected via peptide linkers. The peptide linkers may be, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more amino acids in length. In some embodiments, a linker comprises 5 or more amino acids. In some embodiments, a linker comprises 7-17 amino acids. The linker may be flexible or rigid.
[0049] In some embodiments a zinc finger array may have the sequence:
[0050] SRPGERPFQCRICMRNFSXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXHXXT H[linker]FQCRICMRNFSXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXHXXTH [linker]PFQCRICMRNFSXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXHXXTH LRGS (SEQ ID NO: 761), or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, where “XXXXXXX” represents the amino acids of the ZF recognition helix, which confers DNA-binding specificity upon the zinc finger; each X may be independently chosen. In the above sequence, “XX” in italics may be TR, LR or LK, and “[linker]” represents a linker sequence. In some embodiments, the linker sequence is TGSQKP (SEQ ID NO: 762); this linker may be used when sub-sites targeted by the ZFs are adjacent. In some embodiments, the linker sequence is TGGGGSQKP (SEQ ID NO: 763); this linker may be used when there is a base between the sub- sites targeted by the zinc fingers. The two indicated linkers may be the same or different.
[0051] ZFP domains herein may contain arrays of two or more adjacent ZFs that are directly adjacent to one another (e.g., separated by a short (canonical) linker sequence), or are separated by longer, flexible or structured polypeptide sequences. In some embodiments, directly adjacent fingers bind to contiguous nucleic acid sequences, i.e., to adjacent trinucleotides / triplets. In some embodiments, adjacent fingers cross-bind between each other’s respective target triplets, which may help to strengthen or enhance the recognition of the target sequence, and leads to the binding of overlapping sequences. In some embodiments, distant ZFs within the ZFP domain may recognize (or bind to) non-contiguous nucleotide sequences.
[0052] The F1-F6 amino acid sequences of the ZF may be placed within the ZF framework sequence of SEQ ID NO: 761, or within any other ZF framework known in the art.
[0053] C. TALES
[0054] In some embodiments, the DNA-binding domain of an epigenetic editor described herein comprises a transcription activator-like effector (TALE) domain. The DNA-binding domain of a TALE comprises a highly conserved sequence of about 33-34 amino acids, with a repeat variable di-residue (RVD) at positions 12 and 13 that is central to the recognition of specific nucleotides. TALEs can be engineered to bind practically any desired DNA sequence. Methods for programming TALEs are known in the art. For example, such methods are described in Carroll et al., Genet Soc Amer. (2011) 188(4):773-82; Miller et al., Nat Biotechnol. (2007) 25(7):778-85; Christian et al., Genetics (2008) 186(2):757-61 ; Li et al., Nucl Acids Res. (2010) 39(l):359-72; and Moscou et al., Science (2009) 326(5959): 1501. D. Other DNA-Binding Domains
[0055] Other DNA-binding domains are contemplated for the epigenetic editors described herein. In some embodiments, the DNA-binding domain comprises an argonaute protein domain, e.g., from Natronobacterium gregoryi (NgAgo). NgAgo is a ssDNA-guided endonuclease that is guided to its target site by 5' phosphorylated ssDNA (gDNA), where it produces double-strand breaks. In contrast to Cas9, the NgAgo-gDNA system does not require a protospacer- adjacent motif (PAM). Thus, using a nuclease inactive NgAgo (dNgAgo) can greatly expand the bases that may be targeted. The characterization and use of NgAgo have been described, e.g., in Gao et al., Nat Biotechnol. (2016) 34(7):768-73; Swarts et al., Nature (2014) 507(7491):258-61; and Swarts et al., Nucl Acids Res. (2015) 43(10):5120-9.
[0056] In some embodiments, the DNA-binding domain comprises an inactivated nuclease, for example, an inactivated meganuclease. Additional non-limiting examples of DNA- binding domains include tetracycline-controlled repressor (tetR) DNA-binding domains, leucine zippers, helix-loop-helix (HLH) domains, helix-turn-helix domains, P-sheet motifs, steroid receptor motifs, bZIP domains homeodomains, and AT-hooks.
[0057] E. Configurations
[0058] A fusion protein of an epigenetic editor described herein may have its components structured in different configurations. For example, the DNA-binding domain may be at the C-terminus, the N-terminus, or in between two or more epigenetic effector domains or additional domains. In some embodiments, the DNA-binding domain is at the C-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is at the N-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is linked to one or more nuclear localization signals. In some embodiments, the DNA-binding domain is flanked by an epigenetic effector domain and / or an additional domain on both sides. In some embodiments, where “DBD” indicates DNA-binding domain and “ED” indicates effector domain, the epigenetic editor comprises the configuration of:
[0059] - N’]-[ED1]-[DBD]-[ED2]-[C’
[0060] - N’]-[ED1]-[DBD]-[ED2]-[ED3]-[C’
[0061] - N’]-[ED1]-[ED2]-[DBD]-[ED3]-[C’ or
[0062] - N’]-[ED1]-[ED2]-DBD]-[ED3]-[ED4]-[C’. In particular embodiments, an epigenetic editor described herein comprises SEQ ID NO: 765 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[0063] In particular embodiments, an epigenetic editor described herein comprises SEQ ID NO: 766, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[0064] II. Guide Polynucleotides
[0065] Epigenetic editors described herein that comprise a polynucleotide guided DNA- binding domain may also include a guide polynucleotide that is capable of forming a complex with the DNA-binding domain. The guide polynucleotide may comprise RNA, DNA, or a mixture of both. For example, where the polynucleotide guided DNA-binding domain is a CRISPR-associated protein domain, the guide polynucleotide may be a guide RNA (gRNA). A “guide RNA” or “gRNA” refers to a nucleic acid that is able to hybridize to a target sequence and direct binding of the CRISPR-Cas complex to the target sequence. Methods of using guide polynucleotide sequences with programmable DNA-binding proteins (e.g., CRISPR-associated protein domains) for site-specific DNA targeting (e.g., to modify a genome) are known in the art.
[0066] A guide polynucleotide sequence (e.g., a gRNA sequence) may comprise two parts: 1) a nucleotide sequence comprising a “targeting sequence” that is complementary to a target nucleic acid sequence (“target sequence”), e.g., to a nucleic acid sequence comprised in a genomic target site; and 2) a nucleotide sequence that binds a polynucleotide guided DNA- binding domain (e.g., a CRISPR-Cas protein domain). The nucleotide sequence in 1) may comprise a targeting sequence that is 100% complementary to a genomic nucleic acid sequence, e.g., a nucleic acid sequence comprised in a genomic target site, and thus may hybridize to the target nucleic acid sequence. The nucleotide sequence in 1) may be referred to as, e.g., a crispr RNA, or crRNA. The nucleotide sequence in 2) may be referred to as a scaffold sequence of a guide nucleic acid, e.g., a tracrRNA, or an activating region of a guide nucleic acid, and may comprise a stem-loop structure. Parts 1) and 2) as described above may be fused to form one single guide (e.g., a single guide RNA, or sgRNA), or may be on two separate nucleic acid molecules. In some embodiments, a guide polynucleotide comprises parts 1) and 2) connected by a linker. In some embodiments, a guide polynucleotide comprises parts 1) and 2) connected by a non-nucleic acid linker, for example, a peptide linker or a chemical linker. Part 2 (the scaffold sequence) of a guide polynucleotide as described herein may be, for example, as described in Jinek et al., Science (2012) 337:816-21; U.S. Patent Publication 2016 / 0208288; or U.S. Patent Publication 2016 / 0200779. Variants of part 2) are also contemplated by the present disclosure. For example, the tetraloop and stem loop of a gRNA scaffold (tracrRNA) sequence may be modified to include RNA aptamers, which can be bound by specific protein domains. In some embodiments, such modified gRNAs can be used to facilitate the recruitment of repressive or activating domains fused to the proteininteracting RNA aptamers.
[0067] A gRNA as provided herein typically comprises a targeting domain and a binding domain. The targeting domain (also termed “targeting sequence”) may comprise a nucleic acid sequence that binds to a target site, e.g., to a genomic nucleic acid molecule within a cell. The target site may be a double-stranded DNA sequence comprising a PAM sequence as well as the target sequence, which is located on the same strand as, and directly adjacent to, the PAM sequence. The targeting domain of the gRNA may comprise an RNA sequence that corresponds to the target sequence, i.e., it resembles the sequence of the target domain, sometimes with one or more mismatches, but typically comprising an RNA sequence instead of a DNA sequence. The targeting domain of the gRNA thus may base pair (in full or partial complementarity) with the sequence of the double- stranded target site that is complementary to the target sequence, and thus with the strand complementary to the strand that comprises the PAM sequence. It will be understood that the targeting domain of the gRNA typically does not include a sequence that resembles the PAM sequence. It will further be understood that the location of the PAM may be 5’ or 3’ of the target sequence, depending on the nuclease employed. For example, the PAM is typically 3’ of the target sequence for Cas9 nucleases, and 5’ of the target sequence for Casl2a nucleases. For an illustration of the location of the PAM and the mechanism of gRNA binding to a target site, see, e.g., Figure 1 of Vanegas et al., Fungal Biol Biotechnol. (2019) 6:6, which is incorporated by reference herein. For additional illustration and description of the mechanism of gRNA targeting of an RNA-guided nuclease to a target site, see Fu et al., Nat Biotechnol (2014) 32(3):279-84 and Sternberg et al., Nature (2014) 507(7490):62-7, each incorporated herein by reference.
[0068] In some embodiments, the targeting domain sequence comprises between 17 and 30 nucleotides and corresponds fully to the target sequence (i.e., without any mismatch nucleotides). In some embodiments, however, the targeting domain sequence may comprise one or more, but typically not more than 4, mismatches, e.g., 1, 2, 3, or 4 mismatches. As the targeting domain is part of gRNA, which is an RNA molecule, it will typically comprise ribonucleotides, while the DNA targeting domain will comprise deoxyribonucleotides. An exemplary illustration of a Cas9 target site, comprising a 22 nucleotide target domain, and an NGG PAM sequence, as well as of a gRNA comprising a targeting domain that fully corresponds to the target sequence (and thus base pairs with full complementarity with the DNA strand complementary to the strand comprising the target sequence and PAM) is provided below:
[0069] [ target domain ( DNA) ] [ PAM ]
[0070] 5 ' -N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-G-G-3 ' ( DNA) 3 ' -N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-C-C-5 ' ( DNA)
[0071] I I I I I I I I I I I I I I I I I I I I I I
[0072] 5 ' -N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N- [ gRNA scaffold] -3 ' ( RNA) [ target ing domain ( RNA) ] [ binding domain ]
[0073] An exemplary illustration of a Casl2a target site, comprising a 22 nucleotide target domain, and a TTN PAM sequence, as well as of a gRNA comprising a targeting domain that fully corresponds to the target sequence (and thus base pairs with full complementarity with the DNA strand complementary to the strand comprising the target sequence and PAM) is provided below:
[0074] [ PAM ] [ target domain ( DNA) ]
[0075] 5 ' -T-T-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-3 ' ( DNA) 3 ' -A-A-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-5 ' ( DNA)
[0076] I I I I I I I I I I I I I I I I I I I I I I
[0077] 5 ' - [ gRNA scaffold] -N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-N-3 ' ( RNA) [ binding domain ] [ target ing domain ( RNA) ]
[0078] While not wishing to be bound by theory, at least in some embodiments, it is believed that the length and complementarity of the targeting domain with the target sequence contributes to specificity of the interaction of the gRNA / Cas9 molecule complex with a target nucleic acid. In some embodiments, the targeting domain of a gRNA provided herein is 5 to 50 nucleotides in length. In some embodiments, the targeting domain is 15 to 25 nucleotides in length. In some embodiments, the targeting domain is 18 to 22 nucleotides in length. In some embodiments, the targeting domain is 19-21 nucleotides in length. In some embodiments, the targeting domain is 15 nucleotides in length. In some embodiments, the targeting domain is 16 nucleotides in length. In some embodiments, the targeting domain is 17 nucleotides in length. In some embodiments, the targeting domain is 18 nucleotides in length. In some embodiments, the targeting domain is 19 nucleotides in length. In some embodiments, the targeting domain is 20 nucleotides in length. In some embodiments, the targeting domain is 21 nucleotides in length. In some embodiments, the targeting domain is 22 nucleotides in length. In some embodiments, the targeting domain is 23 nucleotides in length. In some embodiments, the targeting domain is 24 nucleotides in length. In some embodiments, the targeting domain is 25 nucleotides in length. In certain embodiments, the targeting domain fully corresponds, without mismatch, to a target sequence provided herein, or a part thereof. In some embodiments, the targeting domain of a gRNA provided herein comprises 1 mismatch relative to a target sequence provided herein. In some embodiments, the targeting domain comprises 2 mismatches relative to the target sequence. In some embodiments, the target domain comprises 3 mismatches relative to the target sequence.
[0079] Methods for designing, selecting, and validating gRNAs are described herein and known in the art. Software tools can be used to optimize the gRNAs corresponding to a target DNA sequence, e.g., to minimize total off-target activity across the genome. For example, DNA sequence searching algorithms can be used to identify a target sequence in crRNAs of a gRNA for use with Cas9. Exemplary gRNA design tools include the ones described in Bae et al., Bioinformatics (2014) 30:1473-5.
[0080] Guide polynucleotides (e.g., gRNAs) described herein may be of various lengths. In some embodiments, the length of the spacer or targeting sequence depends on the CRISPR- associated protein component of the epigenetic editor system used. For example, Cas proteins from different bacterial species have varying optimal targeting sequence lengths. Accordingly, the spacer sequence may comprise, e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more than 50 nucleotides in length. In some embodiments, the spacer comprises 10-24, 11-20, 11-16, 18-24, 19-21, or 20 nucleotides in length. In some embodiments, a guide polynucleotide (e.g., gRNA) is from 15-100 (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50) nucleotides in length and comprises a spacer sequence of at least 10 (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50) contiguous nucleotides complementary to the target sequence. In some embodiments, a guide polynucleotide described herein may be truncated, e.g., by 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50 or more nucleotides. Table 1 and Table 2 list target sequences for F8 guide RNAs.
[0081] Table 1. F8 target sequences for promoter region
[0082] Table 2. F8 target sequences for downstream promoter region
[0083] In certain embodiments, the 3’ end of the target sequence is immediately adjacent to a PAM sequence (e.g., a canonical PAM sequence such as NGG for SpCas9). The degree of complementarity between the targeting sequence of the guide polynucleotide (e.g., the spacer sequence of a gRNA) and the target sequence may be at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In particular embodiments, the targeting and the target sequence may be 100% complementary. In other embodiments, the targeting sequence and the target sequence may contain, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches.
[0084] A guide polynucleotide (e.g., gRNA) may be modified with, for example, chemical alterations and synthetic modifications. A modified gRNA, for instance, can include an alteration or replacement of one or both of the non-linking phosphate oxygens and / or of one or more of the linking phosphate oxygens in the phosphodiester backbone linkage, an alteration of the ribose sugar (e.g., of the 2’ hydroxyl on the ribose sugar), an alteration of the phosphate moiety, modification or replacement of a naturally occurring nucleobase, modification or replacement of the ribose-phosphate backbone, modification of the 3’ end and / or 5’ end of the oligonucleotide, replacement of a terminal phosphate group or conjugation of a moiety, cap, or linker, or any combination thereof.
[0085] In some embodiments, one or more ribose groups of the gRNA may be modified. Examples of chemical modifications to the ribose group include, but are not limited to, 2’-O- methyl (2’-0Me), 2’ -fluoro (2’-F), 2’ -deoxy, 2’-O-(2-methoxyethyl) (2’ -MOE), 2’-NH2, 2’- O-allyl, 2’-O-ethylamine, 2’-O-cyanoethyl, 2’-O-acetalester, or a bicyclic nucleotide such as locked nucleic acid (LNA), 2’-(5-constrained ethyl (S-cEt)), constrained MOE, or 2’-0,4’-C- aminomethylene bridged nucleic acid (2’,4’-BNANC). 2’-O-methyl modification and / or 2’- fluoro modification may increase binding affinity and / or nuclease stability of the gRNA oligonucleotides.
[0086] In some embodiments, one or more phosphate groups of the gRNA may be chemically modified. Examples of chemical modifications to a phosphate group include, but are not limited to, a phosphorothioate (PS), phosphonoacetate (PACE), thiophosphonoacetate (thioPACE), amide, triazole, phosphonate, and phosphotriester modification. In some embodiments, a guide polynucleotide described herein may comprise one, two, three, or more PS linkages at or near the 5’ end and / or the 3’ end; the PS linkages may be contiguous or noncontiguous.
[0087] In some embodiments, the gRNA herein comprises a mixture of ribonucleotides and deoxyribonucleotides and / or one or more PS linkages.
[0088] In some embodiments, one or more nucleobases of the gRNA may be chemically modified. Examples of chemically modified nucleobases include, but are not limited to, 2- thiouridine, 4-thiouridine, N6-methyladenosine, pseudouridine, 2,6-diaminopurine, inosine, thymidine, 5-methylcytosine, 5-substituted pyrimidine, isoguanine, isocytosine, and nucleobases with halogenated aromatic groups. Chemical modifications can be made in the spacer region, the tracr RNA region, the stem loop, or any combination thereof. Any tracr sequence known in the art is contemplated for a gRNA described herein. In some embodiments, a gRNA described herein has a tracr sequence shown in Table 3 below, or a tracr sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the tracr sequence shown below (“SEQ” means SEQ ID NO).
[0089] Table 3. Exemplary TRACR Sequences
[0090] In some embodiments, the gRNA herein is provided to the cell directly (e.g., through an RNP complex together with the CRIS PR-associated protein domain). In some embodiments, the gRNA is provided to the cell through an expression vector (e.g., a plasmid vector or a viral vector) introduced into the cell, where the cell then expresses the gRNA from the expression vector. Methods of introducing gRNAs and expression vectors into cells are well known in the art.
[0091] III. Effector Domains
[0092] Epigenetic editors described herein include one or more effector protein domains (also “epigenetic effector domains,” or “effector domains,” as used herein) that effect epigenetic modification of a target gene. An epigenetic editor with one or more effector domains may modulate expression of a target gene without altering its nucleobase sequence. In some embodiments, an effector domain described herein may provide repression or silencing of expression of a target, e.g., by repressing transcription or by modifying or remodeling chromatin. Such effector domains are also referred to herein as “repression domains,” “repressor domains,” or “epigenetic repressor domains.” Non-limiting examples of chemical modifications that may be mediated by effector domains include methylation, demethylation, acetylation, deacetylation, phosphorylation, SUMOylation and / or ubiquitination of DNA or histone residues.
[0093] In some embodiments, an effector domain of an epigenetic editor described herein may make histone tail modifications, e.g., by adding or removing active marks on histone tails.
[0094] In some embodiments, an effector domain of an epigenetic editor described herein may comprise or recruit a transcription-related protein, e.g., a transcription repressor. The transcription-related protein may be endogenous or exogenous.
[0095] In some embodiments, an effector domain of an epigenetic editor described herein may, for example, comprise a protein that directly or indirectly blocks access of a transcription factor to the gene of interest harboring the target sequence.
[0096] An effector domain may be a full-length protein or a fragment thereof that retains the epigenetic effector function (a “functional domain”). Functional domains that are capable of modulating (e.g., repressing) gene expression can be derived from a larger protein. For example, functional domains that can reduce target gene expression may be identified based on sequences of repressor proteins. Amino acid sequences of gene expression-modulating proteins may be obtained from available genome browsers, such as the UCSD genome browser or Ensembl genome browser. Protein annotation databases such as UniProt or Pfam can be used to identify functional domains within the full protein sequence. As a starting point, the largest sequence, encompassing all regions identified by different databases, may be tested for gene expression modulation activity. Various truncations then may be tested to identify the minimal functional unit.
[0097] Variants of effector domains described herein are also contemplated by the present disclosure. A variant may, for example, refer to a polypeptide with at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to a wildtype effector domain described herein. In some embodiments, the variant retains at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the epigenetic effector function of the wildtype effector domain.
[0098] In some embodiments, an effector domain described herein may comprise a fusion of two or more effector domains (e.g., KOX1 KRAB and ZIM3). The effector domain may, for example, comprise a fusion of 2, 3, 4, 5, 6, 7, 8, 9, or 10 effector domains, such as effector domains described herein. In certain embodiments, an effector domain comprises a fusion of a truncated form of an effector domain and a second effector domain. In certain embodiments, an effector domain comprises a fusion of the truncated forms of two effector domains (e.g., fusions of the N- and C-terminal portions of the two effector domains).
[0099] In some embodiments, an epigenetic editor described herein may comprise 1 effector domain, 2 effector domains, 3 effector domains, 4 effector domains, 5 effector domains, 6 effector domains, 7 effector domains, 8 effector domains, 9 effector domains, 10 effector domains, or more. In certain embodiments, the epigenetic editor comprises one or more fusion proteins (e.g., one, two, or three fusion proteins), each with one or more effector domains (e.g., one, two, or three effector domains) linked to a DNA-binding domain. In some embodiments, the effector domains may induce a combination of epigenetic modifications, e.g., transcription repression and DNA methylation, DNA methylation and histone deacetylation, DNA methylation and histone demethylation, DNA methylation and histone methylation, DNA methylation and histone phosphorylation, DNA methylation and histone ubiquitylation, DNA methylation, and histone SUMOylation.
[0100] In certain embodiments, an effector domain described herein (e.g., DNMT3L) is encoded by a nucleotide sequence as found in the native genome (e.g., human or murine) for that effector domain. In other embodiments, an effector domain described herein is encoded by a nucleotide sequence that has been codon-optimized for optimal expression in human cells.
[0101] Effector domains described herein may include, for example, transcriptional repressors, DNA methyltransferases, and / or histone modifiers, as further detailed below.
[0102] A. Transcriptional Repressors
[0103] In some embodiments, an epigenetic effector domain described herein mediates repression of a target gene’s expression (e.g., transcription). The effector domain may comprise, e.g., a Kriippel-associated box (KRAB) repressor domain, a Repressor Element Silencing Transcription Factor (REST) repressor domain, a KRAB -associated protein 1 (KAP1) domain, a MAD domain, a FKHR (forkhead in rhabdosarcoma gene) repressor domain, an EGR-1 (early growth response gene product- 1) repressor domain, an ets2 repressor factor repressor domain (ERD), a MAD smSIN3 interaction domain (SID), a WRPW motif of the hairy-related basic helix-loop-helix (bHLH) repressor proteins, an HP1 alpha chromo-shadow repressor domain, an HP1 beta repressor domain, or any combination thereof. The effector domain may recruit one or more protein domains that repress expression of the target gene, e.g., through a scaffold protein. In some embodiments, the effector domain may recruit or interact with a scaffold protein domain that recruits a PRMT protein, a HD AC protein, a SETDB 1 protein, or a NuRD protein domain. In some embodiments, the effector domain comprises a functional domain derived from a zinc finger repressor protein, such as a KRAB domain. KRAB domains are found in approximately 400 human ZFP-based transcription factors. Descriptions of KRAB domains may be found, for example, in Ecco et al., Development (2017) 144(15):2719-29 and Lambert et al., Cell (2018) 172:650-65.
[0104] In certain embodiments, the effector domain comprises a repressor domain (e.g., KRAB) derived from KOX1 / ZNF10, KOX8 / ZNF708, ZNF43, ZNF184, ZNF91, HPF4, HTF10, or HTF34. In some embodiments, the effector domain comprises a repressor domain (e.g., KRAB) derived from ZIM3, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZNF680, ZNF41, ZNF189, ZNF528, ZNF543, ZNF554, ZNF140, ZNF610, ZNF264, ZNF350, ZNF8, ZNF582, ZNF30, ZNF324, ZNF98, ZNF669, ZNF677, ZNF596, ZNF214, ZNF37, ZNF34, ZNF250, ZNF547, ZNF273, ZNF354, ZFP82, ZNF224, ZNF33, ZNF45, ZNF175, ZNF595, ZNF184, ZNF419, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF566, ZNF729, ZIM2, ZNF254, ZNF764, ZNF785, or any combination thereof. For example, the repressor domain may be a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627. In particular embodiments, the repressor domain is a ZIM3 KRAB domain. In further embodiments, the effector domain is derived from a human protein, e.g., a human ZIM3, a human KOX1, a human ZFP28, or a human ZN627.
[0105] Sequences of exemplary effector domains that may reduce or silence target gene expression, or protein sequences that contain them, are provided in Table 4 below (“SEQ” means SEQ ID NO). Further examples of repressors and transcriptional repressor domains can be found, e.g., in PCT Patent Publication WO 2021 / 226077 and Tycko et al., Cell (2020) 183(7):2020-35, each of which is incorporated herein by reference in its entirety.
[0106] Table 4. Exemplary Effector Domains That May Reduce or Silence Gene Expression
[0107] A functional analog of any one of the above-listed proteins, i.e., a molecule having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more) of the protein’s transcription factor function) is encompassed by the present disclosure. For example, the functional analog may be an isoform or a variant of the above-listed protein, e.g., containing a portion of the above protein with or without additional amino acid residues and / or containing mutations relative to the above protein. In some embodiments, the functional analog has a sequence identity that is at least 75, 80, 85, 90, 95, 98, or 99% to one of the sequences listed in Table 4. Homologs, orthologs, and mutants of the above-listed proteins are also contemplated.
[0108] In certain embodiments, an epigenetic editor described herein comprises a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627, and / or an effector domain derived from KAP1, MECP2, HPla, HPlb, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2, optionally wherein the parental protein is a human protein. In particular embodiments, an epigenetic editor described herein comprises a domain derived from KOX1, ZIM3, ZFP28, and / or ZN627, optionally wherein the parental protein is a human protein. In certain embodiments, the epigenetic editor may comprise a KRAB domain derived from KOX1 (ZNF10), e.g., a human KOX1. In certain embodiments, the epigenetic editor may comprise a KRAB domain derived from ZIM3 (ZNF657 or ZNF264), e.g., a human ZIM3. In certain embodiments, the epigenetic editor may comprise a KRAB domain derived from ZFP28, e.g., a human ZFP28. In certain embodiments, the epigenetic editor may comprise a KRAB domain derived from ZN627, e.g., a human ZN627. In certain embodiments, an epigenetic editor described herein may comprise a CDYL2, e.g., a human CDYL2, and / or a TOX domain (e.g., a human TOX domain) in combination with a KOX1 KRAB domain (e.g., a human KOX1 KRAB domain).
[0109] In certain embodiments, an epigenetic effector described herein comprises a repressor domain derived from KOX1 / ZNF10 (SEQ ID NO: 92). For example, the repressor domain may comprise the sequence of SEQ ID NO: 92, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 92.
[0110] In certain embodiments, an epigenetic effector described herein comprises a repressor domain derived from KOX1 / ZNF10, as shown in Table 5 below: Table 5. Exemplary Effector Domains Derived from KOX1 / ZNF10
[0111] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 568, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 568.
[0112] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 569, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 569.
[0113] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 570, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 570.
[0114] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 571, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 571.
[0115] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 572, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 572.
[0116] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 573, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 573.
[0117] In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 574, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 574. In particular embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO: 575, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 575.
[0118] B. Transcriptional Activators
[0119] In some embodiments, an epigenetic effector domain described herein mediates activation of a target gene’s expression (e.g., transcription). In some embodiments, an epigenetic editor for activation of a target as described herein may recruit a domain that is associated with changes in DNA methylation. In some embodiments, an epigenetic effector catalyzes the removal of DNA methylation. A non-limiting example of a domain that is associated with changes in DNA methylation is the ten eleven translocation (Tet) enzyme family. In some embodiments, Tet enzymes oxidize 5-methylcytosines (5mCs) and promote locus- specific reversal of DNA methylation. Non-limiting examples of Tet enzymes are human Tetl (UniProt ID: Q8NFU7) and human Tet2 (UniProt ID: Q6N021). The Tetl protein has been characterized in the art and has at least one catalytic domain. In some embodiments of the present disclosure, fusion proteins comprise the catalytic domain (CD) of Tetl.
[0120] Some exemplary and non-limiting activator effector proteins, domains, truncations, trimmings, or combinations of effector domains that may activate target gene expression through various mechanisms that are contemplated by the present disclosure are OCT4 (POU5F1), GATA4, PU.l (SPI1), EOMES, PAX6 (FVH1), FOXA1 (HNF3A), FOXA2 (HNF3B), FOXD3 (HFH2), CEBPalpha, CEBPbeta, Satbl, KLF4, HNF4 alpha, HNF1 alpha, GATA2, GATA6, PBX1, EBF1, HOXB2, FOXN1, FOXR2, KLF14, SOX7, SPDYE4, CSRNP1, ATF6, CITED2-TAD, CITED 1 -TAD, C3orf62-TAD, CBX-C (CBX2), FAM22F- 23, KLF6-1, SPDYE4-3, Med21, Med25, ZN473_SND, FOXO3_SND, FOXO1_SND, MYBA_SND, MYB_SND, NCOA2_SND, SMCA2_SND, NCOA3_SND, ZN597_SND, APBB1_SND, ANM2_SND, CXXC1_SND, CRTC2_SND, NOTC2_SND, p53 (gene: TP53), PPARgamma, BZLF1, BanP, GCN5_HAT, KDM4A_JMJN_C, KDM4D_FL, KDM6A_JMJC, KDM6B_JMJC, MLL3_SET, MLL4_SET, MLL4_SET2, Dpy-30, ASH1L_SET, ASH2L, Brd9, ZNF473_SND (trimmed), Vp64, P65-RTA (PR), HSF1, P300, SSI 8, VPH, FOXP3, Vp64, GADD45a, GADD45b, GADD45g, SOX2, MYODI (MYF3), GATA3, PAX3, PAX7 (HUP1), ASCL1, FOXO1 (FKHR), Rapl, Rebl, Cbfl, HNF4gamma, NEURODI, NEUROD2, ILS1, KLF15, KLF6, ATMIN, CHOP (DDIT3), C3orf62, KLF7- TAD, ATF6-TAD, KIBRA_SND, c-Myc, BRD2, Med6, KDM4A_JMJC, KDM4B_JMJC, KDM4C_JMJC, KDM4D_JMJC, SMYD5_H316L_C318A, and BRGl_Trunc. Alternative names used in the art, potential amino acid changes, trimmings, and truncations are indicated.
[0121] C. DNA Methyltransferases
[0122] In some embodiments, an effector domain of an epigenetic editor described herein alters target gene expression through DNA modification, such as methylation. Highly methylated areas of DNA tend to be less transcriptionally active than less methylated areas. DNA methylation occurs primarily at CpG sites (shorthand for “C-phosphate-G-” or “cytosine-phosphate-guanine” sites). Many mammalian genes have promoter regions near or including CpG islands (nucleic acid regions with a high frequency of CpG dinucleotides).
[0123] An effector domain described herein may be, e.g., a DNA methyltransferase (DNMT), or a catalytic domain thereof, or may be capable of recruiting a DNA methyltransferase. DNMTs encompass enzymes that catalyze the transfer of a methyl group to a DNA nucleotide, such as canonical cytosine-5 DNMTs that catalyze the addition of methyl groups to genomic DNA (e.g., DNMT3A, DNMT3B, and DNMT3C). This term also encompasses non-canonical family members that do not catalyze methylation themselves but that recruit (including activate) catalytically active DNMTs; a non-limiting example of such a DNMT is DNMT3L. See, e.g., Lyko, Nat Review (2018) 19:81-92. Unless otherwise indicated, a DNMT domain may refer to a polypeptide domain derived from a catalytically active DNMT (e.g., DNMT3A, and DNMT3B) or from a catalytically inactive DNMT (e.g., DNMT3L). A DNMT may repress expression of the target gene through the recruitment of repressive regulatory proteins. In some embodiments, the methylation is at a CG (or CpG) dinucleotide sequence. In some embodiments, the methylation is at a CHG or CHH sequence, where H is any one of A, T, or C.
[0124] In some embodiments, a DNMT described herein can be an animal DNMT (e.g., a mammalian DNMT), a plant DNMT, a fungal DNMT, or a bacterial DNMT. A bacterial DNMT can be obtained from a bacterial species (e.g., a coccus bacterium, bacillus bacterium, spiral bacterium, or an intracellular, gram-positive, or gram-negative bacterium. In certain embodiments, the bacterial species is Mycoplasmatales bacterium, Mycoplasma marinum, or Spiroplasma chinense. In certain embodiments, the bacterial species is not M. penetrans, S. monbiae, H. parainfluenzae, A. luteus, H. aegyptius, H. haemolyticus, Moraxella, E. coli, T. aquaticus, C. crescentus, or C. difficile. In certain embodiments, an epigenetic editor described herein comprises a DNMT domain comprising SEQ ID NO: 604, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 604. In certain embodiments, an epigenetic editor described herein comprises a DNMT domain comprising SEQ ID NO: 605, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 605. In certain embodiments, an epigenetic editor described herein comprises a DNMT domain comprising SEQ ID NO: 606, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 606.
[0125] In certain embodiments, DNMTs comprised in or recruited by the epigenetic editors described herein may include, e.g., DNMT3A, DNMT3B, and / or DNMT3C. In some embodiments, the DNMT is a mammalian (e.g., human or murine) DNMT. In particular embodiments, the DNMT is DNMT3A (e.g., human DNMT3A). In certain embodiments, an epigenetic editor described herein recruits a DNMT comprising a DNMT3A domain comprising SEQ ID NO: 577, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 577. In certain embodiments, an epigenetic editor described herein recruits a DNMT comprising a DNMT3A domain comprising SEQ ID NO: 578, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 578.
[0126] In some embodiments, an effector domain described herein may be a DNMT-like domain. As used herein a “DNMT-like domain” is a regulatory factor of DNMT that may activate or recruit other DNMT domains but does not itself possess methylation activity. In some embodiments, the DNMT-like domain is a mammalian (e.g., human or mouse) DNMT- like domain. In certain embodiments, the DNMT-like domain is DNMT3L, which may be, for example, human DNMT3L or mouse DNMT3L. In certain embodiments, an epigenetic editor described herein comprises a DNMT3L domain comprising SEQ ID NO: 584, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 584. In certain embodiments, an epigenetic editor herein comprises a DNMT3L domain comprising SEQ ID NO: 585, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 585. In certain embodiments, an epigenetic editor described herein comprises a DNMT3L domain comprising SEQ ID NO: 586, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 586. In certain embodiments, an epigenetic editor described herein comprises a DNMT3L domain comprising SEQ ID NO: 587, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 587. In some embodiments, the DNMT3L domain may have, e.g., a mutation corresponding to that at position D226 (such as D226V), Q268 (such as Q268K), or both (numbering according to SEQ ID NO: 584).
[0127] Table 6 below provides exemplary DNMTs that may be part of an epigenetic effector domain described herein or from which an effector domain of an epigenetic editor described herein may be derived.
[0128] Table 6. Exemplary DNMT Sequences
[0129] A functional analog of any one of the above-listed proteins, i.e., a molecule having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more) of the protein’s DNA methylation function or recruiting function) is encompassed by the present disclosure. For example, the functional analog may be an isoform or a variant of the above-listed protein, e.g., containing a portion of the above protein with or without additional amino acid residues and / or containing mutations relative to the above protein. In some embodiments, the functional analog has a sequence identity that is at least 75, 80, 85, 90, 95, 98, or 99% to one of the sequences listed in Table 6. In some embodiments, the effector domain herein comprises only the functional domain (or functional analog thereof), e.g., the catalytic domain or recruiting domain, of an abovelisted protein. In some embodiments, the effector domain herein comprises one or more epigenetic effector domains selected from Table 6, or functional homologs, orthologs, or variants thereof.
[0130] As used herein, a DNMT domain (e.g., a DNMT3A domain or a DNMT3L domain) refers to a protein domain that is identical to the parental protein (e.g., a human or murine DNMT3A or DNMT3L) or a functional analog thereof (e.g., having a functional fragment, such as a catalytic fragment or recruiting fragment, of the parental protein; and / or having mutations that improve the activity of the DNMT protein).
[0131] An epigenetic editor herein may effect methylation at, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more CpG dinucleotide sequences in the target gene or chromosome. The CpG dinucleotide sequences may be located within or near the target gene in CpG islands, or may be located in a region that is not a CpG island. A CpG island generally refers to a nucleic acid sequence or chromosome region that comprises a high frequency of CpG dinucleotides. For example, a CpG island may comprise at least 50% GC content. The CpG island may have a high observed-to-expected CpG ratio, for example, an observed-to-expected CpG ratio of at least 60%. As used herein, an observed-to-expected CpG ratio is determined by Number of CpG * (sequence length) / (Number of C * Number of G). In some embodiments, the CpG island has an observed-to-expected CpG ratio of at least 60%, 70%, 80%, 90% or more. A CpG island may be a sequence or region of, e.g., at least 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, or 800 nucleotides. In some embodiments, only 1, or less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, or 50 CpG dinucleotides are methylated by the epigenetic editor.
[0132] In some embodiments, an epigenetic editor herein effects methylation at a hypomethylated nucleic acid sequence, i.e., a sequence that may lack methyl groups on the 5- methyl cytosine nucleotides (e.g., in CpG) as compared to a standard control. Hypomethylation may occur, for example, in aging cells or in cancer (e.g., early stages of neoplasia) relative to a younger cell or non-cancer cell, respectively.
[0133] In some embodiments, an epigenetic editor described herein induces methylation at a hypermethylated nucleic acid sequence.
[0134] In some embodiments, methylation may be introduced by the epigenetic editor at a site other than a CpG dinucleotide. For example, the target gene sequence may be methylated at the C nucleotide of CpA, CpT, or CpC sequences. In some embodiments, an epigenetic editor comprises a DNMT3A domain and effects methylation at CpG, CpA, CpT, CpC sequences, or any combination thereof. In some embodiments, an epigenetic editor comprises a DNMT3A domain that lacks a regulatory subdomain and only maintains a catalytic domain. In some embodiments, the epigenetic editor comprising a DNMT3A catalytic domain effects methylation exclusively at CpG sequences. In some embodiments, an epigenetic editor comprising a DNMT3A domain that comprises a mutation, e.g. a R836A or R836Q mutation (numbering according to SEQ ID NO: 577), has higher methylation activity at CpA, CpC, and / or CpT sequences as compared to an epigenetic editor comprising a wildtype DNMT3A domain.
[0135] D. Histone Modifiers
[0136] In some embodiments, an effector domain of an epigenetic editor herein mediates histone modification. Histone modifications play a structural and biochemical role in gene transcription, such as by formation or disruption of the nucleosome structure that binds to the histone and prevents gene transcription. Histone modifications may include, for example, acetylation, deacetylation, methylation, phosphorylation, ubiquitination, SUMOylation and the like, e.g., at their N-terminal ends (“histone tails”). These modifications maintain or specifically convert chromatin structure, thereby controlling responses such as gene expression, DNA replication, DNA repair, and the like, which occur on chromosomal DNA. Post-translational modification of histones is an epigenetic regulatory mechanism and is considered essential for the genetic regulation of eukaryotic cells. Recent studies have revealed that chromatin remodeling factors such as SWI / SNF, RSC, NURF, NRD, and the like, which facilitate transcription factor access to DNA by modifying the nucleosome structure; histone acetyltransferases (HATs) that regulate the acetylation state of histones; and histone deacetylases (HDACs), act as important regulators.
[0137] In particular, the unstructured N-termini of histones may be modified by acetylation, deacetylation, methylation, ubiquitylation, phosphorylation, SUMOylation, ribosylation, citrullination O-GlcNAcylation, crotonylation, or any combination thereof. For example, histone acetyltransferases (HATs) utilize acetyl-CoA as a cofactor and catalyze the transfer of an acetyl group to the epsilon amino group of the lysine side chains. This neutralizes the lysine’s positive charge and weakens the interactions between histones and DNA, thus opening the chromosomes for transcription factors to bind and initiate transcription. Acetylation of K14 and K9 lysines of histone H3 by histone acetyltransferase enzymes may be linked to transcriptional competence in humans. Lysine acetylation may directly or indirectly create binding sites for chromatin-modifying enzymes that regulate transcriptional activation. On the other hand, histone methylation of lysine 9 of histone H3 may be associated with heterochromatin, or transcriptionally silent chromatin.
[0138] In certain embodiments, an effector domain of an epigenetic editor described herein comprises a histone methyltransferase domain. The effector domain may comprise, for example, a D0T1L domain, a SET domain, a SUV39H1 domain, a G9a / EHMT2 protein domain, an EZH1 domain, an EZH2 domain, a SETDB1 domain, or any combination thereof. In particular embodiments, the effector domain comprises a histone-lysine-N- methyltransferase SETDB1 domain.
[0139] In some embodiments, the effector domain comprises a histone deacetylase protein domain. In certain embodiments, the effector domain comprises a HD AC family protein domain, for example, a HDAC1, HDAC3, HDAC5, HDAC7, or HDAC9 protein domain. In particular embodiments, the effector domain comprises a nucleosome remodeling and deacetylase complex (NURD), which removes acetyl groups from histones.
[0140] E. Other Effector Domains
[0141] In some embodiments, the effector domain comprises a tripartite motif containing protein (TRIM28, TIFl-beta, or KAP1). In certain embodiments, the effector domain comprises one or more KAP1 proteins. A KAP1 protein in an epigenetic editor herein may form a complex with one or more other effector domains of the epigenetic editor or one or more proteins involved in modulation of gene expression in a cellular environment. For example, KAP1 may be recruited by a KRAB domain of a transcriptional repressor. A KAP1 protein domain may interact with or recruit one or more protein complexes that reduces or silences gene expression. In some embodiments, KAP1 interacts with or recruits a histone deacetylase protein, a histone-lysine methyltransferase protein, a chromatin remodeling protein, and / or a heterochromatin protein. For example, a KAP1 protein domain may interact with or recruit a heterochromatin protein 1 (HP1) protein, a SETDB1 protein, an HD AC protein, and / or a NuRD protein complex component. In some embodiments, a KAP1 protein domain interacts with or recruits a ZFP90 protein (e.g., isoform 2 of ZFP90), and / or a FOXP3 protein. An exemplary KAP1 amino acid sequence is shown in SEQ ID NO: 632.
[0142] In some embodiments, the effector domain comprises a protein domain that interacts with or is recruited by one or more DNA epigenetic marks. For example, the effector domain may comprise a methyl CpG binding protein 2 (MECP2) protein that interacts with methylated DNA nucleotides in the target gene (which may or may not be at a CpG island of the target gene). An MECP2 protein domain in an epigenetic editor described herein may induce condensed chromatin structure, thereby reducing or silencing expression of the target gene. In some embodiments, an MECP2 protein domain in an epigenetic editor described herein may interact with a histone deacetylase (e.g. HD AC), thereby repressing or silencing expression of the target gene. In some embodiments, an MECP2 protein domain in an epigenetic editor described herein may block access of a transcription factor or transcriptional activator to the target sequence, thereby repressing or silencing expression of the target gene. An exemplary MECP2 amino acid sequence is shown in SEQ ID NO: 633.
[0143] Also contemplated as effector domains for the epigenetic editors described herein are, e.g., a chromoshadow domain, a ubiquitin-2 like Rad60 SUMO-like (Rad60-SLD / SUMO) domain, a chromatin organization modifier domain (Chromo) domain, a Y af2 / RYBP C- terminal binding motif domain (YAF2_RYBP), a CBX family C-terminal motif domain (CBX7_C), a zinc finger C3HC4 type (RING finger) domain (ZF-C3HC4_2), a cytochrome b5 domain (Cyt-b5), a helix-loop-helix domain (HLH), a helix-hairpin-helix motif domain (e.g., HHH_3), a high mobility group box domain (HMG-box), a basic leucine zipper domain (e.g., bZIP_l or bZIP_2), a Myb_DNA-binding domain, a homeodomain, a MYM-type zinc finger with FCS sequence domain (ZF-FCS), an interferon regulatory factor 2-binding protein zinc finger domain (IRF-2BP1_2), an SSX repressor domain (SSXRD), a B-box-type zinc finger domain (ZF-B_box), a CXXC zinc finger domain (ZF-CXXC), a regulator of chromosome condensation 1 domain (RCC1), an SRC homology 3 domain (SH3_9), a sterile alpha motif domain (SAM_1), a sterile alpha motif domain (SAM_2), a sterile alpha motif / Pointed domain (SAM_PNT), a Vestigial / Tondu family domain (Vg_Tdu), a LIM domain, an RNA recognition motif domain (RRM_1), a paired amphipathic helix domain (PAH), a proteasomal ATPase OB C-terminal domain (Prot_ATP_ID_OB), a nervy homology 2 domain (NHR2), a hinge domain of cleavage stimulation factor subunit 2 (CSTF2_hinge), a PPAR gamma N-terminal region domain (PPARgamma_N), a CDC48 N- terminal domain (CDC48_2), a WD40 repeat domain (WD40), a Fipl motif domain (Fipl), a PDZ domain (PDZ_6), a Von Willebrand factor type C domain (VWC), a NAB conserved region 1 domain (NCD1), an SI RNA-binding domain (SI), an HNF3 C-terminal domain (HNF_C), a Tudor domain (Tudor_2), a histone-like transcription factor (CBF / NF-Y) and archaeal histone domain (CBFD_NFYB_HMF), a zinc finger protein domain (DUF3669), an EGF-like domain (cEGF), a GATA zinc finger domain (GATA), a TEA / ATTS domain (TEA), a phorbol esters / diacylglycerol binding domain (Cl-1), polycomb-like MTF2 factor 2 domain (Mtf2_C), a transactivation domain of FOXO protein family (FOXO-TAD), a homeobox KN domain (Homeobox_KN), a BED zinc finger domain (ZF-BED), a zinc finger of C3HC4-type RING domain (ZF-C3HC4_4), a RAD51 interacting motif domain (RAD5 l_interact), a p55-binding region of a nethyl-CpG-binding domain protein MBD (MBDa), a Notch domain, a Raf-like Ras-binding domain (RBD), a Spin / Ssty family domain (Spin-Ssty), a PHD finger domain (PHD_3), a Low-density lipoprotein receptor domain class A (Ldl_recept_a), a CS domain, a DM DNA-binding domain, and a QLQ domain.
[0144] In some embodiments, the effector domain is a protein domain comprising a YAF2_RYBP domain or homeodomain or any combination thereof. In certain embodiments, the homeodomain of the YAF2_RYBP domain is a PRD domain, an NKL domain, a HOXL domain, or a LIM domain. In particular embodiments, the YAF2_RYBP domain may comprise a 32 amino acid Yaf2 / RYBP C-terminal binding motif domain (32 aa RYBP).
[0145] In some embodiments, the effector domain comprises a protein domain selected from a group consisting of SUMO3 domain, Chromo domain from M phase phosphoprotein 8 (MPP8), chromoshadow domain from Chromobox 1 (CBX1), and SAM_1 / SPM domain from Scm Polycomb Group Protein Homolog 1 (SCMH1).
[0146] In some embodiments, the effector domain comprises an HNF3 C-terminal domain (HNF_C). The HNF_C domain may be from FOXA1 or FOXA2. In certain embodiments, the HNF_C domain comprises an EH1 (engrailed homology 1) motif.
[0147] In some embodiments, the effector domain may comprise an interferon regulatory factor 2-binding protein zinc finger domain (IRF-2BP1_2), a Cyt-b5 domain from DNA repair factor HERC2 E3 ligase, a variant SH3 domain (SH3_9) from Bridging Integrator 1 (BINI), an HMG-box domain from transcription factor TOX or ZF-C3HC4_2 RING finger domain from the polycomb component PCGF2, a Chromodomain-helicase-DNA binding protein 3 (CHD3) domain, or a ZNF783 domain. V. Epigenetic Editors
[0148] The present disclosure provides epigenetic editors (also referred to herein as epigenetic editing systems) for repressing or activating expression of a gene of interest, e.g., by directing epigenetic modification(s) to a target sequence in the gene of interest.
[0149] The DNA-binding domain (in concert with a guide polynucleotide such as one described herein, where the DNA-binding domain is a polynucleotide guided DNA-binding domain) directs the effector domain to epigenetically modify a target sequence in the gene of interest. In some embodiments, the epigenetic editor comprises an effector domain that represses gene expression, thus recruiting the effector domain to the target sequence in the gene of interest, resulting in repression of gene expression (silencing) of the gene of interest. In some embodiments, the epigenetic editor comprises an effector domain that activates gene expression or relieves gene repression, e.g., by removing DNA methyl marks, thus recruiting the effector domain to the target sequence in the gene of interest, resulting in activation of gene expression (activation) of the gene of interest. In some embodiments, the repression or activation of the gene of interest is durable and inheritable across cell generations. In some embodiments the repression or activation of the gene of interest is temporary. In some embodiments the repression or activation of the gene of interest is reversible.
[0150] In particular embodiments, an epigenetic editor provided herein comprises one or more fusion proteins that comprises (1) a DNA-binding domain and (2) an effector domain, e.g., an epigenetic effector domain. . In particular embodiments, an epigenetic editor provided herein comprises one or more fusion proteins that comprises (1) a DNA-binding domain (2) two or more effector domains, e.g., two or more epigenetic effector domains. In particular embodiments, an epigenetic editor provided herein comprises one or more fusion proteins that comprises (1) a DNA-binding domain and (2) two or more effector domains, e.g., two or more epigenetic effector domains, wherein the two or more epigenetic effector domains are different effector domains, e.g., a DNA methyltransferase domain and a histone modifier.
[0151] In some embodiments of the present disclosure, an epigenetic editor composition comprises a fusion protein, comprising a Sa or Nme2 Cas9 nuclease-dead RNA-guided nuclease fused to a TET1 catalytic domain and a guide RNA. In some embodiments, the TET1 catalytic domain is positioned on the C terminus of the nuclease-dead RNA-guided nuclease. In some embodiments, the TET1 catalytic domain is positioned on the N terminus of the nuclease-dead RNA-guided nuclease. In some embodiments, the fusion protein comprises an SaCas9 nuclease-dead RNA-guided nuclease. In some embodiments, the fusion protein comprises the sequence of SEQ ID NOs: 798, 800, 802, or 804.
[0152] In some embodiments, the fusion protein comprises an Nme2 nuclease-dead RNA- guided nuclease. In some embodiments, the fusion protein comprises the sequence of SEQ ID NOs: 845 or 847. In some embodiments, the guide RNA is targeted to or binds to the sequence of SEQ ID NOs: 856, 857, 867, 868, or 870.
[0153] In some embodiments of the present disclosure comprising a fusion protein and guide RNA, the composition further comprises at least one additional guide RNA. In some particular’ embodiments, the composition comprises three additional guide RNAs (in addition to the required one guide).
[0154] In some embodiments of the present disclosure, an epigenetic editor composition comprises a fusion protein comprising a DNA-binding ZF domain fused to a TET1 catalytic domain. In some embodiments, the composition further comprises at least one additional fusion protein comprising a DNA-binding ZF domain fused to a TET1 catalytic domain, In some particular embodiments, the composition comprises three additional fusion proteins comprising a DNA-binding ZF domain fused to a TET1 catalytic domain (in addition to the required one protein).
[0155] In some embodiments of the present disclosure, an epigenetic editor composition comprises a fusion protein comprising a Sa Cas9 nuclease-dead RNA-guided nuclease with an activation domain fused to its C terminus and a guide RNA. In some embodiments, the activation domain comprises the sequence of any of SEQ ID NOs: 634-741.
[0156] Exemplary epigenetic editing and approaches are described in Nunez JK, Chen J, Pommier GC, et al. Genome- wide programmable transcriptional memory by CRISPR- based epigenome editing. Cell. 2021;184(9):2503-2519. el 7; Amabile A, Migliara A, Capasso P, et al. inheritable Silencing of Endogenous Genes by Hit-and-Run Targeted Epigenetic Editing. Cell, 2016; 167(l):219-232.el4; and Maeder ML, Angstman JF, Richardson ME, et al. Targeted DNA demethylation and activation of endogenous genes using programmable TALE-TET1 fusion proteins. Nat Biotechnol. 2013:31(12): 1137-1142, all of which are hereby incorporated by reference in their entirety.
[0157] In some embodiments, the epigenetic editor comprises a single fusion protein. In some embodiments, the epigenetic editor consists of a single fusion protein. In some embodiments, the epigenetic editor comprises two or more fusion proteins, wherein each fusion protein comprises, independently, (1) a DNA-binding domain and (2) an effector domain. A fusion protein described herein may further comprise one or more linkers (e.g., peptide linkers), detectable tags, nuclear localization signals (NLSs), or any combination thereof.
[0158] As used herein, a “fusion protein” refers to a chimeric protein in which two or more coding sequences (e.g., for a DNA-binding domain and / or an effector domain) are covalently joined, typically via a peptide bond.
[0159] In some embodiments, an epigenetic editor described herein comprises a DNA binding domain and an epigenetic effector domain. In some embodiments, an epigenetic editor described herein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10, or more effector (e.g., repression / repressor) domains, which may be identical or different. In certain embodiments, two or more of said effector domains function synergistically. Combinations of effector domains may comprise DNA methylation domains, histone deacetylation domains, histone methylation domains, and / or scaffold domains that recruit any of the above. For example, an epigenetic editor described herein may comprise one or more transcriptional repressor domains (e.g., a KRAB domain such as KOX1, ZIM3, ZFP28, or ZN627 KRAB) in combination with one or more DNA methylation domains (e.g., a DNMT domain) and / or recruiter domain (e.g., a DNMT3L domain). Such an epigenetic editor may comprise, for instance, a KRAB domain, a DNMT3A domain, and a DNMT3L domain. In some embodiments, the epigenetic editor further comprises an additional effector domain (e.g., a KAP1, MECP2, HPlb, CBX8, CDYL2, TOX, TOX3, TOX4, EED, RBBP4, RCOR1, or SCML2 domain). In some embodiments, the additional effector domain is a CDYL2, TOX, TOX3, TOX4, or HPla domain. For example, an epigenetic editor described herein may comprise a CDYL2 and / or a TOX domain in combination with a KRAB domain (e.g., a KOX1 KRAB domain).
[0160] A. Linkers
[0161] A fusion protein as described herein may comprise one or more linkers that connect components of the epigenetic editor. A linker may be a peptide or non-peptide linker.
[0162] In some embodiments, one or more linkers utilized in an epigenetic editor provided herein is a peptide linker, i.e., a linker comprising a peptide moiety. A peptide linker can be any length applicable to the epigenetic editor fusion proteins described herein. In some embodiments, the linker can comprise a peptide between 1 and 250 (e.g., between 1 and 80) amino acids. In some embodiments, the linker comprises from 1 to 5, 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 80, 1 to 100, 1 to 150, 1 to 200, 1 to 250, 5 to 10, 5 to 20, 5 to 30, 5 to 40, 5 to 60, 5 to 80, 5 to 100, 5 to 150, 5 to 200, 5 to 250, 10 to 20, 10 to 30, 10 to 40, 10 to 50, 10 to 60, 10 to 80, 10 to 100, 10 to 150, 10 to 200, 10 to 250, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 100, 20 to 150, 20 to 200, 20 to 250, 30 to 40, 30 to 50, 30 to 60, 30 to 80, 30 to 100, 30 to 150, 30 to 200, 30 to 250, 40 to 50, 40 to 60, 40 to 80, 40 to 100, 40 to 150, 40 to 200, 40 to 250, 50 to 60, 50 to 80, 50 to 100, 50 to 150, 50 to 200, 50 to 250, 60 to 80, 60 to 100, 60 to 150, 60 to 200, 60 to 250, 80 to 100, 80 to 150, 80 to 200, 80 to 250, 100 to 150, 100 to 200, 100 to 250, 150 to 200, or 150 to 250 amino acids in length. Longer or shorter linkers are also contemplated. In some embodiments, the peptide linker is
[0163] 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 25, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acids in length. For example, the peptide linker may be
[0164] 4, 5, 16, 20, 23, 24, 27, 32, 40, 64, 92, or 104 amino acids in length. The peptide linker may be a flexible or rigid linker. In particular embodiments, the peptide linker comprises the amino acid sequence of any one of SEQ ID NOs: 742-748 and 775-779 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto.
[0165] In certain embodiments, the peptide linker is an XTEN linker. Such a linker may comprise part of the XTEN sequence (Schellenberger et al., Nat Biotechnol (2009) 27(1): 1186-90), an unstructured hydrophilic polypeptide consisting only of residues G, S, P, T, E, and A. The term “XTEN” as used herein refers to a recombinant peptide or polypeptide lacking hydrophobic amino acid residues. XTEN linkers typically are unstructured and comprise a limited set of natural amino acids. Fusion of XTEN to proteins alters its hydrodynamic properties and reduces the rate of clearance and degradation of the fusion protein. These XTEN fusion proteins are produced using recombinant technology, without the need for chemical modifications, and degraded by natural pathways. The XTEN linker may be, for example, 5, 10, 16, 20, 26, or 80 amino acids in length. In some embodiments, the XTEN linker is 16 amino acids in length. In some embodiments, the XTEN linker is 80 amino acids in length. In certain embodiments, the XTEN linker may be XTEN10, XTEN16, XTEN20, or XTEN80. In certain embodiments, the XTEN linker may comprise the amino acid sequence of any one of SEQ ID NOs: 752-757 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In particular embodiments, the XTEN linker comprises the amino acid sequence of SEQ ID NO: 752. In particular embodiments, the XTEN linker comprises the amino acid sequence of SEQ ID NO: 757.
[0166] In some embodiments, one or more linkers utilized in an epigenetic editor provided herein is a non-peptide linker. For example, the linker may be a carbon bond, a disulfide bond, or carbon-heteroatom bond. In certain embodiments, the linker is a carbon-nitrogen bond of an amide linkage. In certain embodiments, the linker is a cyclic or acyclic, substituted or unsubstituted, or branched or unbranched aliphatic or heteroaliphatic linker.
[0167] In some embodiments, one or more linkers utilized in an epigenetic editor provided herein is polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). The linker may comprise, for example, a monomer, dimer, or polymer of aminoalkanoic acid; an aminoalkanoic acid (e.g., glycine, ethanoic acid, alanine, beta- alanine, 3-aminopropanoic acid, 4-aminobutanoic acid, 5-pentanoic acid, etc.); a monomer, dimer, or polymer of aminohexanoic acid (Ahx); or a polyethylene glycol moiety (PEG); or an aryl or heteroaryl moiety. In certain embodiments, the linker may be based on a carbocyclic moiety (e.g., cyclopentane or cyclohexane) or a phenyl ring. The linker may include functionalized moieties to facilitate attachment of a nucleophile (e.g., thiol, amino) from the peptide to the linker. Any electrophile may be used as part of the linker. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, alkyl halides, aryl halides, acyl halides, and isothiocyanates.
[0168] Various linker lengths and flexibilities can be employed between any two components of an epigenetic editor (e.g., between an effector domain (e.g., a repressor domain) and a DNA-binding domain (e.g., a Cas9 domain), between a first effector domain and a second effector domain, etc.). The linkers may range from very flexible linkers, such as glycine / serine-rich linkers, to more rigid linkers, in order to achieve the optimal length for effector domain activity for the specific application. In some embodiments, the more flexible linkers are glycine / serine-rich linkers (GS-rich linkers), where more than 45% (e.g., more than 48, 50, 55, 60, 70, 80, or 90%) of the residues are glycine or serine residues. Nonlimiting examples of the GS-rich linkers are (GGGGS)n (SEQ ID NO: 778), (G)n (SEQ ID NO: 992), and W linker (SEQ ID NO: 637). In some embodiments, the more rigid linkers are in the form of the form (EAAAK)n (SEQ ID NO: 779), (SGGS)n (SEQ ID NO: 745), and (XP)n(SEQ ID NO: 993)). In the aforementioned formulae of flexible and rigid linkers, n may be any integer between 1 and 30. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises a (GGS)n (SEQ ID NO: 994) motif, wherein n is 1, 3, or 7. In some embodiments, the linker comprises a (GGGGS)n motif, wherein n is 4 (SEQ ID NO: 750).
[0169] In some embodiments, a linker in an epigenetic editor described herein comprises a nuclear localization signal, for example, with the amino acid sequence of any one of SEQ ID NOs: 758-763. In some embodiments, a linker in an epigenetic editor described herein comprises an expression tag, e.g., a detectable tag such as a green fluorescent protein.
[0170] B. Nuclear Localization Signals
[0171] A fusion protein described herein may comprise one or more nuclear localization signals, and in certain embodiments, may comprise two or more nuclear localization signals. For example, the fusion protein may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more nuclear localization signals. As used herein, a “nuclear localization signal” (NLS) is an amino acid sequence that directs proteins to the nucleus. In certain embodiments, the NLS may be an SV40 NLS (e.g., with the amino acid sequence of SEQ ID NO: 758). The fusion protein may comprise an NLS at its N-terminus, C-terminus, or both, and / or an NLS may be embedded in the middle of the fusion protein (e.g., at the N- or C- terminus of a DNA-binding domain or an effector domain).
[0172] In some embodiments, the fusion protein may comprise two NLSs. The fusion protein may comprise two NLSs at its N-terminus or C-terminus. The fusion protein may comprise one NLS located at its N-terminus and one NLS embedded in the middle of the fusion protein, or one NLS located at its C-terminus and one NLS embedded in the middle of the fusion protein. The fusion protein may comprise two NLSs embedded in the middle of the fusion protein.
[0173] In some embodiments, the fusion protein may comprise four NLSs. The fusion protein may comprise at least two (e.g., two, three, or four) NLSs at its N-terminus or C- terminus. The fusion protein may comprise at least one (e.g., one, two, three, or four) NLSs embedded in the middle of the fusion protein. In particular embodiments, the fusion protein may comprise two NLSs at its N-terminus and two NLSs at its C-terminus.
[0174] An NLS described herein may be an endogenous NLS sequence. In certain embodiments, an NLS described herein comprises the amino acid sequence of any one of SEQ ID NOs: 758-763, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the selected sequence. In particular embodiments, the NLS comprises the amino acid sequence of SEQ ID NO: 758. Additional NLSs are known in the art.
[0175] In some embodiments, an epigenetic editor comprising a fusion protein that comprises at least one NLS at the N-terminus and at least one NLS at the C-terminus may increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000%, or more as compared to an epigenetic editor with a corresponding fusion protein that does not have at least one NLS at the N-terminus and at least one NLS at the C-terminus.
[0176] In some embodiments, an epigenetic editor comprising a fusion protein that comprises two NLSs at the N-terminus and two NLSs at the C-terminus may increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000%, or more as compared to an epigenetic editor with a corresponding fusion protein that does not have two NLSs at the N-terminus and two NLSs at the C-terminus.
[0177] C. Tags
[0178] Epigenetic editors provided herein may comprise one or more additional sequences (“tags”) for tracking, detection, and localization of the editors. In some embodiments, the epigenetic editor comprises 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more detectable tags. Each of the detectable tags may be the same or different.
[0179] For example, an epigenetic editor fusion protein may comprise cytoplasmic localization sequences, export sequences, such as nuclear export sequences, or other localization sequences, as well as sequence tags that are useful for solubilization, purification, or detection of the fusion proteins. Suitable protein tags provided herein include, but are not limited to, biotin carboxylase carrier protein (BCCP) tags, myc-tags, calmodulin-tags, FLAG- tags, hemagglutinin (HA)-tags, poly-histidine tags (also referred to as histidine tags or His- tags), maltose binding protein (MBP)-tags, nus-tags, glutathione-S-transferase (GST)-tags, green fluorescent protein (GFP)-tags, thioredoxin-tags, S-tags, Softags (e.g., Softag 1 or Softag 3), strep-tags, biotin ligase tags, FlAsH tags, V5 tags, and SBP-tags. Additional suitable sequences will be apparent to those of skill in the art. Sequences disclosed herein that are presented with tag sequences included are also contemplated without the presented tag sequences.
[0180] VI. Pharmaceutical Compositions
[0181] In one aspect, the present disclosure provides a pharmaceutical composition comprising as an active ingredient (or as the sole active ingredient) one or more epigenetic editors described herein or component(s) (e.g., fusion proteins and / or guide polynucleotides) thereof, or nucleic acid molecule(s) encoding said epigenetic editors or component(s) thereof. For example, a pharmaceutical composition may comprise nucleic acid molecule(s) encoding the fusion protein(s) (and guide polynucleotides, where applicable) of an epigenetic editor described herein. In some embodiments, separate pharmaceutical compositions comprise the fusion protein(s) and the guide polynucleotide(s). A pharmaceutical composition may also comprise cells that have undergone epigenetic modification(s) mediated or induced by an epigenetic editor provided herein.
[0182] Generally, the epigenetic editors described herein or component(s) thereof, or nucleic acid molecule(s) encoding said epigenetic editors or component(s) thereof, of the present disclosure are suitable to be administered as a formulation in association with one or more pharmaceutically acceptable excipient(s), e.g., as described below.
[0183] The term “excipient” is used herein to describe any ingredient other than the compound(s) of the present disclosure. The choice of excipient(s) will to a large extent depend on factors such as the particular mode of administration, the effect of the excipient on solubility and stability, and the nature of the dosage form. As used herein, “pharmaceutically acceptable excipient” includes any and all solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic and absorption delaying agents, and the like that are physiologically compatible. Some examples of pharmaceutically acceptable excipients are water, saline, phosphate buffered saline, dextrose, glycerol, ethanol and the like, as well as combinations thereof. In many cases, it will be preferable to include isotonic agents, for example, sugars, polyalcohols such as mannitol, sorbitol, or sodium chloride in the composition. Additional examples of pharmaceutically acceptable substances are wetting agents or minor amounts of auxiliary substances such as wetting or emulsifying agents, preservatives, or buffers, which enhance the shelf life or effectiveness of the antibody.
[0184] Formulations of a pharmaceutical composition suitable for parenteral administration typically comprise the active ingredient combined with a pharmaceutically acceptable carrier, such as sterile water or sterile isotonic saline. Such formulations may be prepared, packaged, or sold in a form suitable for bolus administration or for continuous administration.
[0185] VII. Delivery Methods
[0186] In some embodiments, the epigenetic editor or its component(s) are introduced to target cells in the form of nucleic acid molecule(s) encoding the epigenetic editor or its component(s); accordingly, the pharmaceutical compositions herein comprise the nucleic acid molecule(s). Such nucleic acid molecule(s) may be, for example, DNA, RNA or mRNA, and / or modified nucleic acid sequence(s) (e.g., with chemical modifications, a 5’ cap, or one or more 3’ modifications). In some embodiments, the nucleic acid molecule(s) may be delivered as naked DNA or RNA, for instance by means of transfection or electroporation, or can be conjugated to molecules (e.g., N-acetylgalactosamine) promoting uptake by target cells. In some embodiments, the nucleic acid molecule(s) may be in nucleic acid expression vector(s), which may include expression control sequences such as promoters, enhancers, transcription signal sequences, transcription termination sequences, introns, polyadenylation signals, Kozak consensus sequences, internal ribosome entry sites (IRES), etc. Such expression control sequences are well known in the art. A vector may also comprise a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization), associated with (e.g., inserted into or fused to) a sequence coding for a protein.
[0187] Examples of vectors include, but are not limited to, plasmid vectors; viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., Murine Leukemia Virus, or spleen necrosis virus, vectors derived from retroviruses such as Rous Sarcoma Virus, Harvey Sarcoma Virus, avian leukosis virus, a lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus); and other recombinant vectors. In certain embodiments, the vector is a plasmid or a viral vector. Viral particles or virus-like particles (VLPs) may also be used to deliver nucleic acid molecule(s) encoding epigenetic editors or component(s) thereof as described herein. For example, “empty” viral particles can be assembled to contain any suitable cargo. Viral vectors and viral particles may also be engineered to incorporate targeting ligands to alter target tissue specificity.
[0188] In certain embodiments, an epigenetic editor as described herein or component(s) thereof are encoded by nucleic acid sequence(s) present in one or more viral vectors, or a suitable capsid protein of any viral vector. Examples of viral vectors include adeno- associated viral vectors (e.g., derived from AAV3, AAV3b, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh8, AAV10, and / or variants thereof); retroviral vectors (e.g., Maloney murine leukemia virus, MML-V), adenoviral vectors (e.g., AD100), lentiviral vectors (e.g., HIV and FIV-based vectors), and herpesvirus vectors (e.g., HSV-2).
[0189] In some embodiments, delivery involves an adeno-associated virus (AAV) vector. AAV vector delivery may be particularly useful where the DNA-binding domain of an epigenetic editor fusion protein is a zinc finger array. Without wishing to be bound by any theory, the smaller size of zinc finger arrays compared to larger DNA-binding domains such as Cas protein domains may allow such a fusion protein to be conveniently packed in viral vectors such as an AAV vector.
[0190] Any AAV serotype, e.g., human AAV serotype, can be used for an AAV vector as described herein, including, but not limited to, AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10), and AAV serotype 11 (AAV11), as well as variants thereof. In some embodiments, an AAV variant has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to a wildtype AAV. In certain embodiments, the AAV variant may be engineered such that its capsid proteins have reduced immunogenicity or enhanced transduction ability in humans. In some instances, one or more regions of at least two different AAV serotype viruses are shuffled and reassembled to generate a chimeric variant. For example, a chimeric AAV may comprise inverted terminal repeats (ITRs) that are of a heterologous serotype compared to the serotype of the capsid. The resulting chimeric AAV can have a different antigenic reactivity or recognition compared to its parental serotypes. In some embodiments, a chimeric variant of an AAV includes amino acid sequences from 2, 3, 4, 5, or more different AAV serotypes.
[0191] Non-viral systems are also contemplated for delivery as described herein. Non-viral systems include, but are not limited to, nucleic acid transfection methods including electroporation, sonoporation, calcium phosphate transfection, microinjection, DNA biolistics, lipid-mediated transfection, transfection through heat shock, compacted DNA- mediated transfection, lipofection, cationic agent-mediated transfection, and transfection with liposomes, immunoliposomes, exosomes, or cationic facial amphiphiles (CFAs). In certain embodiments, one or more mRNAs encoding epigenetic editor fusion proteins as described herein may be co-electroporated with one or more guide polynucleotides (e.g., gRNAs) as described herein. One important category of non-viral nucleic acid vectors is nanoparticles, which can be organic (e.g., lipid) or inorganic (e.g., gold). For instance, organic (e.g. lipid and / or polymer) nanoparticles can be suitable for use as delivery vehicles in certain embodiments of this disclosure.
[0192] In some embodiments, delivery is accomplished using a lipid nanoparticle (LNP). LNP compositions are typically sized on the order of micrometers or smaller and may include a lipid bilayer. In some embodiments, a LNP refers to any particle that has a diameter of less than 1000 nm, 500 nm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, or 25 nm. In some embodiments, a nanoparticle may range in size from 1-1000 nm, 1-500 nm, 1-250 nm, 25- 200 nm, 25-100 nm, 35-75 nm, or 25-60 nm. Nanoparticle compositions encompass lipid nanoparticles (LNPs), liposomes (e.g., lipid vesicles), and lipoplexes.
[0193] An LNP as described herein may be made from cationic, anionic, or neutral lipids. In some embodiments, an LNP may comprise neutral lipids, such as the fusogenic phospholipid l,2-Dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) or the membrane component cholesterol, as helper lipids to enhance transfection activity and nanoparticle stability. In some embodiments, an LNP may comprise hydrophobic lipids, hydrophilic lipids, or both hydrophobic and hydrophilic lipids. Any lipid or combination of lipids that are known in the art can be used to produce an LNP. The lipids may be combined in any molar ratios to produce the LNP. In some embodiments, the LNP is a liver-targeting (e.g., preferentially or specifically targeting the liver) LNP.
[0194] LNP formulations and methods of LNP delivery that can be used will be apparent to those skilled in the art from the present disclosure and the state of the art. Non-limiting exemplary compositions and methods can be found in Shah, R., Eldridge, D., Palombo, E., and Harding, I., Lipid Nanoparticles: Production, Characterization and Stability, Springer, 2015, ISBN- 13 978-3319107103; Mitchell, M.J., Billingsley, M.M., Haley, R.M. et al. Engineering precision nanoparticles for drug delivery, Nat Rev Drug Discov 20, 101-124 (2021); Hou, X., Zaks, T., Langer, R. et al. Lipid nanoparticles for mRNA delivery. Nat Rev Mater 6, 1078-1094 (2021); Lipid-Nanoparticle-Based Delivery of CR1SPR / Cas9 Genome- Editing Components, Pardis Kazemian, Si- Yue Yu, Sarah B. Thomson, Alexandra Birkenshaw, Blair R. Leavitt, and Colin J. D. Ross. Molecular Pharmaceutics 2022 19 (6), 1669-1686; Cullis PR, Hope MJ. Lipid Nanoparticle Systems for Enabling Gene Therapies, Mol Ther. 2017 Jul 5;25(7):1467-1475; Hatit, M.Z.C., Lokugamage, M.P., Dobrowolski, C.N. et al. Species-dependent in vivo mRNA delivery and cellular responses to nanoparticles, Nat. Nanotechnol. 17, 310-318 (2022); Lam, K., Schreiner, P., Leung, A., Stainton, P., Reid, S., Yaworski, E., Lutwyche, P. and Heyes, J. (2023), Optimizing Lipid Nanoparticles for Delivery in Primates, Adv. Mater; Dilliard, S.A., Siegwart, D.J. Passive, active and endogenous organ-targeted lipid and polymer nanoparticles for delivery of genetic drugs, Nat Rev Mater (2023); Kasiewicz, L.N., et.al., Lipid nanoparticles incorporating a GalNAc ligand enable in vivo liver AN GPTL3 editing in wild-type and somatic LDLR knockout non-human primates, bioRxiv 2021.11.08.467731, doi: https: / / doi.org / 10.1101 / 2021.l l.08.467731; Tombacz, I., et.al., Highly efficient CD4+ T cell targeting and genetic recombination using engineered CD4+ cell-homing mRNA-LNPs, Molecular Therapy, Volume 29, Issue 11, 2021, 3293-3304; Cheng, Q., Wei, T., Farbiak, L. et al. Selective organ targeting (SORT) nanoparticles for tissue-specific mRNA delivery and CRISPR-Cas gene editing, Nat. Nanotechnol. 15, 313-320 (2020); Zhang, Y., et.al., Lipids and Lipid Derivatives for RNA Delivery, Chemical
[0195] Reviews 2021 121 (20); Lam, K., et.al, Unsaturated, Trialkyl Ionizable Lipids are Versatile Lipid-N anoparticle Components for Therapeutic and Vaccine Applications, Adv.
[0196] Mater. 2023, 35; Han, X., Zhang, H., Butowska, K. et al. An ionizable lipid toolbox for RNA delivery, Nat Commun 12, 7233 (2021); US Patent No. 9,364,435; US Patent No. 8,058,069; US Patent No. 8,822,66; US Patent No. 8,492,359; US Patent No. 11,141,378; US Patent No. 9,518,272; US Patent No. 9,404,127; US Patent No. 9,006,417; US Patent No. 7,901,708; US Patent No. 9,005,654; US Patent No. 9,878,042; US Patent No. 9,682,139; US Patent No. 8,642,076; US Patent No. 9,593,077; US Patent No. 9,415,109; US Patent No. 9,701,623; US Patent No. 10,369,226; US Patent No. 9,999,673; US Patent No. 9,301,923; US Patent No. 10,342,761; US Patent No. 10,137,201; International Publication No. WO2015199952A1; International Publication No. WO2017075531 Al; International Publication No.
[0197] W02018081480A1; International Publication No. W02016081029A1; European Application No. EP3852911A2; each of which are incorporated herein by reference in their entirety.
[0198] Other methods of delivery to target cells will be known to those skilled in the art and can be used with the compositions of the present disclosure.
[0199] Any type of cell may be targeted for delivery of an epigenetic editor or component(s) thereof as described herein. For example, the cells may be eukaryotic or prokaryotic. In some embodiments, the cells are mammalian (e.g., human) cells. Human cells may include, for example, hepatocytes, biliary epithelial cells (cholangiocytes), stellate cells, Kupffer cells, and liver sinusoidal endothelial cells.
[0200] In some embodiments, an epigenetic editor described herein, or component(s) thereof, are delivered to a host cell for transient expression, e.g., via a transient expression vector. Transient expression of the epigenetic editor or its component(s) may result in prolonged or permanent epigenetic modification of the target gene. For example, the epigenetic modification may be stable for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. 11, or 12 weeks or more; or 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months or more, after introduction of the epigenetic editor into the host cell. The epigenetic modification may be maintained after one or more mitotic and / or meiotic events of the host cell. In particular embodiments, the epigenetic modification is maintained across generations in offspring generated or derived from the host cell. VIII. Therapeutic Uses of Epigenetic Editors
[0201] The present disclosure also provides methods for treating or preventing a condition in a subject, comprising administering to the subject an epigenetic editor or pharmaceutical composition as described herein. The epigenetic editor may effect an epigenetic modification of a target polynucleotide sequence in a target gene associated with a disease, condition, or disorder in the subject, thereby modulating expression of the target gene to treat or prevent the disease, condition, or disorder. In some embodiments, the epigenetic editor reduces the expression of the target gene to an extent sufficient to achieve a desired effect, e.g., a therapeutically relevant effect such as the prevention or treatment of the disease, condition, or disorder.
[0202] “Treat”, “treating” and “treatment” refer to a method of alleviating or abrogating a biological disorder and / or at least one of its attendant symptoms. As used herein, to “alleviate” a disease, disorder or condition means reducing the severity and / or occurrence frequency of the symptoms of the disease, disorder, or condition. Further, references herein to “treatment” include references to curative, palliative and prophylactic treatment. In some embodiments, as compared with an equivalent untreated control, alleviating a symptom may involve reduction of the symptom by at least 3%, 5%, 10%, 20%, 40%, 50%, 60%, 80%, 90%, 95%, 98%, 99%, 99.5%, 99.9%, or 100% as measured by any standard technique.
[0203] In some embodiments, the subject may be a mammal, e.g., a human. In some embodiments, the subject is selected from a non-human primate such as chimpanzee, cynomolgus monkey, or macaque, and other ape and monkey species.
[0204] An epigenetic editor of the present disclosure may be administered in a therapeutically effective amount to a subject with a condition described herein. “Therapeutically effective amount,” as used herein, refers to an amount of the therapeutic agent being administered that will relieve to some extent one or more of the symptoms of the disorder being treated, and / or result in clinical endpoint(s) desired by healthcare professionals. An effective amount for therapy may be measured by its ability to stabilize disease progression and / or ameliorate symptoms in a subject, and preferably to reverse disease progression. The ability of an epigenetic editor of the present disclosure to reduce or silence gene expression may be evaluated by in vitro assays, e.g., as described herein, as well as in suitable animal models that are predictive of the efficacy in humans. Suitable dosage regimens will be selected in order to provide an optimum therapeutic response in each particular situation, for example, administered as a single bolus or as a continuous infusion, and with possible adjustment of the dosage as indicated by the exigencies of each case. An epigenetic editor of the present disclosure may be administered without additional therapeutic treatments, i.e., as a stand-alone therapy (monotherapy). Alternatively, treatment with an epigenetic editor of the present disclosure may include at least one additional therapeutic treatment (combination therapy).
[0205] The epigenetic editors or components thereof (or nucleic acid molecules encoding the epigenetic editors or components thereof) of the present disclosure may be administered by any method accepted in the art, e.g., subcutaneously, intradermally, intratumorally, intranodally, intramuscularly, intravenously, intralymphatically, or intraperitoneally. In particular embodiments, a pharmaceutical composition of the present disclosure is administered intravenously to the subject.
[0206] IV. Definitions
[0207] The term “nucleic acid” as used herein refers to any oligonucleotide or polynucleotide containing nucleotides (e.g., deoxyribonucleotides or ribonucleotides) in either single- or double-strand form, and includes DNA and RNA. “Nucleotides” contain a sugar deoxyribose (DNA) or ribose (RNA), a base, and a phosphate group, and are linked together through the phosphate groups. “Bases” include purines and pyrimidines, which include natural compounds such as adenine, thymine, guanine, cytosine, uracil, inosine, and natural analogs; as well as synthetic derivatives of purines and pyrimidines, which include, but are not limited to, modified versions which place new reactive groups such as amines, alcohols, thiols, carboxylates, alkylhalides, etc. Nucleic acids may contain known nucleotide analogs and / or modified backbone residues or linkages, which may be synthetic, naturally occurring, and non-naturally occurring. Such nucleotide analogs, modified residues, and modified linkages are well known in the art, and may provide a nucleic acid molecule with enhanced cellular uptake, reduced immunogenicity, and / or increased stability in the presence of nucleases.
[0208] As used herein, an “isolated” or “purified” nucleic acid molecule is a nucleic acid molecule that exists apart from its native environment. For example, an “isolated” or “purified” nucleic acid molecule (1) has been separated away from the nucleic acids of the genomic DNA or cellular RNA of its source of origin; and / or (2) does not occur in nature. In some embodiments, an “isolated” or “purified” nucleic acid molecule is a recombinant nucleic acid molecule.
[0209] It will be understood that in addition to the specific proteins and nucleic acid molecules mentioned herein, the present disclosure also contemplates the use of variants, derivatives, homologs, and fragments thereof. A variant of any given sequence may have the specific sequence of residues (whether amino acid or nucleic acid residues) modified in such a manner that the polypeptide or polynucleotide in question substantially retains at least one of its endogenous functions. A variant sequence can be obtained by addition, deletion, substitution, modification, replacement and / or variation of at least one residue present in the naturally occurring sequence (in some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 residues). For specific proteins described herein (e.g., KRAB, dCas9, DNMT3A, and DNMT3L proteins described herein), the present disclosure also contemplates any of the protein’s naturally occurring forms, or variants or homologs that retain at least one of its endogenous functions (e.g., at least 50%, 60%, 70%, 80%, 90%, 85%, 96%, 97%, 98%, or 99% of its function as compared to the specific protein described).
[0210] Some exemplary fusion proteins embraced by the present disclosure are provided herein. It will be appreciated by the skilled artisan, that these exemplary proteins are nonlimiting examples and that additional proteins are within the scope of the present disclosure. For example, where fusion exemplary proteins comprising a specific domain, e.g., a mammalian DNMT3L and / or KRAB domain, such as a human or mouse DNMT3L and / or KRAB domain, are provided, the skilled artisan will be able to ascertain that, in some embodiments, fusion proteins with the same configuration, but with one or more of the mammalian domains substituted for a homologous domain from another mammal, e.g., one or more mouse domains substituted for one or more human domains, are also embraced by the present disclosure. For example, where an exemplary fusion protein is provided that comprises a mouse DNMT3L domain, a fusion protein of the same architecture but with the mouse DNMT3L substituted for a human DNMT3L domain is also embraced.
[0211] As used herein, a homologue of any polypeptide or nucleic acid sequence contemplated herein includes sequences having a certain homology with the wildtype amino acid and nucleic sequence. A homologous sequence may include a sequence, e.g. an amino acid sequence which may be at least 50%, 55%, 65%, 75%, 85%, 90%, 91%, 92%< 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the subject sequence. The term “percent identical” in the context of amino acid or nucleotide sequences refers to the percent of residues in two sequences that are the same when aligned for maximum correspondence. In some embodiments, the length of a reference sequence aligned for comparison purposes is at least 30%, (e.g., at least 40, 50, 60, 70, 80, or 90%, or 100%) of the reference sequence. Sequence identity may be measured using sequence analysis software (for example, Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning degrees of homology to various substitutions, deletions, and / or other modifications. In an exemplary approach to determining the degree of identity, a BLAST program may be used, with a probability score between e-3 and e-100 indicating a closely related sequence.
[0212] The percent identity of two nucleotide or polypeptide sequences is determined by, e.g., BLAST® using default parameters (available at the U.S. National Library of Medicine’s National Center for Biotechnology Information website). In some embodiments, the length of a reference sequence aligned for comparison purposes is at least 30%, (e.g., at least 40, 50, 60, 70, 80, or 90%) of the reference sequence.
[0213] It will be understood that the numbering of the specific positions or residues in polypeptide sequences depends on the particular protein and numbering scheme used. Numbering might be different, e.g., in precursors of a mature protein and the mature protein itself, and differences in sequences from species to species may affect numbering. One of skill in the art will be able to identify the respective residue in any homologous protein and in the respective encoding nucleic acid by methods well known in the art, e.g., by sequence alignment and determination of homologous residues.
[0214] The term “modulate” or “alter” refers to a change in the quantity, degree, or extent of a function. Lor example, an epigenetic editor as described herein may modulate the activity of a promoter sequence by binding to a motif within the promoter, thereby inducing, enhancing, or suppressing transcription of a gene operatively linked to the promoter sequence. As other examples, an epigenetic editor as described herein may block RNA polymerase from transcribing a gene, or may inhibit translation of an mRNA transcript. The terms “inhibit,” “repress,” “suppress,” “silence” and the like, when used in reference to an epigenetic editor or a component thereof as described herein, refers to decreasing or preventing the activity (e.g., transcription) of a nucleic acid sequence (e.g., a target gene) or protein relative to the activity of the nucleic acid sequence or protein in the absence of the epigenetic editor or component thereof. The term may include partially or totally blocking activity, or preventing or delaying activity. The inhibited activity may be, e.g., 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% less than that of a control, or may be, e.g., at least 1.5-fold, 2-fold, 3-fold, 4- fold, 5-fold, or 10-fold less than that of a control.
[0215] The term “about” or “approximately” means within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within one or more than one standard deviation, per the practice in the given value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” should be assumed to mean an acceptable error range for the particular value.
[0216] Ranges provided herein are understood to be shorthand for all of the values within the range. For example, a range of 1 to 50 is understood to include any number, combination of numbers, or sub-range from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50, as well as all intervening decimal values between the aforementioned integers such as, for example, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, and 1.9. With respect to sub-ranges, “nested sub-ranges” that extend from either end point of the range are specifically contemplated. For example, a nested sub-range of an exemplary range of 1 to 50 may comprise 1 to 10, 1 to 20, 1 to 30, and 1 to 40 in one direction, or 50 to 40, 50 to 30, 50 to 20, and 50 to 10 in the other direction.
[0217] Unless otherwise defined herein, scientific and technical terms used in connection with the present disclosure shall have the meanings that are commonly understood by those of ordinary skill in the art. Exemplary methods and materials are described below, although methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure. In case of conflict, the present specification, including definitions, will control. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. Throughout this specification and embodiments, the words “have” and “comprise,” or variations such as “has,” “having,” “comprises,” or “comprising,” will be understood to imply the inclusion of a stated integer or group of integers but not the exclusion of any other integer or group of integers. Unless otherwise indicated, the recitation of a listing of elements herein includes any of the elements singly or in any combination. The recitation of an embodiment herein includes that embodiment as a single embodiment, or in combination with any other embodiment(s) herein. All publications, patents, patent applications, and other references mentioned herein are incorporated by reference in their entirety. To the extent that references incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material. Although a number of documents are cited herein, this citation does not constitute an admission that any of these documents forms part of the common general knowledge in the art.
[0218] According to the present disclosure, back-references in the dependent claims are meant as short-hand writing for a direct and unambiguous disclosure of each and every combination of claims that is indicated by the back-reference. Further, headers herein are created for ease of organization and are not intended to limit the scope of the claimed methods and compositions in any manner.
[0219] Some protein sequences, e.g., some fusion protein sequences, provided herein include a peptide tag, e.g., a His6 tag, or a DYKDDDDK (SEQ ID NO: 1114) tag, which are useful for detection and / or purification of tagged proteins, but do not affect protein function. These peptide tags, and additional suitable peptide tags, are well known to those of skill in the art. It will be apparent to the person of skill in the art that the disclosed tags can be substituted for other suitable peptide tags, and that fusion proteins of the same or highly similar sequence, but not including such peptide tags, e.g., from which the peptide tag has been cleaved or which are created without a peptide tag, are suitable for carrying out embodiments of the present disclosure as well.
[0220] In order that the present disclosure may be better understood, the following examples are set forth. These examples are for purposes of illustration only and are not to be construed as limiting the scope of the present disclosure in any manner.
[0221] SEQUENCES
[0222] The SEQ ID NOs (SEQ) of nucleotide (nt) and amino acid (aa) sequences described in the present disclosure are listed below.
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
Claims
CLAIMSWhat is claimed is:
1. A system for activating transcription of a human F8 gene in a human cell, comprising a) one or more fusion proteins that collectively comprise a transcriptional activating domain, or an effector domain that catalyzes the removal of DNA methylation each domain being linked to a DNA-binding domain that binds to a target region in the human F8 promoter or downstream promoter sequence, or b) one or more nucleic acid molecules encoding the one or more fusion proteins, wherein the system does not generate a DNA break in the F8 gene.
2. The system of claim 1, wherein the DNA- binding domain comprises a dead CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain.
3. The system of claim 2, wherein the DNA-binding domain comprises a dCas9 domain and the system further comprises (i) one or more guide RNAs comprising a sequence that binds to any one of the sequences in Table 1 or Table 2, or (ii) nucleic acid molecules coding for the one or more guide RNAs.
4. A human cell comprising the system of any one of claims 1-3, or progeny of the cell.
5. A human cell modified by the system of any one of claims 1-3, or progeny of the cell.
6. A pharmaceutical composition comprising the system of any one of claims 1-3 and a pharmaceutically acceptable excipient.
7. A pharmaceutical composition comprising human cells of claim 4 or 5 and a pharmaceutically acceptable excipient.
8. A method of treating a subject in need thereof, comprising administering the system of any one of claims 1-3, human cells of claim 4 or 5, or the pharmaceutical composition of claim 6 or 7 to the subject.
9. The method of claim 8, wherein the subject has or is suspected of having hemophilia.
10. An epigenetic ally modified human cell comprising an F8 genomic locus, wherein the epigenetically modified human cell is of a cell type that, in its naturally-occurring (wildtype) state, is characterized by expression of F8, wherein the F8 genomic locus comprises two alleles, and wherein both alleles of the F8 genomic locus comprise a naturally occurring genomic sequence associated with expression of functional F8, and wherein the epigenetically modified human cell is characterized by active or increased expression (as compared to a wildtype cell of the same cell type) of F8.
Citation Information
Patent Citations
Promoter for cell-specific gene expression and uses thereof
US20190269796A1
Gene editing using a modified closed-ended DNA (CEDNA)
US20220290186A1
Compositions and methods for epigenetic editing
WO2022140577A2