Compositions and methods for epigenetic regulation of B2M expression

Through the epigenetic editing system, fusion proteins or nucleic acid molecules of DNMT and transcriptional repressor domains are used to repress the transcription of human B2M genes, solving the risk problems of traditional genetic engineering methods and achieving safe and effective gene regulation.

CN119948164APending Publication Date: 2025-05-06CHROMA MEDICINE INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202380061420.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2023-06-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional genetic engineering strategies perform permanent cell manipulation at the genomic level, with risks such as chromosomal translocation, undesired nucleotide insertions and deletions at the target site, and off-target mutations, requiring a safe and effective method of genetically engineered immune cells.

Method used

An epigenetic editing system is provided, which can be used to repress the transcription of human B2M genes by fusion proteins or nucleic acid molecules, including DNA methyltransferase (DNMT) domains and transcription repressor domains.

Benefits of technology

The system is able to safely and effectively regulate the expression of B2M genes, reduce allopathic reactivity, reduce the risk of permanent genome manipulation, and achieve lasting heritable silencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948164A_ABST
    Figure CN119948164A_ABST
Patent Text Reader

Abstract

Disclosed herein are compositions and methods comprising an epigenetic editor for epigenetic modification of B2M, as well as nucleic acids and vectors encoding them. Also disclosed are cells epigenetically modified by the epigenetic editor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 355,061, filed on June 23, 2023, entitled “Compositions and Methods for Epigenetically Regulating B2M Expression,” under 35 U.S.C. §119(e), the entire disclosure of which is incorporated herein by reference in its entirety.

[0003] References to Electronic Sequence Listings

[0004] The contents of the electronic Sequence Listing (C169870008WO00-SEQ-AXW.xml; size: 1,879,683 bytes; and creation date: June 23, 2023) are incorporated herein by reference in its entirety. Background Art

[0005] Adoptive cell therapy using genetically engineered immune cells has emerged as a promising approach for treating cancer, infection, autoimmune diseases, and other conditions. However, conventional genetic engineering strategies typically rely on permanent manipulation of cells at the genomic level, which is associated with certain risks, including, for example, chromosomal translocations, insertions and deletions of undesired nucleotides at target sites, and off-target mutations. There remains a need for efficient and safe methods for genetically engineering immune cells. SUMMARY OF THE INVENTION

[0007] The present disclosure provides systems and compositions for epigenetic modification (herein "epigenetic editors" or "epigenetic editing systems"), and methods of using the same to generate epigenetic modifications at B2M, including in host cells and organisms.

[0008] In some aspects, the present disclosure provides a system for suppressing transcription of a human B2M gene in a human cell, optionally a human T lymphocyte or a human NK cell, comprising

[0009] a) one or more fusion proteins, which together comprise

[0010] a DNA methyltransferase (DNMT) domain and / or a domain that recruits DNMTs, optionally wherein the DNMT domain and / or the recruiter domain comprises a DNMT3A domain and / or a DNMT3L domain, and optionally wherein the recruited DNMT is DNMT3A, and

[0011] Transcriptional repressor domain,

[0012] Each domain is linked to a DNA binding domain that binds to a target region in the human B2M gene, wherein the target region comprises one or more sequences selected from the group consisting of: SEQ ID NO: 700-740, 744, 747-749, 752, 753, 757, 758, 760-806, 812-822, 825, 827, 830, 833, 834, 839-841, 843-845, 849, 851-853, 855, 864, 866-877, 879-883 , 891-896, 898-900, 903-914, 922, 923, 925-927, 934, 936, 943-947, 949, 951-962, 975-981, 983, 985, 987-989, 995, 997-999, 1003-1005, and 1007-1011; or

[0013] b) one or more nucleic acid molecules encoding one or more fusion proteins,

[0014] Optionally wherein the system does not generate DNA breaks in the B2M gene.

[0015] In some embodiments, the DNA binding domain comprises a dead CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain. For example, the DNA binding domain can comprise a dCas9 domain, and the system can further comprise (i) one or more guide RNAs (e.g., comprising SEQ ID NOs: 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 1162-1167, 1169-1170, 1181-1183, 1184-1185, 1186-1187, 1188-1189, 1190-1191, 1201-1202, 1203-1204, 1205-1206, 1207-1208 171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276 and 1278-1282), or (ii) a nucleic acid molecule encoding one or more guide RNAs.

[0016] In some embodiments, the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) two guide RNAs comprising SEQ ID NOs: 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 1162-1167, 1169 -1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276 and 1278-1282, or (ii) a nucleic acid molecule encoding two guide RNAs.

[0017] In some embodiments, the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) three guide RNAs comprising SEQ ID NOs: 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 1162-1167, 1169 -1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276 and 1278-1282, or (ii) a nucleic acid molecule encoding three guide RNAs.

[0018] In some aspects, the present disclosure provides a system for suppressing transcription of a human B2M gene in a human cell, optionally a human T lymphocyte or a human NK cell, comprising

[0019] a) A fusion protein comprising:

[0020] DNMT3A domain,

[0021] DNMT3L domain,

[0022] DNA binding domain, and

[0023] Transcriptional repressor domain, or

[0024] b) a nucleic acid molecule encoding a fusion protein,

[0025] Optionally wherein the system does not generate DNA breaks in the B2M gene.

[0026] In some embodiments, the DNA binding domain comprises a dead CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain. For example, the DNA binding domain can comprise a dCas9 domain, and the system can further comprise (i) one or more guide RNAs (e.g., comprising any one of SEQ ID NOs: 1012-1282), or (ii) a nucleic acid molecule encoding one or more guide RNAs.

[0027] In certain embodiments, the dCas domain comprises a dCas9 sequence, such as a sequence that is at least 90% identical to SEQ ID NO: 12 or 13.

[0028] In some embodiments, the DNA binding domain binds to a target sequence in SEQ ID NO: 1283 or 1284.

[0029] In some embodiments, the DNA binding domain comprises a ZFP domain that targets a nucleotide sequence selected from SEQ ID NOs: 700-740.

[0030] In some embodiments, the DNMT3A domain comprises a sequence that is at least 90% identical to SEQ ID NO:574 or 575.

[0031] The DNMT3L domain can comprise, for example, a sequence having at least 90% identity to a sequence selected from SEQ ID NOs: 578-581. In some embodiments, the DNMT3L domain comprises a sequence having at least 90% identity to a sequence selected from SEQ ID NOs: 582-603. In some embodiments, the DNMT3L domain comprises a sequence having at least 90% identity to a sequence selected from SEQ ID NOs: 601-603.

[0032] In some embodiments, the transcriptional repressor domain comprises a sequence having at least 90% identity to a sequence selected from SEQ ID NO: 33-570. In certain embodiments, the transcriptional repressor domain is a KRAB domain derived from KOX1, ZIM3, ZFP28 or ZN627. The KRAB domain may comprise, for example, a sequence having at least 90% identity to a sequence selected from SEQ ID NO: 89, 116, 245 and 255. In some embodiments, the transcriptional repressor domain comprises a fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1KRAB, and optionally comprises an amino acid sequence of SEQ ID NO: 571 or 572. In certain embodiments, the transcriptional repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1 or SCML2.

[0033] In some embodiments, the system comprises

[0034] a) a fusion protein comprising a DNMT3A domain, a DNMT3L domain, a transcription repressor domain and a DNA binding domain,

[0035] Optionally, wherein one or both of the DNMT3A domain and the DNMT3L domain are human, and

[0036] Optionally, wherein the DNA binding domain is a death CRISPR Cas domain or a ZFP domain; or

[0037] b) A nucleic acid molecule encoding a fusion protein.

[0038] In certain embodiments, the fusion protein comprises a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA binding domain, a third peptide linker, and a transcription repressor domain from the N-terminus to the C-terminus. For example, the fusion protein may comprise a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a first nuclear localization signal (NLS), a DNA binding domain, a second NLS, a third peptide linker, and a transcription repressor domain from the N-terminus to the C-terminus. The fusion protein may comprise a first NLS, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA binding domain, a third peptide linker, a transcription repressor domain, and a second NLS from the N-terminus to the C-terminus. The fusion protein may comprise a first and a second NLS, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA binding domain, a third peptide linker, a transcription repressor domain, and a third and a fourth NLS from the N-terminus to the C-terminus. In a specific embodiment, the transcriptional repressor domain is a KRAB domain, such as human KOX1, ZFP28, ZN627 or ZIM3 KRAB domain. In a specific embodiment, one or both of the second and third peptide linkers are XTEN linkers, which can be selected from XTEN80 (e.g., SEQ ID NO: 643) and XTEN16 (e.g., SEQ ID NO: 638), for example, wherein the second peptide linker is XTEN80, and the third peptide linker is XTEN16.

[0039] In some embodiments, the fusion protein can comprise a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a dSpCas9 domain, a second NLS, an XTEN16 peptide linker, and a human KOX1 KRAB domain from N-terminus to C-terminus. In certain embodiments, the fusion protein comprises SEQ ID NO: 658 or a sequence at least 90% identical thereto.

[0040] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a ZFP domain, a second NLS, an XTEN16 linker, and a human KOX1 KRAB domain. In certain embodiments, the fusion protein comprises SEQ ID NO: 659 or a sequence at least 90% identical thereto.

[0041] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and a third and a fourth NLS. In a specific embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 660, or a sequence at least 90% identical thereto.

[0042] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and a third and a fourth NLS.

[0043] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZFP28KRAB domain, and a third and a fourth NLS. In a specific embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 661, or a sequence at least 90% identical thereto.

[0044] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZFP28 KRAB domain, and a third and a fourth NLS.

[0045] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZN627KRAB domain, and a third and a fourth NLS. In a specific embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 662, or a sequence at least 90% identical thereto.

[0046] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZN627 KRAB domain, and a third and a fourth NLS.

[0047] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and a third and a fourth NLS. In a specific embodiment, the fusion protein may comprise an amino acid sequence of SEQ ID NO: 663, or a sequence at least 90% identical thereto, or an amino acid sequence of SEQ ID NO: 667, or a sequence at least 90% identical thereto.

[0048] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and a third and a fourth NLS.

[0049] In some embodiments, at least one NLS in the fusion proteins described herein is a SV40 NLS (eg, SEQ ID NO: 644).

[0050] In some embodiments, the system comprises:

[0051] a) a first fusion protein comprising a first DNA binding domain and comprising or recruiting a DNMT3A domain,

[0052] a second fusion protein comprising a second DNA binding domain and comprising or recruiting a DNMT3L domain, and

[0053] a third fusion protein comprising a third DNA binding domain and comprising or recruiting a transcriptional repressor domain; or

[0054] b) one or more nucleic acid molecules encoding the fusion protein.

[0055] The present disclosure also provides a human cell comprising the system described herein, or a descendant of the cell. In some embodiments, the cell is a T lymphocyte or a NK cell.

[0056] The disclosure also provides a human cell modified (optionally ex vivo) by the system described herein, or a progeny of the cell. In some embodiments, the cell is a T lymphocyte or a NK cell.

[0057] The present disclosure also provides a pharmaceutical composition comprising a system described herein and a pharmaceutically acceptable excipient. In some embodiments, the composition comprises a lipid nanoparticle (LNP) comprising the system, and / or the DNA binding domain is a dCas domain, and the LNP further comprises one or more gRNAs.

[0058] The present disclosure also provides a pharmaceutical composition comprising the human cells described herein and a pharmaceutically acceptable excipient.

[0059] The present disclosure also provides methods of treating a patient in need thereof, comprising administering to the patient (eg, intravenously) a system, human cell, or pharmaceutical composition described herein. In some embodiments, the patient suffers from cancer or an autoimmune disease.

[0060] The present disclosure also provides the systems, human cells or pharmaceutical compositions described herein for use in treating a patient in need thereof, such as in the methods described herein.

[0061] The present disclosure also provides for use of a system or human cell as described herein in the preparation of a medicament for treating a patient in need thereof, such as in a method as described herein.

[0062] The present disclosure also provides articles of manufacture and kits comprising the systems or human cells described herein.

[0063] Other features, purposes and advantages of the present invention are apparent in the following detailed description. However, it should be understood that the detailed description, although showing embodiments and embodiments of the present invention, is only given in an illustrative rather than a limiting manner. According to the detailed description, various changes and modifications within the scope of the present invention will be apparent to those skilled in the art. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a scatter plot showing relative B2M expression (y-axis) in cells treated with the CRISPR-off epigenetic editing system (DNMT3A-DNMT3L-dCas9-KRAB). The distance from the gRNA target site to the B2M transcription start site (TSS) is shown on the x-axis. The best performing guide was selected from the screen (marked "yes"). Empty dots: not selected. Empty triangle shape: wild-type Cas9. Dark dots: selected.

[0065] Figure 2A-2C Flow cytometry plots of DON008 without gRNA are shown ( Figure 2A ) and flow cytometry of DON008 containing RNA102 and RNA964 ( Figure 2B ). Figure 2A and 2B The gating strategy for the flow cytometry workflow is shown. Figure 2C The B2M multiplex screen revealed 27 RNA-guide pairs with significant B2M silencing.

[0066] Figure 3A-3BHeatmaps indicating day 6 results are shown. The outer edge reports the distance between the gRNA binding sequence and the B2M TSS. The inner heatmap reports the percentage of B2M+ cells observed after treatment with the relevant gRNA.

[0067] Figure 4 Shown is B2M silencing of guide RNA pairs and single guides in human T cells over time. As shown, the top silenced pair from day 6 remains silent at day 20.

[0068] Figure 5A-5B Exon-level differences in B2M expression between WTCas9 and CRISPR-Off are shown. The results show that CRISPR-Off reduces B2M isoform / exon expression more robustly than WTCas9.

[0069] Figure 6A-6B Hybridization capture methylation analysis of B2M double-stranded CRISPR-Off is shown. Fig. 6A Summarizing the test conditions, Figure 6B Methylation observed upstream of the B2M locus is shown.

[0070] Figure 7A-7B Robust B2M CpG methylation in the sorted B2M-negative population is shown.

[0071] Figure 8 Shown is B2M CpG methylation achieved with the RNA138 / 949 duplex compared to no gRNA and RNA104 / 988.

[0072] Figures 9A-9B Shown is a comparison of B2M levels measured in fresh cells with various effectors.

[0073] Figures 10A-10C Shown is a comparison of B2M levels determined in frozen cells with various effectors.

[0074] Fig.11A -B is related to the B2M silencing efficiency under different serum conditions. Fig.11A Shown are the gating strategies used for flow cytometric analysis of the effects of 5% HuS versus 10% HuS on B2M silencing after nucleofection and the durability of B2M silencing. Fig. 11B The time course of B2M silencing is shown, demonstrating that serum percentage did not differ in silencing efficiency.

[0075] Figures 12A-12C The results for day 6 are shown, comparing the silencing efficiency of day 2 nucleofection with that of day 3 nucleofection ( Fig. 12A ), transduction efficiency of nuclear transfection on day 3 ( Fig. 12B) and day 6 silencing results obtained from CAR+ or CAR- cells ( Fig. 12C ).

[0076] Figures 13A-13B IDT gRNA batch comparison is shown. Fig.13A The gating strategy used to test three different batches of 2 B2M guides is shown. Fig. 13B Epigenetic silencing of B2M at day 7 is shown.

[0077] Fig.14A -B shows a graph reporting the results of a dose-response experiment using the dual B2M guide. Fig.14A Shown is the dose response of B2M silencing at day 6 after nucleofection. Fig. 14B Shown is the dose response of B2M silencing at day 13 after nucleofection.

[0078] Figures 15A-15B Shown are responses of allogeneic healthy donor CD8+ T cells to mock-modified or B2M-silenced or B2M-multitargeted T cells observed using a mixed lymphocyte co-culture assay.

[0079] Fig.16A -C shows B2M silencing using guide pairs and triplets. DETAILED DESCRIPTION OF THE INVENTION

[0081] The present disclosure provides an epigenetic editor for suppressing human B2M gene expression.By changing the expression of B2M, the editor herein can be used to generate allogeneic cells (e.g., T cells, NK cells, etc.) with reduced alloreactivity. Unless otherwise indicated, "B2M" (italics) refers to human B2M genes in this article.Human B2M gene sequences can be found in Ensembl accession number ENSG00000166710.Compared with other genome engineering methods, the epigenetic editor of the present invention has several advantages, including reversibility, reduced risk of chromosome translocation and persistent heritable silencing.

[0082] In some embodiments, the region of targeted epigenetic regulation in human B2M gene is about 2kb long and about + / -1kb of B2MTSS. In certain embodiments, the region has the nucleotide sequence of SEQ ID NO:1284 (as shown below). In some embodiments, the B2M region of targeting is about 1kb long and about + / -500bp of B2M TSS. In certain embodiments, the targeting region has the nucleotide sequence of SEQ ID NO:1283 (as shown below). B2M TSS is located at #chr15:55039548 of genome GRCh38.

[0083]

[0084]

[0085] In some embodiments, the length of the targeting site can be 10 to 50 bp (e.g., 10 to 40, 10 to 30, 10 to 20, 15 to 30, 15 to 25, or 15 to 20 bp). In some embodiments, the targeting strand in the targeting region is the sense strand of the gene. In other embodiments, the targeting strand in the targeting region is the antisense strand of the gene.

[0086] In some embodiments, an epigenetic editor as described herein may comprise one or more fusion proteins, wherein each fusion protein comprises a DNA binding domain connected to one or more effector domains for epigenetic modification. In certain embodiments, wherein the DNA binding domain is a DNA binding domain of a polynucleotide guide, the epigenetic editor may further comprise one or more guide polynucleotides. The DNA binding domain, effector domain, and guide polynucleotide of an epigenetic editor as described herein may be selected from, for example, those described below, in any functional combination.

[0087] The epigenetic editor described herein can be transiently expressed in a host cell, or can be integrated into the genome of a host cell; the present disclosure also contemplates such cells and their offspring. Transiently expressed and integrated epigenetic editors or their components can achieve stable epigenetic modifications. For example, after the epigenetic editor described herein is introduced into a host cell, the target gene in the host cell can be stably or permanently repressed or silenced. In some embodiments, compared to the expression level in the absence of an epigenetic editor, the expression of the target gene is reduced or silenced for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or the entire life cycle of a cell or a subject carrying a cell. Epigenetic modifications can be inherited by the offspring of the host cell into which the epigenetic editor is introduced.

[0088] I. DNA binding domain

[0089] Epigenetic editors described herein may include one or more DNA binding domains that guide the effector domain of the epigenetic editor to a target sequence within or near the B2M gene locus. DNA binding domains as described herein may be, for example, DNA binding domains of polynucleotide guides, zinc finger proteins (ZFP) domains, transcription activator-like effector (TALE) domains, meganuclease DNA binding domains, etc. Examples of DNA binding domains can be found in U.S. Patent No. 11,162,114, which is incorporated herein by reference in its entirety.

[0090] In some embodiments, the DNA binding domain described herein is encoded by its native coding sequence.In other embodiments, the DNA binding domain is encoded by a nucleotide sequence that has been codon optimized for optimal expression in human cells.

[0091] A. DNA Binding Domain of Polynucleotide Guides

[0092] In some embodiments, the DNA binding domain herein can be a protein domain guided to a target site in a B2M gene locus by a guide nucleic acid sequence (e.g., a guide RNA sequence). In certain embodiments, the protein domain can be derived from a CRISPR-associated nuclease, such as a class I or class II CRISPR-associated nuclease. In some embodiments, the protein domain can be derived from a Cas nuclease, such as a type II, type IIA, type IIB, type IIC, type V, or type VI Cas nuclease. In certain embodiments, the protein domain can be derived from a class II Cas nuclease, selected from Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas14a, Cas14b, Cas14c, CasX, CasY, CasPhi, C2c4, C2c8, C2c9, C2c10, Csy1, Csy2, Csy3, Cse 1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4 and homologs and modified versions thereof. "Derived from" is used to mean that the protein domain comprises the complete polypeptide sequence of the parent protein, or comprises a variant thereof (e.g., having amino acid residue deletions, insertions and / or substitutions). The variant retains the desired function of the parent protein (e.g., the ability to form a complex with a guide nucleic acid sequence and a target DNA).

[0093] In some embodiments, the CRISPR-associated protein domain can be a Cas9 domain described herein. For example, Cas can refer to a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence similarity with a wild-type Cas9 polypeptide described herein. In some embodiments, the wild-type polypeptide is Cas9 from Streptococcus pyogenes (NCBI reference number NC_002737.2 (SEQ ID NO: 1)) and / or Cas9 of UniProt reference number Q99ZW2 (SEQ ID NO: 2). In some embodiments, the wild-type polypeptide is Cas9 from Staphylococcus aureus (SEQ ID NO: 3). In some embodiments, the CRISPR-associated protein domain is a Cpf1 domain or protein, or a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to a wild-type Cpf1 polypeptide described herein (e.g., Cpf1 from Francisella novicida (UniProt reference U2UMQ6 or SEQ ID NO:4). In certain embodiments, the CRISPR-associated protein domain can be a modified form of a wild-type protein comprising one or more amino acid residue changes, such as deletions, insertions, or substitutions; a fusion or chimera; or any combination thereof.

[0094] The structures of Cas9 sequences and variant Cas9 orthologs of various organisms have been described. Exemplary organisms from which the Cas9 domains herein may be derived include, but are not limited to, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus spp., Staphylococcus aureus, Listeria innocua, Lactobacillus carreii, Francisella novicida, Worilia succinica, Wardersartella, Gammaproteus, Neisseria meningitidis, Campylobacter jejuni, Pasteurella multocida, Bacillus succinica, Rhodospirilla rubrum, Nocardia dasonvillei, Streptomyces viridis, Streptomyces viridis, Streptococcus spp. Molds, pink chain cysts, acidocaloric acid Bacillus, pseudomycobacterium, selenium-reducing Bacillus, Siberian microbacterium, Lactobacillus delbrueckii, Lactobacillus salivarius, Lactobacillus buchneri, Treponema denticola, marine micro-shivering cyanobacteria, Burkholderia, naphthotrophic mononas, polar mononas, crocodile algae, cyanobacteria, aeruginosa, Synechococcus, Arabinus acetoaceticus, degenergen ammonia-producing bacteria, thermocellulosus, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Fingoldia grandis, Thermotolerant halophilic anaerobic bacteria, Pelotomaculum thermopropionium, Acidithiobacillus thermophilus, Acidithiobacillus ferrooxidans, Metachromatium viniferum, Oceanobacillus spp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas marine, Leptotrichia, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spp., Arthrospira maxima, Spirulina platensis, Arthrospira spp., Sphingomyelia spp., Microcoleus prototypically, Oscillatoria cyanobacteria, Petrotoga mobilis, Thermostomium africanum, Streptococcus pasteurianus, Neisseria griseus, Campylobacter salsa, Corynebacterium parvum, Corynebacterium diphtheriae, and Acaryochloris marina. Cas9 sequences also include those from the organisms and loci disclosed in Chylinski et al., RNA Biol. (2013) 10(5):726-37.

[0095] In some embodiments, the Cas9 domain is from Streptococcus pyogenes (spCas9). In some embodiments, the Cas9 domain is from Staphylococcus aureus (saCas9).

[0096] Other Cas domains are also contemplated for use in the epigenetic editors herein. These include, for example, CasX (Cas12E) (e.g., SEQ ID NO: 5), CasY (Cas12d) (e.g., SEQ ID NO: 6), (CasPhi) (e.g., SEQ ID NO: 7), Cas12f1 (Cas14a) (e.g., SEQ ID NO: 8), Cas12f2 (Cas14b) (e.g., SEQ ID NO: 9), Cas12f3 (Cas14c) (e.g., SEQ ID NO: 10), and C2c8 (e.g., SEQ ID NO: 11).

[0097] For epigenetic editing, a nuclease-derived protein domain (e.g., Cas9 or Cpf1 domain) can have reduced nuclease activity or no nuclease activity by mutation, so that the protein domain does not cut DNA or has reduced DNA cutting activity, while retaining the ability to complex with a guide nucleic acid sequence (e.g., guide RNA) and target DNA. For example, compared to a wild-type domain, the nuclease activity may be reduced by at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or 99%. In some embodiments, the CRISPR-associated protein domain described herein is catalytically inactive ("dead"). Examples of such domains include, for example, dCas9 ("dead" Cas9), dCpf1, ddCpf1, dCasPhi, ddCas12a, dLbCpf1, and dFnCpf1. For example, compared to wild-type Cas9, the dCas9 protein domain may include one, two or more mutations that eliminate its nuclease activity. It is known that the DNA cleavage domain of Cas9 includes two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cuts the chain complementary to the gRNA, while the RuvC1 subdomain cuts the non-complementary chain. Mutations in these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A (in RuvC1) and H840A (in HNH) completely inactivate the nuclease activity of SpCas9. Similarly, SaCas9 can be inactivated by mutations D10A and N580A. In some embodiments, dCas9 includes at least one mutation that reduces or eliminates nuclease activity in the HNH subdomain and / or the RuvC1 subdomain. In some embodiments, dCas9 includes only the RuvC1 subdomain or only the HNH subdomain. It should be understood that any mutation that inactivates the RuvC1 and / or HNH domains may be included in the dCas9 herein, for example, insertions, deletions, or single or multiple amino acid substitutions in the RuvC1 domain and / or the HNH domain.

[0098] In some embodiments, the dCas9 protein herein comprises a mutation at position D10 (e.g., D10A), H840 (e.g., H840A), or both corresponding to the wild-type SpCas9 sequence, as numbered in the sequence provided by UniProt Accession No. Q99ZW2 (SEQ ID NO: 2). In a specific embodiment, the dCas9 comprises the amino acid sequence of dSpCas9 (D10A and H840A) (SEQ ID NO: 12).

[0099] In some embodiments, a dCas9 protein as described herein comprises a mutation at position D10 (e.g., D10A), N580 (e.g., N580A), or both corresponding to a wild-type SaCas9 sequence (e.g., SEQ ID NO: 3). In a specific embodiment, dCas9 comprises the amino acid sequence of dSaCas9 (D10A and N580A) (SEQ ID NO: 13).

[0100] Based on the present disclosure and the knowledge in the art, additional suitable mutations that inactivate Cas9 will be apparent to those skilled in the art and are within the scope of the present disclosure. Such mutations may include, but are not limited to, D839A, N863A, and / or K603R in SpCas9. The present disclosure contemplates any mutation that reduces or eliminates the nuclease activity of any Cas9 described herein (e.g., a mutation corresponding to any Cas9 mutation described herein).

[0101] Compared to wild-type Cpf1, the dCpf1 protein domain may comprise one, two or more mutations that reduce or eliminate its nuclease activity. The Cpf1 protein has a RuvC-like endonuclease domain similar to the RuvC domain of Cas9 but without the HNH endonuclease domain, and the N-terminus of Cpf1 does not have the α-helical recognition lobe of Cas9. In some embodiments, dCpf1 comprises one or more mutations corresponding to positions D917A, E1006A or D1255A numbered in the sequence of Francisella novicida Cpf1 protein (FnCpf1; SEQ ID NO: 4). In certain embodiments, the dCpf1 protein comprises a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A, or a corresponding mutation of any one of the Cpf1 amino acid sequences described herein. In some embodiments, dCpf1 comprises the D917A mutation. In a specific embodiment, dCpf1 comprises the amino acid sequence of dFnCpf1 (SEQ ID NO: 14).

[0102] Further nuclease-inactive CRISPR-associated protein domains contemplated herein include those from, for example, dNmeCas9 (e.g., SEQ ID NO: 15), dCjCas9 (e.g., SEQ ID NO: 16), dSt1Cas9 (e.g., SEQ ID NO: 17), dSt3Cas9 (e.g., SEQ ID NO: 18), dLbCpf1 (e.g., SEQ ID NO: 19), dAsCpf1 (e.g., SEQ ID NO: 20), denAsCpf1 (e.g., SEQ ID NO: 21), dHFAsCpf1 (e.g., SEQ ID NO: 22), dRVRAsCpf1 (e.g., SEQ ID NO: 23), dRRAsCpf1 (e.g., SEQ ID NO: 24), dCasX (e.g., SEQ ID NO: 25), and dCasPhi (e.g., SEQ ID NO: 26).

[0103] In some embodiments, the Cas9 domain described herein can be a high-fidelity Cas9 domain, for example, comprising one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA to confer increased target binding specificity. In certain embodiments, the high-fidelity Cas9 domain can be nuclease-inactivated as described herein.

[0104] The CRISPR-associated protein domains described herein can recognize the protospacer adjacent motif (PAM) sequence in the target gene. A "PAM" sequence is typically a 2 to 6 bp DNA sequence immediately following the sequence targeted by the CRISPR-associated protein domain. The PAM sequence is necessary for CRISPR protein binding and cutting, but is not part of the target sequence. The CRISPR-associated protein domain can recognize a naturally occurring or typical PAM sequence, or can have a changed PAM specificity. CRISPR-associated protein domains that bind to atypical PAM sequences have been described in the art. For example, Kleinstiver et al., Nature (2015) 523 (7561): 481-5 and Kleinstiver et al., Nat Biotechnol. (2015) 33: 1293-8 describe Cas9 domains that bind to atypical PAM sequences. Such Cas9 domains may include, for example, those from "VRER" SpCas9, "EQR" SpCas9, "VQR" SpCas9, "SpG Cas9", "SpRYCas9", and "KKH" SaCas9. Nuclease-inactivated versions of these Cas9 domains are also contemplated, such as nuclease-inactivated VRER SpCas9 (e.g., SEQ ID NO: 27), nuclease-inactivated EQR SpCas9 (e.g., SEQ ID NO: 28), nuclease-inactivated VQR SpCas9 (e.g., SEQ ID NO: 29), nuclease-inactivated SpG Cas9 (e.g., SEQ ID NO: 30), nuclease-inactivated SpRY Cas9 (e.g., SEQ ID NO: 31), and nuclease-inactivated KKH SaCas9 (e.g., SEQ ID NO: 32). Another example is the Cas9 of Francisella novicida, engineered to recognize 5'-YG-3' (where "Y" is a pyrimidine).

[0105] Other suitable CRISPR-associated proteins, orthologs, and variants (including nuclease-inactive variants and sequences) will be apparent to those skilled in the art based on this disclosure.

[0106] Guide RNAs that can be used in conjunction with the CRISPR-associated protein domains herein are further described in Section II below.

[0107] B. Zinc finger protein domain

[0108] In some embodiments, the DNA binding domain of the epigenetic editor described herein comprises a zinc finger protein (ZFP) domain (or "ZF domain" as used herein). ZFP is a protein having at least one zinc finger, and binds to DNA in a sequence-specific manner. "Zinc finger" (ZF) or "zinc finger motif" (ZF motif) refers to a polypeptide domain comprising a beta-beta-alpha (ββα)-protein fold stabilized by zinc ions. ZF binds two to four base pairs of nucleotides, typically three or four base pairs (continuous or non-continuous). Each ZF typically comprises about 30 amino acids. The ZFP domain may contain a plurality of ZFs in series contact with its target nucleic acid sequence. Tandem arrays of ZFs may be engineered to generate artificial ZFPs that bind to desired nucleic acid targets. ZFPs can be rationally designed using a database comprising triplet (or quadruple) nucleotide sequences and single ZF amino acid sequences, wherein each triplet or quadruple nucleotide sequence is associated with one or more amino acid sequences of ZFs that bind to a specific triplet or quadruple sequence. See, for example, U.S. Patents 6,453,242, 6,534,261, and 8,772,453.

[0109] ZFPs are widely present in eukaryotic cells and may belong to, for example, the C2H2 class, CCHC class, PHD class, or RING class. An exemplary motif that characterizes one class of these proteins (the C2H2 class) is -Cys-(X) 2-4 -Cys-(X) 12 -His-(X) 3-5 -His- (SEQ ID NO: 657), wherein X is any independently selected amino acid. In some embodiments, the ZFP domain herein may comprise a ZF array comprising consecutive C2H2-ZFs each contacting three or more consecutive nucleotides.

[0110] The ZFP domain of the epigenetic editor described herein may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more ZFs. The ZFP domain may include an array of double-finger or three-finger units, such as 3, 4, 5, 6, 7, 8, 9 or 10 or more units, wherein each unit binds to a subsite in a target sequence. In some embodiments, a ZFP domain comprising at least three ZFs recognizes a target DNA sequence of 9 or 10 nucleotides. In some embodiments, a ZFP domain comprising at least four ZFs recognizes a target DNA sequence of 12 to 14 nucleotides. In some embodiments, a ZFP domain comprising at least six ZFs recognizes a target DNA sequence of 18 to 21 nucleotides.

[0111] In some embodiments, the ZFs in the ZFP domains described herein are connected via peptide linkers. The length of the peptide linker can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more amino acids. In some embodiments, the linker comprises 5 or more amino acids. In some embodiments, the linker comprises 7-17 amino acids. The linker can be flexible or rigid.

[0112] In some embodiments, the zinc finger array may have the following sequence:

[0113] SRPGERPFQCRICMRNFSXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXXHXXTH[connector]FQCRICMRNFSXXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXHXXTH[connector]PFQCRICMRNFSXXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXXHXXTHLRGS(SEQ ID NO:650),

[0114] Or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical thereto, wherein "XXXXXXX" represents the amino acids of the ZF recognition helix that confer DNA binding specificity on the zinc finger; each X can be independently selected. In the above sequence, italicized "XX" can be TR, LR or LK, and "[linker]" represents a linker sequence. In some embodiments, the linker sequence is TGSQKP (SEQ ID NO: 651); this linker can be used when the subsites targeted by the ZF are adjacent. In some embodiments, the linker sequence is TGGGGSQKP (SEQ ID NO: 652); this linker can be used when there are bases between the subsites targeted by the zinc finger. The two indicated linkers can be the same or different. In some embodiments, the length of the linker sequence is a minimum of 5 amino acids. In some embodiments, the length of the linker sequence is a maximum of 250 amino acids.

[0115] The ZFP domains herein can contain an array of two or more adjacent ZFs that are directly adjacent to each other (e.g., separated by a short (typical) linker sequence), or separated by a longer, flexible or structured polypeptide sequence. In some embodiments, directly adjacent fingers bind to a continuous nucleic acid sequence, i.e., to adjacent trinucleotides / triplets. In some embodiments, adjacent fingers are cross-linked between each other's respective target triplets, which can help to strengthen or enhance the recognition of the target sequence and result in the binding of overlapping sequences. In some embodiments, long-range ZFs within a ZFP domain can recognize (or bind to) non-contiguous nucleotide sequences.

[0116] Exemplary B2M target sequences are shown in Table 1 below.

[0117] Table 1 ZFP target sequences within B2M

[0118]

[0119]

[0120] In some embodiments, the ZFP domain of the epigenetic editor of the invention binds to a target sequence selected from any one of SEQ ID NOs: 700-740. The ZF may comprise a ZF framework sequence of SEQ ID NO: 650, or any other ZF framework known in the art.

[0121] C.TALE

[0122] In some embodiments, the DNA binding domain of the epigenetic editor described herein comprises a transcription activator-like effector (TALE) domain. The DNA binding domain of TALE comprises a highly conserved sequence of about 33-34 amino acids, with repeated variable di-residues (RVD) at positions 12 and 13, which are essential for recognizing specific nucleotides. TALEN can be engineered to actually bind to any desired DNA sequence. Methods for programming TALE are known in the art. For example, such methods are described in Carroll et al., Genet Soc Amer. (2011) 188(4):773-82; Miller et al., Nat Biotechnol. (2007) 25(7):778-85; Christian et al., Genetics (2008) 186(2):757-61; Li et al., Nucl Acids Res. (2010) 39(1):359-72; and Moscou et al., Science (2009) 326(5959):1501.

[0123] D. Other DNA binding domains

[0124] Other DNA binding domains of epigenetic editors described herein are contemplated. In some embodiments, the DNA binding domain comprises, for example, an argonaute protein domain from Halobacterium grisea (NgAgo). NgAgo is a ssDNA-guided endonuclease that is guided to its target site by 5' phosphorylated ssDNA (gDNA), where it produces a double-strand break. In contrast to Cas9, the NgAgo-gDNA system does not require a pre-spacer adjacent motif (PAM). Therefore, the bases that can be targeted can be greatly amplified using nuclease-inactivated NgAgo (DNNgAgo). Characterization and use of NgAgo have been described in, for example, Gao et al., Nat Biotechnol. (2016) 34(7):768-73; Swarts et al., Nature (2014) 507(7491):258-61; and Swarts et al., Nucl Acids Res. (2015) 43(10):5120-9.

[0125] In some embodiments, the DNA binding domain comprises an inactivated nuclease, e.g., an inactivated meganuclease. Additional non-limiting examples of DNA binding domains include tetracycline-controlled repressor (tetR) DNA binding domains, leucine zippers, helix-loop-helix (HLH) domains, helix-turn-helix domains, β-sheet motifs, steroid receptor motifs, bZIP domain homology domains, and AT hooks.

[0126] II. Guide polynucleotide

[0127] The epigenetic editors of the DNA binding domains comprising polynucleotide guides described herein may also include guide polynucleotides capable of forming a complex with the DNA binding domains. The guide polynucleotides may include RNA, DNA, or a mixture of the two. For example, in the case where the DNA binding domain of the polynucleotide guide is a CRISPR-associated protein domain, the guide polynucleotide may be a guide RNA (gRNA). "Guide RNA" or "gRNA" refers to a nucleic acid that can hybridize with a target sequence and guide the CRISPR-Cas complex to bind to the target sequence. Methods for using guide polynucleotide sequences with programmable DNA binding proteins (e.g., CRISPR-associated protein domains) for site-specific DNA targeting (e.g., to modify a genome) are known in the art.

[0128] A guide polynucleotide sequence (e.g., a gRNA sequence) can include two parts: 1) a nucleotide sequence comprising a "targeting sequence" that is complementary to a target nucleic acid sequence ("target sequence"), e.g., complementary to a nucleic acid sequence contained in a genomic target site; and 2) a nucleotide sequence of a DNA binding domain (e.g., a CRISPR-Cas protein domain) that binds to a polynucleotide guide. The nucleotide sequence in 1) can include a targeting sequence that is 100% complementary to a genomic nucleic acid sequence, e.g., a nucleic acid sequence contained in a genomic target site, and can therefore hybridize to a target nucleic acid sequence. The nucleotide sequence in 1) can be referred to as, for example, a crispr RNA or crRNA. The nucleotide sequence in 2) can be referred to as a scaffold sequence of a guide nucleic acid, e.g., a tracrRNA, or an activation region of a guide nucleic acid, and can include a stem-loop structure. Portions 1) and 2) as described above can be fused to form a single guide (e.g., a single guide RNA or sgRNA), or can be on two separate nucleic acid molecules. In some embodiments, the guide polynucleotide comprises portions 1) and 2) connected by a linker. In some embodiments, the guide polynucleotide comprises portions 1) and 2) connected by a non-nucleic acid linker (eg, a peptide linker or a chemical linker).

[0129] Part 2 (scaffold sequence) of a guide polynucleotide as described herein can be, for example, as described in Jinek et al., Science (2012) 337: 816-21; U.S. Patent Publication 2016 / 0208288; or U.S. Patent Publication 2016 / 0200779. The present disclosure also contemplates variants of part 2). For example, the tetraloop and stem loop of the gRNA scaffold (tracrRNA) sequence can be modified to include an RNA aptamer that can be bound by a specific protein domain. In some embodiments, such modified gRNAs can be used to promote the recruitment of a repression or activation domain fused to an RNA aptamer that interacts with a protein.

[0130] The gRNA provided herein generally comprises a targeting domain and a binding domain. The targeting domain (also referred to as "targeting sequence") may comprise a nucleic acid sequence that is bound to a target site (e.g., to a genomic nucleic acid molecule in a cell). The target site may be a double-stranded DNA sequence comprising a PAM sequence and a target sequence, which is located on the same chain as the PAM sequence and is directly adjacent to it. The targeting domain of a gRNA may comprise an RNA sequence corresponding to a target sequence, that is, it is similar to the sequence of the target domain, sometimes with one or more mismatches, but generally comprises an RNA sequence rather than a DNA sequence. Therefore, the targeting domain of a gRNA may be paired with a sequence base of a double-stranded target site complementary to the target sequence (completely or partially complementary), and therefore with a chain base pairing complementary to the chain comprising the PAM sequence. It should be understood that the targeting domain of a gRNA generally does not include a sequence similar to a PAM sequence. It should be further understood that, depending on the nuclease employed, the position of the PAM may be 5' or 3' of the target sequence. For example, PAM is generally located at 3' of the target sequence of the Cas9 nuclease and 5' of the target sequence of the Cas12a nuclease. For a description of the location of the PAM and the mechanism by which gRNA binds to the target site, see, e.g., Vanegas et al., Fungal Biol Biotechnol. (2019) 6:6. Figure 1 , which is incorporated herein by reference. For additional descriptions and descriptions of the mechanism of gRNA targeting of RNA-guided nucleases to target sites, see Fu et al., Nat Biotechnol (2014) 32(3):279-84 and Sternberg et al., Nature (2014) 507(7490):62-7, each of which is incorporated herein by reference.

[0131] In some embodiments, the targeting domain sequence comprises 17 to 30 nucleotides and corresponds completely to the target sequence (i.e., without any mismatched nucleotides). However, in some embodiments, the targeting domain sequence may comprise one or more, but typically no more than 4 mismatches, such as 1, 2, 3, or 4 mismatches. Since the targeting domain is part of the gRNA (which is an RNA molecule), it will typically comprise ribonucleotides, while the DNA targeting domain will comprise deoxyribonucleotides.

[0132] An exemplary illustration of a Cas9 target site comprising a 22 nucleotide target domain and an NGG PAM sequence, and a gRNA comprising a targeting domain that corresponds exactly to the target sequence (and thus is perfectly complementary base paired with a DNA strand that is complementary to the strand comprising the target sequence and the PAM) is provided below:

[0133] [Target domain (DNA)][PAM]

[0134] 5'-NNNNNNNNNNNNNNNNNNNNN-NNGG-3'(DNA)

[0135] 3'-NNNNNNNNNNNNNNNNNNNNN-NNCC-5'(DNA)

[0136] ||||||||||||||||||||||

[0137] 5'-NNNNNNNNNNNNNNNNNNNNN-N-[gRNA scaffold]-3'(RNA)

[0138] [Targeting domain (RNA)][Binding domain]

[0139] An exemplary illustration of a Cas12a target site comprising a 22 nucleotide target domain and a TTN PAM sequence, and a gRNA comprising a targeting domain that completely corresponds to the target sequence (and thus completely complementarily base pairs with a DNA strand that is complementary to the strand comprising the target sequence and the PAM) is provided below:

[0140] [PAM][Target domain (DNA)]

[0141] 5'-TTNNNNNNNNNNNNNNNNNNN-NNNN-3'(DNA)

[0142] 3'-AANNNNNNNNNNNNNNNNNNNN-NNNN-5'(DNA)

[0143] ||||||||||||||||||||||

[0144] 5'-[gRNA scaffold]-NNNNNNNNNNNNNNNNNNNN-N-3'(RNA)

[0145] [Binding domain][Targeting domain (RNA)]

[0146] Although not wishing to be bound by theory, at least in some embodiments, it is believed that the length of the targeting domain and the complementarity with the target sequence contribute to the specificity of the gRNA / Cas9 molecular complex interacting with the target nucleic acid. In some embodiments, the length of the targeting domain of the gRNA provided herein is 5 to 50 nucleotides. In some embodiments, the length of the targeting domain is 15 to 25 nucleotides. In some embodiments, the length of the targeting domain is 18 to 22 nucleotides. In some embodiments, the length of the targeting domain is 19-21 nucleotides. In some embodiments, the length of the targeting domain is 15 nucleotides. In some embodiments, the length of the targeting domain is 16 nucleotides. In some embodiments, the length of the targeting domain is 17 nucleotides. In some embodiments, the length of the targeting domain is 18 nucleotides. In some embodiments, the length of the targeting domain is 19 nucleotides. In some embodiments, the length of the targeting domain is 20 nucleotides. In some embodiments, the length of the targeting domain is 21 nucleotides. In some embodiments, the length of the targeting domain is 22 nucleotides. In some embodiments, the length of the targeting domain is 23 nucleotides. In some embodiments, the length of the targeting domain is 24 nucleotides. In some embodiments, the length of the targeting domain is 25 nucleotides. In certain embodiments, the targeting domain corresponds completely to the target sequence or a portion thereof provided herein, without mismatches. In some embodiments, the targeting domain of the gRNA provided herein comprises 1 mismatch relative to the target sequence provided herein. In some embodiments, the targeting domain comprises 2 mismatches relative to the target sequence. In some embodiments, the target domain comprises 3 mismatches relative to the target sequence.

[0147] Methods for designing, selecting and verifying gRNA are described herein, and these methods are known in the art. Software tools can be used to optimize the gRNA corresponding to the target DNA sequence, for example, to minimize the total off-target activity of the entire genome. For example, a DNA sequence search algorithm can be used to identify the target sequence in the crRNA of the gRNA for use with Cas9. Exemplary gRNA design tools include Bae et al., Bioinformatics (2014) 30: tools described in 1473-5.

[0148] Guide polynucleotides (e.g., gRNA) described herein can have various lengths. In some embodiments, the length of the spacer or targeting sequence depends on the CRISPR-associated protein components of the epigenetic editor system used. For example, Cas proteins from different bacterial species have different optimal targeting sequence lengths. Therefore, the length of the spacer sequence includes, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In some embodiments, the length of the spacer includes 10-24, 11-20, 11-16, 18-24, 19-21 or 20 nucleotides. In some embodiments, the length of the guide polynucleotide (e.g., gRNA) is 15-100 (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 ) nucleotides and comprises a spacer sequence of at least 10 (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50) consecutive nucleotides that are complementary to the target sequence. In some embodiments, the guide polynucleotides described herein can be truncated, for example, by 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more nucleotides.

[0149] In certain embodiments, the 3' end of the B2M target sequence is adjacent to a PAM sequence (e.g., a typical PAM sequence, such as NGG of SpCas9). The degree of complementarity between the targeting sequence of the guide polynucleotide (e.g., the spacer sequence of the gRNA) and the target sequence can be at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100%. In specific embodiments, the targeting sequence and the target sequence can be 100% complementary. In other embodiments, the targeting sequence and the target sequence may contain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 mismatches.

[0150] Guide polynucleotides (e.g., gRNAs) can be modified, for example, with chemical alterations and synthetic modifications. For example, a modified gRNA can include an alteration or replacement of one or both non-linked phosphate oxygens and / or one or more linked phosphate oxygens in a phosphodiester backbone bond, an alteration of a ribose (e.g., a 2' hydroxyl on a ribose), an alteration of a phosphate moiety, a modification or replacement of a naturally occurring nucleobase, a modification or replacement of a ribose-phosphate backbone, a modification of the 3' end and / or the 5' end of an oligonucleotide, a replacement or portion of a terminal phosphate group, conjugation of a cap or linker, or any combination thereof.

[0151] In some embodiments, one or more ribose groups of gRNA can be modified. Examples of chemical modifications of ribose groups include, but are not limited to, 2'-O-methyl (2'-OMe), 2'-fluorine (2'-F), 2'-deoxy, 2'-O-(2-methoxyethyl) (2'-MOE), 2'-NH2, 2'-O-allyl, 2'-O-ethylamine, 2'-O-cyanoethyl, 2'-O-acetal esters, or bicyclic nucleotides such as locked nucleic acids (LNA), 2'-(5-restricted ethyl (S-cEt)), restricted MOE or 2'-0,4'-C-aminomethylene bridged nucleic acids (2', 4'-BNANC). 2'-O-methyl modifications and / or 2'-fluorine modifications can increase the binding affinity and / or nuclease stability of gRNA oligonucleotides.

[0152] In some embodiments, one or more phosphate groups of the gRNA may be chemically modified. Examples of chemical modifications of phosphate groups include, but are not limited to, phosphorothioate (PS), phosphonoacetate (PACE), thiophosphonoacetate (thioPACE), amide, triazole, phosphonate, and phosphotriester modifications. In some embodiments, the guide polynucleotides described herein may comprise one, two, three, or more PS bonds at or near the 5' end and / or the 3' end; the PS bonds may be continuous or non-continuous.

[0153] In some embodiments, the gRNA herein comprises a mixture of ribonucleotides and deoxyribonucleotides and / or one or more PS bonds.

[0154] In some embodiments, one or more nucleobases of gRNA can be chemically modified. Examples of chemically modified nucleobases include, but are not limited to, 2-thiouridine, 4-thiouridine, N6-methyladenosine, pseudouridine, 2,6-diaminopurine, inosine, thymidine, 5-methylcytosine, 5-substituted pyrimidine, isoguanine, isocytosine, and nucleobases with halogenated aromatic groups. Chemical modification can be performed in the spacer, tracr RNA region, stem loop, or any combination thereof.

[0155] Table 2 below lists exemplary gRNA target sequences for human B2M epigenetic modification, and the coordinates of the starting position of the targeted site on human chromosome 15 (SEQ represents SEQ ID NO). The table also shows the distance from the starting coordinates to the TSS coordinates in the B2M gene. Table 3 lists exemplary targeting sequences of gRNA.

[0156] Table 2 Exemplary target sequences of gRNA targeting B2M

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168] Table 3 Exemplary targeting sequences of gRNA targeting B2M

[0169]

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179] In some embodiments, the target region of the guide RNA targeting B2M comprises one or more sequences selected from the group consisting of SEQ ID NO: 700-740, 744, 747-749, 752, 753, 757, 758, 760-806, 812-822, 825, 827, 830, 833, 834, 839-841, 843-845, 849, 851-853, 855, 864, 866-877, 879-88 3, 891-896, 898-900, 903-914, 922, 923, 925-927, 934, 936, 943-947, 949, 951-962, 975-981, 983, 985, 987-989, 995, 997-999, 1003-1005, and 1007-1011. In some embodiments, the guide RNA targeting B2M comprises SEQ IDNO:1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1 104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 11 Any one of 62-1167, 1169-1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276 and 1278-1282.

[0180] Any tracr sequence known in the art is contemplated for use in the gRNAs described herein. In some embodiments, the gRNAs described herein have a tracr sequence as shown in Table 4 below, or a tracr sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% identical to the tracr sequence shown below (SEQ represents SEQ ID NO).

[0181] Table 4 Exemplary TRACR sequences

[0182]

[0183] In some embodiments, the gRNA herein is provided directly to the cell (e.g., via an RNP complex and a CRISPR-associated protein domain). In some embodiments, the gRNA is provided to the cell by an expression vector (e.g., a plasmid vector or a viral vector) introduced into the cell, wherein the cell then expresses the gRNA from the expression vector. Methods for introducing gRNA and expression vectors into cells are well known in the art.

[0184] III. Effector domain

[0185] The epigenetic editors described herein include one or more effector protein domains (also referred to as "epigenetic effector domains" or "effector domains" as used herein) that affect the epigenetic modification of the target gene. An epigenetic editor with one or more effector domains can regulate the expression of a target gene without changing its nucleobase sequence. In some embodiments, the effector domains described herein can provide repression or silencing of expression of a target gene (such as B2M), for example, by repressing transcription or by modifying or remodeling chromatin. Such effector domains are also referred to herein as "repressor domains", "repressor domains" or "epigenetic repressor domains". Non-limiting examples of chemical modifications that can be mediated by effector domains include methylation, demethylation, acetylation, deacetylation, phosphorylation, SUMOylation and / or ubiquitination of DNA or histone residues.

[0186] In some embodiments, the effector domain of an epigenetic editor described herein can perform a histone tail modification, for example, by adding or removing an active tag on the histone tail.

[0187] In some embodiments, the effector domain of the epigenetic editor described herein can include or recruit transcription-related proteins, such as transcription repressors. The transcription-related proteins can be endogenous or exogenous.

[0188] In some embodiments, the effector domain of an epigenetic editor described herein can, for example, comprise a protein that directly or indirectly blocks access of a transcription factor to a gene of interest containing a target sequence.

[0189] The effector domain can be a full-length protein or a fragment thereof ("functional domain") that retains epigenetic effector function. Functional domains that can regulate (e.g., repress) gene expression can be derived from larger proteins. For example, functional domains that can reduce target gene expression can be identified based on the sequence of repressor proteins. The amino acid sequence of gene expression regulatory proteins can be obtained from available genome browsers, such as the UCSD genome browser or the Ensembl genome browser. Protein annotation databases such as UniProt or Pfam can be used to identify functional domains within the full protein sequence. As a starting point, the gene expression regulatory activity of the largest sequence covering all regions identified by different databases can be tested. Various truncations can then be tested to identify the minimum functional unit.

[0190] The present disclosure also contemplates variants of the effector domains described herein. For example, a variant refers to a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity and / or sequence similarity to a wild-type effector domain described herein. In a specific embodiment, the variant retains at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the epigenetic effector function of the wild-type effector domain.

[0191] In some embodiments, the effector domain described herein may comprise a fusion of two or more effector domains (e.g., KOX1, KRAB, and ZIM3). The effector domain may, for example, comprise a fusion of 2, 3, 4, 5, 6, 7, 8, 9, or 10 effector domains, such as the effector domains described herein. In certain embodiments, the effector domain comprises a fusion of a truncated form of the effector domain and a second effector domain. In certain embodiments, the effector domain comprises a fusion of truncated forms of two effector domains (e.g., a fusion of the N-terminal and C-terminal portions of the two effector domains).

[0192] In some embodiments, the epigenetic editor described herein may include 1 effector domain, 2 effector domains, 3 effector domains, 4 effector domains, 5 effector domains, 6 effector domains, 7 effector domains, 8 effector domains, 9 effector domains, 10 effector domains or more. In certain embodiments, the epigenetic editor comprises one or more fusion proteins (e.g., one, two or three fusion proteins), each fusion protein having one or more effector domains (e.g., one, two or three effector domains) connected to a DNA binding domain. In some embodiments, the effector domain can induce a combination of epigenetic modifications, for example, transcriptional repression and DNA methylation, DNA methylation and histone deacetylation, DNA methylation and histone demethylation, DNA methylation and histone methylation, DNA methylation and histone phosphorylation, DNA methylation and histone ubiquitination, DNA methylation and histone SUMOylation.

[0193] In certain embodiments, the effector domains described herein (e.g., DNMT3A and / or DNMT3L) are encoded by nucleotide sequences found in the native genome (e.g., human or mouse) of the effector domain. In other embodiments, the effector domains described herein are encoded by nucleotide sequences that have been codon-optimized for optimal expression in human cells.

[0194] The effector domains described herein may include, for example, transcriptional repressors, DNA methyltransferases, and / or histone modifiers, as described in further detail below.

[0195] A. Transcriptional repressors

[0196] In some embodiments, the epigenetic effector domain described herein mediates the repression of target gene expression (e.g., transcription). For example, the effector domain may include a Krüppel-associated box (KRAB) repression domain, a repressor element silencing transcription factor (REST) ​​repression domain, a KRAB-associated protein 1 (KAP1) domain, a MAD domain, a FKHR (forkhead gene in rhabdomyosarcoma gene) repression domain, an EGR-1 (early growth response gene product-1) repression domain, an ets2 repressor repression domain (ERD), a MAD smSIN3 interaction domain (SID), a WRPW motif of a hair-related basic helix-loop-helix (bHLH) repressor protein, an HP1α chromo-shadow repression domain, an HP1β repression domain, or any combination thereof. The effector domain may recruit one or more protein domains that repress target gene expression, such as through a scaffold protein. In some embodiments, the effector domain can recruit or interact with a scaffold protein domain that recruits a PRMT protein, an HDAC protein, a SETDB1 protein, or a NuRD protein domain.

[0197] In some embodiments, the effector domain comprises a functional domain derived from a zinc finger repressor protein, such as a KRAB domain. The KRAB domain is found in approximately 400 human ZFP-based transcription factors. A description of the KRAB domain can be found, for example, in Ecco et al., Development (2017) 144 (15): 2719-29 and Lambert et al., Cell (2018) 172: 650-65.

[0198] In certain embodiments, the effector domain comprises a repressor domain (e.g., KRAB) derived from KOX1 / ZNF10, KOX8 / ZNF708, ZNF43, ZNF184, ZNF91, HPF4, HTF10, or HTF34. In some embodiments, the effector domain comprises a repressor domain (e.g., KRAB) derived from ZIM3, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZNF680, ZNF41, ZNF189, ZNF528, ZNF543, ZNF554, ZNF140, ZNF610, ZNF264, ZNF350, ZNF8, ZNF582, ZNF30, ZNF324, ZNF98, ZNF669, ZNF677, ZNF596, ZNF214, ZNF37, Z NF34, ZNF250, ZNF547, ZNF273, ZNF354, ZFP82, ZNF224, ZNF33, ZNF45, ZNF175, ZNF595, ZNF184, ZNF419, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF566, ZNF729, ZIM2, ZNF254, ZNF764, ZNF785, or any combination thereof, a repressor domain (e.g., KRAB). For example, the repressor domain is a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627. In a specific embodiment, the repressor domain is a ZIM3 KRAB domain. In further embodiments, the effector domain is derived from a human protein, such as human ZIM3, human KOX1, human ZFP28, or human ZN627.

[0199] Sequences of exemplary effector domains that can reduce or silence target gene expression or protein sequences containing them are provided in Table 5 below (SEQ represents SEQ ID NO). Further examples of repressor and transcriptional repressor domains can be found in, for example, PCT Patent Publication WO 2021 / 226077 and Tycko et al., Cell (2020) 183 (7): 2020-35, each of which is incorporated herein by reference in its entirety.

[0200] Table 5 Exemplary effector domains that can reduce or silence gene expression

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212] The present disclosure encompasses functional analogs of any of the proteins listed above, i.e., molecules having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the transcription factor function of the protein). For example, a functional analog can be an isoform or variant of the protein listed above, e.g., containing a portion of the above protein with or without additional amino acid residues and / or containing a mutation relative to the above protein. In some embodiments, a functional analog has at least 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity with one of the sequences listed in Table 5. Homologues, orthologues, and mutants of the above-listed proteins are also contemplated.

[0213] In certain embodiments, the epigenetic editor described herein comprises a KRAB domain derived from KOX1, ZIM3, ZFP28 or ZN627, and / or an effector domain derived from KAP1, MECP2, HP1a, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1 or SCML2, optionally wherein the parent protein is a human protein. In a specific embodiment, the epigenetic editor described herein comprises a domain derived from KOX1, ZIM3, ZFP28 and / or ZN627, optionally wherein the parent protein is a human protein. In certain embodiments, the epigenetic editor may include a KRAB domain derived from KOX1 (ZNF10), such as human KOX1. In certain embodiments, the epigenetic editor may include a KRAB domain derived from ZIM3 (ZNF657 or ZNF264), such as human ZIM3. In certain embodiments, the epigenetic editor may include a KRAB domain derived from ZFP28, such as human ZFP28. In certain embodiments, the epigenetic editor may include a KRAB domain derived from ZN627, such as human ZN627. In certain embodiments, the epigenetic editor described herein may include CDYL2, such as human CDYL2, and / or a TOX domain (e.g., human TOX domain) in combination with a KOX1 KRAB domain (e.g., human KOX1 KRAB domain).

[0214] In certain embodiments, the epigenetic effectors described herein comprise a repressor domain derived from KOX1 / ZNF10 (SEQ ID NO: 89). For example, the repressor domain may comprise the sequence of SEQ ID NO: 89, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 89.

[0215] In certain embodiments, the epigenetic effectors described herein comprise a repressor domain derived from KOX1 / ZNF10, as shown in Table 6 below:

[0216] Table 6 Exemplary effector domains derived from KOX1 / ZNF10

[0217] protein Protein sequence KOX1 / ZNF10 KRAB 1 SEQ ID NO:565 KOX1 / ZNF10KRAB 2 SEQ ID NO:566 KOX1 / ZNF10KRAB 3 SEQ ID NO:567 KOX1 / ZNF10 (aa 11-72) SEQ ID NO:568 KOX1 / ZNF10 (aa 11-108) SEQ ID NO:569 KOX1 / ZNF10 variants SEQ ID NO:570 KOX1 KRAB-ZIM3 chimera SEQ ID NO:571 ZIM3-KOX1 KRAB chimera SEQ ID NO:572

[0218] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:565, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:565.

[0219] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:566, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:566.

[0220] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:567, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:567.

[0221] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:568, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:568.

[0222] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:569, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:569.

[0223] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:570, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:570.

[0224] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:571, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:571.

[0225] In specific embodiments, the repressor domain may comprise the amino acid sequence of SEQ ID NO:572, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO:572.

[0226] B. DNA methyltransferase

[0227] In some embodiments, the effector domain of the epigenetic editor described herein changes the expression of the target gene by DNA modification, such as methylation. The transcriptional activity of highly methylated regions of DNA is often lower than that of regions with lower methylation. DNA methylation mainly occurs at CpG sites (abbreviation for "C-phospho-G-" or "cytosine-phospho-guanine" sites). Many mammalian genes have promoter regions near or including CpG islands (nucleic acid regions with high frequency CpG dinucleotides).

[0228] The effector domain described herein can be, for example, a DNA methyltransferase (DNMT) or its catalytic domain, or can be capable of recruiting a DNA methyltransferase. DNMT covers enzymes that catalyze the transfer of methyl groups to DNA nucleotides, such as typical cytosine-5 DNMTs (e.g., DNMT1, DNMT3A, DNMT3B, and DNMT3C) that catalyze the addition of methyl groups to genomic DNA. The term also covers atypical family members that do not catalyze methylation per se but recruit (including activation) catalytically active DNMTs; a non-limiting example of such DNMTs is DNMT3L. See, for example, Lyko, Nat Review (2018) 19: 81-92. Unless otherwise indicated, a DNMT domain can refer to a polypeptide domain derived from a catalytically active DNMT (e.g., DNMT1, DNMT3A, and DNMT3B) or derived from a non-catalytically active DNMT (e.g., DNMT3L). DNMT can suppress the expression of a target gene by recruiting repressive regulatory proteins. In some embodiments, methylation is at a CG (or CpG) dinucleotide sequence. In some embodiments, the methylation is at a CHG or CHH sequence, where H is any of A, T, or C.

[0229] In some embodiments, DNMT described herein can be animal DNMT (e.g., mammalian DNMT), plant DNMT, fungal DNMT or bacterial DNMT. Bacterial DNMT can be obtained from bacterial species (e.g., cocci, bacillus, helicobacter or intracellular, gram-positive or gram-negative bacteria). In certain embodiments, the bacterial species is mycoplasma, marine mycoplasma or Spiroplasma chinense. In certain embodiments, the bacterial species is not penetrating mycoplasma, S.monbiae, Haemophilus parainfluenzae, Arthrobacter luteus, Haemophilus aegypti, Haemophilus hemolyticus, Moraxella, Escherichia coli, Thermus aquaticus, Crescent bacillus or Clostridium difficile. In certain embodiments, the epigenetic editor described herein comprises a DNMT domain containing SEQ ID NO: 601 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 601. In certain embodiments, an epigenetic editor described herein comprises a DNMT domain comprising SEQ ID NO: 602, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 602. In certain embodiments, an epigenetic editor described herein comprises a DNMT domain comprising SEQ ID NO: 603, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 603.

[0230] In certain embodiments, the DNMT in the epigenetic editor described herein may include, for example, DNMT1, DNMT3A, DNMT3B and / or DNMT3C. In some embodiments, the DNMT is a mammalian (e.g., human or mouse) DNMT. In a specific embodiment, the DNMT is DNMT3A (e.g., human DNMT3A). In certain embodiments, the epigenetic editor described herein comprises a DNMT3A domain containing SEQ ID NO: 574 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 574. In certain embodiments, the epigenetic editor described herein comprises a DNMT3A domain comprising SEQ ID NO: 575, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 575. In some embodiments, the DNMT3A domain may have, for example, a mutation at position H739 (e.g., H739A or H739E), R771 (e.g., R771L), and / or R836 (e.g., R836A or R836Q), or any combination thereof (numbered according to SEQ ID NO: 574).

[0231] In some embodiments, the effector domain described herein can be a DNMT-like domain. As used herein, a "DNMT-like domain" is a regulatory factor of DNMT, which can activate or recruit other DNMT domains, but does not have methylation activity itself. In some embodiments, the DNMT-like domain is a mammalian (e.g., human or mouse) DNMT-like domain. In certain embodiments, the DNMT-like domain is DNMT3L, which can be, for example, human DNMT3L or mouse DNMT3L. In certain embodiments, the epigenetic editor described herein comprises a DNMT3L domain containing SEQ ID NO: 578 or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to SEQ ID NO: 578. In certain embodiments, the epigenetic editor herein comprises a DNMT3L domain comprising SEQ ID NO: 579, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 579. In certain embodiments, the epigenetic editor described herein comprises a DNMT3L domain comprising SEQ ID NO: 580, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 580. In certain embodiments, the epigenetic editor described herein comprises a DNMT3L domain comprising SEQ ID NO: 581, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 581. In some embodiments, the DNMT3L domain can have, for example, mutations corresponding to positions D226 (e.g., D226V), Q268 (e.g., Q268K), or both (numbered according to SEQ ID NO: 578).

[0232] In certain embodiments, the epigenetic editor herein may include both DNMT and DNMT-like effector domains. For example, the epigenetic editor may include a DNMT3A-3L domain, wherein DNMT3A and DNMT3L may be covalently linked. In other embodiments, the epigenetic editor described herein may include an effector domain that includes only a DNMT3A domain (e.g., human DNMT3A), or only a DNMT-like domain (e.g., DNMT3L, which may be human or mouse DNMT3L).

[0233] Table 7 below provides exemplary DNMTs that can be part of the epigenetic effectors described herein, or from which the effector domains of the epigenetic editors described herein can be derived.

[0234] Table 7 Exemplary DNMT sequences

[0235]

[0236]

[0237]

[0238]

[0239] The present disclosure encompasses functional analogs of any of the above proteins, i.e., molecules having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the DNA methylation function or recruitment function of the protein). For example, a functional analog may be an isoform or variant of the above-listed proteins, e.g., containing a portion of the above-listed proteins with or without additional amino acid residues and / or containing mutations relative to the above-listed proteins. In some embodiments, a functional analog has at least 75%, 80%, 85%, 90%, 95%, 98% or 99% sequence identity with one of the sequences listed in Table 7. In some embodiments, the effector domains herein comprise only the functional domains (or functional analogs thereof) of the above-listed proteins, e.g., catalytic domains or recruitment domains. In some embodiments, the effector domains herein comprise one or more epigenetic effector domains selected from Table 7, or functional homologs, orthologs, or variants thereof.

[0240] As used herein, a DNMT domain (e.g., a DNMT3A domain or a DNMT3L domain) refers to a protein domain that is the same as a parent protein (e.g., human or mouse DNMT3A or DNMT3L) or a functional analog thereof (e.g., having a functional fragment of the parent protein, such as a catalytic fragment or a recruitment fragment; and / or having a mutation that improves the activity of the DNMT protein).

[0241] The epigenetic editor herein can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900 or 1000 or more CpG dinucleotide sequences to achieve methylation. The CpG dinucleotide sequence can be located in or near the target gene in the CpG island, or can be located in a region that is not a CpG island. CpG islands generally refer to nucleic acid sequences or chromosome regions containing high-frequency CpG dinucleotides. For example, a CpG island may contain at least 50% GC content. CpG islands can have a high observed to expected CpG ratio, for example, an observed to expected CpG ratio of at least 60%. As used herein, the observed to expected CpG ratio is determined by the number of CpG* (sequence length) / (number of C*number of G). In some embodiments, the CpG island has an observed to expected CpG ratio of at least 60%, 70%, 80%, 90% or more. The CpG island can be, for example, a sequence or region of at least 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750 or 800 nucleotides. In some embodiments, only 1 or less than 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40 or 50 CpG dinucleotides are methylated by an epigenetic editor.

[0242] In some embodiments, the epigenetic editors herein achieve methylation at hypomethylated nucleic acid sequences, i.e., sequences that may lack methyl groups on 5-methylcytosine nucleotides (e.g., in CpG) compared to standard controls. For example, hypomethylation may occur in senescent cells or cancer (e.g., early stages of neoplasia) relative to young cells or non-cancerous cells, respectively.

[0243] In some embodiments, the epigenetic editors described herein induce methylation at a hypermethylated nucleic acid sequence.

[0244] In some embodiments, methylation can be introduced by an epigenetic editor at sites other than CpG dinucleotides. For example, the target gene sequence can be methylated at the C nucleotide of a CpA, CpT or CpC sequence. In some embodiments, the epigenetic editor includes a DNMT3A domain and realizes methylation at a CpG, CpA, CpT, CpC sequence or any combination thereof. In some embodiments, the epigenetic editor includes a DNMT3A domain that lacks a regulatory subdomain and only maintains a catalytic domain. In some embodiments, the epigenetic editor comprising a DNMT3A catalytic domain realizes methylation only at a CpG sequence. In some embodiments, compared with an epigenetic editor comprising a wild-type DNMT3A domain, an epigenetic editor comprising a DNMT3A domain has a higher methylation activity at a CpA, CpC and / or CpT sequence, and the DNMT3A domain includes a mutation, such as an R836A or R836Q mutation (according to SEQ ID NO: 574 numbering).

[0245] C. Histone modifiers

[0246] In some embodiments, the effector domain of the epigenetic editor herein mediates histone modification. Histone modification plays a structural and biochemical role in gene transcription, such as by forming or destroying a nucleosome structure that binds to histones and prevents gene transcription. Histone modification can include, for example, acetylation, deacetylation, methylation, phosphorylation, ubiquitination, SUMOylation, etc. at its N-terminus ("histone tail"). These modifications maintain or specifically transform chromatin structure, so as to control responses such as gene expression, DNA replication, DNA repair, etc. that occur on chromosome DNA. The post-translational modification of histones is an epigenetic regulatory mechanism, and is considered to be necessary for eukaryotic cell genetic regulation. Recent studies have shown that chromatin remodeling factors such as SWI / SNF, RSC, NURF, NRD, etc., promote transcription factors to approach DNA by modifying nucleosome structure; histone acetyltransferase (HAT), which regulates histone acetylation state; and histone deacetylase (HDAC), as an important regulatory factor.

[0247] In particular, the unstructured N-terminal of histone can be modified by acetylation, deacetylation, methylation, ubiquitination, phosphorylation, SUMOylation, ribosylation, citrullination, O-glycosylation, crotonylation or any combination thereof. For example, histone acetyltransferase (HAT) utilizes acetyl-CoA as a cofactor and catalyzes the transfer of acetyl groups to the epsilon amino group of lysine side chains. This neutralizes the positive charge of lysine, and weakens the interaction between histone and DNA, thereby opening chromosomes so that transcription factors can be combined and start transcription. The acetylation of K14 and K9 lysine of histone H3 by histone acetyltransferase may be related to people's transcriptional capacity. Lysine acetylation can directly or indirectly produce the binding site of the chromatin modifying enzyme regulating transcriptional activation. On the other hand, the histone methylation of lysine 9 of histone H3 may be related to heterochromatin or transcriptional silent chromatin.

[0248] In certain embodiments, the effector domain of the epigenetic editor described herein comprises a histone methyltransferase domain. The effector domain may comprise, for example, a DOT1L domain, a SET domain, a SUV39H1 domain, a G9a / EHMT2 protein domain, an EZH1 domain, an EZH2 domain, a SETDB1 domain, or any combination thereof. In a specific embodiment, the effector domain comprises a histone-lysine-N-methyltransferase SETDB1 domain.

[0249] In some embodiments, the effector domain comprises a histone deacetylase protein domain. In certain embodiments, the effector domain comprises an HDAC family protein domain, for example, HDAC1, HDAC3, HDAC5, HDAC7 or HDAC9 protein domain. In a specific embodiment, the effector domain comprises a nucleosome remodeling and deacetylase complex (NURD), which removes acetyl groups from histones.

[0250] D. Other effector domains

[0251] In some embodiments, the effector domain comprises a triple motif containing a protein (TRIM28, TIF1-β or KAP1). In certain embodiments, the effector domain comprises one or more KAP1 proteins. The KAP1 protein in the epigenetic editor herein can form a complex with one or more other effector domains of the epigenetic editor or one or more proteins involved in gene expression regulation in the cell environment. For example, KAP1 can be recruited by the KRAB domain of a transcriptional repressor. The KAP1 protein domain can interact with one or more protein complexes that reduce or silence gene expression or recruit one or more protein complexes that reduce or silence gene expression. In some embodiments, KAP1 interacts with histone deacetylase proteins, histone-lysine methyltransferase proteins, chromatin remodeling proteins and / or heterochromatin proteins or recruits histone deacetylase proteins, histone-lysine methyltransferase proteins, chromatin remodeling proteins and / or heterochromatin proteins. For example, the KAP1 protein domain can interact with or recruit heterochromatin protein 1 (HP1) protein, SETDB1 protein, HDAC protein and / or NuRD protein complex components. In some embodiments, the KAP1 protein domain interacts with or recruits ZFP90 protein (e.g., isoform 2 of ZFP90) and / or FOXP3 protein. An exemplary KAP1 amino acid sequence is shown in SEQ ID NO: 629.

[0252] In some embodiments, the effector domain comprises a protein domain that interacts with one or more DNA epigenetic markers or is recruited by one or more DNA epigenetic markers. For example, the effector domain may include a methyl CpG binding protein 2 (MECP2) protein (which may or may not be located at the CpG island of the target gene) that interacts with the methylated DNA nucleotides in the target gene. The MECP2 protein domain in the epigenetic editor described herein can induce a condensed chromatin structure, thereby reducing or silencing the expression of the target gene. In some embodiments, the MECP2 protein domain in the epigenetic editor described herein can interact with histone deacetylase (e.g., HDAC), thereby repressing or silencing the expression of the target gene. In some embodiments, the MECP2 protein domain in the epigenetic editor described herein can block the approach of a transcription factor or a transcription activator to a target sequence, thereby repressing or silencing the expression of the target gene. Exemplary MECP2 amino acid sequences are shown in SEQ ID NO:630.

[0253] Also contemplated are effector domains of the epigenetic editors described herein, e.g., chromoshadow domains, ubiquitin-2-like Rad60 SUMO-like (Rad60-SLD / SUMO) domains, chromatin organization modifier domain (Chromo) domains, YAF2 / RYBP C-terminal binding motif domain (YAF2_RYBP), CBX family C-terminal motif domain (CBX7_C), zinc finger C3HC4 type (RING finger) domain (ZF-C3HC4_2), cytochrome b5 domain (Cyt-b5), helix-loop-helix domain (HLH), helix-hairpin-helix motif domain (e.g., HHH_3), high mobility group box domain (HMG-box), basic leucine zipper domain (e.g.,bZIP_1 or bZIP_2), Myb_DNA binding domain, homeodomain, MYM-type zinc finger with FCS sequence domain (ZF-FCS), interferon regulatory factor 2 binding protein zinc finger domain (IRF-2BP1_2), SSX repression domain (SSXRD), B-box type zinc finger domain (ZF-B_box), CXXC zinc finger domain (ZF-CXXC), chromosome condensation regulatory factor 1 domain (RCC1), SRC homology 3 domain (SH3_9), sterile alpha motif domain (SAM_1), sterile alpha motif domain (SAM_2), sterile alpha motif / tip domain (SAM_PNT), degeneration-like / Tondu family domain (Vg_Tdu), LIM domain, RNA recognition motif domain (RRM_1), paired amphipathic helix domain (PAH), proteasome ATPase OB C-terminal domain (Prot_ATP_ID_OB), neurohomology 2 domain (NHR2), cleavage stimulating factor subunit 2 hinge domain (CSTF2_hinge), PPARγ N-terminal region domain (PPARγ_N), CDC48 N-terminal domain (CDC48_2), WD40 repeat domain (WD40), Fip1 motif domain (Fip1), PDZ domain (PDZ_6), Von Willebrand factor C-type domain (VWC), NAB conserved region 1 domain (NCD1), S1 RNA-binding domain (S1), HNF3 C-terminal domain (HNF_C), Tudor domain (Tudor_2), histone-like transcription factor (CBF / NF-Y) and archaeal histone domain (CBFD_NFYB_HMF), zinc finger protein domain (DUF3669), EGF-like domain (cEGF), GATA zinc finger domain (GATA), TEA / ATTS domain (TEA), phorbol ester / diacylglycerol binding domain (C1-1), polycomb-like MTF2 factor 2 domain (Mtf2_C), transactivation domain of FOXO protein family (FOXO-TAD), homeobox KN domain (Home obox_KN), BED zinc finger domain (ZF-BED), C3HC4 type RING zinc finger domain (ZF-C3HC4_4), RAD51 interaction motif domain (RAD51_interact), p55 binding region of methyl-CpG-binding domain protein MBD (MBDa), Notch domain, Raf-like Ras binding domain (RBD), Spin / Ssty family structure (Spin-Ssty), PHD finger domain (PHD_3), low-density lipoprotein receptor domain class A (Ldl_recept_a), CS domain, DM DNA binding domain and QLQ domain.

[0254] In some embodiments, the effector domain is a protein domain comprising a YAF2_RYBP domain or a homology domain or any combination thereof. In certain embodiments, the homology domain of the YAF2_RYBP domain is a PRD domain, a NKL domain, a HOXL domain or a LIM domain. In a specific embodiment, the YAF2_RYBP domain may comprise a 32 amino acid YAF2 / RYBP C-terminal binding motif domain (32aa RYBP).

[0255] In some embodiments, the effector domain comprises a protein domain selected from the group consisting of a SUMO3 domain, a Chromo domain from M-phase phosphoprotein 8 (MPP8), a chromoshadow domain from Chromobox 1 (CBX1), and a SAM_1 / SPM domain from Scm polycomb group protein homolog 1 (SCMH1).

[0256] In some embodiments, the effector domain comprises the HNF3 C-terminal domain (HNF_C). The HNF_C domain may be from FOXA1 or FOXA2. In certain embodiments, the HNF_C domain comprises an EH1 (engrailed homology 1) motif.

[0257] In some embodiments, the effector domain may include an interferon regulatory factor 2 binding protein zinc finger domain (IRF-2BP1_2), a Cyt-b5 domain from the DNA repair factor HERC2 E3 ligase, a variant SH3 domain (SH3_9) from the bridging integrator 1 (BIN1), an HMG-box domain from the transcription factor TOX, or a ZF-C3HC4_2RING finger domain from the polycomb component PCGF2, a chromosome domain-helicase-DNA binding protein 3 (CHD3) domain, or a ZNF783 domain.

[0258] IV. Epigenetic editors

[0259] Provided herein are epigenetic editors (i.e., epigenetic editing systems) that, for example, use one or more DNA binding domains as described herein and one or more effector domains as described herein (e.g., epigenetic repressor domains) in any combination to direct epigenetic modifications to a target sequence in a gene of interest. A DNA binding domain (consistent with a guide polynucleotide such as a guide polynucleotide as described herein, wherein the DNA binding domain is a DNA binding domain of a polynucleotide guide) directs the effector domain to epigenetically modify the target sequence, resulting in gene repression or silencing, which can be persistent and can be inherited between cell generations. In some aspects, the epigenetic editors described herein can reversibly or irreversibly repress or silence genes in cells.

[0260] In a specific embodiment, the epigenetic editor described herein comprises one or more fusion proteins, each fusion protein comprising (1) a DNA binding domain and (2) an effector domain. The effector domain can be located on one or more fusion proteins contained in the epigenetic editor. For example, a single fusion protein can include all effector domains with a DNA binding domain. Alternatively, the effector domain or a subset thereof can be on a separate fusion protein, each fusion protein having a DNA binding domain (which can be the same or different). The fusion protein described herein may further include one or more linkers (e.g., a peptide linker), a detectable label, a nuclear localization signal (NLS), or any combination thereof. As used herein, "fusion protein" refers to a chimeric protein in which two or more coding sequences (e.g., for a DNA binding domain and / or an effector domain) are directly or indirectly covalently or non-covalently linked.

[0261] In some embodiments, the epigenetic editor described herein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10 or more effector (e.g., repression / repressor) domains, which may be the same or different. In certain embodiments, two or more of the effector domains work together. The combination of effector domains may include a DNA methylation domain, a histone deacetylation domain, a histone methylation domain, and / or a scaffold domain for raising any of the above. For example, the epigenetic editor described herein may include one or more transcriptional repressor domains (e.g., KRAB domains such as KOX1, ZIM3, ZFP28 or ZN627 KRAB) combined with one or more DNA methylation domains (e.g., DNMT domains) and / or a recruitment domain (e.g., DNMT3L domain). This epigenetic editor may include, for example, a KRAB domain, a DNMT3A domain, and a DNMT3L domain. In some embodiments, the epigenetic editor further comprises an additional effector domain (e.g., KAP1, MECP2, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, RBBP4, RCOR1, or SCML2 domain). In some embodiments, the additional effector domain is a CDYL2, TOX, TOX3, TOX4, or HP1a domain. For example, the epigenetic editor described herein can comprise a CDYL2 and / or TOX domain in combination with a KRAB domain (e.g., a KOX1KRAB domain).

[0262] A. Connector

[0263] A fusion protein as described herein may comprise one or more linkers connecting components of an epigenetic editor. The linker may be a peptide linker or a non-peptide linker.

[0264] In some embodiments, one or more linkers used in the epigenetic editors provided herein are peptide linkers, i.e., linkers comprising a peptide portion. The peptide linker can be any length suitable for the epigenetic editor fusion protein described herein. In some embodiments, the linker can comprise a peptide of 1 to 200 (e.g., 1 to 80) amino acids. In some embodiments, the length of the linker comprises 1 to 5, 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 80, 1 to 100, 1 to 150, 1 to 200, 5 to 10, 5 to 20, 5 to 30, 5 to 40, 5 to 60, 5 to 80, 5 to 100, 5 to 150, 5 to 200, 10 to 20, 10 to 30, 10 to 40, 10 to 50, 10 to 60, 10 to 80, 10 to 100, 10 to 150, 10 to 200, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 80, 60 to 100, 60 to 150, 60 to 200, 80 to 100, 80 to 150, 80 to 200, 100 to 150, 100 to 200, or 150 to 200 amino acids. Longer or shorter linkers are also contemplated. In some embodiments, the length of the linker is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 25, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95 or 100 amino acids. For example, the length of the peptide linker can be 4, 5, 16, 20, 24, 27, 32, 40, 64, 92 or 104 amino acids. The peptide linker can be a flexible or rigid linker. In specific embodiments, the peptide linker comprises the amino acid sequence of any one of SEQ ID NOs: 631-637 and 664-666, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical thereto.

[0265] In certain embodiments, the peptide linker is an XTEN linker. Such a linker may comprise a portion of an XTEN sequence (Schellenberger et al., Nat Biotechnol (2009) 27 (1): 1186-90), which is an unstructured hydrophilic polypeptide consisting only of residues G, S, P, T, E and A. As used herein, the term "XTEN" refers to a recombinant peptide or polypeptide lacking hydrophobic amino acid residues. An XTEN linker is typically unstructured and comprises a limited set of natural amino acids. The fusion of XTEN with a protein changes its hydrodynamic properties and reduces the clearance and degradation rate of the fusion protein. These XTEN fusion proteins are produced using recombinant technology, do not require chemical modification, and are degraded by natural pathways. The length of an XTEN linker may be, for example, 5, 10, 16, 20, 26 or 80 amino acids. In some embodiments, the length of an XTEN linker is 16 amino acids. In some embodiments, the length of an XTEN linker is 80 amino acids. In certain embodiments, an XTEN linker may be XTEN10, XTEN16, XTEN20 or XTEN80. In certain embodiments, the XTEN linker can comprise the amino acid sequence of any one of SEQ ID NOs: 638-643, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In a specific embodiment, the XTEN linker comprises the amino acid sequence of SEQ ID NO: 638. In a specific embodiment, the XTEN linker comprises the amino acid sequence of SEQ ID NO: 643.

[0266] In some embodiments, one or more linkers used in the epigenetic editors provided herein are non-peptide linkers, for example, the linker can be a carbon bond, a disulfide bond, or a carbon-heteroatom bond. In certain embodiments, the linker is a carbon-nitrogen bond of an amide bond. In certain embodiments, the linker is a cyclic or non-cyclic, substituted or unsubstituted, or branched or non-branched aliphatic or heteroaliphatic linker.

[0267] In some embodiments, one or more joints used in the epigenetic editors provided herein are polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). The joint may include, for example, a monomer, dimer, or polymer of an aminoalkanoic acid; an aminoalkanoic acid (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-pentanoic acid, etc.); a monomer, dimer, or polymer of aminocaproic acid (Ahx); or a polyethylene glycol moiety (PEG); or an aryl or heteroaryl moiety. In certain embodiments, the joint may be based on a carbocyclic moiety (e.g., cyclopentane or cyclohexane) or a benzene ring. The joint may include a functionalized moiety to facilitate attachment of a nucleophile (e.g., thiol, amino) from a peptide to a joint. Any electrophile may be used as part of a joint. Exemplary electrophiles include, but are not limited to, activated esters, activated amides, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0268] Various linker lengths and flexibilities can be used between any two components of an epigenetic editor (e.g., between an effector domain (e.g., a repressor domain) and a DNA binding domain (e.g., a Cas9 domain), between a first effector domain and a second effector domain, etc.). Linkers can range from very flexible linkers (e.g., glycine / serine-rich linkers) to more rigid linkers to achieve the optimal length for effector domain activity for a particular application. In some embodiments, a more flexible linker is a glycine / serine-rich linker (GS-rich linker), wherein more than 45% (e.g., more than 48%, 50%, 55%, 60%, 70%, 80%, or 90%) of the residues are glycine or serine residues. Non-limiting examples of GS-rich linkers are (GGGGS)n (SEQ ID NO: 1285), (G)n (SEQ ID NO: 1288), and W linkers (SEQ ID NO: 637). In some embodiments, the more rigid linker is of the form (EAAAK)n (SEQ ID NO: 1286), (SGGS)n (SEQ ID NO: 1287), and (XP)n (SEQ ID NO: 1289). In the formulas of the above flexible and rigid linkers, n can be any integer between 1 and 30. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the linker comprises a (GGS)n motif, wherein n is 1, 3, or 7 (SEQ ID NO: 1290). In some embodiments, the linker comprises a (GGGGS)n motif, wherein n is 4 (SEQ ID NO: 636).

[0269] In some embodiments, the linker in the epigenetic editor described herein comprises a nuclear localization signal, e.g., having an amino acid sequence of any one of SEQ ID NOs: 644-649. In some embodiments, the linker in the epigenetic editor described herein comprises an expression tag, e.g., a detectable tag, such as green fluorescent protein.

[0270] B. Nuclear localization signal

[0271] Fusion protein as described herein can include one or more nuclear localization signals, and in certain embodiments, can include two or more nuclear localization signals.For example, fusion protein can include 1,2,3,4,5,6,7,8,9,10 or more nuclear localization signals.As used herein, "nuclear localization signal" (NLS) is the amino acid sequence that guides protein to the nucleus.In certain embodiments, SV40 can be SV40 NLS (for example, with SEQ ID NO:644 amino acid sequence).Fusion protein can include NLS at its N-terminal, C-terminal or both, and / or NLS can be embedded in the middle of fusion protein (for example, at N or C-terminal of DNA binding domain or effector domain).

[0272] In some embodiments, the fusion protein may include two NLSs. The fusion protein may include two NLSs at its N-terminus or C-terminus. The fusion protein may include one NLS at its N-terminus and one NLS embedded in the middle of the fusion protein, or one NLS at its C-terminus and one NLS embedded in the middle of the fusion protein. The fusion protein may include two NLSs embedded in the middle of the fusion protein.

[0273] In some embodiments, the fusion protein can include four NLS. The fusion protein can include at least two (e.g., two, three, or four) NLS at its N-terminus or C-terminus. The fusion protein can include at least one (e.g., one, two, three, or four) NLS embedded in the middle of the fusion protein. In specific embodiments, the fusion protein can include two NLS at its N-terminus or two NLS at its C-terminus.

[0274] The NLS described herein can be an endogenous NLS sequence. In certain embodiments, the NLS described herein comprises the amino acid sequence of any one of SEQ ID NOs: 644-649, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the selected sequence. In a specific embodiment, the NLS comprises the amino acid sequence of SEQ ID NO: 644. Additional NLSs are known in the art.

[0275] In some embodiments, an epigenetic editor comprising a fusion protein comprising at least one NLS at the N-terminus and at least one NLS at the C-terminus can increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000%, or more compared to an epigenetic editor with a corresponding fusion protein without at least one NLS at the N-terminus and at least one NLS at the C-terminus.

[0276] In some embodiments, an epigenetic editor comprising a fusion protein comprising two NLSs at the N-terminus and two NLSs at the C-terminus can increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000% or more compared to an epigenetic editor with a corresponding fusion protein without two NLSs at the N-terminus and two NLSs at the C-terminus.

[0277] C. Labels

[0278] The epigenetic editors provided herein may include one or more additional sequences ("tags") for tracking, detecting, and locating the editor. In some embodiments, the epigenetic editor includes 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more detectable tags. Each detectable tag can be the same or different.

[0279] For example, epigenetic editor fusion protein can include cytoplasm localization sequence, output sequence (such as nuclear export sequence) or other localization sequence, and can be used for dissolving, purifying or detecting sequence tag of fusion protein.Suitable protein tag provided herein includes but is not limited to biotin carboxylase carrier protein (BCCP) tag, myc tag, calmodulin-tag, FLAG tag, hemagglutinin (HA) tag, polyhistidine tag (also referred to as histidine tag or His tag), maltose binding protein (MBP) tag, nus tag, glutathione-S-transferase (GST) tag, green fluorescent protein (GFP) tag, thioredoxin tag, S tag, Softag (for example, Softag 1 or Softag 3), streptococcus tag, biotin ligase tag, FlAsH tag, V5 tag and SBP tag.Other suitable sequence will be apparent to those skilled in the art.

[0280] D. Fusion protein configuration

[0281] The fusion proteins of the epigenetic editors described herein may have their components constructed in different configurations. For example, the DNA binding domain may be at the C-terminus, at the N-terminus, or between two or more epigenetic effector domains or additional domains. In some embodiments, the DNA binding domain is at the C-terminus of the epigenetic editor. In some embodiments, the DNA binding domain is at the N-terminus of the epigenetic editor. In some embodiments, the DNA binding domain is connected to one or more nuclear localization signals. In some embodiments, the DNA binding domain is flanked by epigenetic effector domains and / or additional domains on both sides. In some embodiments, wherein "DBD" indicates a DNA binding domain and "ED" indicates an effector domain, the epigenetic editor comprises the following configurations:

[0282] -N']-[ED1]-[DBD]-[ED2]-[C'

[0283] -N']-[ED1]-[DBD]-[ED2]-[ED3]-[C'

[0284] -N']-[ED1]-[ED2]-[DBD]-[ED3]-[C'

[0285] or

[0286] -N']-[ED1]-[ED2]-DBD]-[ED3]-[ED4]-[C'.

[0287] In some embodiments, the epigenetic editor comprises a DNA binding domain (DBD), a DNA methyltransferase (DNMT) domain, and a transcriptional repressor ("repressor") domain that represses or silences target gene expression. The DBD, DNMT, and transcriptional repressor domains can be any domain as described herein in any combination. The DBD, DNMT domain, and repressor domain can be in any configuration, for example, any one of the domains is at the N-terminus, C-terminus, or in the middle of the fusion protein. In some embodiments, the epigenetic editor comprises a fusion protein having the following configuration:

[0288] N']-[DNMT domain]-[DBD]-[repressor domain]-[C'

[0289] N']-[repressor domain]-[DBD]-[DNMT domain]-[C'

[0290] N']-[DNMT domain]-[repressor domain]-[DBD]-[C'

[0291] or

[0292] N']-[repressor domain]-[DNMT domain]-[DBD]-[C'.

[0293] In some embodiments, the connection structure "]-[" of any one of the epigenetic editor structures is a linker, e.g., a peptide linker; a detectable tag; a peptide bond; a nuclear localization signal; and / or a promoter or regulatory sequence. In the epigenetic editor structure, multiple connection structures "]-[" may be the same, or may each be a different linker, tag, NLS or peptide bond. In some embodiments, the DNMT domain may include any one of the domains in Table 7, or any combination or homolog thereof. In specific embodiments, the DNMT domain includes DNMT3A or a truncated version thereof, DNMT3L or a truncated version thereof, or both. In specific embodiments, the DBD is a DNA binding domain (e.g., dCas9) or a ZFP domain of a non-catalytically active polynucleotide guide. In certain embodiments, the repressor domain includes any one of the domains shown in Table 5 or 6, or any combination or homolog thereof. For example, the repressor domain may be a KRAB domain. In certain embodiments, the repressor domain is a ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1 or SCML2 domain, or a fusion of two of the domains (e.g., a fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1 KRAB). In specific embodiments, the repressor domain is a KRAB domain from ZFP28, ZN627, ZIM3 or KOX1.

[0294] In some embodiments, the base editor comprises a configuration selected from the group consisting of:

[0295] N']-[DNMT3A-DNMT3L]-[DBD]-[repressor]-[C'

[0296] N']-[repressor]-[DBD]-[DNMT3A-DNMT3L]-[C'

[0297] N']-[repressor]-[DBD]-[DNMT3A]-[C'

[0298] N']-[DNMT3A]-[DBD]-[repressor]-[C'

[0299] N']-[repressor]-[DBD]-[DNMT3A]-[DNMT3L]-[C'

[0300] N']-[DNMT3A]-[DNMT3L]-[DBD]-[repressor]-[C'

[0301] N']-[DNMT3A]-[DBD]-[C'

[0302] N']-[DBD]-[DNMT3A]-[C'

[0303] N']-[DNMT3L]-[DBD]-[C'

[0304] N']-[DBD]-[DNMT3L]-[C'

[0305] Wherein [DNMT3A-DNMT3L] indicates that the DNMT3A and DNMT3L domains are directly fused via a peptide bond, and wherein the connecting structure]-[ is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. The DBD, repressor, DNMT3A, and DNMT3L domains can be any domain as described herein in any combination. For example, the DNMT3A and DNMT3L domains can be selected from those in Table 7. In a specific embodiment, the DBD is a CRISPR-associated protein domain (e.g., dCas9) or a ZFP domain; the repressor domain is a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627; the DNMT3A domain is a human DNMT3A domain; and the DNMT3L domain is a human or mouse DNMT3L domain; any combination of these components is also contemplated by the present disclosure.

[0306] In some embodiments, the base editor comprises a configuration selected from the group consisting of:

[0307] N']-[DNMT3A]-[DBD]-[SETDB1]-[C'

[0308] N']-[DNMT3A]-[DNMT3L]-[DBD]-[SETDB1]-[C'

[0309] N']-[DNMT3A-DNMT3L]-[DBD]-[SETDB1]-[C'

[0310] N']-[SETDB1]-[DBD]-[DNMT3A]-[DNMT3L]-[C'

[0311] N']-[SETDB1]-[DBD]-[DNMT3A]-[C'

[0312] Wherein [DNMT3A-DNMT3L] indicates that the DNMT3A and DNMT3L domains are directly fused via a peptide bond, and wherein the connecting structure]-[ is any one of a linker, a detectable tag, an affinity domain, a peptide bond, a nuclear localization signal, a promoter, and / or a regulatory sequence as described herein. The DBD, SETDB1, DNMT3A, and DNMT3L domains can be any domain as described herein in any combination. In a specific embodiment, the DBD is a CRISPR-associated protein domain (e.g., dCas9) or a ZFP domain; the SETDB1 domain is derived from human SETDB1, ZIM3, ZFP28, or ZN627; the DNMT3A domain is a human DNMT3A domain; and the DNMT3L domain is a human or mouse DNMT3L domain; any combination of these components is also contemplated by the present disclosure.

[0313] Specific constructs contemplated herein include:

[0314] DNMT3A-DNMT3L-XTEN80-NLS-dCas9-NLS-XTEN16-KOX1 KRAB (configuration 1),

[0315] DNMT3A-DNMT3L-XTEN80-NLS-ZFP domain-NLS-XTEN16-KOX1KRAB (configuration 2),

[0316] NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS (configuration 3),

[0317] NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1 KRAB-NLS (configuration 4),

[0318] NLS-NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS-NLS (configuration 5), and

[0319] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1KRAB-NLS-NLS (Configuration 6).

[0320] DNMT3L and DNMT3A can be derived from human parent proteins, mouse parent proteins, or any combination thereof. In certain embodiments, DNMT3L and DNMT3A are derived from mouse and human parent proteins (mDNMT3L and hDNMT3A), respectively. In certain embodiments, DNMT3L and DNMT3A are both derived from human parent proteins (hDNMT3L and hDNMT3A). In some embodiments, dCas9 is dSpCas9. In some embodiments, KOX1 is human KOX1. Any one of configurations 1-6 is also contemplated, in which the KOX1 KRAB domain is replaced by a ZFP28, ZN627, or ZIM3 KRAB domain. In some embodiments, ZFP28, ZN627, and ZIM3 are human ZFP28, ZN627, and ZIM3, respectively. In a specific embodiment, the fusion construct may have the following configurations:

[0321] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-KOX1KRAB-NLS-NLS (configuration 7),

[0322] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-KOX1KRAB-NLS-NLS (configuration 8),

[0323] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZFP28KRAB-NLS-NLS (configuration 9),

[0324] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-ZFP28KRAB-NLS-NLS (configuration 10),

[0325] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZN627KRAB-NLS-NLS (configuration 11),

[0326] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-ZN627KRAB-NLS-NLS (configuration 12),

[0327] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZIM3KRAB-NLS-NLS (configuration 13), or

[0328] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP domain-XTEN16-ZIM3KRAB-NLS-NLS (Configuration 14).

[0329] In specific embodiments, the fusion constructs described herein can have Configuration 1 and comprise SEQ ID NO: 658, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In the following SEQ ID NO: 658, the XTEN linker is Underlined , W connector is Bold, with Underline and italics , NLS sequences are in bold, DNMT3A sequences are in italics, and DNMT3L sequences are Underlined and italic of , the dCas9 domain is bold and italicized, and the KOX1 KRAB domain is Underlined and bold :

[0330]

[0331]

[0332]

[0333] In specific embodiments, the fusion constructs described herein can have Configuration 2 and comprise SEQ ID NO: 659, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In the following SEQ ID NO: 659, the XTEN linker is Underlined , W connector is Bold, with Underline and italics , the NLS sequence is bold and underlined, the DNMT3A sequence is italicized, and the DNMT3L sequence is Take it down Strikethrough and italics , the ZFP domain is in bold, and the KOX1 KRAB domain is Underlined and bold The variable amino acids indicated by X are the amino acids of the DNA recognition helix of the zinc finger, and the italicized XX can be TR, LR, or LK.

[0334]

[0335]

[0336] In certain embodiments, the six "XXXXXXX" regions in SEQ ID NO: 659 comprise amino acid sequences that form zinc fingers. In the above sequences, [linker] represents a linker sequence. In some embodiments, one or both linker sequences may be TGSQKP (SEQ ID NO: 651). In some embodiments, one or both linker sequences may be TGGGGSQKP (SEQ ID NO: 652). In some embodiments, one linker sequence may have the amino acid sequence of SEQ ID NO: 651, and the other linker sequence may have the amino acid sequence of SEQ ID NO: 652.

[0337] In specific embodiments, the fusion construct described herein can have configuration 7 and comprise SEQ ID NO:660, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical thereto.

[0338] In specific embodiments, the fusion construct described herein can have configuration 9 and comprise SEQ ID NO:661, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical thereto.

[0339] In specific embodiments, the fusion construct described herein can have configuration 11 and comprise SEQ ID NO:662, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical thereto.

[0340] In specific embodiments, the fusion construct described herein can have configuration 13 and comprise SEQ ID NO:663, or a sequence at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical thereto.

[0341] In some embodiments, the fusion construct described herein (e.g., a fusion construct of any one of Configurations 1-14) is within an expression construct comprising a WPRE sequence, a polyadenylation site, or both. In certain embodiments, the WPRE sequence is in the 3' non-coding region. In certain embodiments, the WPRE sequence is upstream of the polyadenylation site. In specific embodiments, the expression construct comprises a fusion construct (e.g., any one of Configurations 1-14) and a WPRE sequence in the 3' non-coding region upstream of the polyadenylation site.

[0342] A variety of fusion proteins can be used to achieve activation or suppression of a target gene or multiple target genes. For example, an epigenetic editor fusion protein comprising a DNA binding domain (e.g., a dCas9 domain) and an effector domain can be co-delivered with two or more guide polynucleotides (e.g., gRNA), and each guide polynucleotide targets different target DNA sequences. The target sites of two DNA binding domains can be identical or adjacent to each other, or spaced such as about 100 base pairs, about 200 base pairs, about 300 base pairs, about 400 base pairs, about 500 base pairs, or about 600 or more base pairs. In addition, when targeting double-stranded DNA such as an endogenous locus, guide polynucleotides can target the same or different chains (one or more targeting positive chains and / or one or more targeting negative chains).

[0343] In some embodiments, an epigenetic editor targeting B2M is used in combination with an epigenetic editor targeting TRAC, TRBC, CIITA, PDCD1, TIM-3, TIGIT, LAG3, CTLA4, AAVS1, CCR5, TET2, TGFBR2, A2AR, CISH, PTPN11, PTPN6, PTPA, PTPN2, JUNB, TOX, TOX2, NR4A1, NR4A2, NR4A3, MAP4K1, REL, IRF4, DGKA, PIK3CD, HLA-A, USP16, DCK, FAS, or any combination thereof.

[0344] V. Target Sequence

[0345] The epigenetic editors in this article can be directed to target sequences in B2M to achieve epigenetic modification of the B2M gene.

[0346] As used herein, a "target sequence", "target site" or "target region" is a nucleic acid sequence present in a gene of interest; in some cases, the target sequence may be outside the gene of interest but in the vicinity thereof, where methylation of the target sequence or binding of a repressor represses expression of the gene. In some embodiments, the target sequence may be a hypomethylated or hypermethylated nucleic acid sequence.

[0347] The target sequence can be in any part of the target gene. In some embodiments, the target sequence is a part of the non-coding sequence of the gene or near it. In some embodiments, the target sequence is a part of the gene exon. In some embodiments, the target sequence is a part of the transcriptional regulatory sequence (such as a promoter or enhancer) of the gene or near it. In some embodiments, the target sequence is adjacent to, overlaps with, or covers a CpG island. In certain embodiments, the target sequence is located within about 3000, 2900, 2800, 2700, 2600, 2500, 2400, 2300, 2200, 2100, 2000, 1900, 1800, 1700, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, or 100 base pairs (bp) of the B2M TSS. In certain embodiments, the target sequence is within 500 bp of the B2M TSS. In certain embodiments, the target sequence is within 1000 bp of the B2M TSS.

[0348] In some embodiments, a target sequence can be hybridized to a guide polynucleotide sequence (e.g., a gRNA) complexed with a fusion protein comprising a DNA binding domain of a polynucleotide guide (e.g., a CRISPR protein such as dCas9) and an effector domain. The guide polynucleotide sequence can be designed to have complementarity with the target sequence, or to have identity with the opposite strand of the target sequence. In some embodiments, the guide polynucleotide comprises a spacer sequence that is approximately 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a protospacer sequence in the target sequence. In a specific embodiment, the guide polynucleotide comprises a spacer sequence that is 100% identical to a protospacer sequence in the target sequence.

[0349] In some embodiments, wherein the DNA binding domain of the epigenetic editor described herein is a zinc finger array, the target sequence can be recognized by the zinc finger array.

[0350] In some embodiments, where the DNA binding domain of an epigenetic editor described herein is a TALE, the target sequence can be recognized by the TALE.

[0351] The target sequence described herein can be specific to a copy of the target gene, or can be specific to an allele of the target gene. Therefore, epigenetic modification and its expression regulation can be specific to a copy or an allele of the target gene. For example, an epigenetic editor can suppress the expression of a specific copy (e.g., a copy associated with a disease or disorder, or a copy carrying a mutation associated with a disease or disorder) of a target sequence recognized by a DNA binding domain.

[0352] In some embodiments, the target B2M genomic region can be within the sequence shown in SEQ ID NO:1283 or 1284.

[0353] VI. Epigenetic modification

[0354] Epigenetic editors described herein can carry out sequence-specific epigenetic modification (e.g., chemically modified changes) to target genes containing target sequences. This epigenetic regulation may be safer and more reversible than regulation due to gene editing (e.g., generation using DNA double-strand breaks). In some embodiments, epigenetic regulation can reduce or silence target genes. In some embodiments, modification is at a specific site of the target sequence. In some embodiments, modification is at a specific allele of the target gene. Therefore, epigenetic modification may cause the expression of a copy of a target gene containing a specific allele to be regulated (e.g., reduced), while the expression of another copy of the target gene is not affected. In some embodiments, specific alleles are associated with diseases, conditions, or illnesses.

[0355] In some embodiments, epigenetic modification reduces or eliminates the transcription of a target gene carrying a target sequence. In some embodiments, epigenetic modification reduces or eliminates the transcription of a copy of a target gene carrying a specific allele identified by an epigenetic editor. In some embodiments, the epigenetic editor reduces the level of a protein encoded by the target gene or eliminates its expression. In some embodiments, the epigenetic editor reduces the level of a protein encoded by a copy of a target gene carrying a specific allele identified by an epigenetic editor or eliminates its expression. The target B2M gene can be epigenetically modified in vitro, in vitro or in vivo.

[0356] The effector domain of the epigenetic editor described herein can change (for example, deposit or remove) the chemical modification at the nucleotide of the target gene or the histone associated with the target gene. The chemical modification can be changed at a single nucleotide or a single histone, or can be changed at 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000 or more nucleotides.

[0357] In some embodiments, the effector domain of the epigenetic editor described herein can change the CpG dinucleotides within the target gene. In some embodiments, compared to the original state of the gene or the gene in comparable cells not contacted with the epigenetic editor, all CpG dinucleotides within 2000, 1500, 1000, 500 or 200bp of the target sequence flank (e.g., in the change site described herein) are changed according to the modification type described herein. In some embodiments, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, or more CpG dinucleotides are altered compared to the original state of the gene or the gene in a comparable cell that has not been contacted with the epigenetic editor. In some embodiments, at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the CpG dinucleotides are altered compared to the original state of the gene or the gene in comparable cells that have not been contacted with an epigenetic editor. In some embodiments, a single CpG dinucleotide is altered compared to the original state of the gene or the gene in comparable cells that have not been contacted with an epigenetic editor.

[0358] The effector domain of the epigenetic editor described herein can change the histone modification state of the histone associated with or combined with the target gene.For example, the effector domain can be deposited on one or more lysine residues of the histone tail of the histone associated with the target gene to modify.In some embodiments, the effector domain can cause the deacetylation of one or more histone tails of the histone associated with the target gene, thereby reducing or silencing the expression of the target gene.In some embodiments, the histone modification state is a methylation state.For example, the effector domain can cause H3K9, H3K27 or H4K20 methylation (for example, one or more of H3K9me2, H3K9me3, H3K27me2, H3K27me3 and H4K20me3 methylation) at one or more histone tails associated with the target gene, thereby reducing or silencing the expression of the target gene.

[0359] In some embodiments, compared to the original state of the chromosome or the chromosome in comparable cells not contacted with the epigenetic editor, all histone tails of the histone bound to the DNA nucleotides within 2000, 1500, 1000, 500 or 200bp of the target sequence flank are changed according to the modification type described herein. In some embodiments, compared to the original state of the chromosome or the chromosome in comparable cells not contacted with the epigenetic editor, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120 or more histone tails of the binding histones are changed. In some embodiments, compared to the original state of the chromosome or the chromosome in comparable cells not contacted with the epigenetic editor, at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98% or 99% of the histone tails of the binding histones are changed. For example, compared to the original state of the chromosome or the chromosome in comparable cells not contacted with the epigenetic editor, a single histone tail of the binding histone may be changed. As another example, compared to the original state of the chromosome or the chromosome in comparable cells not contacted with the epigenetic editor, a single binding histone octamer may be changed.

[0360] The chemical modification deposited at the target gene DNA nucleotide or histone residue can be located at or adjacent to the target sequence in the target gene. In some embodiments, the effector domain of the epigenetic editor described herein changes the chemical modification state of the nucleotides or histone tails bound to nucleotides 100-200, 200-300, 300-400, 400-55, 500-600, 600-700 or 700-800 nucleotides 5' or 3' of the target sequence in the target gene. In some embodiments, the effector domain alters the chemical modification state of nucleotides or histone tails that bind to nucleotides within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 nucleotides flanking a target sequence. As used herein, "flanking" refers to the nucleotide positions from the 5' to the 5' end and from the 3' to the 3' end of a particular sequence (e.g., a target sequence).

[0361] In some embodiments, the effector domain mediates or induces a chemical modification change in a nucleotide or histone tail that binds to a nucleotide that is away from the target sequence. This modification can start near the target sequence and can then extend to one or more nucleotides in the target gene that are away from the target sequence. For example, the effector domain can initiate a change in the chemical modification state of one or more nucleotides or one or more histone residues bound to one or more nucleotides within 10, 20, 30, 400, 500 nucleotides flanking the target sequence, and the change in the chemical modification state can extend to at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, 3000 or more nucleotides from the target sequence in the target gene, upstream or downstream of the target sequence. In certain embodiments, chemical modification can start from less than 2, 3, 5, 10, 20, 30, 40, 50 or 100 nucleotides in the target gene and extend to at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000 or more nucleotides in the target gene. In some embodiments, chemical modification extends to nucleotides in the entire target gene. Additional proteins or transcription factors, for example, transcription repressors, methyltransferases, or transcriptional regulatory scaffold proteins, can participate in the expansion of chemical modification. Alternatively, an epigenetic editor may be involved alone.

[0362] In some embodiments, compared to a control cell, a control tissue, or a control subject (e.g., in the absence of an epigenetic editor), the epigenetic editor described herein reduces the expression of the target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99% or more, as measured by transcription of the target gene in a cell, tissue, or subject. In some embodiments, compared to a control cell, a control tissue, or a control subject, the epigenetic editor described herein reduces the expression of a copy of the target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99% or more, as measured by transcription of a copy of the target gene in a cell, tissue, or subject. In certain embodiments, the copy of the target gene contains a specific sequence or allele recognized by the epigenetic editor. In a specific embodiment, the epigenetic modified copy encodes a functional protein, and thus the epigenetic editor disclosed herein can reduce or eliminate the expression and / or function of the protein. For example, compared to a control cell, a control tissue or a control subject, the epigenetic editor described herein can reduce the expression and / or function of a protein encoded by a target gene in a cell, tissue or subject by at least 3 times, at least 4 times, at least 5 times, at least 6 times, at least 7 times, at least 8 times, at least 9 times, at least 10 times, at least 11 times, at least 12 times, at least 13 times, at least 14 times, at least 15 times, at least 20 times, at least 25 times, at least 30 times, at least 35 times, at least 40 times, at least 45 times, at least 50 times, at least 60 times, at least 70 times, at least 80 times, at least 90 times or at least 100 times.

[0363] Modulation of target gene expression can be determined by determining any parameter that is indirectly or directly affected by target gene expression. Such parameters include, for example, changes in RNA or protein levels; changes in protein activity; changes in product levels; changes in downstream gene expression; changes in transcription or activity of reporter genes such as, for example, luciferase, CAT, β-galactosidase, or GFP; changes in signal transduction; changes in phosphorylation and dephosphorylation; changes in receptor-ligand interactions; changes in second messengers such as, for example, cGMP, cAMP, IP3, and Ca 2+The concentration of the gene or proteins can be measured by a method such as a change in the concentration of the gene or proteins; a change in cell growth; a change in angiogenesis; and / or a change in any functional effect of gene expression. The measurement can be performed in vitro, in vivo and / or ex vivo, and can be prepared by conventional methods, for example, measurement of RNA or protein levels, measurement of RNA stability and / or identification of downstream or reporter gene expression. The readout can be performed by, for example, chemiluminescence, fluorescence, colorimetric reaction, antibody binding, inducible markers, ligand binding assays, changes in intracellular second messengers such as cGMP and inositol triphosphate (IP3), changes in intracellular calcium levels; cytokine release, etc.

[0364] Methods for determining the expression level of a gene (e.g., a target of an epigenetic editor) can include, for example, by reverse transcription PCR, quantitative RT-PCR, droplet digital PCR (ddPCR), RNA blotting, RNA sequencing, DNA sequencing (e.g., sequencing of complementary deoxyribonucleic acid (cDNA) obtained from RNA) to determine the transcript level of a gene; Next-generation (Next-Gen) sequencing, nanopore sequencing, pyrophosphate sequencing or nanostring sequencing. Protein levels expressed from genes can be determined, for example, by Western blotting, enzyme-linked immunosorbent assay, mass spectrometry, immunohistochemistry or flow cytometry analysis. Gene expression product levels can be normalized relative to internal standards, such as total messenger ribonucleic acid (mRNA) or the expression level of a specific gene (e.g., housekeeping gene).

[0365] In some embodiments, the effect of the epigenetic editor in regulating the expression of the target gene can be checked using a reporter system. For example, the epigenetic editor can be designed to target a reporter gene encoding a reporter protein (such as a fluorescent protein). The expression of the reporter gene in such a model system can be monitored by, for example, flow cytometry, fluorescence activated cell sorting (FACS) or fluorescence microscopy. In some embodiments, a cell group can be transfected with a vector containing a reporter gene. A vector can be constructed so that when the vector is transfected with cells, the reporter gene is expressed. Suitable reporter genes include genes encoding fluorescent proteins, such as green, yellow, cherry, cyan or orange fluorescent proteins. The cell group carrying the reporter system can be transfected with DNA, mRNA or a vector encoding an epigenetic editor targeting a reporter gene.

[0366] VII. Epigenetically modified cells

[0367] In one aspect, the present disclosure provides cells that have been modified using one or more epigenetic editors described herein. In some embodiments, a nucleic acid molecule encoding the epigenetic editor or its components is administered to a cell. Any type of cell can be modified as described herein. Cells can be modified in vitro, in vivo or ex vivo. Cells suitable for modification can be obtained from patients or healthy donors.

[0368] In some aspects, the cell is an immune cell. Immune cells can include T cells, B cells, natural killer (NK) cells, dendritic cells, and monocytes / macrophages. In some aspects, the cell is an α / β T cell. In some aspects, the cell is a γ / δ T cell. In some embodiments, the cell is a cytotoxic T cell, for example, a CD8 + Cytotoxic T cells. In some embodiments, the cells are T helper cells, e.g., CD4 + In some aspects, the cell is a T helper cell. In some aspects, the cell is a regulatory T cell. In some aspects, the cell is a NK cell. In some aspects, the cell is a dendritic cell. In some aspects, the cell is a macrophage.

[0369] In some aspects, the cell is a stem cell. "Stem cell" refers to an undifferentiated cell that is capable of producing more stem cells of the same type indefinitely and from which other specialized cells can be produced by differentiation. Adult stem cells are usually multipotent, while induced or embryo-derived stem cells are pluripotent.

[0370] In some aspects, the cell is a progenitor cell. "Progenitor cell" refers to a cell that is capable of differentiating to form one or more types of cells, but has limited self-renewal in vitro and in vivo.

[0371] In some embodiments, the cells are capable of differentiating into the above-mentioned immune cells. The cells can be, for example, embryonic stem cells (ESCs), hematopoietic stem cells (HSCs), hematopoietic progenitor cells (HPCs), or hematopoietic stem and progenitor cells (HSPCs). "Hematopoietic stem and progenitor cells" or "HSPCs" refer to cells that express the antigen marker CD34 (CD34 + In a specific embodiment, the term "HSPC" refers to cells that express the antigenic marker CD34 (CD34 + ) and the absence of lineage (lin) markers. + and / or Lin - The cell population includes hematopoietic stem cells and hematopoietic progenitor cells.

[0372] In some embodiments, the cell is an induced pluripotent stem cell (iPSC) reprogrammed from a somatic cell (eg, a T cell).

[0373] In some embodiments, the cells are obtained from umbilical cord blood of a healthy donor. In some embodiments, the cells are obtained from adult peripheral blood or mobilized from bone marrow of a healthy donor.

[0374] In some embodiments, the cell as described herein is modified by a method comprising transfecting cells with a system comprising (a) one or more epigenetic editors described herein, or (b) a nucleic acid molecule encoding the epigenetic editor. In certain embodiments, the modified cell is a T cell. In some embodiments, the modified T cell expresses one or more epigenetic editors, which can selectively reduce or silence the expression of one or more target genes in the cell. In a specific embodiment, the target gene is B2M. In some embodiments, the T cell is modified ex vivo. In some embodiments, the modified T cell may further express an engineered TCR or CAR for at least one antigen expressed on the surface of a target cell (e.g., a malignant or infected cell). In some embodiments, the modified T cell does not express at least one gene encoding an endogenous TCR component. In a specific embodiment, the modified T cell is non-allogeneic reactive. In a specific embodiment, the modified T cell is particularly suitable for allogeneic transplantation.

[0375] VIII. Pharmaceutical composition

[0376] In one aspect, the present disclosure provides a pharmaceutical composition comprising as an active ingredient (or as the only active ingredient) one or more epigenetic editors described herein or components thereof (e.g., fusion proteins and / or guide polynucleotides), or nucleic acid molecules encoding the epigenetic editors or components thereof. For example, a pharmaceutical composition may comprise a nucleic acid molecule (and a guide polynucleotide, where applicable) encoding a fusion protein of an epigenetic editor described herein. In some embodiments, a separate pharmaceutical composition comprises a fusion protein and a guide polynucleotide.

[0377] In some embodiments, the present disclosure provides a pharmaceutical composition comprising as an active ingredient (or as the only active ingredient) a cell that has undergone epigenetic modification mediated or induced by (a) one or more epigenetic editors provided herein, for example, wherein a nucleic acid molecule encoding the epigenetic editor is administered to the cell ex vivo.

[0378] Typically, the epigenetic editors or components thereof described herein, nucleic acid molecules encoding the epigenetic editors or components thereof, or cells modified by the epigenetic editors of the present disclosure are suitable for administration as a formulation in combination with one or more pharmaceutically acceptable excipients, for example, as described below.

[0379] The term "excipient" is used herein to describe any ingredient other than the compounds of the present disclosure. The choice of excipient will depend to a large extent on factors such as the specific mode of administration, the effect of the excipient on solubility and stability, and the nature of the dosage form. As used herein, "pharmaceutically acceptable excipients" include any and all physiologically compatible solvents, dispersion media, coatings, antibacterial and antifungal agents, isotonic agents, and absorption delaying agents. Some examples of pharmaceutically acceptable excipients are water, saline, phosphate buffered saline, dextrose, glycerol, ethanol, etc., and combinations thereof. In many cases, it is preferred to include isotonic agents in the composition, such as sugars, polyols such as mannitol, sorbitol, or sodium chloride. Another example of a pharmaceutically acceptable substance is a wetting agent or a small amount of auxiliary substance, such as a wetting agent or emulsifier, preservative, or buffer, which enhances the shelf life or effectiveness of the antibody.

[0380] The preparation of the pharmaceutical composition suitable for parenteral administration generally comprises the active ingredient in combination with a pharmaceutically acceptable carrier (e.g., sterile water or sterile isotonic saline). Such preparations can be prepared, packaged or sold in a form suitable for push injection or continuous administration. The pharmaceutical composition described herein can be administered to a subject, for example, subcutaneously, intradermally, intratumorally, intranodally, intramuscularly, intravenously, intralymphatically or intraperitoneally. In a specific embodiment, the pharmaceutical composition of the present disclosure is administered intravenously to a subject.

[0381] IX. Delivery Method

[0382] In some embodiments, the epigenetic editor or its components are introduced into the target cell in the form of a nucleic acid molecule encoding the epigenetic editor or its components; therefore, the pharmaceutical composition herein comprises a nucleic acid molecule. Such nucleic acid molecules may be, for example, DNA, RNA or mRNA, and / or a modified nucleic acid sequence (e.g., with chemical modification, 5' cap or one or more 3' modifications). In some embodiments, the nucleic acid molecule may be delivered as naked DNA or RNA, for example, by transfection or electroporation, or may be conjugated with a molecule (e.g., N-acetylgalactosamine) that promotes uptake by the target cell. In some embodiments, the nucleic acid molecule may be in a nucleic acid expression vector, which may include an expression control sequence, such as a promoter, an enhancer, a transcription signal sequence, a transcription termination sequence, an intron, a polyadenylation signal, a Kozak consensus sequence, an internal ribosome entry site (IRES), etc. Such expression control sequences are well known in the art. The vector may also include a sequence encoding a signal peptide (e.g., for nuclear localization, nucleolar localization, or mitochondrial localization), which is associated with a sequence encoding a protein (e.g., inserted or fused).

[0383] Examples of vectors include, but are not limited to, plasmid vectors; viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retrovirus (e.g., murine leukemia virus or spleen necrosis virus, derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus); and other recombinant vectors. In certain embodiments, the vector is a plasmid or viral vector. Viral particles or virus-like particles (VLPs) can also be used to deliver nucleic acid molecules encoding epigenetic editors or components thereof as described herein. For example, "empty" viral particles can be assembled to contain any suitable cargo. Viral vectors and viral particles can also be engineered to incorporate targeting ligands to change target tissue specificity.

[0384] In certain embodiments, an epigenetic editor or a component thereof as described herein is encoded by a nucleic acid sequence present in one or more viral vectors or a suitable capsid protein of any viral vector. Examples of viral vectors include adeno-associated viral vectors (e.g., derived from AAV3, AAV3b, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh8, AAV10, and / or variants thereof); retroviral vectors (e.g., Maloney murine leukemia virus, MML-V), adenoviral vectors (e.g., AD100), lentiviral vectors (e.g., vectors based on HIV and FIV), and herpes virus vectors (e.g., HSV-2).

[0385] In some embodiments, delivery involves an adeno-associated virus (AAV) vector. When the DNA binding domain of the epigenetic editor fusion protein is a zinc finger array, AAV vector delivery may be particularly useful. Without wishing to be bound by any theory, the smaller size of the zinc finger array can allow this fusion protein to be conveniently packaged in a viral vector such as an AAV vector compared to a larger DNA binding domain such as a Cas protein domain.

[0386] Any AAV serotype, such as a human AAV serotype, can be used for the AAV vectors as described herein, including but not limited to AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10) and AAV serotype 11 (AAV11) and variants thereof. In some embodiments, the AAV variant has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% amino acid sequence identity with wild-type AAV. In certain embodiments, the AAV variant can be engineered so that its capsid protein has reduced immunogenicity or enhanced transduction ability in humans. In some cases, one or more regions of at least two different AAV serotype viruses are rearranged and recombined to generate chimeric variants. For example, a chimeric AAV may comprise an inverted terminal repeat (ITR) that is a heterologous serotype compared to the serotype of the capsid. The resulting chimeric AAV may have different antigenic reactivity or recognition compared to its parental serotype. In some embodiments, chimeric variants of AAV include amino acid sequences from 2, 3, 4, 5 or more different AAV serotypes.

[0387] Non-viral systems are also contemplated for delivery as described herein.Non-viral systems include, but are not limited to, nucleic acid transfection methods, including electroporation, sonoporation, calcium phosphate transfection, microinjection, DNA gene guns, lipid-mediated transfection, transfection by heat shock, transfection mediated by compressed DNA, lipofection, cationic agent-mediated transfection, and transfection with liposomes, immunoliposomes, exosomes, or cationic surface amphiphiles (CFA).In certain embodiments, one or more mRNAs encoding epigenetic editor fusion proteins as described herein can be co-electroporated with one or more guide polynucleotides (e.g., gRNA) as described herein.An important category of non-viral nucleic acid vectors is nanoparticles, which can be organic (e.g., lipids) or inorganic (e.g., gold).For example, in certain embodiments of the present disclosure, organic (e.g., lipids and / or polymers) nanoparticles may be suitable for use as delivery vehicles.

[0388] In some embodiments, lipid nanoparticles (LNP) are used to deliver. The size of LNP compositions is usually in micrometer or smaller order of magnitude, and can include lipid bilayer. In some embodiments, LNP refers to any particle with a diameter less than 1000nm, 500nm, 250nm, 200nm, 150nm, 100nm, 75nm, 50nm or 25nm. In some embodiments, the size range of nanoparticles can be 1-1000nm, 1-500nm, 1-250nm, 25-200nm, 25-100nm, 35-75nm or 25-60nm. Nanoparticle compositions include lipid nanoparticles (LNP), liposomes (e.g., lipid vesicles) and lipid complexes.

[0389] LNP as described herein can be made of cationic lipids, anionic lipids or neutral lipids. In some embodiments, LNP can include neutral lipids, such as fusogenic phospholipids 1,2-dioleoyl-sn-glycero-3-phosphoethanolamine (DOPE) or membrane component cholesterol, as auxiliary lipids to enhance transfection activity and nanoparticle stability. In some embodiments, LNP can include hydrophobic lipids, hydrophilic lipids or both hydrophobic and hydrophilic lipids. Any lipid or lipid combination known in the art can be used to produce LNP. Lipid can be combined to produce LNP with any molar ratio. In some embodiments, LNP is a LNP targeting T cells (e.g., preferentially or specifically targeting T cells).

[0390] X. Epigenetic editors and therapeutic uses of modified cells

[0391] The present disclosure also provides a method for treating or preventing a condition in a subject, comprising administering to the subject a) one or more epigenetic editors as described herein, b) a nucleic acid molecule encoding an epigenetic editor, c) a cell modified by an epigenetic editor, or d) a pharmaceutical composition comprising any one of a)-c).

[0392] In one aspect, the epigenetic editor can achieve epigenetic modification of a target polynucleotide sequence in a target gene associated with a disease, condition, or disorder of a subject, thereby regulating the expression of the target gene to treat or prevent the disease, condition, or disorder. In some embodiments, the epigenetic editor reduces the expression of the target gene to a sufficient degree to achieve a desired effect, e.g., a therapeutically relevant effect, such as preventing or treating a disease, condition, or disorder.

[0393] In one aspect, cells (e.g., allogeneic cells) modified by one or more epigenetic editors disclosed herein can be administered as a drug to a subject suffering from a disease, condition or illness, so as to treat a disease, condition or illness. In some embodiments, the subject is administered allogeneic T cells, which have been epigenetically modified as described herein, for example, with reduced or silent B2M expression. In some embodiments, the modified T cells further express engineered TCR or CAR directly for at least one antigen expressed on the surface of a target cell (e.g., malignant or infected cell). In some embodiments, the modified T cells do not express at least one gene encoding endogenous TCR components.

[0394] In some embodiments, the subject can be a mammal, such as a human. In some embodiments, the subject is selected from a non-human primate, such as a chimpanzee, a cynomolgus monkey, a monkey or a macaque, and other apes and monkey species.

[0395] XI. definition

[0396] As used herein, the term "nucleic acid" refers to any oligonucleotide or polynucleotide containing nucleotides (e.g., deoxyribonucleotides or ribonucleotides) in single-stranded or double-stranded form, and includes DNA and RNA. "Nucleotide" contains sugar deoxyribose (DNA) or ribose (RNA), bases and phosphate groups, and is linked together by the phosphate groups. "Bases" include purines and pyrimidines, which include natural compounds such as adenine, thymine, guanine, cytosine, uracil, inosine and natural analogs; and synthetic derivatives of purines and pyrimidines, including but not limited to modified forms of placing new reactive groups such as amines, alcohols, thiols, carboxylates, halogenated alkyls, etc. Nucleic acids may contain known nucleotide analogs and / or modified backbone residues or bonds, which may be synthetic, naturally occurring and non-naturally occurring. Such nucleotide analogs, modified residues and modified bonds are well known in the art, and can provide nucleic acid molecules with enhanced cellular uptake, reduced immunogenicity and / or increased stability in the presence of nucleases.

[0397] As used herein, an "isolated" or "purified" nucleic acid molecule is a nucleic acid molecule that is free from its natural environment. For example, an "isolated" or "purified" nucleic acid molecule (1) has been separated from the nucleic acid of its source, genomic DNA or cellular RNA; and / or (2) does not exist in nature. In some embodiments, an "isolated" or "purified" nucleic acid molecule is a recombinant nucleic acid molecule.

[0398] It should be understood that in addition to the specific proteins and nucleic acid molecules mentioned herein, the present disclosure also contemplates the use of variants, derivatives, homologs and fragments thereof. Variants of any given sequence may have a specific residue sequence (whether amino acids or nucleic acid residues) modified in a manner that the polypeptide or polynucleotide in question substantially retains at least one of its endogenous functions. Variant sequences can be obtained by adding, deleting, replacing, modifying, substituting and / or mutating at least one residue (in some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15 or 20 residues) present in a naturally occurring sequence. For specific proteins described herein (e.g., KRAB, dCas9, DNMT3A and DNMT3L proteins described herein), the present disclosure also contemplates any of the naturally occurring forms of the protein, or variants or homologs retaining at least one of its endogenous functions (e.g., at least 50%, 60%, 70%, 80%, 90%, 85%, 96%, 97%, 98% or 99% of its function compared to the specific protein described).

[0399] As used herein, any polypeptide or homologue of the nucleic acid sequence considered herein includes sequences with certain homology to wild-type amino acids and nucleic acid sequences. Homologous sequences can include sequences at least 50%, 55%, 65%, 75%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% identical to the subject sequence, such as amino acid sequences. In the context of amino acids or nucleotide sequences, the term "identity percentage" refers to the percentage of identical residues in two sequences when the maximum correspondence is compared. In some embodiments, the length of the reference sequence compared for comparison purposes is at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80% or 90% or 100%) of the reference sequence. Sequence identity can be measured using sequence analysis software (e.g., sequence analysis software packages, BLAST, BESTFIT, GAP or PILEUP / PRETTYBOX programs of Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705). Such software matches identical or similar sequences by assigning homology to various substitutions, deletions and / or other modifications. In an exemplary method for determining the degree of identity, the BLAST program can be used, where probability scores between e-3 and e-100 indicate closely related sequences.

[0400] For example, you can The percentage identity of two nucleotide or polypeptide sequences is determined using default parameters (available on the website of the National Center for Biotechnology Information at the U.S. National Library of Medicine). In some embodiments, the length of the reference sequence aligned for comparison purposes is at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or 90%) of the reference sequence.

[0401] It will be appreciated that the numbering of specific positions or residues in a polypeptide sequence depends on the specific protein and numbering scheme used. For example, the numbering of a precursor of a mature protein and the mature protein itself may be different, and sequence differences between species may also affect the numbering. One skilled in the art will be able to identify any homologous proteins and respective residues in the respective encoding nucleic acids by methods well known in the art, for example, by sequence alignment and determination of homologous residues.

[0402] The term "regulate" or "change" refers to the change in the amount, degree or degree of function. For example, an epigenetic editor as described herein can regulate the activity of a promoter sequence by binding to a motif in a promoter, thereby inducing, enhancing or repressing the transcription of a gene operably connected to a promoter sequence. As other examples, an epigenetic editor as described herein can block RNA polymerase transcription of a gene, or can inhibit the translation of an mRNA transcript. When referring to the use of an epigenetic editor as described herein or its components, the terms "inhibit", "repress", "suppress", "silence", etc. refer to the activity of a nucleic acid sequence or protein in the absence of an epigenetic editor or its components, reducing or preventing the activity (e.g., transcription) of a nucleic acid sequence (e.g., a target gene) or a protein. The term may include partially or completely blocking activity, or preventing or delaying activity. The inhibited activity can be, for example, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% less than the activity of a control, or can be, for example, at least 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, or 10-fold less than the activity of a control.

[0403] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one skilled in the art, which will depend in part on how the value is measured or determined, for example, the limitations of the measurement system. For example, depending on the practice of a given value, "about" can mean within one or more than one standard deviation. When a particular value is described in the application and claims, the term "about" should be construed to mean an acceptable error range for that particular value unless otherwise indicated.

[0404] The scope provided herein should be understood as the abbreviation of all values ​​in the scope. For example, the scope of 1 to 50 should be understood as including any number, combination of numbers or subranges selected from the group consisting of the following: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49 or 50, and all intermediate decimal values ​​between the aforementioned integers, for example, 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8 and 1.9. About subrange, " nested subrange " extending from any end point of scope is particularly considered. For example, nested sub-ranges of the exemplary range of 1 to 50 may include 1 to 10, 1 to 20, 1 to 30, and 1 to 40 in one direction, or 50 to 40, 50 to 30, 50 to 20, and 50 to 10 in the other direction.

[0405] Unless otherwise defined herein, scientific and technical terms used in conjunction with the present disclosure shall have the meanings commonly understood by those of ordinary skill in the art. Exemplary methods and materials are described below, although methods and materials similar or equivalent to the methods and materials described herein may also be used in the practice or testing of the present disclosure. In the event of a conflict, this specification, including definitions, shall prevail. In addition, unless the context otherwise requires, singular terms shall include plural numbers, and plural terms shall include singular numbers. Throughout this specification and embodiments, the words "having" and "including" or variants, such as "having", "having", "including" or "including" shall be understood to mean including a specified integer or integer group, but do not exclude any other integer or integer group. Unless otherwise stated, the narration of the element list herein includes any single or any combination of elements. The narration of the embodiment herein includes the embodiment as a single embodiment, or in combination with any other embodiment herein. All publications, patents, patent applications, and other references mentioned herein are incorporated by reference in their entirety. To the extent that the references incorporated by reference contradict the disclosure contained in the specification, this specification is intended to replace and / or take precedence over any such contradictory materials. Although a number of documents are cited herein, this citation does not constitute an admission that any of these documents forms part of the common general knowledge in the art.

[0406] In accordance with the present disclosure, back-references in dependent claims are meant as shorthand to directly and unambiguously disclose every claim combination indicated by the back-reference. Additionally, the headings herein are created for organizational convenience and are not intended to limit the scope of the claimed invention in any way.

[0407] In order to better understand the present disclosure described, the following examples are set forth. These examples are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure in any way. Example

[0408] Example 1: Design and synthesis of fusion protein

[0409] A fusion protein ("CRISPR-off") comprising dCas9, DNMT3A, DNMT3L and KOX1KRAB was generated. From N-terminus to C-terminus, the protein has the following functional domains and linker: huDNMT3A-Linker-huDNMT3L-XTEN80-NLS-dspCas9-NLS-XTEN16-huKOX1KRAB (SEQ ID NO: 658). The CRISPR-off plasmid construct is described in et al., Cell(2021)184(9):2503-19.

[0410] ZF fusion proteins ("ZF-off") comprising DNMT3A, 3L and KOX1 KRAB were also generated. These fusion proteins had the following general structure: huDNMT3A-linker-huDNMT3L-XTEN80-NLS-ZFP domain-NLS-XTEN16-huKOX1 KRAB (SEQ ID NO: 659).

[0411] Example 2: Selection of the B2Mg region for RNA targeting

[0412] Using the Benchling gRNA platform for humans (GRCh38), gRNAs targeting the genomic region within 1 kb of the TSS of the human B2M gene were computationally designed. First, gRNAs containing poly-TTTT sequences were discarded. CasOFFinder (Bae et al., Bioinformatics (2014) 30 (10): 1473-5) was used for gRNA off-target analysis. If the gRNA matched multiple positions in the target genome, it would be discarded.

[0413] The final set of 258 gRNA sequences were screened in GripTiteTM HEK 293 cells for primary screening. DNA plasmids containing the coding sequences of gRNAs under the control of the U6 promoter were ordered from a supplier.

[0414] Example 3: Selection of ZFP target sites and design of ZFP

[0415] A library of two-finger ZFPs (2F units), each recognizing a 6bp DNA site, was used to design a larger six-finger ZFP array targeting an 18bp DNA binding site. The source of the 2F unit is a set of three-finger zinc finger proteins that have been selected to bind to specific targets using the bacteria-2-hybrid (B2H) selection system (Hurt et al., PNAS (2003) 100: 12271-6; Maeder et al., Mol Cell (2008) 31 (2): 294-301). A list of targetable DNA sites was created by generating all possible triplet combinations of the 6bp binding sites represented in the library and allowing 0 or 1bp between the 6bp target sites. To identify ZF target sites in human B2M, sequences within 1 kb of TSS (human (GRCh38)) were queried according to the list.

[0416] For each identified ZF target site, multiple ZF proteins can be designed. The design of the six recognition helices used to generate the complete protein is accomplished by selecting 2F units and considering multiple factors, such as the binding preferences of known zinc finger proteins, the frequency of selecting amino acids at positions -1, 2, 3, and 6 to bind to the desired target base in the B2H selection system, avoiding the selection of amino acids at positions -1, 2, 3, and 6 in B2H to bind to multiple different bases, and maintaining context dependence by matching flanking bases as much as possible. The complete ZF sequence is derived from the naturally occurring Zif268 protein, and the selected recognition helices are maintained in the sequence context in which they are selected in B2H (finger 1-2 or finger 2-3 of Zif268).

[0417] The 2F units were linked by the linker TGSQKP (SEQ ID NO: 651), where the 6 bp binding sites were contiguous and by the linker TGGGGSQKP (SEQ ID NO: 652), where 1 bp separated the 6 bp binding sites. A final set of 280 ZFPs were selected for primary screening that targeted 41 different DNA regions within 1 kb of the B2M TSS (chr15: 44711517) and had no other perfect matches in the genome (GRCh38) (Table 1).

[0418] Example 4: GripTite TM Guide RNA Screening in HEK 293MSR Cells

[0419] This example describes studies in which the efficacy of gRNA targeting B2M in HEK 293 cells (human embryonic kidney cells) was screened.

[0420] Introduction of gRNA+CRISPR-Off into HEK 293 cells

[0421] Six 96-well plates (Sigma-Aldrich) were seeded with 20,000 GripTite cells per well in the appropriate cell culture medium. TM 293MSR cells (Thermo Fisher, catalog number R79507). These cells are derived from human embryonic kidney cells (HEK293). The cells were allowed to grow for 24 hours in a 37°C incubator under 5% CO2. 25ng of gRNA encoding DNA fragment and 50ng of CRISPR-off encoding plasmid were resuspended in DPBS buffer (Thermo Fisher, catalog number 14190144). In addition, 10ng of EF1a: Puromycin resistance plasmid (PLA015) was added to the transfection mixture to achieve a total payload of 85ng DNA.

[0422] By adding the resuspended components to the Mirus Transfection reagent (Mirus, catalog number MIR2300) is used to produce transfection mixture. The transfection mixture is added to a total of six screening plates in duplicate. Wild-type (WT) CRISPR Cas9 has two different TSS adjacent gRNAs (positive control), CRISPR-off without gRNA (negative control), CRISPR-off with non-B2M locus targeting gRNA (negative control) and only empty vector (negative control) are also part of this experiment. Cells are passaged twice a week, treated with trypsin and Versene, and then separated into fresh culture medium in a new culture plate.

[0423] β2M flow cytometry

[0424] At 6, 13, and 20 days after transfection, transfected GripTite TM293 MSR cells were treated with trypsin and Versene and washed with PBS containing 2% FBS. The cells were then stained for 20 minutes at 4°C with a 1:300 dilution of PE-conjugated anti-human β2M antibody (BioLegend, catalog number 395704) and Zombie Violet Fixable Viability Dye (BioLegend, catalog number 423113), prepared in advance according to the manufacturer's recommendations and diluted 1:1000 in PBS containing 2% FBS. The stained cells were washed and incubated in fixation buffer (BioLegend, catalog number 420801) for 20 minutes. The cells were then washed before being placed on an Agilent Novocyte Penteon flow cytometer, which can collect up to 20,000 live cell events per well. The screening conditions were compared to negative (no gRNA) control expression levels to assess % silencing.

[0425] result

[0426] Relative B2M expression levels in cells transfected with one of the 258 tested gRNAs are shown in Figure 1 and Table 8. The best performing B2M gRNA was Figure 1 The expression of B2M in the no-gRNA control experiment was also quantified. The smooth fit of the entire screen indicated an efficient gRNA silencing pattern centered around the TSS of B2M, as shown in Figure 1 shown.

[0427] After treatment with multiple gRNA candidates, robust silencing of the B2M gene was observed, resulting in reduced B2M expression and only 30%-40% of β2M-positive cells were observed.

[0428] Table 8 Targeting domain sequences of the best performing gRNAs targeting B2M

[0429]

[0430]

[0431]

[0432]

[0433]

[0434]

[0435]

[0436]

[0437]

[0438]

[0439]

[0440]

[0441]

[0442]

[0443]

[0444]

[0445]

[0446]

[0447]

[0448]

[0449] The 172 best performing gRNAs (i.e., with the best β2M protein knockdown efficiency) from the above primary screening were aligned as single guide RNAs (sgRNAs) for further follow-up studies in the mRNA / sgRNA format.

[0450] Example 5: gRNA screening and confirmation in primary T cells

[0451] This example describes studies in which gRNAs were screened in human primary T cells.

[0452] Using EasySep TMHuman T cell separation kit (StemCell Technologies, catalog number 17951) is separated from human leukocyte removal product (StemCell Technologies, catalog number 70500) T cells. T cells are thawed and activated. Before nuclear transfection, T cells are thawed, washed, and supplemented with 5% human AB serum (Gemini Bio-Product, catalog number 100-512), 2mM L-alanyl-L-glutamine, 5ng / mL IL-7 and 5ng / mL IL-15 complete T cell culture medium (X-VIVO15 culture medium; Lonza, catalog number BEBP04-744Q), using Dynabeads people T activator CD3 / CD28 (Thermo Fisher, catalog number 11131D) for T cell expansion and activation, at 37 ° C and 5% CO2, with 3: 1 beads and cell quantity ratio stimulation for about 48 hours. The beads were then magnetically removed from the culture and the T cells were cultured in fresh complete T cell medium for approximately 24 hours. T cells were then nucleofected at 2E5 cells / well using the P3 primary cell 96-well Nucleofector kit (Lonza, catalog number V4SP-3960) and an Amaxa 4D nucleofector (Lonza) with pulse code EO115, using 2.5 μg CRISPR-off mRNA (TriLink) plus 2.5 μg sgRNA (IDT).

[0453] After nucleofection, T cells were resuspended in complete T cell medium and maintained by changing the medium twice a week and passaged as needed. TM The cells were restimulated with human CD3 / CD28 T cell activator (StemCell Technologies, catalog number 10991).

[0454] Cell surface β2M protein expression on live T cells was assessed by flow cytometry at days 6, 13, and 20 after nucleofection. No mRNA, CRISPR-off mRNA plus non-B2M targeting sgRNA, CRISPR-off mRNA without gRNA, WT Cas9 mRNA plus exon targeting sgRNA, stain only (no mRNA or gRNA), isotype (no mRNA or gRNA), and unstained (no mRNA or gRNA) controls were also run on each screening plate.

[0455] β2M flow cytometry assays were performed as described in Example 5. Test samples were compared to negative (CRISPR-off mRNA without sgRNA) control expression levels to assess % silencing.

[0456] Example 6: ZF screening in primary T cells

[0457] This example describes studies in which ZFP domains targeting different genomic regions of the B2M gene will be screened in human primary T cells.

[0458] T cells were isolated from human leukapheresis products and stored cryogenically. Prior to nuclear transfection, T cells were thawed and stimulated with CD3 / CD28 beads in complete T cell culture medium for approximately 48 hours at 37°C and 5% CO2. The beads were then magnetically removed from the culture and the T cells were cultured in fresh complete T cell culture medium. T cells were nuclear transfected with ZF-off mRNA using a Lonza Amaxa 4D nuclear transfection device. After nuclear transfection, T cells were resuspended in complete T cell culture medium and maintained by replacing the culture medium and isolating cells as needed twice a week. On day 13 after nuclear transfection, cells were re-stimulated with a soluble CD3 / CD28 T cell activator. On days 6, 13, and 20 after nuclear transfection, cell surface B2M protein expression on live T cells was assessed by flow cytometry. No mRNA, non-B2M-targeting ZF-off mRNA, WT Cas9 mRNA plus exon-targeting gRNA, stain only, isotype, and unstained controls were also run on each screening plate.

[0459] β2M flow cytometry assays were performed as described in Example 5. Screening conditions were compared to negative (non-B2M-targeting ZF) control expression levels to assess % silencing The following ZF constructs were tested:

[0460]

[0461]

[0462]

[0463]

[0464]

[0465]

[0466]

[0467]

[0468]

[0469]

[0470]

[0471]

[0472]

[0473]

[0474]

[0475]

[0476]

[0477]

[0478]

[0479]

[0480]

[0481]

[0482]

[0483]

[0484]

[0485]

[0486]

[0487]

[0488]

[0489]

[0490]

[0491]

[0492]

[0493]

[0494]

[0495]

[0496]

[0497]

[0498]

[0499]

[0500]

[0501]

[0502]

[0503]

[0504]

[0505] Example 7: Full specificity screening of constructs in primary human T cells

[0506] The specificity of CRISPR-off and ZF-off constructs to silence B2M was tested in human primary T cells. The readouts to assess specificity were RNAseq, methylation arrays, and whole-genome bisulfite sequencing assays. Genome-wide expression and methylation changes after epigenetic editing compared to negative controls were analyzed.

[0507] Example 8: CpG methylation pattern

[0508] CpG methylation patterns in primary human T cells treated with CRISPR-off or ZF-off were investigated. Hybridization capture assays were performed on bisulfite-treated DNA to investigate methylation patterns at CpG sites induced by CRISPR-off or ZF-off in a 1 kb region surrounding the B2M TSS.

[0509] Example 9: Screening Follow-up and Hit Validation

[0510] Reconfirm the best hits from the gRNA and ZF-off screens by repeating the screening experimental conditions and adjusting the dose of CRISPR-off mRNA+sgRNA or ZF-off mRNA up and down by several half-logs as appropriate to establish dose-response curves. Select the gRNA and ZF-off mRNA that exhibit the best potency and long-term durability characteristics for downstream candidate development.

[0511] Example 10: Allogeneic functional assay of primary T cells

[0512] Allogeneic healthy donor CD8 + T cell responses to mock-modified or B2M-silenced T cells were assessed via mixed lymphocyte co-culture assays and / or cytotoxicity assays.

[0513] Allogeneic healthy donor CD8 + T cell proliferation and / or activation, as measured by flow cytometry, of cell dye dilution and expression of cell surface activation markers. Responses to B2M-silenced cells are expected to be reduced relative to responses to mock-modified cells, indicating that allogeneic healthy donor CD8 + T cell proliferation and activation were reduced. In addition, flow cytometry viability dye staining or cell viability imaging analysis was used to assess the difference between allogeneic healthy donor CD8 + Modified T cell death after co-incubation with T cells. + In the presence of T cells, B2M-silenced T cells are expected to preferentially survive compared to mock-modified T cells.

[0514] Example 11: Guide RNA screening in primary T cells using CRISPR-off constructs

[0515] B2M single guide rescreening was performed in primary T cells using 172 guide RNAs (as shown in Table 9 below) and mRNA encoding fusion protein construct 15. Annotation of the amino acid sequence of fusion protein construct 15 is shown below. The results are shown in Table 9 below.

[0516] Ten guides showed greater than 20% silencing, and 18 guides showed greater than 10% silencing. RNA988 provided 40% silencing.

[0517] Annotation of the amino acid sequence of the fusion protein configuration 15

[0518]

[0519] Table 9 was measured on the 6th day after administration, using different gRNAs targeting B2M in primary human T cells, and the normalized percentage of B2M+ cells in primary human T cell populations was treated with CRISPR-off epigenetic repressors. Data from two replicates ("Plate 1" and "Plate 2"), as well as the weighted average of two replicates, are shown. The corresponding gRNA start position on chromosome 15 (GRCh38) is also provided.

[0520]

[0521]

[0522]

[0523]

[0524]

[0525]

[0526]

[0527]

[0528]

[0529]

[0530]

[0531]

[0532]

[0533]

[0534]

[0535]

[0536]

[0537]

[0538]

[0539]

[0540]

[0541]

[0542]

[0543]

[0544]

[0545]

[0546]

[0547]

[0548]

[0549]

[0550]

[0551]

[0552]

[0553]

[0554]

[0555]

[0556]

[0557]

[0558]

[0559]

[0560] Example 12: B2M dual guide screening

[0561] To improve silencing robustness and durability, assays were performed using administration of both guides to the same cells. This example describes a study in which gRNA pairs were screened in human primary T cells.

[0562] Using EasySep TMHuman T cell separation kit (StemCell Technologies, catalog number 17951) is separated from human leukocyte removal product (StemCell Technologies, catalog number 70500) T cells. T cells are thawed and activated. Before nuclear transfection, T cells are thawed, washed, and supplemented with 5% human AB serum (Gemini Bio-Product, catalog number 100-512), 2mM L-alanyl-L-glutamine, 5ng / mL IL-7 and 5ng / mL IL-15 complete T cell culture medium (X-VIVO15 culture medium; Lonza, catalog number BEBP04-744Q), using Dynabeads people T activator CD3 / CD28 (Thermo Fisher, catalog number 11131D) for T cell expansion and activation, at 37 ° C and 5% CO2, with 3: 1 beads and the number of cells than stimulated for about 48 hours. The beads were then magnetically removed from the culture and the T cells were cultured in fresh complete T cell medium for approximately 24 hours. T cells were then nucleofected at 2E5 cells / well using the P3 primary cell 96-well Nucleofector kit (Lonza, catalog number V4SP-3960) and an Amaxa 4D nucleofector (Lonza) with pulse code EO115, using 2.5 μg CRISPR-off mRNA (TriLink) plus 2.5 μg sgRNA (IDT).

[0563] After nucleofection, T cells were resuspended in complete T cell medium and maintained by changing the medium twice a week and passaged as needed. TM The cells were restimulated with human CD3 / CD28 T cell activator (StemCell Technologies, catalog number 10991).

[0564] Cell surface β2M protein expression on live T cells was assessed by flow cytometry at days 6, 13, and 20 after nucleofection. No mRNA, CRISPR-off mRNA plus non-B2M targeting sgRNA, CRISPR-off mRNA without gRNA, WT Cas9 mRNA plus exon targeting sgRNA, stain only (no mRNA or gRNA), isotype (no mRNA or gRNA), and unstained (no mRNA or gRNA) controls were also run on each screening plate.

[0565] β2M flow cytometry assays were performed as described in Example 5. The gating strategy was as follows Figure 2A (without gRNA) and Figure 2B(RNA102 & RNA964) were shown. The test samples were compared with the negative (CRISPR-off mRNA without sgRNA) control expression levels to assess the silencing %. The results are shown in Figure 2C shown.

[0566] Figure 3A-3B Shown are the percentage of B2M+ positive cells observed after administration of different guide RNA pairs, as well as the distance from the guide RNA binding site to the B2M TSS. Figure 4 B2M silencing by six guide RNA pairs measured at days 6, 13, and 20. All six guide RNA pairs reduced B2M expression at each time point compared to the no gRNA control.

[0567] Example 13: B2M CpG methylation pattern

[0568] CpG methylation patterns in primary human T cells treated with CRISPR-off were investigated. Hybridization capture assays were performed on bisulfite-treated DNA to investigate methylation patterns at CpG sites induced by CRISPR-off in a 1 kb region surrounding the B2M TSS.

[0569] B2M was silenced with two sets of dual guide combinations (RNA138 / 949 and RNA104 / 988). Samples were sorted at day 14 after nucleofection; pure B2M negative (B2M-) and B2M positive (B2M+) cell populations were sent for methylation analysis. More than 99% of the sorted B2M+ cells were positive for B2M, and less than 1% of the B2M- cells were positive for B2M. After sorting, the B2M samples were then restimulated with PMA / ionomycin or left in standard medium; after incubating these samples to observe silencing, the restimulated and control samples were also sent for hybrid capture methylation analysis.

[0570] The experimental procedures for each sample are outlined as follows Fig. 6A shown. Figure 6B The methylation patterns under each condition around the B2M locus are shown. Figure 7A-7B As shown, robust B2M CpG methylation was observed in the sorted B2M-negative population. Figure 8 As shown, extensive B2M CpG methylation was achieved with the combination of RNA138 / 949 and RNA104 / 988.

[0571] Example 14: B2M Silencing in Fresh and Frozen and Multiple Effector / Guide Conditions

[0572] Fresh primary human T cells were transfected with various combinations of effectors (FP13 or FP11a) and / or RNA (guide 1, guide 2, Milan TRACR, or US TRACR) or with WT Cas9 ( Fig. 9A ). Six days after transfection, the percentage of T cells expressing B2M was measured. Silencing was achieved when the effector and guide were combined, but not when the effector and guide were used alone. B2M expression in transfected T cells was also measured at day 14 after transfection ( Fig. 9B ). When the effector and guide are combined, silencing is maintained over time.

[0573] The same transfection was repeated in previously frozen primary human T cells ( Fig. 10A ). Similarly, silencing can be achieved when effectors and guides are combined. Fig. 10B and DON24, Fig. 10C ) were compared in primary human T cells (transfection was performed as described above). When effectors and guides were combined, the effectiveness of B2M silencing varied between donors, with DON24 ( Fig. 10C ) showed stronger B2M silencing in all effector / guide combinations. Transfections were performed with fusion protein 13 and fusion protein 11a along with gRNA, WT Cas9, or a no gRNA control, and B2M expression was assessed at 6, 12, 20, 28, and 35 days post-transfection. Silencing was achieved when effectors and guides were combined, and previously frozen T cells retained higher B2M silencing over time.

[0574] Example 15: B2M Silencing at Various Serum Concentrations

[0575] B2M silencing was measured over time in primary human T cells under different serum conditions (5% vs. 10% human serum) after transfection with B2M-silencing gRNA, WT Cas9, or no gRNA control. An exemplary gating strategy for B2M expression measurement is shown in Fig.11A There was no difference in B2M silencing under different culture conditions at any time point after transfection ( Fig. 11B ).

[0576] Example 16: B2M silencing under multiple conditions with multiple targets at multiple transduction times

[0577] Chimeric antigen receptor (CAR) is transduced into T cells treated with silencing gRNA and can affect gRNA silencing efficacy, CAR expression or both. In order to determine whether the above-mentioned gRNA is the case, on the 2nd or 3rd day after thawing, primary human T cells from donors DON001, DON006, DON020, DON023 are nuclear transfected. On the 1st, 2nd or 3rd day after thawing, B cell maturation antigen (BCMA) CAR is also used to transduce T cells. With 6 different gRNAs combined with 2.5 μg fusion protein 11a, transfected T cells. When combined with BCMA CAR transduction, nuclear transfection with gRNA on the 3rd day after thawing leads to more robust B2M silencing, as shown by the reduction of B2M, HLA-DR and CD3 expression. In addition, compared with the 3rd day, when BCMA CAR is transduced on the 1st or 2nd day after thawing, B2M, HLA-DR and CD3 expression remain low. Different pairs of gRNAs showed different B2M silencing abilities ( Fig. 12A ).

[0578] The transduction efficiency of BCMA CAR varied depending on the day of transduction. T cells transduced with BMCA CAR on day 1 or day 3 after thawing resulted in greater CAR expression than transduction on day 2 after thawing ( Fig. 12B Although silencing B2M with gRNA was more effective in CAR- cells, as measured by B2M, HLA-DR, and CD3 expression, B2M silencing was also effective in CAR+ cells ( Fig. 12C ), indicating that B2M-silenced human T cells can also express the transduced CAR.

[0579] Example 17: B2M silencing using gRNA from multiple production batches

[0580] To determine whether internucleofection variation was caused by gRNA quality, primary human T cells from donors DON006 and DON023 were transfected with different batches of gRNA. Three batches of two B2M silencing gRNAs were tested in pairs and combined with 2.5 μg of effector fusion protein 11a. An exemplary gating strategy for nucleofected T cells is shown in Fig.13A Although all gRNA pairs resulted in significant silencing of B2M expression, there were slight batch-dependent differences in the silencing efficiency of gRNAs 7 days after nucleofection ( Fig. 13B ).

[0581] Example 18: B2M dual-guide dose response assay

[0582] The dose response of 12 guide pairs was determined at two points. 2.5 μg of fusion protein 11a was used, and the starting dose of each sgRNA was 2.5 μg. Fig.14A) and Day 13 ( Fig. 14B ) reaction was observed.

[0583] Example 19: Allogeneic functional assay of primary T cells

[0584] Allogeneic healthy donor CD8 + T cell responses to mock-modified or B2M-silenced T cells were assessed via mixed lymphocyte co-culture assay.

[0585] Allogeneic healthy donor CD8 + T cell proliferation and / or activation, such as measured by flow cytometry, cell dye dilution and expression of cell surface activation markers. Allogeneic T cells responded less to B2M-silenced cells, resulting in CD8 + and CD4 + T cell proliferation was reduced (measured by the 7-day CellTrace Violet dilution assay), and activation was observed relative to the response of unmodified cells (measured by cell surface staining CD25 expression). Results are shown in Figures 15A-15B .

[0586] Using EasySep TM T cells were isolated from human leukapheresis products (StemCell Technologies, catalog number 70500) using the Human T Cell Isolation Kit (StemCell Technologies, catalog number 17951) and cultured in Cryo Cryopreserved in CS10 freezing medium (Biolife Solutions, catalog number 210502). Prior to nucleofection, T cells were thawed, washed, and grown in complete T cell medium (ImmunoCult TM-XFT cell expansion medium; StemCell Technologies, catalog number 10981), using Dynabeads human T activator CD3 / CD28 for T cell expansion and activation (Thermo Fisher, catalog number 11131D), at 37°C and 5% CO2, with a 1:1 ratio of beads to cells for about 72 hours. The beads were then magnetically removed from the culture, and T cells were then nuclear transfected with 2.5 μg CRISPR-Off mRNA plus 2.5 μg sgRNA (IDT) at 2E5 cells / well using the P3 primary cell 96-well Nucleofector kit (Lonza, catalog number V4SP-3960) and an Amaxa4D nucleofector (Lonza) with pulse code EO115.

[0587] After nucleofection, T cells were resuspended in complete T cell medium and maintained by changing the medium and passage as needed twice a week. On day 8 after nucleofection, B2M-silenced cells were sorted and cultured again until the day of the assay. On the day of the assay, unedited and B2M-silenced T cells were treated with 50 μg / ml mitomycin C for 30 minutes at 37°C, then washed, and subsequently stained with 0.5 μM CFSE in PBS for 3 minutes at room temperature, then washed. Allogeneic PBMCs were thawed and stained with CellTrace Violet (CTV) by incubation in 10 mM CTV in PBS for 10 minutes at 37°C, then washed. T cells and PBMCs were co-incubated for 7 days at a 1:1 T cell: PBMC ratio in T cell medium without added cytokines. At the assay endpoint, cell surface expression of CD3, CD4, CD8, and CD25 was assessed by flow cytometry of co-cultured samples. Proliferation of CD8+ and CD4+ T cells within allogeneic PBMCs was assessed by analyzing the CFSE-CD3+CD8+ or CFSE-CD3+CD4+ cell populations and quantifying the frequency of CTV dilution. Activation of CD8+ and CD4+ T cells within allogeneic PBMCs was assessed by analyzing the CFSE-CD3+CD8+ or CFSE-CD3+CD4+ cell populations and quantifying the frequency of CD25 cell surface expression.

[0588] Example 27: B2M three-guide screening

[0589] To improve silencing robustness and durability, assays were performed using administration of three guides to the same cells. This example describes studies in which gRNA triplets were screened in human primary T cells (FIGURE 16A-C).

[0590] Using EasySep TMHuman T cell separation kit (StemCell Technologies, catalog number 17951) is separated from human leukocyte removal product (StemCell Technologies, catalog number 70500) T cells. T cells are thawed and activated. Before nuclear transfection, T cells are thawed, washed, and supplemented with 5% human AB serum (Gemini Bio-Product, catalog number 100-512), 2mM L-alanyl-L-glutamine, 5ng / mL Il-7 and 5ng / mL IL-15 complete T cell culture medium (X-VIVO15 culture medium; Lonza, catalog number BEBP04-744Q), using Dynabeads people T activator CD3 / CD28 (Thermo Fisher, catalog number 11131D) for T cell expansion and activation, at 37 ° C and 5% CO2, with 3: 1 beads and the number of cells than stimulated for about 48 hours. The beads were then magnetically removed from the culture and the T cells were cultured in fresh complete T cell medium for approximately 24 hours. T cells were then nucleofected at 2E5 cells / well using a P3 primary cell 96-well Nucleofector kit (Lonza, catalog number V4SP-3960) and an Amaxa 4D nucleofector (Lonza) with pulse code EO115, with 2.5 μg CRISPR-off mRNA (TriLink) plus a total of 2.5 μg sgRNA (IDT) (split into two or three guides).

[0591] After nucleofection, T cells were resuspended in complete T cell medium and maintained by changing the medium twice a week and passaged as needed. TM The cells were restimulated with human CD3 / CD28 T cell activator (StemCell Technologies, catalog number 10991).

[0592] Cell surface β2M protein expression on live T cells was assessed by flow cytometry at days 6, 13, and 20 after nucleofection. No mRNA, CRISPR-off mRNA plus non-B2M targeting sgRNA, CRISPR-off mRNA without gRNA, WT Cas9 mRNA plus exon targeting sgRNA, stain only (no mRNA or gRNA), isotype (no mRNA or gRNA), and unstained (no mRNA or gRNA) controls were also run on each screening plate.

[0593] β2M flow cytometry assays were performed as described in Example 5. Test samples were compared to negative (CRISPR-off mRNA without sgRNA) control expression levels to assess % silencing. Results are shown in Figures 16A-16B .

[0594] sequence

[0595] Listed below are the SEQ ID NOs (SEQ) of the nucleotide (nt) and amino acid (aa) sequences described in the present disclosure.

[0596]

[0597]

[0598]

[0599]

[0600]

[0601]

[0602]

[0603]

[0604]

[0605]

[0606]

[0607]

[0608]

[0609]

[0610]

[0611]

[0612]

[0613]

[0614]

[0615]

[0616]

[0617]

[0618]

[0619]

[0620]

[0621]

[0622]

[0623]

[0624]

[0625]

[0626]

[0627]

[0628]

[0629]

[0630]

[0631]

[0632]

[0633]

[0634]

[0635]

[0636]

[0637]

[0638]

[0639]

[0640]

[0641]

[0642]

[0643]

[0644]

[0645]

[0646]

[0647]

[0648]

[0649]

[0650]

[0651]

[0652]

[0653]

[0654]

[0655]

[0656]

[0657]

[0658]

[0659]

[0660]

[0661]

[0662]

[0663]

[0664]

[0665]

[0666]

[0667]

[0668]

[0669]

[0670]

[0671]

[0672]

[0673]

[0674]

[0675]

[0676]

[0677]

[0678]

[0679]

[0680]

[0681]

[0682]

[0683]

[0684]

[0685]

[0686]

[0687]

[0688]

[0689]

[0690]

[0691]

[0692]

[0693]

[0694]

[0695]

[0696]

[0697]

[0698]

[0699]

[0700]

[0701]

[0702]

[0703]

[0704]

[0705]

[0706]

[0707]

[0708]

[0709]

[0710]

[0711]

[0712]

[0713]

[0714]

[0715]

[0716]

[0717]

[0718]

[0719]

[0720]

[0721]

[0722]

[0723]

[0724]

[0725]

[0726]

[0727]

[0728]

[0729]

[0730]

[0731]

[0732]

[0733]

[0734]

[0735]

[0736]

[0737]

[0738]

[0739]

[0740]

[0741]

[0742]

[0743]

[0744]

[0745]

[0746]

[0747]

[0748]

[0749]

[0750]

[0751]

[0752]

[0753]

[0754]

[0755]

[0756]

[0757]

[0758]

[0759]

[0760]

[0761]

[0762]

[0763]

[0764]

[0765]

[0766]

[0767]

[0768]

[0769]

[0770]

[0771]

[0772]

[0773]

[0774]

[0775]

[0776]

[0777]

[0778]

[0779]

[0780]

[0781]

[0782]

[0783]

[0784]

[0785]

[0786]

[0787]

[0788]

[0789]

[0790]

[0791]

[0792]

[0793]

[0794]

[0795]

[0796]

[0797]

[0798]

[0799]

[0800]

[0801]

[0802]

[0803]

[0804]

[0805]

[0806]

[0807]

[0808]

[0809]

[0810]

[0811]

[0812]

[0813]

[0814]

[0815]

[0816]

[0817]

[0818]

[0819]

[0820]

[0821]

[0822]

[0823]

[0824]

[0825]

[0826]

[0827]

[0828]

[0829]

[0830]

[0831]

[0832]

[0833]

[0834]

[0835]

[0836]

[0837]

[0838]

[0839]

[0840]

[0841]

[0842]

[0843]

[0844]

[0845]

[0846]

[0847]

[0848]

[0849]

Claims

1. A system for suppressing transcription of a human B2M gene in a human cell, optionally a human T lymphocyte or a human NK cell, the system comprising a) one or more fusion proteins, which together comprise a DNA methyltransferase (DNMT) domain and / or a domain that recruits DNMTs, optionally wherein the DNMT domain and / or the recruiter domain comprises a DNMT3A domain and / or a DNMT3L domain, and optionally wherein the recruited DNMT is DNMT3A, and Transcriptional repressor domain, Each domain is linked to a DNA binding domain that binds to a target region in the human B2M gene, wherein the target region comprises one or more sequences selected from the group consisting of SEQ ID NO: 700-740, 744, 747-749, 752, 753, 757, 758, 760-806, 812-822, 825, 827, 830, 833, 834, 839-841, 843-845, 849, 851-853, 855, 864, 866-877, 879-883 , 891-896, 898-900, 903-914, 922, 923, 925-927, 934, 936, 943-947, 949, 951-962, 975-981, 983, 985, 987-989, 995, 997-999, 1003-1005, and 1007-1011, or b) one or more nucleic acid molecules encoding one or more fusion proteins, wherein the system does not generate DNA breaks in the B2M gene.

2. The system of claim 1, wherein the DNA binding domain comprises a death CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain.

3. The system of claim 2, wherein the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) one or more guide RNAs comprising SEQ ID NOs: 710, 741-747, 749-759, 770-780, 782-1007, 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 116 2-1167, 1169-1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276, 1278-1282 and 1735-1737, or (ii) a nucleic acid molecule encoding the one or more guide RNAs.

4. The system of claim 2 or 3, wherein the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) two guide RNAs comprising SEQ ID NOs: 710, 741-747, 749-759, 770-780, 782-1007, 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 115 any two of 162-1167, 1169-1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276, 1278-1282 and 1735-1737, or (ii) a nucleic acid molecule encoding the two guide RNAs.

5. The system of claim 2 or 3, wherein the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) three guide RNAs comprising SEQ ID NOs: 710, 741-747, 749-759, 770-780, 782-1007, 1015, 1018-1020, 1023, 1024, 1028, 1029, 1031-1077, 1083-1093, 1096, 1098, 1101, 1104, 1105, 1110-1112, 1114-1116, 1120, 1122-1124, 1126, 1135, 1137-1148, 1150-1154, 115 any three of 162-1167, 1169-1171, 1174-1185, 1193, 1194, 1196-1198, 1205, 1207, 1214-1218, 1220, 1222-1233, 1246-1252, 1254, 1256, 1258-1260, 1266, 1268-1270, 1274-1276, 1278-1282 and 1735-1737, or (ii) a nucleic acid molecule encoding the three guide RNAs.

6. A system for suppressing transcription of human B2M gene in human cells, optionally human T lymphocytes or human NK cells, the system comprising a) A fusion protein comprising: DNMT3A domain, DNMT3L domain, DNA binding domain, and Transcriptional repressor domain, or b) a nucleic acid molecule encoding a fusion protein, wherein the system does not generate DNA breaks in the B2M gene.

7. The system of claim 6, wherein the DNA binding domain comprises a death CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain.

8. The system of claim 7, wherein the DNA binding domain comprises a dCas9 domain, and the system further comprises (i) one or more guide RNAs comprising any one of SEQ ID NOs: 1012-1282, or (ii) a nucleic acid molecule encoding the one or more guide RNAs.

9. The system of any one of claims 2, 3, 4, 5, 7, and 8, wherein the dCas domain comprises a dCas9 sequence, optionally a sequence having at least 90% identity to SEQ ID NO: 12 or 13.

10. The system of any one of claims 1-9, wherein the DNA binding domain binds to a target sequence in SEQ ID NO: 1283 or 1284.

11. The system of claim 2 or 7, wherein the ZFP domain targets a nucleotide sequence selected from the group consisting of SEQ ID NOs: 700-740.

12. The system of any one of claims 1-11, wherein the DNMT3A domain comprises a sequence that is at least 90% identical to SEQ ID NO: 574 or 575.

13. The system of any one of claims 1-12, wherein the DNMT3L domain comprises a sequence that is at least 90% identical to a sequence selected from SEQ ID NOs: 578-581.

14. The system of any one of claims 1-12, wherein the DNMT3L domain comprises a sequence that is at least 90% identical to a sequence selected from SEQ ID NOs: 582-603.

15. The system of any one of claims 1-5 and 7-11, wherein the DNMT domain comprises a sequence that is at least 90% identical to a sequence selected from SEQ ID NOs: 601-603.

16. The system of any one of claims 1-15, wherein the transcriptional repressor domain comprises a sequence that is at least 90% identical to a sequence selected from SEQ ID NOs: 33-570.

17. The system of any one of claims 1-15, wherein the transcriptional repressor domain comprises a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627.

18. The system of claim 17, wherein the KRAB domain comprises a sequence that is at least 90% identical to a sequence selected from the group consisting of SEQ ID NOs: 86, 116, 245, and 255.

19. The system of any one of claims 1-15, wherein the transcriptional repressor domain comprises a fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1 KRAB, and optionally comprises the amino acid sequence of SEQ ID NO: 571 or 572.

20. The system of any one of claims 1-15, wherein the transcriptional repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2.

21. The system of any one of claims 1-20, wherein the system comprises: a) a fusion protein comprising a DNMT3A domain, a DNMT3L domain, a transcription repressor domain and a DNA binding domain, Optionally wherein one or both of said DNMT3A domain and said DNMT3L domain are human, and Optionally wherein the DNA binding domain is a death CRISPR Cas domain or a ZFP domain; or b) a nucleic acid molecule encoding the fusion protein.

22. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, the DNMT3A domain, a first peptide linker, the DNMT3L domain, a second peptide linker, the DNA binding domain, a third peptide linker, and the transcriptional repressor domain.

23. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, the DNMT3A domain, a first peptide linker, the DNMT3L domain, a second peptide linker, a first nuclear localization signal (NLS), the DNA binding domain, a second NLS, a third peptide linker, and the transcriptional repressor domain.

24. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first nuclear localization signal (NLS), the DNMT3A domain, a first peptide linker, the DNMT3L domain, a second peptide linker, the DNA binding domain, a third peptide linker, the transcriptional repressor domain, and a second NLS.

25. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second nuclear localization signal (NLS), the DNMT3A domain, a first peptide linker, the DNMT3L domain, a second peptide linker, the DNA binding domain, a third peptide linker, the transcriptional repressor domain, and third and fourth NLS.

26. The system of any one of claims 21-25, wherein the transcriptional repressor domain comprises a KRAB domain, optionally a human KOX1, ZFP28, ZN627 or ZIM3 KRAB domain.

27. The system of any of claims 22-26, wherein one or both of the second peptide linker and the third peptide linker are XTEN linkers, optionally selected from XTEN80 and XTEN16, and further optionally wherein the second peptide linker is XTEN80 and the third peptide linker is XTEN16.

28. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a dSpCas9 domain, a second NLS, an XTEN16 peptide linker, and a human KOX1 KRAB domain.

29. The system of claim 28, wherein the fusion protein comprises SEQ ID NO: 658 or a sequence at least 90% identical thereto.

30. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a ZFP domain, a second NLS, an XTEN16 linker, and a human KOX1 KRAB domain.

31. The system of claim 30, wherein the fusion protein comprises SEQ ID NO: 659 or a sequence at least 90% identical thereto.

32. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and a third and a fourth NLS.

33. The system of claim 32, wherein the fusion protein comprises SEQ ID NO: 660 or a sequence at least 90% identical thereto.

34. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and a third and a fourth NLS.

35. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZFP28 KRAB domain, and a third and a fourth NLS.

36. The system of claim 35, wherein the fusion protein comprises SEQ ID NO: 661 or a sequence at least 90% identical thereto.

37. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZFP28 KRAB domain, and a third and a fourth NLS.

38. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZN627 KRAB domain, and a third and a fourth NLS.

39. The system of claim 38, wherein the fusion protein comprises SEQ ID NO: 662 or a sequence at least 90% identical thereto.

40. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZN627 KRAB domain, and a third and a fourth NLS.

41. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and a third and a fourth NLS.

42. The system of claim 41, wherein the fusion protein comprises SEQ ID NO: 663 or a sequence at least 90% identical thereto or SEQ ID NO: 667 or a sequence at least 90% identical thereto.

43. The system of claim 21, wherein the fusion protein comprises, from N-terminus to C-terminus, a first and a second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and a third and a fourth NLS.

44. The system of any one of claims 23-43, wherein at least one of the NLSs is an SV40 NLS.

45. The system of any one of claims 1-5 and 9-20, wherein the system comprises: a) a first fusion protein comprising a first DNA binding domain and comprising or recruiting said DNMT3A domain, a second fusion protein comprising a second DNA binding domain and comprising or recruiting the DNMT3L domain, and a third fusion protein comprising a third DNA binding domain and comprising or recruiting the transcriptional repressor domain; or b) one or more nucleic acid molecules encoding the fusion protein.

46. ​​A human cell comprising the system of any one of claims 1-45, or a descendant of said cell, optionally wherein said cell is a T lymphocyte or a NK cell.

47. A human cell modified by the system of any one of claims 1-45, or a descendant of said cell, optionally wherein said cell is a T lymphocyte or a NK cell, optionally wherein said cell is modified ex vivo.

48. A pharmaceutical composition comprising the system of any one of claims 1-45 and a pharmaceutically acceptable excipient, optionally wherein The composition comprises a lipid nanoparticle (LNP) comprising the system, and / or The DNA binding domain is a dCas domain, and the LNP further comprises one or more gRNAs.

49. A pharmaceutical composition comprising the human cell of claim 46 or 47 and a pharmaceutically acceptable excipient.

50. A method of treating a patient in need thereof, the method comprising administering to the patient the system of any one of claims 1-45, the human cell of claim 46 or 47, or the pharmaceutical composition of claim 48 or 49.

51. The method of claim 50, wherein the patient has cancer or an autoimmune disease.

52. The system of any one of claims 1-45, the human cell of claim 46 or 47, or the pharmaceutical composition of claim 48 or 49, for use in treating a patient in need thereof, optionally in the method of claim 50 or 51.

53. Use of the system of any one of claims 1-45 or the human cell of claim 46 or 47 in the preparation of a medicament for treating a patient in need thereof, optionally in the method of claim 50 or 51.

Citation Information

Patent Citations

  • RNA-guided nucleases and active fragments and variants thereof and methods of use

    US11162114B2

  • Delivery system for functional nucleases

    US20160200779A1

  • Switchable cas9 nucleases and uses thereof

    US20160208288A1

  • Selection of sites for targeting by zinc finger proteins and methods of designing zinc finger proteins to bind to preselected sites

    US6453242B1

  • Regulation of endogenous gene expression in cells using zinc finger proteins

    US6534261B1