Compositions and methods for epigenetic regulation of pcsk9 expression

By using a fusion protein of DNA methyltransferase and transcriptional repressor domains in human cells, combined with the CRISPR-Cas and ZFP systems, the PCSK9 gene was targeted and regulated, overcoming the double-strand break risk of traditional genetic editing and achieving safe and durable reduction of PCSK9 expression for the treatment of hypercholesterolemia and cardiovascular disease.

CN121889501APending Publication Date: 2026-04-17CHROMA MEDICINE INC
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHROMA MEDICINE INC
Filing Date
2024-03-06
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies pose risks of double-strand DNA breaks and genotoxicity when targeting and regulating PCSK9 expression. There is a need to develop a safe epigenetic modification method to reduce PCSK9 expression and thus treat hypercholesterolemia and cardiovascular disease.

Method used

A fusion protein containing DNA methyltransferase and transcriptional repressor domains is used to target the PCSK9 gene through a DNA-binding domain to achieve epigenetic silencing. Specifically, it includes the CRISPR-Cas system and zinc finger protein (ZFP) domains, which bind to specific guide RNA or gRNA to regulate the transcription of the PCSK9 gene.

Benefits of technology

It achieves a safe and durable reduction of PCSK9 expression in human cells, reduces the concentration of LDL particles in the blood, lowers the risk of cardiovascular disease, and avoids the double-strand break risk of traditional genetic editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121889501A_ABST
    Figure CN121889501A_ABST
Patent Text Reader

Abstract

The present application relates to compositions and methods comprising an epigenetic editor for epigenetic modification of PCSK9, as well as nucleic acids and vectors encoding the epigenetic editor. Also disclosed are cells epigenetically modified by the epigenetic editor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 488,736, filed March 6, 2023, entitled “Compositions and methods for epigenetic regulation of PCSK9 expression”, filed under 35 USC §119(e), the entire disclosure of which is incorporated herein by reference.

[0003] Reference to electronic sequence listing

[0004] The contents of the electronic serial number (C169870005WO01-SEQ-AXW.xml; size: 1,866,473 bytes; and creation date: March 6, 2024) are incorporated herein by reference in their entirety. Background Technology

[0005] For over a decade, genome editing has been considered a promising treatment for genetic diseases. However, manipulation at the DNA level using traditional genetic editors remains risky, given the potential for unintended double-strand DNA breaks, heterogeneous repair (including large and small insertions and deletions at intended sites), and toxicity. In contrast, targeted epigenetic modifications offer the potential to alter gene expression without causing double-strand break-induced genotoxicity.

[0006] One promising candidate for epigenetic silencing is the proprotein convertase subtilisin / kexin type 9 (PCSK9) gene. PCSK9 is a key target for treating heart disease (a leading cause of death worldwide) (Berberich et al., Nature Rev Cardiol. (2019) 16(1):9-20). The human PCSK9 gene, located on chromosome 1, shares approximately 94% and 80% homology with its cynomolgus monkey and mouse counterparts, respectively. The gene has CpG islands in its promoter region and is isolated from other genes and features cis-regulatory characteristics. PCSK9 protein is primarily produced by the liver.

[0007] In humans, PCSK9 plays a crucial role in regulating the circulating levels of low-density lipoprotein (LDL) particles due to its binding to the LDL receptor (LDLR). LDLR reduces the circulating concentration of LDL particles by mediating endocytosis and degradation within cells. In the absence of PCSK9, or if the interaction between PCSK9 and LDLR is blocked, the rate of LDLR recirculation to the cell surface increases, and the recirculated LDLR protein continues to remove LDL particles from the extracellular fluid (Tombling et al., Atherosclerosis (2021) 330:52-60). Conversely, when endocytosed LDLR binds to PCSK9, LDLR degrades along with the LDL particles it carries. Clinical and genetic studies have established that circulating LDL contributes to atherosclerotic cardiovascular disease (Ference et al., Eur Heart J (2017) 38:2459-72). Furthermore, loss-of-function mutations in PCSK9 are associated with low LDL levels (Zhao et al., Am J Hum Genet. (2006) 79(3):514-23). ​​Genetic or pharmacological reductions in PCSK9 can decrease cardiovascular events (Ference et al., N Engl J Med. (2016) 375(22):2144-53; Sabatine et al., N Engl J Med. (2017) 376(18):1713-22). Decreased PCSK9 expression contributes to increased LDLR recirculation, which leads to lower blood LDL particle concentrations.

[0008] Given the crucial role of PCSK9 in the pathogenesis of hypercholesterolemia and cardiovascular disease, there is a need for novel and improved therapies that target PCSK9 expression. Invention Overview

[0010] This disclosure provides systems and compositions for epigenetic modification (referred to herein as "epigenetic editors" or "epigenetic editing systems"), and methods for using them to generate epigenetic modifications at PCSK9 (including in host cells and organisms).

[0011] In some aspects, this disclosure provides a system for inhibiting the transcription of the human PCSK9 gene in human cells (optionally human hepatocytes), the system comprising

[0012] a) One or more fusion proteins, which together contain

[0013] DNA methyltransferase (DNMT) domain and / or domain recruiting DNMT, optionally wherein the DNMT domain and / or recruiting domain includes a DNMT3A domain and / or a DNMT3L domain, and optionally wherein the recruited DNMT is DNMT3A, and

[0014] Transcription repressor domain,

[0015] Each domain is linked to a DNA-binding domain that binds to a target region in the human PCSK9 gene; or

[0016] b) One or more nucleic acid molecules that encode one or more fusion proteins.

[0017] In some embodiments, the DNA-binding domain binds to a target sequence in SEQ ID NO: 1488 or 1489. In some embodiments, the DNA-binding domain targets the fusion protein to one or more sequences selected from the PCSK9 gene in SEQ ID NO: 700-747 and 1036-1261.

[0018] In some embodiments, the DNA-binding domain comprises a death CRISPR Cas (dCas) domain, a ZFP domain, or a TALE domain. For example, the DNA-binding domain may comprise a dCas9 domain, and the system may further comprise (i) one or more guide RNAs (e.g., any one of SEQ ID NO: 1262-1487), or (ii) a nucleic acid molecule encoding one or more guide RNAs. In some embodiments, the dCas domain comprises a dCas9 sequence, such as a sequence having at least 90% identity with SEQ ID NO: 12 or 13.

[0019] In some embodiments, the fusion protein comprises a death CRISPR Cas (dCas) domain, and the system comprises one or more PCSK9 binding guide RNAs (gRNAs) provided herein. In some embodiments, the system comprises a single gRNA. In some embodiments, the system comprises two gRNAs. In some embodiments, the system comprises three gRNAs. In some embodiments, the system comprises four gRNAs. In some embodiments, the system comprises five or more gRNAs. In some embodiments, the system comprises gRNAs selected from the gRNAs provided in Table 2. In some embodiments, the system comprises gRNAs selected from the gRNAs provided in Table 7. In some embodiments, the system comprises gRNAs selected from the gRNAs provided in Table 8. In some embodiments, the system comprises sgRNAs selected from the gRNAs provided in Table 10. In some embodiments, the system comprises gRNAs selected from the gRNAs provided in Table 12.

[0020] In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA009 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA009. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA003 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA003. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA093 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA093. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA011 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA011. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA007 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA007. In some embodiments, the system comprises a gRNA containing the gRNA target sequence of gRNA077 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA077. In some embodiments, the system comprises a gRNA containing the gRNA target sequence of gRNA113 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA113. In some embodiments, the system comprises a gRNA containing the gRNA target sequence of gRNA004 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA004. In some embodiments, the system comprises a gRNA containing the gRNA target sequence of gRNA008 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA008. In some embodiments, the system comprises a gRNA containing the gRNA target sequence of gRNA012 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA012. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA111 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA111. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA005 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA005. In some embodiments, the system comprises a gRNA containing the gRNA targeting sequence of gRNA013 in Table 10, or a gRNA that binds to the same target domain sequence as gRNA013.

[0021] In some embodiments, the system comprises gRNA g041 as provided in Table 12, or gRNA that binds to the target domain sequence of g041. In some embodiments, the system comprises gRNA g049 as provided in Table 12, or gRNA that binds to the target domain sequence of g049. In some embodiments, the system comprises gRNA g056 as provided in Table 12, or gRNA that binds to the target domain sequence of g056. In some embodiments, the fusion protein comprises the fusion protein disclosed in Example 12. In some embodiments, the system comprises the fusion protein as provided in Example 12 and gRNA as provided in Table 10. In some embodiments, the system comprises the fusion protein as provided in Example 12 and gRNA as provided in Table 12.

[0022] In some embodiments, the system comprises fusion protein 9, variant 1 (Example 12), and gRNA g041. In some embodiments, the system comprises fusion protein 9 variant 2 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 10 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 11 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 12 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 13 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 14 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 15 (Example 12) and gRNA g041. In some embodiments, the system comprises fusion protein 9 variant 1 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 9 variant 2 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 10 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 11 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 12 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 13 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 14 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 15 (Example 12) and gRNA g049. In some embodiments, the system comprises fusion protein 9 variant 1 (Example 12) and gRNAs g041 and g049. In some embodiments, the system comprises fusion protein 9 variant 2 (Example 12) and gRNAs g041 and g049. In some embodiments, the system comprises fusion protein 10 (Example 12) and gRNAs g041 and g049. In some embodiments, the system includes fusion protein 11 (Example 12) and gRNAs g041 and g049. In some embodiments, the system includes fusion protein 12 (Example 12) and gRNAs g041 and g049. In some embodiments, the system includes fusion protein 13 (Example 12) and gRNAs g041 and g049. In some embodiments, the system includes fusion protein 14 (Example 12) and gRNAs g041 and g049. In some embodiments, the system includes fusion protein 15 (Example 12) and gRNAs g041 and g049.

[0023] In some embodiments, the DNA-binding domain includes a ZFP domain that targets a nucleotide sequence selected from SEQ ID NO: 700-747. In some embodiments, the ZFP domain sequentially includes the F1-F6 amino acid sequence of any one of ZF001 to ZF048 as shown in Table 1.

[0024] In some implementations, the DNMT3A domain contains a sequence that is at least 90% identical to SEQ ID NO: 574 or 575.

[0025] The DNMT3L domain may contain, for example, a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 578-581. In some embodiments, the DNMT3L domain contains a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 582-603. In some embodiments, the DNMT3L domain contains a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 601-603.

[0026] In some embodiments, the transcriptional repressor domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 33-570. In some embodiments, the transcriptional repressor domain is a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627. The KRAB domain may comprise, for example, a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 89, 116, 245, and 255. In some embodiments, the transcriptional repressor domain comprises a fusion of the N- and C-terminal regions of ZIM3 and KOX1 KRABs, and optionally comprises the amino acid sequence of SEQ ID NO: 571 or 572. In some embodiments, the transcriptional repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2.

[0027] In some implementations, the system includes

[0028] a) A fusion protein comprising a DNMT3A domain, a DNMT3L domain, a transcriptional repressor domain, and a DNA-binding domain.

[0029] Optionally, one or both of the DNMT3A and DNMT3L domains are human, and

[0030] Optionally, the DNA-binding domain is a death CRISPR-Cas domain or a ZFP domain; or

[0031] b) Nucleic acid molecules that encode fusion proteins.

[0032] In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA-binding domain, a third peptide linker, and a transcriptional repressor domain. For example, the fusion protein may comprise, from N-terminus to C-terminus, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a first nuclear localization signal (NLS), a DNA-binding domain, a second NLS, a third peptide linker, and a transcriptional repressor domain. Alternatively, the fusion protein may comprise, from N-terminus to C-terminus, a first nuclear NLS, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA-binding domain, a third peptide linker, a transcriptional repressor domain, and a second NLS. Another fusion protein may comprise, from N-terminus to C-terminus, first and second NLS, a DNMT3A domain, a first peptide linker, a DNMT3L domain, a second peptide linker, a DNA-binding domain, a third peptide linker, a transcriptional repressor domain, and third and fourth NLS. In a particular embodiment, the transcriptional repressor domain is a KRAB domain, such as the human KOX1, ZFP28, ZN627, or ZIM3 KRAB domain. In a particular embodiment, one or both of the second and third peptide linkers are XTEN linkers, which may be selected from XTEN80 (e.g., SEQ ID NO: 643) and XTEN16 (e.g., SEQ ID NO: 638), for example, where the second peptide linker is XTEN80 and the third peptide linker is XTEN16.

[0033] In some embodiments, the fusion protein may include, from the N-terminus to the C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a dSpCas9 domain, a second NLS, an XTEN16 peptide linker, and a human KOX1 KRAB domain. In some embodiments, the fusion protein comprises SEQ ID NO: 658 or a sequence at least 90% identical thereto.

[0034] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a ZFP domain, a second NLS, an XTEN16 linker, and a human KOX1 KRAB domain. In some embodiments, the fusion protein comprises SEQ ID NO: 659 or at least 90% identical to it, optionally wherein the ZFP comprises, in sequence, the F1-F6 amino acid sequences of any one of ZF001 to ZF048 as shown in Table 1.

[0035] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and third and fourth NLS. In a particular embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 660 or a sequence at least 90% identical thereto.

[0036] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and third and fourth NLS.

[0037] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZFP28KRAB domain, and third and fourth NLS. In a particular embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 661 or a sequence at least 90% identical thereto.

[0038] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZFP28 KRAB domain, and third and fourth NLS.

[0039] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZN627KRAB domain, and third and fourth NLS. In a particular embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 662 or a sequence at least 90% identical thereto.

[0040] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZN627 KRAB domain, and third and fourth NLS.

[0041] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and third and fourth NLS. In a particular embodiment, the fusion protein may comprise the amino acid sequence of SEQ ID NO: 663 or a sequence at least 90% identical thereto.

[0042] In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and third and fourth NLS.

[0043] In some embodiments, at least one NLS in the fusion protein described herein is an SV40 NLS (e.g., SEQ ID NO: 644).

[0044] In some implementations, the system includes:

[0045] a) A first fusion protein comprising a first DNA-binding domain and comprising or recruiting a DNMT3A domain.

[0046] The second fusion protein contains a second DNA-binding domain and includes or recruits a DNMT3L domain, and

[0047] A third fusion protein, comprising a third DNA-binding domain and containing or recruiting a transcriptional repressor domain; or

[0048] b) One or more nucleic acid molecules that encode fusion proteins.

[0049] This disclosure also provides human cells comprising the system described herein, or descendants of such cells. In some embodiments, the cells are hepatocytes.

[0050] This disclosure also provides pharmaceutical compositions comprising the systems described herein and pharmaceutically acceptable excipients.

[0051] This disclosure also provides methods for treating patients in need, comprising administering (e.g., intravenously) the system or pharmaceutical composition described herein to the patient. In some embodiments, the patient has heart disease; has elevated low-density lipoprotein cholesterol (LDL-C) or hypercholesterolemia; is at risk of developing myocardial infarction, stroke, or unstable angina; and / or has primary hyperlipidemia (e.g., heterozygous familial hypercholesterolemia (HeFH) or homozygous familial hypercholesterolemia (HoFH)).

[0052] This disclosure also provides systems or pharmaceutical compositions described herein for treating patients in need, such as those described herein.

[0053] This disclosure also provides the use of the system described herein in the preparation of medicaments for treating patients in need, for example, in the methods described herein.

[0054] This disclosure also provides articles and kits that include the systems described herein.

[0055] Other features, objects, and advantages of the methods and compositions disclosed herein will become apparent in the following detailed description. However, it should be understood that while this detailed description indicates embodiments and implementations of the disclosed methods and compositions, it is given by way of illustration only and not limitation. Various changes and modifications within the scope of this disclosure will become apparent to those skilled in the art from the detailed description. Attached Figure Description

[0056] Figure 1 This is a diagram showing the predicted binding sites of ZF protein and computationally designed gRNA on the PCSK9 gene.

[0057] Figure 2 This is a scatter plot showing the relative PCSK9 expression (y-axis) on day 7 in cells treated with CRISPR-off (DNMT3A-3L-dCas9-KRAB). The genomic distance from the gRNA target site to the PCSK9 TSS is shown on the x-axis.

[0058] Figure 3 This is a diagram showing the overlap between the first 40 gRNAs and the PCSK9 gene.

[0059] Figure 4A This is a bar graph showing the levels of PCSK9 secreted on days 7 and 28 after treatment with the indicated gRNA. The dashed line indicates silencing achieved through wild-type (WT) Cas9.

[0060] Figure 4B This is a scatter plot showing the correlation between PCSK9 mRNA expression and PCSK9 protein secretion in cells treated with gRNA. CRISPRi (dCas9-KRAB) represents the dCas9-KRAB fusion protein.

[0061] Figure 5 It is a line graph showing the silencing of PCSK9 after treatment with CRISPRi (dCas9-KRAB), CRISPR-off (DNMT3A-3L-dCas9-KRAB), and indicated gRNA.

[0062] Figure 6 This is a linear graph showing the PCSK9 secretion in cells treated with CRISPRoff and simvastatin compared to cells treated with the CRISPRoff system alone.

[0063] Figure 7 This is a bar graph showing the reduction in PCSK9 secretion in Huh7 liver cancer cells treated with CRISPRoff and a given gRNA.

[0064] Figure 8 This is a scatter plot showing the activity and toxicity of 247 PCSK9 targeting ZF protein. Relative PCSK9 expression is shown on the x-axis, and the corresponding cell counts relative to pUC and off-target controls are shown on the y-axis. The diagonal line represents a 1:1 correlation between relative PCSK9 expression and cell count.

[0065] Figure 9 This is a scatter plot showing the relative PCSK9 expression (y-axis) and the corresponding target genome distance (x-axis) relative to the PCSK9 transcription start site (TSS) in cells treated with the ZF-off (DNMT3A-3L-ZF-KRAB) construct.

[0066] Figure 10 This diagram shows the entire human PCSK9 locus (67.5 kb), flanked by upstream and downstream genomic regions of 35.5 kb and 7 kb, respectively, which were introduced into and expressed in transgenic mice. These transgenic mice express human PCSK9 under the control of their own (human) endogenous promoter.

[0067] Figure 11A A schematic diagram of a fusion protein construct with a variant NLS configuration is shown. Figure 11B A schematic diagram of another fusion protein construct with a variant KRAB domain is shown.

[0068] Figure 12A-12B This indicates the use of 6.25 ng RNA ( Figure 12A ) or 2.5 ng RNA ( Figure 12B A graph showing the percentage of PCSK9 protein levels measured after treatment with fusion protein constructs having multiple NLS layouts in HeLa cells. Human and mouse DNMT3L sequences are indicated as h3L and m3L, respectively.

[0069] Figure 13 The figure shows that the construct with 2X NLS is 3 times more effective than CRISPR-off in silencing mPcsk9 in Hepa1-6 cells.

[0070] Figure 14AThis is a graph showing that the construct with 2X NLS is more effective at silencing mPcsk9 in Huh7 cells than CRISPR-off. Figure 14B-14C It shows the 5th day ( Figure 14B ) and the 15th day ( Figure 14C Figure 1 shows that different doses of the construct with 2X NLS are more effective than CRISPR-off in silencing mPcsk9 in Huh7 cells.

[0071] Figure 15 The graph shows that 2XNLS provides improved performance in multiple ZF cells in Huh7 cells in a CRISPR-off-like form where dCas9 is replaced by zinc fingers.

[0072] Figure 16 This is a set of diagrams showing that methylation of the CTLA4 promoter with bacterial DNMT protein can induce epigenetic silencing at a locus.

[0073] Figure 17 This is a set of graphs showing the methylation profiles at the VIM3 locus of cells treated on day 30 with different constructs carrying a bacterial DNA methyltransferase fused to dCas9. Samples treated with M. SssI were methylated by 20%.

[0074] Figure 18 This is a set of diagrams showing the methylation profiles captured by hybridization at the CLTA locus in cells on day 29, comparing the dCas9 fusion form of M.SssI with mouse DNMT3A / 3L.

[0075] Figures 19A-19D This indicates that 0.5 ng of effector DNA was used with CLTA-GFP as a marker. Figure 19A ), using 3ng effector DNA with GFP as a marker ( Figure 19B ) and using 0.5 ng effector DNA with GFP as a marker ( Figure 19C A set of diagrams showing the epigenetic silencing activity of the alternative KRAB domain against CRISPR-off during testing. Figure 19D The results are shown after 30 days using various nanogram amounts of effector DNA.

[0076] Figures 20A-20B This shows the experimental timeline used to test PCSK9 silencing in transgenic mice expressing human PCSK9 (hPCSK9). Figure 20A ) and experimental design ( Figure 20B The diagram is as described in Example 11.

[0077] Figure 21A and 21B It displays within 42 days ( Figure 21A ) and 84 days Figure 21B A set of graphs showing the silencing of PCSK9 in transgenic mice expressing human PCSK9, as measured intracellularly. Human PCSK9 levels in mouse blood were measured by ELISA at indicated time points as described in Example 11.

[0078] Figure 22 This is a graph showing the experimental timeline used to conduct dose-response experiments, evaluating PCSK9 silencing in transgenic mice expressing human PCSK9, as described in Example 11.

[0079] Figures 23A-23B This is a set of graphs showing the results of dose-response experiments evaluating PCSK9 silencing in transgenic mice expressing human PCSK9. The constructs tested included CRISPR-Off (…). Figure 23A ) and ZF-Off ( Figure 23B )Constructor.

[0080] Figure 24 Results of the partial hepatectomy (PHx) experiment are shown, demonstrating the persistence of PCSK9 epigenetic silencing in mice, as described in Example 17.

[0081] Figure 25 Multiple effectors exhibited good specificity in Huh-7 cells. Huh-7 cells were treated with different repressor fusion proteins (fusion proteins 11 and 13, respectively) and PCSK9-targeting guide RNA 041. Differential gene expression was assessed by RNAseq of treated and untreated cells, and the values ​​were visualized in a volcano plot. No significant off-target effects were detected.

[0082] Figures 26A-26C The first five guide sequences demonstrating robust PCSK9 silencing in PXB cells are shown. Fresh human hepatocytes were isolated from the PXB mouse model. Long-term stability and function of the hepatocytes, as well as robust PCSK9 secretion, were confirmed.

[0083] Cells were treated with the epigenetic repressor PLA2628 (fusion protein 12 as provided in Example 12) and controls, as shown in the illustration. PCSK9 secretion was measured and plotted against a negative control (PLA2628 plus a non-PCSK9-targeting gRNA). All gRNAs were specifically measured by RNAseq on day 14 post-delivery. Volcano plots of the two exemplary RNAs evaluated are shown on the right (volcano plots of other RNAs are not shown). No significant off-target effects were observed for any of the gRNAs tested.

[0084] Figure 27The observed levels of PCSK9 secreted in PXB cells on day 14 are shown. Invention Details

[0086] This disclosure provides an epigenetic editor for regulating PCSK9 gene expression. By altering PCSK9 expression, the systems, compositions, and methods described herein can be used to treat conditions such as hypercholesterolemia (e.g., heterozygous familial hypercholesterolemia (HeFH), homozygous familial hypercholesterolemia (HoFH), familial hypercholesterolemia (HF), or established atherosclerotic cardiovascular disease (ASCVD)) or renal insufficiency (RI). Unless otherwise stated, “PCSK9” herein refers to human PCSK9. The human PCSK9 gene sequence can be found in Ensembl accession number ENSG00000169174. This epigenetic editor offers several advantages over other genome engineering methods, including reversibility, reduced risk of chromosomal translocations, and durable, heritable silencing.

[0087] In some embodiments, the region in the human PCSK9 gene targeted for epigenetic regulation is approximately 2 kb long and located approximately + / - 1 kb of the PCSK9 TSS. In some embodiments, this region has the nucleotide sequence of SEQ ID NO: 1488. In some embodiments, the targeted PCSK9 region is approximately 1069 bp long and located approximately + / - 500 bp of the PCSK9 TSS. In some embodiments, the targeted region has the nucleotide sequence of SEQ ID NO: 1489. The PCSK9 TSS is located at #chr1:55039548 in GRCh38 of the genome.

[0088] In some embodiments, the epigenetic editor as described herein may comprise one or more fusion proteins, each fusion protein comprising a DNA-binding domain linked to one or more effector domains for epigenetic modification. In some embodiments, wherein the DNA-binding domain is a polynucleotide-guided DNA-binding domain, the epigenetic editor may further comprise one or more guide polynucleotides. The DNA-binding domain, effector domain, and guide polynucleotide of the epigenetic editor as described herein may be selected in any functional combination from those described below, for example.

[0089] The epigenetic editors described herein can be transiently expressed in host cells or integrated into the genome of host cells; this disclosure also considers such cells and their progeny. Transiently expressed and integrated epigenetic editors or components thereof can achieve stable epigenetic modifications. For example, after the introduction of the epigenetic editors described herein into a host cell, target genes in the host cell can be stably or permanently repressed or silenced. In some embodiments, the expression of target genes is reduced or silenced for at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 2 months, at least 3 months, at least 4 months, at least 5 months, at least 6 months, at least 1 year, at least 2 years, or the entire lifespan of the cell or the subject carrying the cell, compared to the expression level without the epigenetic editor. The epigenetic modification can be inherited by the progeny of the host cell in which the epigenetic editor has been introduced.

[0090] The epigenetic editor disclosed herein can be introduced into patients who require it (e.g., human patients), for example, into the patient's hepatocytes, bile epithelial cells (cholangiocartilages), astrocytes, Kupffer cells, and sinusoidal endothelial cells.

[0091] I. DNA-binding domain

[0092] The epigenetic editor described herein may include one or more DNA-binding domains that guide the effector domains of the epigenetic editor to target sequences within or near the PCSK9 gene locus. The DNA-binding domains described herein may be, for example, polynucleotide-guided DNA-binding domains, zinc finger protein (ZFP) domains, transcription activator-like effector (TALE) domains, large-scale nuclease DNA-binding domains, etc. Examples of DNA-binding domains can be found in U.S. Patent 11,162,114, which is incorporated herein by reference in its entirety.

[0093] In some embodiments, the DNA-binding domain described herein is encoded by its natural coding sequence. In other embodiments, the DNA-binding domain is encoded by a nucleotide sequence that has been codon-optimized for optimal expression in human cells.

[0094] A. Polynucleotide-guided DNA-binding domain

[0095] In some embodiments, the DNA-binding domain described herein may be a protein domain guided by a guide nucleic acid sequence (e.g., a guide RNA sequence) to a target site in the PCSK9 gene locus. In some embodiments, the protein domain may be derived from a CRISPR-associated nuclease, such as a class I or class II CRISPR-associated nuclease. In some embodiments, the protein domain may be derived from a Cas nuclease, such as type II, type IIA, type IIB, type IIC, type V, or type VI Cas nucleases. In some embodiments, the protein domain may be derived from a class II Cas nuclease selected from Cas1, Cas1B, Cas2, Cas3, etc. Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Cas14a, Cas14b, Cas14c, CasX, CasY, CasPhi, C2c4, C2c8, C2c9, C2c10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx1S, Csf1, Csf2, CsO, Csf4 and their homologues and modified forms. "Derived from" is used to mean that a protein domain contains the complete polypeptide sequence of the parent protein, or contains a variant of it (e.g., with amino acid residue deletions, insertions, and / or substitutions). The variant retains the desired function of the parent protein (e.g., the ability to form a complex with the guide nucleic acid sequence and the target DNA).

[0096] In some embodiments, the CRISPR-related protein domain may be the Cas9 domain described herein. For example, Cas may refer to a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to the wild-type Cas9 polypeptide described herein. In some embodiments, the wild-type polypeptide is Cas9 from Streptococcus pyogenes (NCBI reference number NC_002737.2 (SEQ ID NO: 1)) and / or UniProt reference number Q99ZW2 (SEQ ID NO: 2). In some embodiments, the wild-type polypeptide is Cas9 from Staphylococcus aureus (SEQ ID NO: 3). In some embodiments, the CRISPR-associated protein domain is a Cpf1 domain or protein, or a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to the wild-type Cpf1 polypeptide described herein (e.g., Cpf1 from Francisella (UniProt reference number U2UMQ6 or SEQ ID NO: 4)). In some embodiments, the CRISPR-associated protein domain may be a modified form of the wild-type protein comprising one or more amino acid residue variations, such as deletions, insertions, or substitutions; fusions or chimeras; or any combination thereof.

[0097] The Cas9 sequences and structures of variant Cas9 orthologs in various organisms have been described. Exemplary organisms from which the Cas9 domains described herein can be derived include, but are not limited to, *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Streptococcus* species, *Staphylococcus aureus*, *Listeria monocytogenes*, *Lactobacillus gasseri*, *Francisella catarrhalis*, *Wallinnella succinate*, *Wallinnella succinate*, *Gamma-Proteus*, *Neisseria meningitidis*, *Campylobacter jejuni*, *Pasteurella multocida*, *Cibotium succinate*, *Rhodospirillum rubrum*, *Nocardia dasoflavones*, *Streptomyces prostidycinus*, *Streptomyces viridans*, and *Streptomyces succinate*. *Streptomyces chromatids*, *Streptomyces roseum*, *Bacillus cyclophosphamide*, *Bacillus pseudomycosis*, *Bacillus selenogenes*, *Microbacterium siberianum*, *Lactobacillus delbrueckii*, *Lactobacillus salivarius*, *Lactobacillus bruchelospermum*, *Treponema denticulatum*, marine micro-oscillating cyanobacteria, Burkholderia bacteria, *Polymonas naphthalene-eating*, *Polymonas* species, *Cyclocarya var. var. vara*, *Cyanobacteria* species, *Microcystis aeruginosa*, *Synechococcus* species, *Halophytes arabatii*, *Ammonia-producing bacteria*, *Bretschneidera sinensis*, *Candidatus ... Desulforudis, Clostridium botulinum, Clostridium difficile, Clostridium falciparum, Glyfungoldella, thermophilic anaerobic halophilic bacteria, thermophilic propionic acid anaerobic enterobacteria, thermophilic acidophilic thiobacillus, ferrooxidizing thiobacillus, purple sulfur bacteria, marine bacilli, halophilic nitrosococci, Watson nitrosococci, marine pseudoalternaria, clustered fibrillary bacilli, long-leaved methanalobacteria, variable anemones, foamy nodule, Nostoc, Spirulina macrophylla, Spirulina platensis, Spirulina genus, Linbiella genus, Prototype microsheath algae, Oscillatoria genus, petroleum algae, African thermocrystal bacilli, Streptococcus pastoris, Neisseria grayi, Campylobacter gullii, Cleaner-eating bacteria, Corynebacterium diphtheriae, and marine cyanobacteria. The Cas9 sequence also includes those from organisms and loci disclosed in Chylinski et al., RNA Biol. (2013) 10(5):726-37.

[0098] In some embodiments, the Cas9 domain is derived from Streptococcus pyogenes (SpCas9). In some embodiments, the Cas9 domain is derived from Staphylococcus aureus (SaCas9).

[0099] Other Cas domains were also considered for use in the epigenetic editor of this paper. These include, for example, those from CasX (Cas12E) (e.g., SEQ ID NO: 5), CasY (Cas12d) (e.g., SEQ ID NO: 6), Casφ (CasPhi) (e.g., SEQ ID NO: 7), Cas12f1 (Cas14a) (e.g., SEQ ID NO: 8), Cas12f2 (Cas14b) (e.g., SEQ ID NO: 9), Cas12f3 (Cas14c) (e.g., SEQ ID NO: 10), and C2c8 (e.g., SEQ ID NO: 11).

[0100] For epigenetic editing, nuclease-derived protein domains (e.g., Cas9 or Cpf1 domains) can be mutated to have reduced or no nuclease activity, such that the protein domain does not cleave DNA or has reduced DNA cleavage activity, while retaining the ability to complex with guide nucleic acid sequences (e.g., guide RNA) and target DNA. For example, nuclease activity may be reduced by at least 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 99% compared to the wild-type domain. In some embodiments, the CRISPR-related protein domains described herein are catalytically inactive (“dead”). Examples of such domains include, for example, dCas9 (“dead” Cas9), dCpf1, ddCpf1, dCasPhi, ddCas12a, dLbCpf1, and dFnCpf1. For example, the dCas9 protein domain may contain one, two, or more mutations that eliminate its nuclease activity compared to wild-type Cas9. The DNA cleavage domain of Cas9 is known to comprise two subdomains: the HNH nuclease subdomain and the RuvC1 subdomain. The HNH subdomain cleaves the strand complementary to the gRNA, while the RuvC1 subdomain cleaves the non-complementary strand. Mutations within these subdomains can silence the nuclease activity of Cas9. For example, mutations D10A (in RuvC1) and H840A (in HNH) completely inactivate the nuclease activity of SpCas9. Similarly, SaCas9 can be inactivated by mutations D10A and N580A. In some embodiments, dCas9 contains at least one mutation in the HNH subdomain and / or the RuvC1 subdomain that reduces or eliminates nuclease activity. In some embodiments, dCas9 contains only the RuvC1 subdomain or only the HNH subdomain. It should be understood that any mutation that inactivates the RuvC1 and / or HNH domains can be included in the dCas9 described herein, such as insertions, deletions, or single or multiple amino acid substitutions in the RuvC1 and / or HNH domains.

[0101] In some embodiments, the dCas9 protein described herein contains a mutation at position D10 (e.g., D10A), H840 (e.g., H840A), or both of the wild-type SpCas9 sequence (SEQ ID NO: 2), as numbered in the sequence provided in UniProt accession number Q99ZW2. In a particular embodiment, dCas9 comprises the amino acid sequence of dSpCas9 (D10A and H840A) (SEQ ID NO: 12).

[0102] In some embodiments, the dCas9 protein as described herein contains a mutation at position D10 (e.g., D10A), N580 (e.g., N580A), or both of the wild-type SaCas9 sequence (e.g., SEQ ID NO: 3). In a particular embodiment, dCas9 contains the amino acid sequence of dSaCas9 (D10A and N580A) (SEQ ID NO: 13).

[0103] Based on this disclosure and knowledge in the art, other suitable mutations that inactivate Cas9 will be apparent to those skilled in the art and are within the scope of this disclosure. Such mutations may include, but are not limited to, D839A, N863A, and / or K603R in SpCas9. This disclosure considers any mutation that reduces or eliminates the nuclease activity of any Cas9 described herein (e.g., a mutation corresponding to any Cas9 mutation described herein).

[0104] Compared to wild-type Cpf1, the dCpf1 protein domain may contain one, two, or more mutations that reduce or eliminate its nuclease activity. The Cpf1 protein has a RuvC-like endonuclease domain similar to that of Cas9 but without the HNH endonuclease domain, and the N-terminus of Cpf1 lacks the α-helical recognition leaflet of Cas9. In some embodiments, dCpf1 contains one or more mutations corresponding to positions D917A, E1006A, or D1255A in the sequence of the novel culprit Francisella Cpf1 protein (FnCpf1; SEQ ID NO: 4). In some embodiments, the dCpf1 protein contains a mutation corresponding to D917A, E1006A, D1255A, D917A / E1006A, D917A / D1255A, E1006A / D1255A, or D917A / E1006A / D1255A, or any of the amino acid sequences of Cpf1 described herein. In some embodiments, dCpf1 contains the D917A mutation. In a particular embodiment, dCpf1 contains the amino acid sequence of dFnCpf1 (SEQ ID NO:14).

[0105] Other nuclease-inactivating CRISPR-related protein domains considered in this paper include those from, for example, dNmeCas9 (e.g., SEQ ID NO: 15), dCjCas9 (e.g., SEQ ID NO: 16), dSt1Cas9 (e.g., SEQ ID NO: 17), dSt3Cas9 (e.g., SEQ ID NO: 18), dLbCpf1 (e.g., SEQ ID NO: 19), dAsCpf1 (e.g., SEQ ID NO: 20), denAsCpf1 (e.g., SEQ ID NO: 21), dHFAsCpf1 (e.g., SEQ ID NO: 22), dRVRAsCpf1 (e.g., SEQ ID NO: 23), dRRAsCpf1 (e.g., SEQ ID NO: 24), dCasX (e.g., SEQ ID NO: 25), and dCasPhi (e.g., SEQ ID NO: 26).

[0106] In some embodiments, the Cas9 domain described herein may be a high-fidelity Cas9 domain, for example, comprising one or more mutations that reduce the electrostatic interaction between the Cas9 domain and the sugar-phosphate backbone of DNA to confer increased target binding specificity. In some embodiments, the high-fidelity Cas9 domain may be nuclease-inactivated as described herein.

[0107] The CRISPR-associated protein domains described herein recognize protospacer adjacent motif (PAM) sequences in target genes. A PAM sequence is typically a 2- to 6 bp DNA sequence immediately following the sequence targeted by the CRISPR-associated protein domain. The PAM sequence is required for CRISPR protein binding and cleavage but is not part of the target sequence. CRISPR-associated protein domains can recognize naturally occurring or typical PAM sequences, or they can have altered PAM specificity. CRISPR-associated protein domains that bind atypical PAM sequences have been described in the art. For example, the Cas9 domain that binds atypical PAM sequences has been described in Kleinstiver et al., Nature (2015) 523(7561):481-5 and Kleinstiver et al., Nat Biotechnol. (2015) 33:1293-8. Such Cas9 domains may include, for example, those from “VRER”SpCas9, “EQR”SpCas9, “VQR”SpCas9, “SpG Cas9”, “SpRYCas9”, and “KKH”SaCas9. Nuclease-inactivated versions of these Cas9 domains are also considered, such as nuclease-inactivated VRER SpCas9 (e.g., SEQ ID NO: 27), nuclease-inactivated EQR SpCas9 (e.g., SEQ ID NO: 28), nuclease-inactivated VQRSpCas9 (e.g., SEQ ID NO: 29), nuclease-inactivated SpG Cas9 (e.g., SEQ ID NO: 30), nuclease-inactivated SpRY Cas9 (e.g., SEQ ID NO: 31), and nuclease-inactivated KKH SaCas9 (e.g., SEQ ID NO: 32). Another example is the Cas9 of the novel culprit *Francis*, which has been engineered to recognize 5'-YG-3' (where “Y” is pyrimidine).

[0108] Based on this disclosure, other suitable CRISPR-related proteins, orthologs, and variants (including nuclease-inactivated variants and sequences) will be apparent to those skilled in the art.

[0109] Guide RNAs that can be used in conjunction with the CRISPR-related protein domains described in this article are further described in Section II below.

[0110] B. Zinc finger protein domain

[0111] In some implementations, the DNA-binding domain of the epigenetic editor described herein comprises a zinc finger protein (ZFP) domain (or, as used herein, a “ZF domain”). A ZFP is a protein having at least one zinc finger and binding to DNA in a sequence-specific manner. A “zinc finger” (ZF) or “zinc finger motif” (ZF motif) refers to a polypeptide domain comprising a folded β-β-α (ββα) protein stable by zinc ions. A ZF binds two to four base pairs of nucleotides, typically three or four base pairs (continuous or discontinuous). Each ZF typically contains approximately 30 amino acids. A ZFP domain may contain multiple ZFs that form tandem contacts with their target nucleic acid sequences. Tandem arrays of ZFs can be engineered to generate artificial ZFPs that bind to desired nucleic acid targets. ZFPs can be rationally designed using a database containing triplet (or quadruple) nucleotide sequences and individual ZF amino acid sequences, wherein each triplet or quadruple nucleotide sequence is associated with one or more amino acid sequences of a ZF that binds to a specific triplet or quadruple sequence. See, for example, U.S. Patents 6,453,242, 6,534,261 and 8,772,453.

[0112] ZFPs are widely distributed in eukaryotic cells and can belong to, for example, the C2H2 class, CCHC class, PHD class, or RING class. An exemplary motif characterizing one of these classes (C2H2 class) is -Cys-(X). 2-4 -Cys-(X) 12 -His-(X) 3-5 -His-(SEQ ID NO: 657), where X is any independently chosen amino acid. In some embodiments, the ZFP domain herein may comprise a ZF array comprising consecutive C2H2-ZFs, each contacting three or more consecutive nucleotides.

[0113] The ZFP domain of the epigenetic editor described herein may include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more ZFs. The ZFP domain may include an array of bidigital or tridigital units, for example, 3, 4, 5, 6, 7, 8, 9 or 10 or more units, wherein each unit binds to a subsite in the target sequence. In some embodiments, a ZFP domain containing at least three ZFs recognizes a target DNA sequence of 9 or 10 nucleotides. In some embodiments, a ZFP domain containing at least four ZFs recognizes a target DNA sequence of 12 to 14 nucleotides. In some embodiments, a ZFP domain containing at least six ZFs recognizes a target DNA sequence of 18 to 21 nucleotides.

[0114] In some embodiments, the ZF in the ZFP domain described herein is linked via a peptide linker. The length of the peptide linker can be, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more amino acids. In some embodiments, the linker contains 5 or more amino acids. In some embodiments, the linker contains 7-17 amino acids. The linker can be flexible or rigid.

[0115] In some implementations, the zinc finger array can have a sequence:

[0116] SRPGERPFQCRICMRNFSXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXXHXXTH[connector]FQCRICMRNFSXXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXHXXTH[connector]PFQCRICMRNFSXXXXXXXXHXXTHTGEKPFQCRICMRNFSXXXXXXXXHXXTHLRGS (SEQ ID NO: 650),

[0117] Or a sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the one indicated by the zinc finger, wherein “XXXXXXX” represents the amino acid of the ZF-recognized helix that imparts DNA binding specificity to the zinc finger; each X can be chosen independently. In the above sequence, the italicized “XX” can be TR, LR, or LK, and “[Linker]” indicates the linker sequence. In some embodiments, the linker sequence is TGSQKP (SEQ ID NO: 651); this linker can be used when the subsites targeted by the ZF are adjacent. In some embodiments, the linker sequence is TGGGGSQKP (SEQ ID NO: 652); this linker can be used when there are bases between the subsites targeted by the zinc finger. The two indicated linkers can be the same or different.

[0118] The ZFP domain described herein may contain arrays of two or more adjacent ZFs, which are directly adjacent to each other (e.g., separated by short (typical) linker sequences) or separated by longer, flexible, or structured polypeptide sequences. In some embodiments, directly adjacent fingers bind to consecutive nucleic acid sequences, i.e., to adjacent trinucleotides / triads. In some embodiments, adjacent fingers are cross-linked between their respective target triads, which can help enhance or strengthen the recognition of target sequences and lead to binding of overlapping sequences. In some embodiments, distant ZFs within the ZFP domain can recognize (or bind) discontinuous nucleotide sequences.

[0119] The amino acid sequences of the ZF DNA recognition helix of the exemplary ZFP domain and its PCSK9 target sequence are shown in Table 1 below, where the numbers in parentheses represent SEQ ID NO:

[0120] Table 1. ZF sequences of exemplary ZFP domains targeting PCSK9

[0121]

[0122]

[0123]

[0124]

[0125]

[0126] In some embodiments, the ZFP domain of the epigenetic editor of this disclosure binds to a target sequence selected from any one of SEQ ID NO: 700-747. In a further embodiment, the ZFP domain sequentially comprises the F1-F6 amino acid sequences of any one of ZF001-ZF048 as shown in Table 1. The F1-F6 amino acid sequences may be placed within the ZF frame sequence of SEQ ID NO: 650, or within any other ZF frame known in the art.

[0127] C.TALE

[0128] In some embodiments, the DNA-binding domain of the epigenetic editor described herein comprises a transcription activator-like effector (TALE) domain. The DNA-binding domain of a TALE contains a highly conserved sequence of approximately 33-34 amino acids, with repeating variable di-residues (RVDs) at positions 12 and 13, which are crucial for recognizing specific nucleotides. TALEs can be engineered to practically bind to any desired DNA sequence. Methods for programming TALEs are known in the art. For example, this method is described in Carroll et al., Genet Soc Amer. (2011) 188(4):773-82; Miller et al., Nat Biotechnol. (2007) 25(7):778-85; Christian et al., Genetics (2008)186(2):757-61; Li et al., Nucl Acids Res. (2010) 39(1):359-72; and Moscou et al., Science (2009) 326(5959):1501.

[0129] D. Other DNA-binding domains

[0130] Other DNA-binding domains have been considered for use in the epigenetic editor described herein. In some embodiments, the DNA-binding domain comprises, for example, the argonaute protein domain (NgAgo) from *Haloxybacterium gargearii*. NgAgo is an ssDNA-guided endonuclease that is guided by 5' phosphorylated ssDNA (gDNA) to its target site, where it produces a double-strand break. Unlike Cas9, the NgAgo-gDNA system does not require a protospacer adjacent motif (PAM). Therefore, using nuclease-inactivated NgAgo (dNgAgo) can greatly amplify the number of bases that can be targeted. The characterization and uses of NgAgo have been described, for example, in Gao et al., Nat Biotechnol. (2016) 34(7):768-73; Swarts et al., Nature (2014) 507(7491):258-61; and Swarts et al., Nucl Acids Res. (2015) 43(10):5120-9.

[0131] In some implementations, the DNA-binding domain contains an inactivated nuclease, for example, an inactivated broad-spectrum nuclease. Other non-limiting examples of DNA-binding domains include tetracycline-controlled repressor (tetR) DNA-binding domains, leucine zippers, helical-loop-helical (HLH) domains, helical-turn-helical domains, β-sheet motifs, steroid receptor motifs, bZIP domains, homology domains, and AT-hook.

[0132] II. Guide polynucleotides

[0133] The epigenetic editor comprising a polynucleotide guide DNA-binding domain described herein may also include guide polynucleotides capable of forming complexes with the DNA-binding domain. Guide polynucleotides may comprise RNA, DNA, or a mixture of both. For example, in the case where the polynucleotide guide DNA-binding domain is a CRISPR-associated protein domain, the guide polynucleotide may be guide RNA (gRNA). "Guide RNA" or "gRNA" refers to a nucleic acid capable of hybridizing to a target sequence and directing the binding of the CRISPR-Cas complex to the target sequence. Methods using guide polynucleotide sequences with programmable DNA-binding proteins (e.g., CRISPR-associated protein domains) for site-specific DNA targeting (e.g., to modify the genome) are known in the art.

[0134] A guide polynucleotide sequence (e.g., a gRNA sequence) may comprise two parts: 1) a nucleotide sequence containing a “target sequence” complementary to a target nucleic acid sequence (“target sequence”), for example, complementary to a nucleic acid sequence contained in a genomic target site; and 2) a nucleotide sequence binding a polynucleotide guide DNA-binding domain (e.g., a CRISPR-Cas protein domain). The nucleotide sequence in 1) may contain a target sequence that is 100% complementary to a genomic nucleic acid sequence (e.g., a nucleic acid sequence contained in a genomic target site) and therefore can hybridize with the target nucleic acid sequence. The nucleotide sequence in 1) may be referred to, for example, crripRNA or crRNA. The nucleotide sequence in 2) may be referred to as a scaffold sequence of the guide nucleic acid, such as tracrRNA, or an activation region of the guide nucleic acid, and may contain a stem-loop structure. Parts 1) and 2) as described above may be fused to form a single guide (e.g., a single guide RNA or sgRNA), or may be on two separate nucleic acid molecules. In some embodiments, the guide polynucleotide comprises parts 1) and 2) linked by a linker. In some implementations, the guide polynucleotide includes portions 1) and 2) linked by non-nucleic acid linkers (e.g., peptide linkers or chemical linkers).

[0135] The guide polynucleotide portion 2 (scaffold sequence) described herein can be, for example, as described in Jinek et al., Science (2012) 337:816-21; U.S. Patent Publication 2016 / 0208288; or U.S. Patent Publication 2016 / 0200779. Variations of portion 2) are also contemplated in this disclosure. For example, the tetraloop and stem loop of the gRNA scaffold (tracrRNA) sequence can be modified to include RNA aptamers that can be bound by specific protein domains. In some embodiments, such modified gRNA can be used to promote the recruitment of repressor or activator domains that facilitate the fusion of RNA aptamers with proteins.

[0136] The gRNAs described herein typically contain a targeting domain and a binding domain. The targeting domain (also called the “target sequence”) may contain a nucleic acid sequence that binds to a target site (e.g., a genomic nucleic acid molecule within a cell). The target site may be a double-stranded DNA sequence containing the PAM sequence as well as the target sequence, which is located on the same strand as and directly adjacent to the PAM sequence. The targeting domain of the gRNA may contain an RNA sequence corresponding to the target sequence; that is, it is a sequence similar to the target domain, sometimes with one or more mismatches, but typically contains an RNA sequence rather than a DNA sequence. Therefore, the targeting domain of the gRNA may pair with (completely or partially complementary) the bases of the double-stranded target site sequence complementary to the target sequence, and thus with the bases of the strand complementary to the strand containing the PAM sequence. It should be understood that the targeting domain of the gRNA typically does not include a sequence similar to the PAM sequence. It should be further understood that the position of the PAM can be either the 5' or 3' of the target sequence, depending on the nuclease used. For example, the PAM is typically located at the 3' of the target sequence of the Cas9 nuclease and the 5' of the target sequence of the Cas12a nuclease. For an explanation of the location of PAM and the mechanism by which gRNA binds to its target site, see, for example, Vanegas et al., Fungal Biol Biotechnol. (2019) 6:6. Figure 1 For further explanation and description of the mechanism by which gRNA targets RNA-guided nucleases to their target sites, see Fu et al., Nat Biotechnol (2014) 32(3):279-84 and Sternberg et al., Nature (2014) 507(7490):62-7, each of which is incorporated herein by reference.

[0137] In some embodiments, the target domain sequence contains 17 to 30 nucleotides and corresponds exactly to the target sequence (i.e., there are no mismatched nucleotides). However, in some embodiments, the target domain sequence may contain one or more, but typically no more than four mismatches, for example, 1, 2, 3, or 4 mismatches. Since the target domain is part of gRNA (which is an RNA molecule), it typically contains ribonucleotides, while the DNA target domain will contain deoxyribonucleotides.

[0138] The following provides an exemplary illustration of a Cas9 target site, comprising a 22-nucleotide target domain and an NGG PAM sequence, and a gRNA containing a target domain that perfectly corresponds to the target sequence (and therefore contains base pairs that are perfectly complementary to the DNA strand containing the target sequence and PAM):

[0139] [Target domain (DNA)][PAM]

[0140] 5'-NNNNNNNNNNNNNNNNNNNNN-NNGG-3' (DNA)

[0141] 3'-NNNNNNNNNNNNNNNNNNNNN-NNCC-5' (DNA)

[0142] | | | | | | | | | | | | | | | | | | | | | |

[0143] 5'-NNNNNNNNNNNNNNNNNNNNN-N-[gRNA scaffold]-3' (RNA)

[0144] [Targeting domain (RNA)][Binding domain]

[0145] The following provides an exemplary description of a CaSL2a target site, which comprises a 22-nucleotide target domain and a TTN PAM sequence, as well as a gRNA containing a target domain that perfectly corresponds to the target sequence (and therefore contains base pairs that are perfectly complementary to the DNA strand containing the target sequence and PAM):

[0146] [PAM][Target domain (DNA)]

[0147] 5'-TTNNNNNNNNNNNNNNNNNNN-NNNN-3' (DNA)

[0148] 3'-AANNNNNNNNNNNNNNNNNNNN-NNNN-5' (DNA)

[0149] | | | | | | | | | | | | | | | | | | | | | |

[0150] 5'-[gRNA Scaffold]-NNNNNNNNNNNNNNNNNNNNN-N-3' (RNA)

[0151] [Binding domain][Targeting domain (RNA)]

[0152] While not wishing to be bound by theory, it is believed, at least in some embodiments, that the length of the targeting domain and its complementarity with the target sequence contribute to the specificity of the interaction between the gRNA / Cas9 molecular complex and the target nucleic acid. In some embodiments, the targeting domain of the gRNA provided herein is 5 to 50 nucleotides long. In some embodiments, the targeting domain is about 15 to 25 nucleotides long. In some embodiments, the targeting domain is about 18 to 22 nucleotides long. In some embodiments, the targeting domain is about 19-21 nucleotides long. In some embodiments, the targeting domain is 15 nucleotides long. In some embodiments, the targeting domain is 16 nucleotides long. In some embodiments, the targeting domain is 17 nucleotides long. In some embodiments, the targeting domain is 18 nucleotides long. In some embodiments, the targeting domain is 19 nucleotides long. In some embodiments, the targeting domain is 20 nucleotides long. In some embodiments, the targeting domain is 21 nucleotides long. In some embodiments, the targeting domain is 22 nucleotides long. In some embodiments, the targeting domain is 23 nucleotides long. In some embodiments, the target domain is 24 nucleotides long. In some embodiments, the target domain is 25 nucleotides long. In some embodiments, the target domain corresponds exactly to the target sequence or a portion thereof provided herein, with no mismatches. In some embodiments, the target domain of the gRNA provided herein contains one mismatch relative to the target sequence provided herein. In some embodiments, the target domain contains two mismatches relative to the target sequence. In some embodiments, the target domain contains three mismatches relative to the target sequence.

[0153] This article describes methods for designing, selecting, and validating gRNAs, and these methods are known in the art. Software tools can be used to optimize gRNAs corresponding to target DNA sequences, for example, to minimize overall off-target activity across the genome. For example, DNA sequence search algorithms can be used to identify target sequences in the crRNA of a gRNA for use with Cas9. Exemplary gRNA design tools include those described in Bae et al., Bioinformatics (2014) 30:1473-5.

[0154] The guide polynucleotides (e.g., gRNAs) described herein can have various lengths. In some embodiments, the length of the spacer or target sequence depends on the CRISPR-related protein component of the epigenetic editor system used. For example, Cas proteins from different bacterial species have different optimal target sequence lengths. Therefore, the length of the spacer sequence can contain, for example, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more nucleotides. In some embodiments, the length of the spacer contains 10-24, 11-20, 11-16, 18-24, 19-21 or 20 nucleotides. In some embodiments, the guide polynucleotide (e.g., gRNA) is 15-100 nucleotides in length (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50) and contains at least 10 nucleotides of the target sequence. (For example, a spacer sequence of 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50) consecutive complementary nucleotides. In some embodiments, the guide polynucleotide described herein may be truncated, for example, by 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 40, 50, or more nucleotides.

[0155] In some embodiments, the 3' end of the PCSK9 target sequence is immediately adjacent to the PAM sequence (e.g., a typical PAM sequence, such as the NGG of SpCas9). The complementarity between the guide polynucleotide's target sequence (e.g., the spacer region sequence of gRNA) and the target sequence can be at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%. In certain embodiments, the target sequence and the target sequence can be 100% complementary. In other embodiments, the target sequence and the target sequence can contain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 mismatches.

[0156] Guide polynucleotides (e.g., gRNA) can be modified, for example, by chemical alterations and synthetic modifications. For example, modified gRNA may include alterations or substitutions of one or two non-linked phosphate groups and / or one or more linked phosphate groups in a phosphodiester backbone, alterations to the ribose (e.g., the 2' hydroxyl group on the ribose), alterations to the phosphate moiety, modifications or substitutions of naturally occurring nucleobases, modifications or substitutions of the ribose phosphate backbone, modifications to the 3' and / or 5' ends of oligonucleotides, substitutions of terminal phosphate groups, or partial, capped, or linker conjugations, or any combination thereof.

[0157] In some implementations, one or more ribosomes of the gRNA may be modified. Examples of chemical modifications to the ribosomes include, but are not limited to, 2'-O-methyl (2'-OMe), 2'-fluoro (2'-F), 2'-deoxy, 2'-O-(2-methoxyethyl) (2'-MOE), 2'-NH2, 2'-O-allyl, 2'-O-ethylamine, 2'-O-cyanoethyl, 2'-O-acetyl ester, or bicyclic nucleotides such as locked nucleic acids (LNA), 2'-(5-restricted ethyl (S-cEt)), restricted MOE, or 2'-O,4'-C-aminomethylene-bridged nucleic acids (2',4'-BNANC). 2'-O-methyl modification and / or 2'-fluoro modification can increase the binding affinity of the gRNA oligonucleotide and / or the stability of nucleases.

[0158] In some embodiments, one or more phosphate groups of the gRNA may be chemically modified. Examples of chemical modification of phosphate groups include, but are not limited to, phosphate thioester (PS), phosphonoacetate (PACE), thiophosphonoacetate (thioPACE), amide, triazole, phosphonate, and phosphate triester modifications. In some embodiments, the guide polynucleotide described herein may contain one, two, three, or more PS bonds at or near the 5' end and / or 3' end; the PS bonds may be continuous or discontinuous.

[0159] In some implementations, the gRNA described herein comprises a mixture of ribonucleotides and deoxyribonucleotides and / or one or more PS bonds.

[0160] In some implementations, one or more nucleotides of the gRNA may be chemically modified. Examples of chemically modified nucleotides include, but are not limited to, 2-thiouridine, 4-thiouridine, N6-methyladenosine, pseudouridine, 2,6-diaminopurine, inosine, thymidine, 5-methylcytosine, 5-substituted pyrimidine, isoguanine, isocytosine, and nucleotides having a halogenated aromatic group. Chemical modification may be performed in the spacer region, the tracr RNA region, the stem-loop, or any combination thereof.

[0161] Table 2 below lists exemplary gRNA target sequences for epigenetic modification of human PCSK9, along with the coordinates of the start and end positions of the target sites on human chromosome 1 (SEQ: SEQ ID NO). The table also shows the distance from the start coordinates to the TSS coordinates within the PCSK9 gene.

[0162] Table 2. Exemplary target sequences of gRNAs targeting PCSK9

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169]

[0170]

[0171]

[0172]

[0173]

[0174] In some implementations, the gRNA described herein does not contain the sequence CCCGCACCUUGGCGCAGCGG (SEQ ID NO:1490).

[0175] Consider any tracr sequence known in the art for use with the gRNA described herein. In some embodiments, the gRNA described herein has a tracr sequence shown in Table 3 below, or a tracr sequence that is at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to the tracr sequence (SEQ:SEQ ID NO) shown below.

[0176] Table 3. Exemplary TRAC sequences

[0177]

[0178] In some embodiments, the gRNA described herein is delivered directly to cells (e.g., via an RNP complex and a CRISPR-associated protein domain). In some embodiments, the gRNA is delivered to cells via an expression vector (e.g., a plasmid vector or a viral vector) introduced into the cells, wherein the cells subsequently express the gRNA from the expression vector. Methods for introducing gRNA and expression vectors into cells are well known in the art.

[0179] III. Effect subdomain

[0180] The epigenetic editors described herein include one or more effector protein domains (also referred to herein as “epigenetic effector domains” or “effector domains”) that influence the epigenetic modification of target genes. Epigenetic editors having one or more effector domains can regulate the expression of target genes without altering their nucleotide sequence. In some embodiments, the effector domains described herein can provide repression or silencing of the expression of a target gene (such as PCSK9), for example, by repressing transcription or by modifying or remodeling chromatin. Such effector domains are also referred herein as “repression domains,” “repressor domains,” or “epigenetic repressor domains.” Non-limiting examples of chemical modifications that can be mediated by effector domains include methylation, demethylation, acetylation, deacetylation, phosphorylation, SUMOylation, and / or ubiquitination of DNA or histone residues.

[0181] In some implementations, the effector domains of the epigenetic editor described herein can be modified with histone tails, for example, by adding or removing an active marker on the histone tail.

[0182] In some implementations, the effector domains of the epigenetic editor described herein may contain or recruit transcription-related proteins, such as transcription repressors. These transcription-related proteins may be endogenous or exogenous.

[0183] In some implementations, the effector domain of the epigenetic editor described herein may, for example, contain proteins that directly or indirectly block transcription factors from approaching genes of interest containing target sequences.

[0184] Effector domains can be full-length proteins or fragments thereof that retain the function of epigenetic effectors (“functional domains”). Functional domains capable of regulating (e.g., repressing) gene expression can be derived from larger proteins. For example, functional domains that can reduce the expression of target genes can be identified based on the sequence of repressor proteins. The amino acid sequences of gene expression regulatory proteins can be obtained from available genome browsers, such as the UCSD Genome Browser or the Ensembl Genome Browser. Protein annotation databases such as UniProt or Pfam can be used to identify functional domains within a whole protein sequence. As a starting point, gene expression regulatory activity can be tested on the largest sequence covering all regions identified by different databases. Various truncations can then be tested to identify the smallest functional units.

[0185] This disclosure also contemplates variants of the effector domains described herein. For example, a variant may refer to a polypeptide having at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity and / or sequence similarity to the wild-type effector domain described herein. In certain embodiments, the variant retains at least about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% of the epigenetic effector function of the wild-type effector domain.

[0186] In some embodiments, the effect subdomains described herein may comprise a fusion of two or more effect subdomains (e.g., KOX1, KRAB, and ZIM3). The effect subdomain may, for example, comprise a fusion of 2, 3, 4, 5, 6, 7, 8, 9, or 10 effect subdomains, such as the effect subdomains described herein. In some embodiments, the effect subdomain comprises a truncated form of an effect subdomain and a fusion of a second effect subdomain. In some embodiments, the effect subdomain comprises a fusion of truncated forms of two effect subdomains (e.g., a fusion of the N-terminal and C-terminal portions of two effect subdomains).

[0187] In some embodiments, the epigenetic editor described herein may comprise one, two, three, four, five, six, seven, eight, nine, ten, or more effector domains. In some embodiments, the epigenetic editor comprises one or more fusion proteins (e.g., one, two, or three fusion proteins), each fusion protein having one or more effector domains (e.g., one, two, or three effector domains) connected to a DNA-binding domain. In some embodiments, the effector domains may induce combinations of epigenetic modifications, such as transcriptional repression and DNA methylation, DNA methylation and histone deacetylation, DNA methylation and histone demethylation, DNA methylation and histone methylation, DNA methylation and histone phosphorylation, DNA methylation and histone ubiquitination, and DNA methylation and histone SUMOylation.

[0188] In some embodiments, the effector domains described herein (e.g., DNMT3A and / or DNMT3L) are encoded by nucleotide sequences targeting that effector domain found in the natural genome (e.g., human or mouse). In other embodiments, the effector domains described herein are encoded by nucleotide sequences that have been codon-optimized for optimal expression in human cells.

[0189] The effector domains described herein may include, for example, transcriptional repressors, DNA methyltransferases, and / or histone modifiers, as further detailed below.

[0190] A. Transcription repressors

[0191] In some embodiments, the epigenetic effector domains described herein mediate the repression of target gene expression (e.g., transcription). For example, the effector domain may include a Krüppel-associated box (KRAB) repressor domain; a repressor element silencing transcription factor (REST) ​​repressor domain; a KRAB-associated protein 1 (KAP1) domain; a MAD domain; an FKHR (forkhead gene in the rhabdomyosarcoma gene) repressor domain; an EGR-1 (early growth response gene product 1) repressor domain; an ets2 repressor repressor domain (ERD); a MAD smSIN3 interaction domain (SID); a WRPW motif of a hair-associated basic helix-loop-helix (bHLH) repressor protein; an HP1α chromatin shading repressor domain; an HP1β repressor domain; or any combination thereof. The effector domain may recruit one or more protein domains that repress the expression of target genes, for example, through scaffold proteins. In some implementations, the effector domain may recruit or interact with the scaffold protein domain, which in turn recruits PRMT, HDAC, SETDB1, or NuRD protein domains.

[0192] In some implementations, the effector domain contains functional domains derived from zinc finger repressors, such as the KRAB domain. The KRAB domain is found in approximately 400 human ZFP-based transcription factors. Descriptions of the KRAB domain can be found, for example, in Ecco et al., Development (2017) 144(15):2719-29 and Lambert et al., Cell (2018) 172:650-65.

[0193] In some embodiments, the effector domain comprises a repressor domain (e.g., KRAB) derived from KOX1 / ZNF10, KOX8 / ZNF708, ZNF43, ZNF184, ZNF91, HPF4, HTF10, or HTF34. In some embodiments, the effector domain comprises a repressor domain derived from ZIM3, ZNF436, ZNF257, ZNF675, ZNF490, ZNF320, ZNF331, ZNF816, ZNF680, ZNF41, ZNF189, ZNF528, ZNF543, ZNF554, ZNF140, ZNF610, ZNF264, ZNF350, ZNF8, ZNF582, ZNF30, ZNF324, ZNF98, ZNF669, ZNF677, ZNF596, ZNF214, ZNF37, Z... The repressor domains (e.g., KRAB) of NF34, ZNF250, ZNF547, ZNF273, ZNF354, ZFP82, ZNF224, ZNF33, ZNF45, ZNF175, ZNF595, ZNF184, ZNF419, ZFP28-1, ZFP28-2, ZNF18, ZNF213, ZNF394, ZFP1, ZFP14, ZNF416, ZNF557, ZNF566, ZNF729, ZIM2, ZNF254, ZNF764, ZNF785, or any combination thereof. For example, the repressor domain may be a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627. In a particular embodiment, the repressor domain is the ZIM3 KRAB domain. In a further embodiment, the effector domain is derived from a human protein, such as human ZIM3, human KOX1, human ZFP28, or human ZN627.

[0194] Sequences of exemplary effector domains that can reduce or silence target gene expression, or protein sequences containing them, are provided in Table 4 below (SEQ: SEQ ID NO). Further examples of repressor and transcriptional repressor domains can be found, for example, in PCT patent publication WO 2021 / 226077 and Tycko et al., Cell (2020) 183(7):2020-35, each of which is incorporated herein by reference in its entirety.

[0195] Table 4. Exemplary effector domains that can reduce or silence gene expression

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204] This disclosure covers functional analogs of any of the proteins listed above, i.e., molecules having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the protein transcription factor function). For example, a functional analog may be an isoform or variant of one of the proteins listed above, such as a portion of the aforementioned protein with or without additional amino acid residues and / or containing a mutation relative to the aforementioned protein. In some embodiments, the functional analog has at least 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with one of the sequences listed in Table 4. Homologs, orthologs, and mutants of the proteins listed above are also considered.

[0205] In some embodiments, the epigenetic editor described herein comprises a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627, and / or an effector domain derived from KAP1, MECP2, HP1a, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2, optionally wherein the parent protein is a human protein. In specific embodiments, the epigenetic editor described herein comprises domains derived from KOX1, ZIM3, ZFP28, and / or ZN627, optionally wherein the parent protein is a human protein. In some embodiments, the epigenetic editor may comprise a KRAB domain derived from KOX1, such as human KOX1 (ZNF10). In some embodiments, the epigenetic editor may comprise a KRAB domain derived from ZIM3, such as human ZIM3 (ZNF657 or ZNF264). In some embodiments, the epigenetic editor may include a KRAB domain derived from ZFP28, such as human ZFP28. In some embodiments, the epigenetic editor may include a KRAB domain derived from ZN627, such as human ZN627. In some embodiments, the epigenetic editor described herein may include a CDYL2, such as human CDYL2, combined with a KOX1 KRAB domain (e.g., human KOX1 KRAB domain), and / or a TOX domain (e.g., human TOX domain).

[0206] In some embodiments, the epigenetic effector described herein comprises a repressor domain derived from KOX1 / ZNF10 (SEQ ID NO: 89). For example, the repressor domain may comprise the sequence of SEQ ID NO: 89, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 89.

[0207] In some embodiments, the epigenetic effectors described herein include repressor domains derived from KOX1 / ZNF10, as shown in Table 5 below:

[0208] Table 5. Exemplary effector subdomains derived from KOX1 / ZNF10

[0209]

[0210] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 565, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 565.

[0211] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 566, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 566.

[0212] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 567, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 567.

[0213] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 568, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 568.

[0214] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 569, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 569.

[0215] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 570, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 570.

[0216] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 571, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 571.

[0217] In a particular embodiment, the repressor domain may contain the amino acid sequence of SEQ ID NO: 572, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 572.

[0218] B. DNA methyltransferase

[0219] In some implementations, the effector domains of the epigenetic editor described herein alter target gene expression through DNA modifications, such as methylation. Transcriptional activity is often lower in highly methylated DNA regions than in less methylated regions. DNA methylation primarily occurs at CpG sites (short for "C-phosphate-G-" or "cytosine-phosphate-guanine" sites). Many mammalian genes have promoter regions (nucleic acid regions with a high frequency of CpG dinucleotides) that are near or include CpG islands.

[0220] The effector domains described herein can be, for example, DNA methyltransferases (DNMTs) or their catalytic domains, or domains capable of recruiting DNA methyltransferases. DNMT encompasses enzymes that catalyze the transfer of methyl groups to DNA nucleotides, such as typical cytosine-5 DNMTs (e.g., DNMT1, DNMT3A, DNMT3B, and DNMT3C) that catalyze the addition of methyl groups to genomic DNA. The term also encompasses atypical family members that do not themselves catalyze methylation but recruit (including activate) catalytically active DNMTs; a non-limiting example of such a DNMT is DNMT3L. See, for example, Lyko, Nat Review (2018) 19:81-92. Unless otherwise stated, a DNMT domain can refer to a polypeptide domain derived from a catalytically active DNMT (e.g., DNMT1, DNMT3A, and DNMT3B) or a non-catalytically active DNMT (e.g., DNMT3L). DNMTs can repress the expression of target genes by recruiting repressive regulatory proteins. In some embodiments, methylation occurs at a CG (or CpG) dinucleotide sequence. In some implementations, methylation is performed at the CHG or CHH sequence, where H is any one of A, T, or C.

[0221] In some embodiments, the DNMT described herein may be an animal DNMT (e.g., mammalian DNMT), a plant DNMT, a fungal DNMT, or a bacterial DNMT. Bacterial DNMTs may be obtained from bacterial species (e.g., cocci, bacilli, spirilla, or intracellular Gram-positive or Gram-negative bacteria). In some embodiments, the bacterial species are mycoplasma bacteria, marine mycoplasma, or Chinese spiroplasma. In some embodiments, the bacterial species are not *Mycoplasma penetratingis*, *S. monbiae*, *Haemophilus parainfluenzae*, *Arthrobacter lutea*, *Haemophilus haemolyticus*, *Haemophilus hemolyticus*, *Moraxella*, *Escherichia coli*, *Thermophyton aquaticus*, *S. crescentis*, or *Clostridium difficile*. In some embodiments, the epigenetic editor described herein comprises a DNMT domain containing the sequence of SEQ ID NO: 601 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 601. In some embodiments, the epigenetic editor described herein comprises a DNMT domain containing a sequence identical to SEQ ID NO: 602 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO: 602. In some embodiments, the epigenetic editor described herein comprises a DNMT domain containing a sequence identical to SEQ ID NO: 603 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO: 603.

[0222] In some embodiments, the DNMT in the epigenetic editor described herein may include, for example, DNMT1, DNMT3A, DNMT3B, and / or DNMT3C. In some embodiments, the DNMT is a mammalian (e.g., human or mouse) DNMT. In a particular embodiment, the DNMT is DNMT3A (e.g., human DNMT3A). In some embodiments, the epigenetic editor described herein comprises a DNMT3A domain containing the sequence of SEQ ID NO: 574 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 574. In some embodiments, the epigenetic editor described herein comprises a DNMT3A domain containing the sequence of SEQ ID NO: 575 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 575. In some implementations, the DNMT3A domain may have mutations at, for example, at positions H739 (such as H739A or H739E), R771 (such as R771L), and / or R836 (such as R836A or R836Q) or any combination thereof (according to SEQ ID NO: 574).

[0223] In some embodiments, the effector domains described herein may be DNMT-like domains. As used herein, a “DNMT-like domain” is a regulator of DNMT that can activate or recruit other DNMT domains but is not itself methylating active. In some embodiments, the DNMT-like domain is a mammalian (e.g., human or mouse) DNMT-like domain. In some embodiments, the DNMT-like domain is DNMT3L, which may be, for example, human DNMT3L or mouse DNMT3L. In some embodiments, the epigenetic editor described herein comprises a DNMT3L domain containing the sequence of SEQ ID NO: 578 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 578. In some embodiments, the epigenetic editor described herein comprises a DNMT3L domain containing a sequence identical to SEQ ID NO: 579 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO: 579. In some embodiments, the epigenetic editor described herein comprises a DNMT3L domain containing a sequence identical to SEQ ID NO: 580 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO: 580. In some embodiments, the epigenetic editor described herein comprises a DNMT3L domain containing a sequence identical to SEQ ID NO: 581 or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of SEQ ID NO: 581. In some implementations, the DNMT3L domain may have mutations (according to SEQ ID NO: 578) corresponding to mutations at positions D226 (e.g., D226V), Q268 (e.g., Q268K), or both.

[0224] In some embodiments, the epigenetic editor described herein may include both DNMT and DNMT-like effector domains. For example, the epigenetic editor may include a DNMT3A-3L domain, wherein DNMT3A and DNMT3L may be covalently linked. In other embodiments, the epigenetic editor described herein may include an effector domain that contains only the DNMT3A domain (e.g., human DNMT3A) or only a DNMT-like domain (e.g., DNMT3L, which may be human or mouse DNMT3L).

[0225] Table 6 below provides exemplary DNMTs that may be part of the epigenetic effector subdomains described herein, or that may be derived from the effector subdomains of the epigenetic editors described herein.

[0226] Table 6. Exemplary DNMT Sequences

[0227]

[0228]

[0229] This disclosure covers any functional analogues of the proteins listed above, i.e., molecules having the same or substantially the same biological function (e.g., retaining 70% or more, 80% or more, 90% or more, 95% or more, or 98% or more of the protein's DNA methylation or recruitment function). For example, a functional analogue may be an isoform or variant of one of the proteins listed above, such as containing a portion of the protein listed above with or without additional amino acid residues and / or containing a mutation relative to the protein listed above. In some embodiments, the functional analogue has at least 75%, 80%, 85%, 90%, 95%, 98%, or 99% sequence identity with one of the sequences listed in Table 6. In some embodiments, the effector domains herein comprise only the functional domains (or functional analogues thereof) of the proteins listed above, such as catalytic or recruitment domains. In some embodiments, the effector domains herein comprise one or more epigenetic effector domains selected from Table 6, or their functional homologs, orthologs, or variants.

[0230] As used herein, a DNMT domain (e.g., the DNMT3A domain or the DNMT3L domain) is a protein domain that is identical to that of the parent protein (e.g., human or mouse DNMT3A or DNMT3L) or its functional analogues (e.g., having a functional fragment of the parent protein, such as a catalytic fragment or a recruitment fragment; and / or having a mutation that improves the activity of the DNMT protein).

[0231] The epigenetic editor described in this paper can achieve methylation at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more CpG dinucleotide sequences in, for example, in a target gene or chromosome. The CpG dinucleotide sequences can be located inside or near the target gene in a CpG island, or they can be located in regions that are not CpG islands. A CpG island typically refers to a nucleic acid sequence or chromosomal region containing a high frequency of CpG dinucleotides. For example, a CpG island may contain at least 50% GC content. CpG islands can have a high CpG observed to expected ratio, for example, a CpG observed to expected ratio of at least 60%. As used herein, the CpG observation to expected ratio is determined by the number of CpGs * (sequence length) / (number of Cs * number of Gs). In some embodiments, CpG islands have a CpG observation to expected ratio of at least 60%, 70%, 80%, 90%, or higher. CpG islands can be, for example, sequences or regions of at least 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, or 800 nucleotides. In some embodiments, only one or fewer CpG dinucleotides are methylated by an epigenetic editor.

[0232] In some implementations, the epigenetic editor described herein performs methylation at demethylated nucleic acid sequences, i.e., sequences that may lack a methyl group on a 5-methylcytosine nucleotide (e.g., in CpG) compared to a standard control. For example, demethylation can occur in senescent cells or cancer cells (e.g., early stages of tumor formation) relative to young cells or non-cancer cells, respectively.

[0233] In some implementations, the epigenetic editor described herein induces methylation at hypermethylated nucleic acid sequences.

[0234] In some embodiments, methylation can be introduced by an epigenetic editor at sites other than CpG dinucleotides. For example, the target gene sequence can be methylated at the C nucleotide of the CpA, CpT, or CpC sequence. In some embodiments, the epigenetic editor contains a DNMT3A domain and achieves methylation at CpG, CpA, CpT, CpC sequences, or any combination thereof. In some embodiments, the epigenetic editor contains a DNMT3A domain lacking a regulatory subdomain and maintaining only the catalytic domain. In some embodiments, an epigenetic editor containing a DNMT3A catalytic domain achieves methylation only at the CpG sequence. In some embodiments, an epigenetic editor containing a DNMT3A domain with mutations, such as R836A or R836Q mutations (according to SEQ ID NO: 574), exhibits higher methylation activity at CpA, CpC, and / or CpT sequences compared to an epigenetic editor containing a wild-type DNMT3A domain.

[0235] C. Histone modifications

[0236] In some implementations, the effector domains of the epigenetic editor described herein mediate histone modifications. Histone modifications play structural and biochemical roles in gene transcription, for example, by forming or disrupting nucleosome structures that bind to histones and prevent gene transcription. Histone modifications can include, for example, at their N-terminal ends (“histone tails”), acetylation, deacetylation, methylation, phosphorylation, ubiquitination, SUMOylation, etc. These modifications maintain or specifically transform chromatin structure, thereby controlling responses occurring on chromosomal DNA, such as gene expression, DNA replication, DNA repair, etc. Post-translational modifications of histones are epigenetic regulatory mechanisms and are considered essential for eukaryotic genetic regulation. Recent studies have revealed chromatin remodeling factors such as SWI / SNF, RSC, NURF, NRD, etc., which facilitate transcription factor access to DNA by modifying nucleosome structures; histone acetyltransferases (HATs) that regulate histone acetylation states; and histone deacetylases (HDACs) as important regulatory factors.

[0237] Specifically, the unstructured N-terminus of histones can be modified by acetylation, deacetylation, methylation, ubiquitination, phosphorylation, SUMOylation, ribosylation, citrullination, O-glycosylation, crotonylation, or any combination thereof. For example, histone acetyltransferases (HATs) utilize acetyl-CoA as a cofactor and catalyze the transfer of an acetyl group to the ε-amino group of the lysine side chain. This neutralizes the positive charge of lysine and weakens the interaction between histone and DNA, thereby opening the chromosome for transcription factor binding and initiation of transcription. Acetylation of histone H3 lysines at K14 and K9 by histone acetyltransferases may be associated with human transcriptional capacity. Lysine acetylation can directly or indirectly create binding sites for chromatin-modifying enzymes that regulate transcriptional activation. On the other hand, histone methylation of lysine 9 in histone H3 may be associated with heterochromatin or transcriptionally silent chromatin.

[0238] In some embodiments, the effector domain of the epigenetic editor described herein comprises a histone methyltransferase domain. The effector domain may comprise, for example, a DOT1L domain, a SET domain, an SUV39H1 domain, a G9a / EHMT2 protein domain, an EZH1 domain, an EZH2 domain, a SETDB1 domain, or any combination thereof. In a particular embodiment, the effector domain comprises a histone-lysine-N-methyltransferase SETDB1 domain.

[0239] In some embodiments, the effector domain comprises a histone deacetylase protein domain. In some embodiments, the effector domain comprises an HDAC family protein domain, such as HDAC1, HDAC3, HDAC5, HDAC7, or HDAC9. In a particular embodiment, the effector domain comprises a nucleosome remodeling and deacetylase complex (NURD) that removes acetyl groups from histones.

[0240] D. Other effect subdomains

[0241] In some embodiments, the effector domain contains a ternary motif comprising a protein (TRIM28, TIF1-β, or KAP1). In some embodiments, the effector domain contains one or more KAP1 proteins. The KAP1 protein in the epigenetic editor described herein can form a complex with one or more other effector domains of the epigenetic editor or with one or more proteins involved in gene expression regulation in the cellular environment. For example, KAP1 can be recruited via the KRAB domain of a transcriptional repressor. The KAP1 protein domain can interact with or recruit one or more protein complexes that reduce or silence gene expression. In some embodiments, KAP1 interacts with or recruits histone deacetylases, histone-lysine methyltransferases, chromatin remodeling proteins, and / or heterochromatin proteins. For example, the KAP1 protein domain can interact with or recruit components of the heterochromatin protein 1 (HP1), SETDB1, HDAC, and / or NuRD protein complexes. In some embodiments, the KAP1 protein domain interacts with or recruits ZFP90 proteins (e.g., isoform 2 of ZFP90) and / or FOXP3 proteins. An exemplary KAP1 amino acid sequence is shown in SEQ ID NO:629.

[0242] In some embodiments, the effector domain comprises a protein domain that interacts with or is recruited by one or more DNA epigenetic markers. For example, the effector domain may comprise a methyl CpG-binding protein 2 (MECP2) protein that interacts with methylated DNA nucleotides in the target gene (which may or may not be located on a CpG island in the target gene). The MECP2 protein domain in the epigenetic editor described herein can induce condensed chromatin structure, thereby reducing or silencing the expression of the target gene. In some embodiments, the MECP2 protein domain in the epigenetic editor described herein can interact with histone deacetylases (e.g., HDACs) to repress or silence the expression of the target gene. In some embodiments, the MECP2 protein domain in the epigenetic editor described herein can block the access of transcription factors or transcription activators to the target sequence, thereby repressing or silencing the expression of the target gene. An exemplary MECP2 amino acid sequence is shown in SEQ ID NO: 630.

[0243] Effector domains of the epigenetic editors described herein are also considered, such as chromatin shadow domains, ubiquitin-2-like Rad60 SUMO-like (Rad60-SLD / SUMO) domains, chromatin organization modifier domains (Chromo) domains, Yaf2 / RYBP C-terminal binding motif domains (YAF2_RYBP), CBX family C-terminal motif domains (CBX7_C), zinc finger C3HC4 type (ring finger) domains (ZF-C3HC4_2), cytochrome b5 domains (Cyt-b5), helical-loop-helical domains (HLH), helical-hairpin-helical motif domains (e.g., HHH_3), high-mobility group box domains (HMG-box), and basic leucine zipper domains (e.g.,bZIP_1 or bZIP_2), Myb_DNA binding domain, homology domain, MYM type zinc finger with FCS sequence domain (zf-FCS), interferon regulator 2 binding protein zinc finger domain (IRF-2BP1_2), SSX repressor domain (SSXRD), B-box type zinc finger domain (ZF-B_box), CXXC zinc finger domain (ZF-CXXC), chromosome condensation regulator 1 domain (RCC1), SRC homology 3 domain (SH3_9), sterile α motif domain (SAM_1), sterile α motif domain (SAM_2), sterile α motif / tip domain (SAM_PNT), degenerate / Tondu family domain (Vg_Tdu), LIM domain, RNA recognition motif domain (RRM_1), pairing amphiphilic helical domain (PAH), proteasome ATPase OB C-terminal domain (Prot_ATP_ID_OB), Neural Homeostasis 2 domain (NHR2), Cleavage Stimulant Subunit 2 Hinge domain (CSTF2_hinge), PPARγ N-terminal domain (PPARγ_N), CDC48 N-terminal domain (CDC48_2), WD40 repeat domain (WD40), Fip1 motif domain (Fip1), PDZ domain (PDZ_6), Von Willebrand factor C-type domain (VWC), NAB conserved region 1 domain (NCD1), S1 RNA-binding domain (S1), HNF3 C-terminal domain (HNF_C), Tudor domain (Tudor_2), histone-like transcription factor (CBF / NF-Y) and archaea histone domain (CBFD_NFYB_HMF), zinc finger protein domain (DUF3669), EGF-like domain (cEGF), GATA zinc finger domain (GATA), TEA / ATTS domain (TEA), phorbol ester / diglyceride binding domain (C1-1), multicomb-like MTF2 factor 2 domain (Mtf2_C), deactivation domain of FOXO protein family (FOXO-TAD), homeobox KN domain (Hom The protein contains the following domains: eobox_KN, BED zinc finger domain (ZF-BED), C3HC4-type cyclic zinc finger domain (ZF-C3HC4_4), RAD51 interacting motif domain (RAD51_interact), p55 binding region of the methyl-CpG-binding domain MBD (MBDa), Notch domain, Raf-like Ras binding domain (RBD), Spin / Ssty family domains (Spin-Ssty), PHD finger domain (PHD_3), low-density lipoprotein receptor class A domain (Ldl_recept_a), CS domain, DM DNA binding domain, and QLQ domain.

[0244] In some embodiments, the effector domain is a protein domain containing the YAF2_RYBP domain or a homologous domain or any combination thereof. In some embodiments, the homologous domain of the YAF2_RYBP domain is a PRD domain, an NKL domain, a HOXL domain, or a LIM domain. In a particular embodiment, the YAF2_RYBP domain may contain a 32-amino acid Yaf2 / RYBP C-terminal binding motif domain (32 aa RYBP).

[0245] In some implementations, the effector domain includes a protein domain selected from the group consisting of: the SUMO3 domain, the cromo domain from M-phase phosphorylated protein 8 (MPP8), the chromatin shadow domain from Chromobox 1 (CBX1), and the SAM_1 / SPM domain from Scm polycomb family homolog 1 (SCMH1).

[0246] In some implementations, the effector domain contains the HNF3 C-terminal domain (HNF_C). The HNF_C domain may be derived from FOXA1 or FOXA2. In some implementations, the HNF_C domain contains the EH1 (Engrailed homology 1) motif.

[0247] In some implementations, the effector domain may include an interferon regulator 2 binding protein zinc finger domain (IRF-2BP1_2), a Cyt-b5 domain from the DNA repair factor HERC2 E3 ligase, a variant SH3 domain (SH3_9) from bridging integrin 1 (BIN1), an HMG-box domain from the transcription factor TOX, or a ZF-C3HC4_2 ring finger domain from the multicomb component PCGF2, a staining domain-helicase-DNA binding protein 3 (CHD3) domain, or a ZNF783 domain.

[0248] IV. Epigenetic Editor

[0249] This article provides an epigenetic editor (i.e., an epigenetic editing system) that, for example, uses in any combination of one or more DNA-binding domains as described herein and one or more effector domains as described herein (e.g., epigenetic repressor domains) to direct epigenetic modifications to target sequences in genes of interest. The DNA-binding domain (working in conjunction with guide polynucleotides as described herein, wherein the DNA-binding domain is a polynucleotide-guided DNA-binding domain) directs the effector domain to epigenetically modify the target sequence, resulting in gene repression or silencing, which can be persistent and heritable across cell generations. In some aspects, the epigenetic editor described herein can reversibly or irreversibly repress or silence genes in cells.

[0250] In certain embodiments, the epigenetic editor described herein comprises one or more fusion proteins, each fusion protein comprising (1) a DNA-binding domain and (2) an effector domain. The effector domain may be on one or more fusion proteins included in the epigenetic editor. For example, a single fusion protein may comprise all effector domains having a DNA-binding domain. Alternatively, effector domains or subsets thereof may be on individual fusion proteins, each having a DNA-binding domain (which may be the same or different). The fusion proteins described herein may further comprise one or more adapters (e.g., peptide adapters), detectable tags, nuclear localization signals (NLS), or any combination thereof. As used herein, “fusion protein” means a chimeric protein in which two or more coding sequences (e.g., for the DNA-binding domain and / or effector domain) are directly or indirectly covalently or non-covalently linked.

[0251] In some embodiments, the epigenetic editor described herein comprises 2, 3, 4, 5, 6, 7, 8, 9, 10 or more effector (e.g., repressor / repressor) domains, which may be the same or different. In some embodiments, two or more of the effector domains function synergistically. Combinations of effector domains may include DNA methylation domains, histone deacetylation domains, histone methylation domains, and / or scaffold domains that recruit any of the above. For example, the epigenetic editor described herein may include one or more transcriptional repressor domains (e.g., KRAB domains such as KOX1, ZIM3, ZFP28, or ZN627 KRAB) combined with one or more DNA methylation domains (e.g., DNMT domains) and / or recruitment domains (e.g., DNMT3L domains). Such an epigenetic editor may include, for example, KRAB domains, DNMT3A domains, and DNMT3L domains. In some embodiments, the epigenetic editor further includes additional effector domains (e.g., KAP1, MECP2, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, RBBP4, RCOR1, or SCML2 domains). In some embodiments, the additional effector domains are CDYL2, TOX, TOX3, TOX4, or HP1a domains. For example, the epigenetic editor described herein may include CDYL2 and / or TOX domains in combination with a KRAB domain (e.g., the KOX1 KRAB domain).

[0252] A. Connector

[0253] As described herein, fusion proteins may contain one or more adapters that connect components to an epigenetic editor. These adapters may be peptide or non-peptide adapters.

[0254] In some embodiments, one or more adapters used in the epigenetic editor provided herein are peptide adapters, i.e., adapters containing a peptide moiety. Peptide adapters can be of any length suitable for the epigenetic editor fusion protein described herein. In some embodiments, the adapter may contain a peptide of 1 to 200 (e.g., 1 to 80) amino acids. In some embodiments, the length of the connector includes 1 to 5, 1 to 10, 1 to 20, 1 to 30, 1 to 40, 1 to 50, 1 to 60, 1 to 80, 1 to 100, 1 to 150, 1 to 200, 5 to 10, 5 to 20, 5 to 30, 5 to 40, 10 to 50, 10 to 60, 10 to 80, 10 to 100, 10 to 150, 10 to 200, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 200, 20 to 30, 20 to 40, 20 to 50, 20 to 60, 20 to 80, 20 to 100, 20 to 150, 20 to 200, 30 to 40, 30 to 50, 30 to 60, 30 to 80, 30 to 100, 30 to 150, 30 to 200, 40 to 50, 40 to 60, 40 to 80, 40 to 100, 40 to 150, 40 to 200, 50 to 60, 50 to 80, 50 to 100, 50 to 150, 50 to 200, 60 to 80, 60 to 100, 60 to 150, 60 to 200, 80 to 100, 80 to 150, 80 to 200, 100 to 150, 100 to 200, or 150 to 200 amino acids. Longer or shorter linkers were also considered. In some embodiments, the peptide linker is 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 25, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amino acids long. For example, the peptide linker can be 4, 5, 16, 20, 24, 27, 32, 40, 64, 92, or 104 amino acids long. The peptide linker can be flexible or rigid. In a particular embodiment, the peptide linker comprises the amino acid sequence of any one of SEQ ID NO: 631-637 and 664-665, or at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0255] In some embodiments, the peptide linker is an XTEN linker. This linker may contain a portion of an XTEN sequence (Schellenberger et al., Nat Biotechnol (2009) 27(1):1186-90), an XTEN sequence being a nonstructured hydrophilic polypeptide consisting only of residues G, S, P, T, E, and A. As used herein, the term “XTEN” refers to a recombinant peptide or polypeptide lacking hydrophobic amino acid residues. XTEN linkers are typically unstructured and contain a limited set of native amino acids. Fusion of an XTEN to a protein alters its hydrodynamic properties and reduces the clearance and degradation rate of the fusion protein. These XTEN fusion proteins are generated using recombinant technologies, require no chemical modification, and are degraded via natural pathways. The length of the XTEN linker can be, for example, 5, 10, 16, 20, 26, or 80 amino acids. In some embodiments, the length of the XTEN linker is 16 amino acids. In some embodiments, the length of the XTEN linker is 80 amino acids. In some embodiments, the XTEN linker can be XTEN10, XTEN16, XTEN20, or XTEN80. In some embodiments, the XTEN adapter may comprise the amino acid sequence of any one of SEQ ID NO: 638-643, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In a particular embodiment, the XTEN adapter comprises the amino acid sequence of SEQ ID NO: 638. In a particular embodiment, the XTEN adapter comprises the amino acid sequence of SEQ ID NO: 643.

[0256] In some embodiments, one or more linkers used in the epigenetic editor provided herein are non-peptide linkers. For example, a linker may be a carbon bond, a disulfide bond, or a carbon-heteroatom bond. In some embodiments, a linker is a carbon-nitrogen bond of an amide bond. In some embodiments, a linker is a cyclic or acyclic, substituted or unsubstituted, or branched or unbranched aliphatic or heteroaliphatic linker.

[0257] In some embodiments, one or more linkers used in the epigenetic editor provided herein are polymeric (e.g., polyethylene, polyethylene glycol, polyamide, polyester, etc.). Linkers may comprise, for example, monomers, dimers, or polymers of aminoalkyl acids; aminoalkyl acids (e.g., glycine, acetic acid, alanine, β-alanine, 3-aminopropionic acid, 4-aminobutyric acid, 5-valeric acid, etc.); monomers, dimers, or polymers of aminohexanoic acid (Ahx); or polyethylene glycol moiety (PEG); or aromatic or heteroaromatic moiety. In some embodiments, linkers may be based on carbocyclic moiety (e.g., cyclopentane or cyclohexane) or benzene ring. Linkers may include functionalized moieties to facilitate the attachment of nucleophilic groups (e.g., thiol, amino) of peptides to the linker. Any electrophilic group may be used as part of the linker. Exemplary electrophilic groups include, but are not limited to, activated esters, activated amides, alkyl halides, aryl halides, acyl halides, and isothiocyanates.

[0258] A variety of linker lengths and flexibility can be employed between any two components of the epigenetic editor (e.g., between an effector domain (e.g., a repressor domain) and a DNA-binding domain (e.g., a Cas9 domain), between a first effector domain and a second effector domain, etc.). Linkers can range from very flexible linkers (e.g., glycine / serine-rich linkers) to more rigid linkers to achieve the optimal length for effector domain activity for a specific application. In some embodiments, the more flexible linker is a glycine / serine-rich linker (GS-rich linker) in which more than 45% (e.g., more than 48%, 50%, 55%, 60%, 70%, 80%, or 90%) of the residues are glycine or serine residues. Non-limiting examples of GS-rich linkers are (GGGGS)n (SEQ ID NO: 664), (G)n, and W linkers (SEQ ID NO: 637). In some embodiments, the more rigid joint is of the form (EAAAK)n (SEQ ID NO: 665), (SGGS)n (SEQ ID NO: 631), and (XP)n. In the formulas for the flexible and rigid joints described above, n can be any integer between 1 and 30. In some embodiments, n is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15. In some embodiments, the joint comprises the (GGS)n motif, where n is 1, 3, or 7. In some embodiments, the joint comprises the (GGGGS)n motif, where n is 4 (SEQ ID NO: 636).

[0259] In some embodiments, the adapter in the epigenetic editor described herein includes a nuclear localization signal, for example, an amino acid sequence having any one of SEQ ID NO: 644-649. In some embodiments, the adapter in the epigenetic editor described herein includes an expression tag, such as a detectable tag, like green fluorescent protein.

[0260] B. Nuclear location signal

[0261] The fusion protein described herein may contain one or more nuclear localization signals, and in some embodiments, may contain two or more nuclear localization signals. For example, the fusion protein may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nuclear localization signals. As used herein, a “nuclear localization signal” (NLS) is an amino acid sequence that guides the protein to the cell nucleus. In some embodiments, the NLS may be an SV40 NLS (e.g., an amino acid sequence having SEQ ID NO: 644). The fusion protein may contain an NLS at its N-terminus, C-terminus, or both, and / or the NLS may be embedded in the middle of the fusion protein (e.g., at the N- or C-terminus of a DNA-binding domain or effector domain).

[0262] In some embodiments, the fusion protein may contain two NLSs. The fusion protein may contain two NLSs at its N-terminus or C-terminus. The fusion protein may contain one NLS located at its N-terminus and one NLS embedded in the middle of the fusion protein, or one NLS located at its C-terminus and one NLS embedded in the middle of the fusion protein. The fusion protein may contain two NLSs embedded in the middle of the fusion protein.

[0263] In some embodiments, the fusion protein may contain four NLSs. The fusion protein may contain at least two (e.g., two, three, or four) NLSs at its N-terminus or C-terminus. The fusion protein may contain at least one (e.g., one, two, three, or four) NLS embedded in the middle of the fusion protein. In a particular embodiment, the fusion protein may contain two NLSs at its N-terminus and two NLSs at its C-terminus.

[0264] The NLS described herein may be an endogenous NLS sequence. In some embodiments, the NLS described herein comprises the amino acid sequence of any one of SEQ ID NO: 644-649, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a selected sequence. In a particular embodiment, the NLS comprises the amino acid sequence of SEQ ID NO: 644. Other NLS are known in the art.

[0265] In some implementations, an epigenetic editor containing a fusion protein (which includes at least one NLS at the N-terminus and at least one NLS at the C-terminus) can increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000%, or more compared to an epigenetic editor having a corresponding fusion protein (which does not have at least one NLS at the N-terminus and at least one NLS at the C-terminus).

[0266] In some implementations, an epigenetic editor containing a fusion protein (which includes two NLS at the N-terminus and two NLS at the C-terminus) can increase the efficiency of the epigenetic editor by at least 5%, at least 10%, at least 15%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 200%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1,000%, at least 5,000%, at least 10,000%, at least 50,000%, at least 100,000%, or more compared to an epigenetic editor having a corresponding fusion protein (which does not have two NLS at the N-terminus and two NLS at the C-terminus).

[0267] C. Labels

[0268] The epigenetic editor provided herein may include one or more additional sequences (“tags”) for tracking, detecting, and locating the editor. In some embodiments, the epigenetic editor includes 1, 2, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more detectable tags. Each detectable tag may be the same or different.

[0269] For example, epigenetic editor fusion proteins may include cytoplasmic localization sequences, output sequences (such as nuclear output sequences), or other localization sequences, as well as sequence tags for dissolving, purifying, or detecting the fusion protein. Suitable protein tags provided herein include, but are not limited to, biotinylate carboxylase carrier protein (BCCP) tags, myc tags, calmodulin tags, FLAG tags, hemagglutinin (HA) tags, polyhistidine tags (also known as histidine tags or His tags), maltose-binding protein (MBP) tags, nus tags, glutathione S-transferase (GST) tags, green fluorescent protein (GFP) tags, thioredoxin tags, S tags, Softag (e.g., Softag 1 or Softag 3), strep tags, biotin ligase tags, FLASH tags, V5 tags, and SBP tags. Other suitable sequences will be apparent to those skilled in the art.

[0270] D. Fusion protein conformation

[0271] The fusion protein of the epigenetic editor described herein can have its components structured in different conformations. For example, the DNA-binding domain can be at the C-terminus, N-terminus, or between two or more epigenetic effector domains or additional domains. In some embodiments, the DNA-binding domain is at the C-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is at the N-terminus of the epigenetic editor. In some embodiments, the DNA-binding domain is linked to one or more nuclear localization signals. In some embodiments, the DNA-binding domain is flanked by epigenetic effector domains and / or additional domains. In some embodiments, wherein “DBD” indicates the DNA-binding domain and “ED” indicates the effector domain, the epigenetic editor comprises the conformation:

[0272] -N']-[ED1]-[DBD]-[ED2]-[C'

[0273] -N']-[ED1]-[DBD]-[ED2]-[ED3]-[C'

[0274] -N']-[ED1]-[ED2]-[DBD]-[ED3]-[C'

[0275] or

[0276] -N']-[ED1]-[ED2]-DBD]-[ED3]-[ED4]-[C'.

[0277] In some embodiments, the epigenetic editor comprises a DNA-binding domain (DBD), a DNA methyltransferase (DNMT) domain, and a transcriptional repression (“repression”) domain that represses or silences the expression of a target gene. The DBD, DNMT, and transcriptional repressor domains can be any of those described herein, in any combination. The DBD, DNMT, and repressor domains can be of any conformation, for example, any of the domains may be located at the N-terminus, C-terminus, or middle of the fusion protein. In some embodiments, the epigenetic editor comprises a fusion protein having the following conformations:

[0278] N']-[DNMT domain]-[DBD]-[repressor domain]-[C'

[0279] N']-[Repressor Domain]-[DBD]-[DNMT Domain]-[C'

[0280] N']-[DNMT domain]-[repressor domain]-[DBD]-[C'

[0281] or

[0282] N']-[Repressor domain]-[DNMT domain]-[DBD]-[C'].

[0283] In some embodiments, the linker “]-[” in any of the epigenetic editor structures is a linker, such as a peptide linker; a detectable tag; a peptide bond; a nuclear localization signal; and / or a promoter or regulatory sequence. In the epigenetic editor structure, multiple linker “]-[” can be the same or can each be a different linker, tag, NLS, or peptide bond. In some embodiments, the DNMT domain may comprise any one of the domains in Table 6 or any combination thereof or homologs thereof. In a particular embodiment, the DNMT domain comprises DNMT3A or a truncated version thereof, DNMT3L or a truncated version thereof, or both. In a particular embodiment, DBD is a non-catalytically active polynucleotide-guided DNA-binding domain (e.g., dCas9) or a ZFP domain. In some embodiments, the repressor domain comprises any one of the domains shown in Table 4 or 5, or any combination thereof or homologs thereof. For example, the repressor domain may be a KRAB domain. In some embodiments, the repressor domain is a ZFP28, ZN627, KAP1, MeCP2, HP1b, CBX8, CDYL2, TOX, Tox3, Tox4, EED, RBBP4, RCOR1, or SCML2 domain, or a fusion of two of these domains (e.g., a fusion of the N and C-terminal regions of ZIM3 and KOX1 KRAB). In specific embodiments, the repressor domain is a KRAB domain derived from ZFP28, ZN627, ZIM3, or KOX1.

[0284] In some implementations, the base editor comprises a configuration selected from the following:

[0285] N']-[DNMT3A-DNMT3L]-[DBD]-[Repressor]-[C'

[0286] N']-[Repressor]-[DBD]-[DNMT3A-DNMT3L]-[C'

[0287] N']-[Repressor]-[DBD]-[DNMT3A]-[C'

[0288] N']-[DNMT3A]-[DBD]-[Repressor]-[C'

[0289] N']-[Repressor]-[DBD]-[DNMT3A]-[DNMT3L]-[C'

[0290] N']-[DNMT3A]-[DNMT3L]-[DBD]-[Repressor]-[C'

[0291] N']-[DNMT3A]-[DBD]-[C'

[0292] N']-[DBD]-[DNMT3A]-[C'

[0293] N']-[DNMT3L]-[DBD]-[C'

[0294] N']-[DBD]-[DNMT3L]-[C'

[0295] Wherein [DNMT3A-DNMT3L] indicates that the DNMT3A and DNMT3L domains are directly fused via a peptide bond, and wherein the linker [-] is any of the adapter, detectable tag, affinity domain, peptide bond, nuclear localization signal, promoter, and / or regulatory sequence as described herein. The DBD, repressor, DNMT3A, and DNMT3L domains can be any of those described herein, in any combination. For example, the DNMT3A and DNMT3L domains can be those selected from Table 6. In a particular embodiment, the DBD is a CRISPR-associated protein domain (e.g., dCas9) or a ZFP domain; the repressor domain is a KRAB domain derived from KOX1, ZIM3, ZFP28, or ZN627; the DNMT3A domain is a human DNMT3A domain; and the DNMT3L domain is a human or mouse DNMT3L domain; any combination of these components is also contemplated in this disclosure.

[0296] In some implementations, the base editor comprises a configuration selected from the following:

[0297] N']-[DNMT3A]-[DBD]-[SETDB1]-[C'

[0298] N']-[DNMT3A]-[DNMT3L]-[DBD]-[SETDB1]-[C'

[0299] N']-[DNMT3A-DNMT3L]-[DBD]-[SETDB1]-[C'

[0300] N']-[SETDB1]-[DBD]-[DNMT3A]-[DNMT3L]-[C'

[0301] N']-[SETDB1]-[DBD]-[DNMT3A]-[C'

[0302] Wherein [DNMT3A-DNMT3L] indicates that the DNMT3A and DNMT3L domains are directly fused via a peptide bond, and wherein the linker [-] is any of the adapter, detectable tag, affinity domain, peptide bond, nuclear localization signal, promoter, and / or regulatory sequence as described herein. The DBD, SETDB1, DNMT3A, and DNMT3L domains can be any of those described herein, in any combination. In a particular embodiment, DBD is a CRISPR-associated protein domain (e.g., dCas9) or a ZFP domain; the SETDB1 domain is derived from human SETDB1, ZIM3, ZFP28, or ZN627; the DNMT3A domain is the human DNMT3A domain; and the DNMT3L domain is the human or mouse DNMT3L domain; any combination of these components is also contemplated in this disclosure.

[0303] The specific constructs considered in this article include:

[0304] DNMT3A-DNMT3L-XTEN80-NLS-dCas9-NLS-XTEN16-KOX1 KRAB (Configuration 1),

[0305] DNMT3A-DNMT3L-XTEN80-NLS-ZFP structural domain-NLS-XTEN16-KOX1 KRAB (configuration 2),

[0306] NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS (Configuration 3),

[0307] NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-KOX1 KRAB-NLS (configuration 4),

[0308] NLS-NLS-DNMT3A-DNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS-NLS (configuration 5), and

[0309] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-KOX1 KRAB-NLS-NLS (configuration 6).

[0310] DNMT3L and DNMT3A can be derived from human parent proteins, mouse parent proteins, or any combination thereof. In some embodiments, DNMT3L and DNMT3A are derived from mouse and human parent proteins (mDNMT3L and hDNMT3A, respectively). In some embodiments, both DNMT3L and DNMT3A are derived from human parent proteins (hDNMT3L and hDNMT3A). In some embodiments, dCas9 is dSpCas9. In some embodiments, KOX1 is human KOX1. Configurations 1-6 are also considered, wherein the KOX1 KRAB domain is replaced by a ZFP28, ZN627, or ZIM3 KRAB domain. In some embodiments, ZFP28, ZN627, and ZIM3 are human ZFP28, ZN627, and ZIM3, respectively. In specific embodiments, the fusion construct can have the following configurations:

[0311] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-KOX1 KRAB-NLS-NLS (Configuration 7),

[0312] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-KOX1 KRAB-NLS-NLS (configuration 8),

[0313] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZFP28 KRAB-NLS-NLS (Configuration 9),

[0314] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-ZFP28 KRAB-NLS-NLS (configuration 10),

[0315] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZN627 KRAB-NLS-NLS (Configuration 11),

[0316] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-ZN627 KRAB-NLS-NLS (configuration 12),

[0317] NLS-NLS-hDNMT3A-hDNMT3L-XTEN80-dCas9-XTEN16-ZIM3 KRAB-NLS-NLS (configuration 13), or

[0318] NLS-NLS-DNMT3A-DNMT3L-XTEN80-ZFP structural domain-XTEN16-ZIM3 KRAB-NLS-NLS (configuration 14).

[0319] In certain embodiments, the fusion construct described herein may comprise the sequence provided below (SEQ ID NO: 1495), or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In SEQ ID NO: 1495 below, the XTEN connector is... underlined The W connector is Bold, underlined and italics The NLS sequence is in bold, the DNMT3A sequence is in italics, and the DNMT3L sequence is in regular font. Underlined and italic The dCas9 domain is in bold and italics, and the KOX1 KRAB domain is... Underlined and bold :

[0320]

[0321] In a particular implementation, the fusion construct described herein may have configuration 2 and contain SEQ ID NO: 659, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In SEQ ID NO: 659 below, the XTEN connector is... underlined The W connector is Bold, underlined and italic of The NLS sequence is in bold and underlined, the DNMT3A sequence is in italics, and the DNMT3L sequence is... Underlined and italic The ZFP domain is in bold, and the KOX1 KRAB domain is... Underlined and bold The variable amino acid represented by Xs is the amino acid of the DNA recognition helix of the zinc finger, and the italicized XX can be TR, LR, or LK.

[0322]

[0323] In some embodiments, the six “XXXXXXX” regions in SEQ ID NO: 659 sequentially contain the F1–F6 amino acid sequences shown in Table 1, for any one of ZF001–ZF048. [Linker] indicates a linker sequence. In some embodiments, one or both linker sequences may be TGSQKP (SEQ ID NO: 651). In some embodiments, one or both linker sequences may be TGGGGSQKP (SEQ ID NO: 652). In some embodiments, one linker sequence may have the amino acid sequence of SEQ ID NO: 651, and the other linker sequence may have the amino acid sequence of SEQ ID NO: 652.

[0324] In certain implementations, the fusion construct described herein may comprise the sequence provided below (SEQ ID NO: 1496), or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical thereto. In SEQ ID NO: 1496 below, the XTEN connector is... underlined The W connector is Bold, underlined and italics The NLS sequence is in bold and underlined, the DNMT3A sequence is in italics, and the DNMT3L sequence is... Underlined and italic body The ZFP domain is in bold, and the KOX1 KRAB domain is... Underlined and bold The variable amino acid represented by Xs is the amino acid of the DNA recognition helix of the zinc finger, and the italicized XX can be TR, LR, or LK.

[0325]

[0326] In some embodiments, the six “XXXXXXX” regions in SEQ ID NO: 1496 sequentially comprise the F1–F6 amino acid sequences shown in Table 1, for any one of ZF001–ZF048. [Linker] indicates a linker sequence. In some embodiments, one or both linker sequences may be TGSQKP (SEQ ID NO: 651). In some embodiments, one or both linker sequences may be TGGGGSQKP (SEQ ID NO: 652). In some embodiments, one linker sequence may have the amino acid sequence of SEQ ID NO: 651, and the other linker sequence may have the amino acid sequence of SEQ ID NO: 652.

[0327] In some embodiments, the fusion protein may further include a Dnmt3A ADD domain, for example, downstream of the Dnmt3A domain sequence disclosed in SEQ ID NO: 658, 659, 1495, or 1496 above. In some embodiments, the ADD sequence is located between the Dnmt3A and Dnmt3L sequences of the fusion protein. In some embodiments, the ADD sequence is located at the C-terminus of the Dnmt3A domain. In some embodiments, the Dnmt3A and ADD sequences are separated by a adapter (e.g., the adapter disclosed herein). In some embodiments, the ADD and Dnmt3L sequences are separated by a adapter (e.g., the adapter disclosed herein). In some embodiments, the ADD domain includes the sequence:

[0328] MAAIPALDPEAEPSMDVILVGSSELSSSVSPGTGRDLIAYEVKANQRNIEDICICCGSLQVHTQHPLFEGGICAPCKDKFLDALFLYDDDGYQSYCSICCSGETLLICGNPDCTRCYCFECVDSLVGPGTSGKVHAMSNWVCYLCLPSSRSGLLQRRRKWRSQLKAFYDRESENPLEMFETVPVWRRQPVRVL SLFEDIKKELTSLGFLESGSDPGQLKHVVDVTDTVRKDVEEWGPFDLVYGATPPLGHTCDRPPSWYLFQFHRLLQYARPKPGSPRPFFWMFVDNLVLNKEDLDVASRFLEMEPVTIPDVHGGSLQNAVRVWSNIPAIRSRHWALVSEEELSLLAQNKQSSKLAAKWPTKLVKNCFLPLREYFKYFSTELTSSL (SEQ ID NO: 1497).

[0329] In a particular implementation, the fusion construct described herein may have configuration 7 and contain SEQ ID NO: 660, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0330] In a particular implementation, the fusion construct described herein may have configuration 9 and contain SEQ ID NO: 661, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0331] In a particular implementation, the fusion construct described herein may have configuration 11 and contain SEQ ID NO: 662, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0332] In a particular implementation, the fusion construct described herein may have configuration 13 and contain SEQ ID NO: 663, or a sequence that is at least 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to it.

[0333] In some embodiments, the fusion constructs described herein (e.g., fusion constructs of any one of configurations 1-14) are contained within an expression construct comprising a WPRE sequence, a polyadenylation site, or both. In some embodiments, the WPRE sequence is located in the 3' untranslated region. In some embodiments, the WPRE sequence is upstream of the polyadenylation site. In a particular embodiment, the expression construct comprises the fusion construct (e.g., any one of configurations 1-14) and the WPRE sequence in the 3' untranslated region upstream of the polyadenylation site.

[0334] In some implementations, the fusion constructs described herein may have the sequence of any one of the fusion proteins 1-12 as shown in Example 12.

[0335] Various fusion proteins can be used to activate or repress one or more target genes. For example, an epigenetic editor fusion protein containing a DNA-binding domain (e.g., a dCas9 domain) and an effector domain can be co-delivered with two or more guide polynucleotides (e.g., gRNA), each targeting a different target DNA sequence. The target sites of the two DNA-binding domains can be the same or adjacent to each other, or spaced apart by, for example, about 100 base pairs, about 200 base pairs, about 300 base pairs, about 400 base pairs, about 500 base pairs, or about 600 or more base pairs. Additionally, when targeting double-stranded DNA such as endogenous gene loci, guide polynucleotides can target the same or different strands (one or more targeting the positive strand and / or one or more targeting the negative strand).

[0336] V. target sequence

[0337] The epigenetic editor described herein can be directed to target sequences in PCSK9 to achieve epigenetic modifications of the PCSK9 gene. As used herein, a “target sequence,” “target site,” or “target region” is a nucleic acid sequence present in the gene of interest; in some cases, the target sequence may be external to but near the gene of interest, wherein methylation of the target sequence or binding of a repressor inhibits the expression of the gene. In some embodiments, the target sequence may be a demethylated or hypermethylated nucleic acid sequence.

[0338] The target sequence can be in any part of the target gene. In some embodiments, the target sequence is part of or near a non-coding sequence of the gene. In some embodiments, the target sequence is part of an exon of the gene. In some embodiments, the target sequence is part of or near a transcriptional regulatory sequence of the gene (such as a promoter or enhancer). In some embodiments, the target sequence is adjacent to, overlaps with, or covers a CpG island. In some embodiments, the target sequence is located within approximately 3000, 2900, 2800, 2700, 2600, 2500, 2400, 2300, 2200, 2100, 2000, 1900, 1800, 1700, 1600, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, 600, 500, 400, 300, 200, or 100 base pairs (bp) flanking the PCSK9 TSS. In some embodiments, the target sequence is within 500 bp flanking the PCSK9 TSS. In some embodiments, the target sequence is within 1000 bp flanking the PCSK9 TSS.

[0339] In some embodiments, the target sequence may hybridize with a guide polynucleotide sequence (e.g., gRNA) complexed with a fusion protein comprising a polynucleotide guide DNA-binding domain (e.g., a CRISPR protein such as dCas9) and an effector domain. The guide polynucleotide sequence may be programmed to be complementary to the target sequence or identical to the opposite strand of the target sequence. In some embodiments, the guide polynucleotide comprises a spacer sequence that is approximately 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the original spacer sequence in the target sequence. In a particular embodiment, the guide polynucleotide comprises a spacer sequence that is 100% identical to the original spacer sequence in the target sequence.

[0340] In some implementations, when the DNA-binding domain of the epigenetic editor described herein is a zinc finger array, the target sequence can be recognized by the zinc finger array.

[0341] In some implementations, when the DNA-binding domain of the epigenetic editor described herein is a TALE, the target sequence can be recognized by the TALE.

[0342] The target sequences described herein can be specific to a single copy of a target gene or to an allele of a target gene. Therefore, epigenetic modifications and their expression regulation can be specific to a single copy of a target gene or to an allele. For example, an epigenetic editor can suppress the expression of a specific copy carrying a target sequence recognized by a DNA-binding domain (e.g., a copy associated with a disease or condition, or a copy carrying a mutation associated with a disease or condition).

[0343] In some implementations, the target PCSK9 genomic region may be located within the sequence shown below (chr1:55038548-55040548), with or without a terminal A:

[0344]

[0345] In some implementations, the target sequence may be GRCh38 Chr1:55039228-55040296, as shown below:

[0346]

[0347] VI. Epigenetic modifications

[0348] The epigenetic editors described herein can perform sequence-specific epigenetic modifications (e.g., chemical modifications) on target genes carrying target sequences. This epigenetic regulation may be safer and more easily reversible than regulation induced by gene editing (e.g., using the generation of DNA double-strand breaks). In some embodiments, epigenetic regulation can reduce or silence the target gene. In some embodiments, the modification is performed at a specific site on the target sequence. In some embodiments, the modification is performed at a specific allele of the target gene. Therefore, epigenetic modification may result in the regulation (e.g., reduction) of the expression of one copy of the target gene containing the specific allele, while the expression of the other copy of the target gene remains unaffected. In some embodiments, the specific allele is associated with a disease, condition, or symptom.

[0349] In some embodiments, epigenetic modification reduces or eliminates transcription of a target gene carrying a target sequence. In some embodiments, epigenetic modification reduces or eliminates transcription of a copy of a target gene carrying a specific allele recognized by an epigenetic editor. In some embodiments, the epigenetic editor reduces the level of a protein encoded by the target gene or eliminates its expression. In some embodiments, the epigenetic editor reduces the level of a protein encoded by a copy of a target gene carrying a specific allele recognized by an epigenetic editor or eliminates its expression. The target PCSK9 gene can be epigenetically modified in vitro, ex vivo, or in vivo.

[0350] The effector domains of the epigenetic editor described herein can alter (e.g., deposit or remove) chemical modifications at nucleotide sites of the target gene or at histone sites associated with the target gene. Chemical modifications can be altered at a single nucleotide or a single histone, or at 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, or more nucleotides.

[0351] In some embodiments, the effector domain of the epigenetic editor described herein can alter CpG dinucleotides within the target gene. In some embodiments, all CpG dinucleotides within 2000, 1500, 1000, 500, or 200 bp flanking the target sequence (e.g., at the alteration sites described herein) are altered according to the modification type described herein, compared to the gene's original state or in a comparable cell that has not been exposed to the epigenetic editor. In some implementations, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 550, 600, 650, 700 or more CpG dinucleotides are altered compared to the gene's original state or in comparable cells that have not been exposed to the epigenetic editor. In some embodiments, at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the CpG dinucleotides are altered compared to the gene's original state or in comparable cells that have not been exposed to the epigenetic editor. In some embodiments, a single CpG dinucleotide is altered compared to the gene's original state or in comparable cells that have not been exposed to the epigenetic editor.

[0352] The effector domain of the epigenetic editor described herein can alter the histone modification state of histones associated with or bound to a target gene. For example, the effector domain can deposit modifications on one or more lysine residues of the histone tail of a histone associated with a target gene. In some embodiments, the effector domain can cause deacetylation of one or more histone tails of a histone associated with a target gene, thereby reducing or silencing the expression of the target gene. In some embodiments, the histone modification state is a methylated state. For example, the effector domain can cause methylation of H3K9, H3K27, or H4K20 at one or more histone tails associated with a target gene (e.g., one or more of H3K9me2, H3K9me3, H3K27me2, H3K27me3, and H4K20me3 methylation), thereby reducing or silencing the expression of the target gene.

[0353] In some embodiments, compared to chromosomes in their original state or in comparable cells not exposed to the epigenetic editor, all histone tails of histones flanking the target sequence within 2000, 1500, 1000, 500, or 200 bp are altered according to the modification type described herein. In some embodiments, compared to chromosomes in their original state or in comparable cells not exposed to the epigenetic editor, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, or more histone tails binding to histones are altered. In some implementations, at least 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% of the histone tails binding histones are altered compared to the original state of the chromosome or the chromosome in comparable cells that have not been exposed to the epigenetic editor. For example, a single histone tail binding histones may be altered compared to the original state of the chromosome or the chromosome in comparable cells that have not been exposed to the epigenetic editor. As another example, a single histone octamer binding histone may be altered compared to the original state of the chromosome or the chromosome in comparable cells that have not been exposed to the epigenetic editor.

[0354] Chemical modifications deposited at nucleotide or histone residues in the target gene DNA can be located at or immediately adjacent to the target sequence in the target gene. In some embodiments, the effector domain of the epigenetic editor described herein alters the chemical modification state of the histone tail, which binds to nucleotides of 100-200, 200-300, 300-400, 400-55, 500-600, 600-700, or 700-800 nucleotides of the target sequence in the target gene, or to nucleotides of 100-200, 200-300, 300-400, 400-55, 500-600, 600-700, or 700-800 nucleotides of the target sequence in the target gene. In some implementations, the effector domain alters the nucleoside within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 nucleotides flanking the target sequence. The chemical modification state of a histone tail that binds to an acid or nucleotide within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 nucleotides flanking a target sequence. As used herein, "flanking" refers to nucleotide positions at the 5' to 5' end and 3' to 3' end of a specific sequence (e.g., the target sequence).

[0355] In some implementations, the effector domain mediates or induces chemical modification changes in nucleotides or histone tails flanking the target sequence. This modification can begin near the target sequence and subsequently extend to one or more nucleotides in the target gene flanking the target sequence. For example, the effector domain can initiate chemical modifications to one or more nucleotides within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nucleotides flanking the target sequence, or to one or more histone residues within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, or 500 nucleotides flanking the target sequence. The modification state is altered, and the chemical modification state can extend to one or more nucleotides at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 2500, or 3000 nucleotides upstream or downstream of the target sequence in the target gene. In some embodiments, the chemical modification can begin at fewer than 2, 3, 5, 10, 20, 30, 40, 50, or 100 nucleotides in the target gene and extend to at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 2000 nucleotides in the target gene. In some embodiments, the chemical modification extends to nucleotides throughout the entire target gene. Other proteins or transcription factors, such as transcriptional repressors, methyltransferases, or transcriptional regulatory scaffold proteins, can participate in the expansion of chemical modifications. Alternatively, they may be involved in epigenetic editors on their own.

[0356] In some embodiments, compared to control cells, control tissues, or control subjects (e.g., in the absence of an epigenetic editor), the epigenetic editor described herein reduces the expression of a target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more, as measured by transcription of the target gene in cells, tissues, or subjects. In some embodiments, compared to control cells, control tissues, or control subjects, the epigenetic editor described herein reduces the expression of a copy of a target gene by at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 99%, or more, as measured by transcription of a copy of the target gene in cells, tissues, or subjects. In some embodiments, the copy of the target gene carries a specific sequence or allele recognized by the epigenetic editor. In certain implementations, the epigenetic modification copies encode functional proteins, and therefore the epigenetic editors disclosed herein can reduce or eliminate protein expression and / or function. For example, compared to control cells, control tissues, or control subjects, the epigenetic editors described herein can reduce the expression and / or function of proteins encoded by target genes in cells, tissues, or subjects by at least 3-fold, at least 4-fold, at least 5-fold, at least 6-fold, at least 7-fold, at least 8-fold, at least 9-fold, at least 10-fold, at least 11-fold, at least 12-fold, at least 13-fold, at least 14-fold, at least 15-fold, at least 20-fold, at least 25-fold, at least 30-fold, at least 35-fold, at least 40-fold, at least 45-fold, at least 50-fold, at least 60-fold, at least 70-fold, at least 80-fold, at least 90-fold, or at least 100-fold.

[0357] Regulation of target gene expression can be determined by identifying any parameters that are indirectly or directly affected by target gene expression. These parameters include, for example, changes in RNA or protein levels; changes in protein activity; changes in product levels; changes in downstream gene expression; changes in the transcription or activity of reporter genes such as luciferase, CAT, β-galactosidase, or GFP; changes in signal transduction; changes in phosphorylation and dephosphorylation; changes in receptor-ligand interactions; and changes in second messengers such as cGMP, cAMP, IP3, and Ca2+. +Changes in concentration; changes in cell growth; changes in angiogenesis; and / or any functional effects of gene expression. Measurements can be performed in vitro, in vivo, and / or ex vivo, and can be performed using conventional methods, such as measurements of RNA or protein levels, measurements of RNA stability, and / or identification of downstream or reporter gene expression. Readouts can be performed, for example, by chemiluminescence, fluorescence, colorimetric reactions, antibody binding, inducible markers, ligand binding assays, changes in intracellular second messengers such as cGMP and inositol trisphosphate (IP3), changes in intracellular calcium levels, and cytokine release.

[0358] Methods for determining gene expression levels (e.g., targets of epigenetic editors) may include, for example, determining gene transcript levels via reverse transcription PCR, quantitative RT-PCR, droplet digital PCR (ddPCR), Northern blotting, RNA sequencing, DNA sequencing (e.g., sequencing of complementary deoxyribonucleic acid (cDNA) obtained from RNA); next-generation sequencing, nanopore sequencing, pyrosequencing, or NanoString sequencing. Protein levels from gene expression may be determined, for example, by Western blotting, enzyme-linked immunosorbent assay (ELISA), mass spectrometry, immunohistochemistry, or flow cytometry. Gene expression product levels may be normalized to internal standards, such as total messenger RNA (mRNA) or the expression levels of specific genes (e.g., housekeeping genes).

[0359] In some implementations, a reporter system can be used to examine the effects of an epigenetic editor in regulating the expression of a target gene. For example, an epigenetic editor can be designed to target a reporter gene encoding a reporter protein, such as a fluorescent protein. The expression of the reporter gene in such a model system can be monitored, for example, by flow cytometry, fluorescence-activated cell sorting (FACS), or fluorescence microscopy. In some implementations, cell populations can be transfected with a vector carrying the reporter gene. Vectors can be constructed such that the reporter gene is expressed when cells are transfected with the vector. Suitable reporter genes include genes encoding fluorescent proteins, such as green, yellow, cherry, cyan, or orange fluorescent proteins. Cell populations carrying a reporter system can be transfected with the DNA, mRNA, or vector of an epigenetic editor encoding a targeted reporter gene.

[0360] VII. Pharmaceutical Composition

[0361] In one aspect, this disclosure provides pharmaceutical compositions comprising one or more epigenetic editors or components thereof (e.g., fusion proteins and / or guide polynucleotides) as an active ingredient (or as the sole active ingredient), or nucleic acid molecules encoding said epigenetic editors or components thereof. For example, the pharmaceutical composition may comprise a nucleic acid molecule encoding a fusion protein (and guide polynucleotides, where applicable) of an epigenetic editor described herein. In some embodiments, the pharmaceutical composition comprises both a fusion protein and a guide polynucleotide. The pharmaceutical composition may also comprise cells that have undergone epigenetic modifications mediated or induced by the epigenetic editors provided herein.

[0362] Generally, the epigenetic editor or components thereof described herein, or the nucleic acid molecules encoding said epigenetic editor or components thereof, are suitable for administration as a formulation in combination with one or more pharmaceutically acceptable excipients, for example, as described below.

[0363] The term "excipient" is used herein to describe any component other than the compounds disclosed herein. The selection of excipients will depend largely on factors such as the specific administration method, the effect of the excipient on solubility and stability, and the nature of the dosage form. As used herein, "pharmaceuticalally acceptable excipients" include any and all physiologically compatible solvents, dispersion media, coatings, antimicrobial and antifungal agents, isotonic agents, and absorption retardants. Some examples of pharmaceutically acceptable excipients are water, saline, phosphate-buffered saline, dextran, glycerol, ethanol, and combinations thereof. In many cases, isotonic agents, such as sugars, polyols such as mannitol, sorbitol, or sodium chloride, are preferably included in the composition. Other examples of pharmaceutically acceptable substances are wetting agents or small amounts of excipients, such as wetting agents or emulsifiers, preservatives, or buffers, which enhance the shelf life or efficacy of antibodies.

[0364] Pharmaceutical compositions suitable for parenteral administration typically contain an active ingredient in combination with a pharmaceutically acceptable carrier (e.g., sterile water or sterile isotonic saline). Such formulations may be prepared, packaged, or marketed in a form suitable for bolus or continuous administration.

[0365] VIII. Delivery method

[0366] In some embodiments, an epigenetic editor or a component thereof is introduced into target cells in the form of a nucleic acid molecule encoding the epigenetic editor or a component thereof; therefore, the pharmaceutical compositions herein comprise nucleic acid molecules. Such nucleic acid molecules may be, for example, DNA, RNA, or mRNA, and / or modified nucleic acid sequences (e.g., having chemical modifications, a 5' cap, or one or more 3' modifications). In some embodiments, the nucleic acid molecules may be delivered as naked DNA or RNA, for example, by transfection or electroporation, or may be conjugated to molecules that promote uptake by target cells (e.g., N-acetylgalactosamine). In some embodiments, the nucleic acid molecules may be in a nucleic acid expression vector, which may include expression control sequences such as promoters, enhancers, transcription signal sequences, transcription termination sequences, introns, polyadenylation signals, Kozak concordant sequences, internal ribosome entry sites (IRES), etc. Such expression control sequences are well known in the art. The vector may also contain a sequence encoding a signal peptide (e.g., for nuclear, nucleolar, or mitochondrial localization) associated with a protein-encoding sequence (e.g., insertion or fusion).

[0367] Examples of vectors include, but are not limited to, plasmid vectors; viral vectors based on vaccinia virus, poliovirus, adenovirus, adeno-associated virus, SV40, herpes simplex virus, human immunodeficiency virus, retroviruses (e.g., murine leukemia virus or spleen necrosis virus, vectors derived from retroviruses such as Rous sarcoma virus, Harvey sarcoma virus, avian leukosis virus, lentivirus, human immunodeficiency virus, myeloproliferative sarcoma virus, and mammary tumor virus); and other recombinant vectors. In some embodiments, the vector is a plasmid or a viral vector. Viral particles or virus-like particles (VLPs) can also be used to deliver nucleic acid molecules encoding epigenetic editors or components thereof as described herein. For example, “empty” viral particles can be assembled to contain any suitable cargo. Viral vectors and viral particles can also be engineered to incorporate targeting ligands to alter target tissue specificity.

[0368] In some embodiments, the epigenetic editor or its components, as described herein, are encoded by nucleic acid sequences present in one or more viral vectors or by suitable capsid proteins of any viral vector. Examples of viral vectors include adeno-associated virus vectors (e.g., derived from AAV3, AAV3b, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAVrh8, AAV10 and / or variants thereof); retroviral vectors (e.g., Moloney murine leukemia virus, MML-V); adenoviral vectors (e.g., AD100); lentiviral vectors (e.g., HIV and FIV-based vectors); and herpesvirus vectors (e.g., HSV-2).

[0369] In some implementations, delivery involves an adeno-associated virus (AAV) vector. AAV vector delivery can be particularly useful when the DNA-binding domain of the epigenetic editor fusion protein is a zinc finger array. Without being bound by any theory, the smaller size of the zinc finger array compared to larger DNA-binding domains such as Cas protein domains allows such fusion proteins to be conveniently packaged into viral vectors such as AAV vectors.

[0370] Any AAV serotype, such as human AAV serotype, can be used in the AAV vectors described herein, including but not limited to AAV serotype 1 (AAV1), AAV serotype 2 (AAV2), AAV serotype 3 (AAV3), AAV serotype 4 (AAV4), AAV serotype 5 (AAV5), AAV serotype 6 (AAV6), AAV serotype 7 (AAV7), AAV serotype 8 (AAV8), AAV serotype 9 (AAV9), AAV serotype 10 (AAV10), and AAV serotype 11 (AAV11) and their variants. In some embodiments, the AAV variant has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity with wild-type AAV. In some embodiments, the AAV variant may be engineered such that its capsid protein has reduced immunogenicity or enhanced transduction capacity in humans. In some cases, one or more regions of at least two different AAV serotypes are rearranged and reassembled to generate chimeric variants. For example, a chimeric AAV may contain an inverted terminal repeat (ITR) sequence that is heterologous to the serotype of the capsid. The resulting chimeric AAV may have different antigenic reactivity or recognition compared to its parental serotype. In some embodiments, the chimeric variant of AAV includes amino acid sequences from 2, 3, 4, 5 or more different AAV serotypes.

[0371] Non-viral systems have also been considered for delivery as described herein. Non-viral systems include, but are not limited to, nucleic acid transfection methods, including electroporation, acoustic transfection, calcium phosphate transfection, microinjection, DNA gene gun transfection, heat shock transfection, compressed DNA-mediated transfection, liposome transfection, cationic reagent-mediated transfection, and transfection using liposomes, immunoliposomes, exosomes, or cationic surface amphiphilic molecules (CFAs). In some embodiments, one or more mRNAs encoding the epigenetic editor fusion protein as described herein may be co-electroplated with one or more guide polynucleotides (e.g., gRNAs) as described herein.

[0372] It can target any type of cell to deliver an epigenetic editor or components thereof as described herein. For example, the cell can be eukaryotic or prokaryotic. In some embodiments, the cell is a mammalian (e.g., human) cell. Human cells may include, for example, hepatocytes, bile duct epithelial cells (cholangiocarcinoma cells), astrocytes, Kupffer cells, and hepatic sinusoidal endothelial cells.

[0373] In some embodiments, for example, the epigenetic editor or a component thereof described herein is delivered to host cells via a transient expression vector for transient expression. Transient expression of the epigenetic editor or a component thereof can lead to prolonged or permanent epigenetic modification of the target gene. For example, after the introduction of the epigenetic editor into the host cell, the epigenetic modification can be stable for at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks or longer; or for 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months or longer. The epigenetic modification can be maintained after one or more mitotic and / or meiotic events in the host cell. In certain embodiments, the epigenetic modification is maintained across generations in offspring generated or derived from the host cell.

[0374] IX. Therapeutic uses of epigenetic editors

[0375] This disclosure also provides methods for treating or preventing a condition in a subject, including administering an epigenetic editor or pharmaceutical composition as described herein to the subject. The epigenetic editor can perform epigenetic modifications on target polynucleotide sequences in target genes associated with a subject's disease, condition, or symptom, thereby regulating the expression of the target gene to treat or prevent the disease, condition, or symptom. In some embodiments, the epigenetic editor reduces the expression of the target gene to a level sufficient to achieve a desired effect, such as a treatment-related effect, like prevention or treatment of a disease, condition, or symptom.

[0376] In some implementations, a system for regulating (e.g., inhibiting) the expression of PCSK9 is administered to a subject, wherein the system comprises (1) a fusion protein of an epigenetic editor as described herein and a guide polynucleotide in the relevant case, or (2) a nucleic acid molecule encoding the fusion protein and the guide polynucleotide in the relevant case.

[0377] "Treatment" refers to a method of reducing or eliminating at least one of a biological symptom and / or its accompanying symptoms. As used herein, "reducing" a disease, symptom, or condition means reducing the severity and / or frequency of the symptoms of the disease, symptom, or condition. Furthermore, "treatment" as used herein includes references to curative, palliative, and preventative treatments. In some implementations, symptom reduction may involve reducing symptoms by at least 3%, 5%, 10%, 20%, 40%, 50%, 60%, 80%, 90%, 95%, 98%, 99%, 99.5%, 99.9%, or 100% compared to an equivalent untreated control, as measured by any standard technique.

[0378] In some implementations, the subjects may be mammals, such as humans. In some implementations, the subjects may be non-human primates, such as chimpanzees, macaques, monkeys, or rhesus monkeys, as well as other ape and monkey species.

[0379] In some implementations, the patient has a condition selected from hypercholesterolemia (e.g., familial hypercholesterolemia, such as heterozygous familial hypercholesterolemia (HeFH) or homozygous familial hypercholesterolemia (HoFH), or established atherosclerotic cardiovascular disease (ASCVD)) or renal insufficiency (RI).

[0380] In some embodiments, the patient to be treated with the epigenetic editor of this disclosure has received prior treatment for the condition to be treated (e.g., hypercholesterolemia (such as HeFH, HoFH, HF, or identified ASCVD) or RI). In other embodiments, the patient has not received such prior treatment. In some embodiments, the patient's prior treatment for the condition (e.g., previous treatment for hypercholesterolemia) has failed.

[0381] The epigenetic editor of this disclosure can be administered to patients suffering from the conditions described herein at a therapeutically effective amount. As used herein, a “therapeutically effective amount” means the amount of therapeutic agent administered that provides some relief from one or more symptoms of the treated condition and / or results in a clinical endpoint desired by a healthcare professional. The effective amount of therapy can be measured by its ability to stabilize disease progression and / or improve symptoms in a patient, and preferably by its ability to reverse disease progression. The ability of the epigenetic editor of this disclosure to reduce or silence PCSK9 expression can be assessed by in vitro assays, such as those described herein, and in suitable animal models that predict efficacy in humans. Appropriate dosing regimens are selected to provide the optimal therapeutic response in each specific situation, e.g., as a single bolus or as a continuous infusion, and with possible adjustments to the dosage as indicated by the emergency status of each case.

[0382] The epigenetic editor of this disclosure can be administered without additional therapeutic treatment, i.e., as a single therapy (monotherapy). Alternatively, treatment with the epigenetic editor of this disclosure may include at least one additional therapeutic treatment (combination therapy). In some embodiments, the additional therapeutic agent is known in the art for treating hypercholesterolemia or RI. The therapeutic agents include, but are not limited to, statins, fibrates, HMG-CoA reductase inhibitors, niacin, bile acid modulators or chelators, cholesterol absorption inhibitors or modulators, CETP inhibitors, MTTP inhibitors, and PPAR agonists.

[0383] The epigenetic editor or a component thereof (or a nucleic acid molecule encoding an epigenetic editor or a component thereof) disclosed herein may be administered by any method acceptable in the art, such as subcutaneously, intradermally, intratumorally, intralymphaticly, intramuscularly, intravenously, intralymphaticly, or intraperitoneally. In a particular embodiment, the pharmaceutical composition of the present disclosure is administered intravenously to a subject.

[0384] X. definition

[0385] As used herein, the term "nucleic acid" refers to any oligonucleotide or polynucleotide containing nucleotides in single-stranded or double-stranded form (e.g., deoxyribonucleotides or ribonucleotides), and includes DNA and RNA. A "nucleotide" contains a sugar, deoxyribose (DNA) or ribose (RNA), a base, and a phosphate group linked together by the phosphate group. A "base" includes purines and pyrimidines, including natural compounds such as adenine, thymine, guanine, cytosine, uracil, inosine, and natural analogs; and synthetic derivatives of purines and pyrimidines, including but not limited to modified forms with novel reactive groups such as amines, alcohols, thiols, carboxylates, haloalkanes, etc. Nucleic acids may contain known nucleotide analogs and / or modified backbone residues or bonds, which may be synthetic, naturally occurring, or unnatural. Such nucleotide analogs, modified residues, and modified bonds are well known in the art and, in the presence of nucleases, can provide nucleic acid molecules with enhanced cellular uptake, reduced immunogenicity, and / or increased stability.

[0386] As used herein, an “isolated” or “purified” nucleic acid molecule is a nucleic acid molecule that exists outside its natural environment. For example, an “isolated” or “purified” nucleic acid molecule (1) has been isolated from the nucleic acid of its source genomic DNA or cellular RNA; and / or (2) does not exist in nature. In some embodiments, an “isolated” or “purified” nucleic acid molecule is a recombinant nucleic acid molecule.

[0387] It should be understood that, in addition to the specific proteins and nucleic acid molecules mentioned herein, this disclosure also considers the use of their variants, derivatives, homologs, and fragments. Any variant of a given sequence may have a specific sequence (whether amino acid or nucleic acid residues) modified in a manner that substantially preserves at least one of the intrinsic functions of the polypeptide or polynucleotide in question. Variant sequences can be obtained by adding, deleting, substituting, modifying, replacing, and / or mutating at least one residue present in a naturally occurring sequence (in some embodiments, no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or 20 residues). For the specific proteins described herein (e.g., the KRAB, dCas9, DNMT3A, and DNMT3L proteins described herein), this disclosure also considers any naturally occurring form of the protein, or variants or homologs that retain at least one of their intrinsic functions (e.g., at least 50%, 60%, 70%, 80%, 90%, 85%, 96%, 97%, 98%, or 99% of their function compared to the specific protein described herein).

[0388] This document provides some exemplary fusion proteins covered by this disclosure. Those skilled in the art will understand that these exemplary proteins are non-limiting examples, and additional proteins are within the scope of this disclosure. For example, in the provision of exemplary fusion proteins containing specific domains (e.g., mammalian DNMT3A, DNMT3L, and / or KRAB domains, such as human or mouse DNMT3A, DNMT3L, and / or KRAB domains), in some embodiments, those skilled in the art will be able to determine that this disclosure also covers fusion proteins having the same conformation, but in which one or more mammalian domains are replaced by homologous domains from another mammal, for example, one or more mouse domains are replaced by one or more human domains. For example, in the provision of exemplary fusion proteins containing a mouse DNMT3L domain, fusion proteins having the same conformation, but in which the mouse DNMT3L is replaced by a human DNMT3L domain, are also covered.

[0389] As used herein, homologs of any polypeptide or nucleic acid sequence considered herein include sequences that have some degree of homology with wild-type amino acid and nucleic acid sequences. Homologous sequences may include sequences that are at least 50%, 55%, 65%, 75%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to the subject sequence, such as amino acid sequences. In the context of amino acid or nucleotide sequences, the term "percentage identical" refers to the percentage of identical residues in the two sequences at maximum correspondence. In some embodiments, the length of the reference sequence for comparison purposes is at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, 90%, or 100%) of the reference sequence. Sequence identity can be measured using sequence analysis software (e.g., the Sequence Analysis Software Package of the Genetics Computer Group, University of Wisconsin Biotechnology Center, 1710 University Avenue, Madison, Wis. 53705, BLAST, BESTFIT, GAP, or PILEUP / PRETTYBOX programs). Such software matches identical or similar sequences by assigning homology grades to various substitutions, deletions, and / or other modifications. In an exemplary method for determining the degree of identity, the BLAST program can be used, where a probability score between e-3 and e-100 indicates closely related sequences.

[0390] For example, the percentage identity of two nucleotide or polypeptide sequences can be determined using BLAST® with default parameters (available on the National Center for Biotechnology Information website of the U.S. National Library of Medicine). In some embodiments, the length of the reference sequence for comparison for comparative purposes is at least 30% (e.g., at least 40%, 50%, 60%, 70%, 80%, or 90%) of the reference sequence.

[0391] It should be understood that the numbering of a specific position or residue in a polypeptide sequence depends on the specific protein and numbering scheme used. Numbering can differ, for example, in the precursor of a mature protein and in the mature protein itself, and species-to-species sequence differences can affect numbering. Those skilled in the art will be able to identify residues in any homologous protein and its respective encoding nucleic acid using methods well known in the art, for example, by sequence alignment and determination of homologous residues.

[0392] The terms “regulation” or “alteration” refer to a change in the quantity, degree, or range of a function. For example, an epigenetic editor as described herein can regulate the activity of a promoter sequence by binding to a motif within the promoter, thereby inducing, enhancing, or repressing transcription of a gene operatively linked to the promoter sequence. As other examples, an epigenetic editor as described herein can prevent RNA polymerase from transcribing a gene or can inhibit the translation of mRNA transcripts. When referring to the use of an epigenetic editor or its components as described herein, the terms “inhibit,” “repress,” “suppress,” “silence,” etc., refer to reducing or preventing the activity (e.g., transcription) of a nucleic acid sequence (e.g., a target gene) or a protein relative to the absence of an epigenetic editor or its components. Terms may include partial or complete blocking of activity, or prevention or delay of activity. The inhibitory activity may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% lower than the control activity, or may be at least 1.5, 2, 3, 4, 5, or 10 times lower than the control activity.

[0393] The terms “about” or “approximately” mean within an acceptable range of error for a particular value as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, for example, limitations of the measurement system. For instance, depending on the practice of a given value, “about” may mean within one or more standard deviations. When a particular value is described in the application and claims, unless otherwise stated, the term “about” should be regarded as indicating an acceptable range of error for that particular value.

[0394] The ranges provided in this document should be understood as abbreviations of all values ​​within that range. For example, the range 1 to 50 should be understood as including any number, combination of numbers, or subrange of the following groups: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50, as well as all intermediate decimal values ​​between the aforementioned integers, such as 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, and 1.9. Regarding subranges, "nested sub-ranges" extending from either end of a range are specifically considered. For example, nested sub-ranges of the exemplary range 1 to 50 may include 1 to 10, 1 to 20, 1 to 30, and 1 to 40 in one direction, or 50 to 40, 50 to 30, 50 to 20, and 50 to 10 in another direction.

[0395] Unless otherwise defined herein, scientific and technical terms used in conjunction with this disclosure shall have the meanings commonly understood by one of ordinary skill in the art. Exemplary methods and materials are described below; however, similar or equivalent methods and materials may be used in practice or testing with respect to those described herein. In case of discrepancies, this specification (including definitions) shall prevail. Furthermore, unless the context otherwise requires, singular terms shall include plural terms, and plural terms shall include singular terms. Throughout this specification and embodiments, the words “have” and “comprise”, or variations such as “has,” “having,” “comprise,” or “comprising”, shall be understood to mean comprising the stated integer or group of integers, but not excluding any other integer or group of integers. Unless otherwise stated, the description of elements listed herein includes any element alone or in any combination. The description of embodiments herein includes embodiments as a single embodiment or in combination with any other embodiments herein. All publications, patents, patent applications, and other references mentioned herein are incorporated herein by reference in their entirety. In the event of any conflict between any references incorporated herein and the disclosure contained herein, this specification is intended to supersede and / or take precedence over any such conflicting material. Although numerous references are cited herein, such citations do not constitute an acknowledgment that any of these references is part of common general knowledge in the art.

[0396] According to this disclosure, backreferences in dependent claims refer to a direct and explicit abbreviation of each or each combination of claims indicated by the backreference. Furthermore, the headings herein are created for ease of organization and are not intended to limit the scope of the claimed methods and compositions in any way.

[0397] Some protein sequences provided herein, such as some fusion protein sequences, include peptide tags, such as His6 tags or DYKDDDDK (SEQ ID NO: 1528) tags, which can be used to detect and / or purify tagged proteins without affecting protein function. These peptide tags and other suitable peptide tags are well known to those skilled in the art. It will be apparent to those skilled in the art that the disclosed tags can replace other suitable peptide tags, and that fusion proteins having the same or highly similar sequences but not including such peptide tags (e.g., peptide tags cleaved from the fusion protein or fusion proteins without peptide tags) are also suitable for embodiments of this disclosure.

[0398] To better understand this disclosure, the following embodiments are described. These embodiments are for illustrative purposes only and should not be construed as limiting the scope of this disclosure in any way. Example

[0399] Example 1: Design and Synthesis of Fusion Proteins

[0400] A fusion protein (“CRISPR-off”) comprising dCas9, DNMT3A, DNMT3L, and KOX1 KRAB was designed and constructed. From the N-terminus to the C-terminus, the protein has the domain DNMT3A-linker-DNMT3L-XTEN80-NLS-dSpCas9-NLS-XTEN16-KOX1 KRAB (SEQ ID NO: 658 and 1495). The CRISPR-off plasmid construct has been described in Nuñez (Nuñez etal., Cell (2021) 184(9):2503-19) and was ordered from Twist Biosciences.

[0401] ZF fusion proteins (“ZF-off”) containing DNMT3A, 3L, and KOX1 KRAB were also constructed. The construct has the general structure DNMT3A-linker-DNMT3L-XTEN80-NLS-ZFP domain-NLS-XTEN16-KOX1Krab (SEQ ID NO: 659 and 1496).

[0402] Example 2: Selection of target PCSK9 sequences for gRNA epigenetic silencing

[0403] The target distance of gRNAs for PCSK9 in humans (GRCh38), mice (mm10), and long-tailed macaques (5.0) to PCSK9 was designed computationally using the Benchling gRNA platform (Benchling (2021), retrieved from benchling.com). gRNAs containing poly-TTTT sequences were discarded first. We performed off-target analyses of the gRNAs using CasOFFinder (Bae et al., Bioinformatics (2014) 30(10):1473-5). gRNAs that matched multiple locations in the corresponding genome constructed for each individual species were discarded.

[0404] Cross-reactive sequence analysis was performed on human PCSK9 gRNA to annotate sequence mismatches with macaque or mouse gRNA sequences. Specifically, gRNA sequence alignment was performed to identify the degree of DNA similarity at each nucleotide, including annotation of guide sequences containing up to zero, one, or two nucleotide mismatches. A final set of 226 gRNA sequences was selected for primary screening of PCSK9 in HeLa cells.

[0405] Example 3: Selection of ZF target sites and design of ZF proteins for epigenetic silencing

[0406] Two-finger ZFP (2F unit) libraries (each recognizing a 6 bp DNA site) were used to design larger six-finger ZFP arrays targeting 18 bp DNA binding sites. The 2F units were derived from a set of three-finger zinc finger proteins that had been selected using a bacterial 2-hybridization (B2H) selection system to bind to specific target sites (Hurt et al., PNAS (2003) 100:12271-6; Maeder et al., Mol Cell (2008) 31(2):294-301). A list of targetable DNA sites was created by generating all possible triplet combinations of the 6 bp binding sites represented in the library, allowing for 0 or 1 bp between the 6 bp target sites. To identify zinc finger target sites within PCSK9, sequences within + / - 1 kb of TSS (human (GRCh38)) were queried against this list. Multiple ZF proteins could be designed for each identified ZF target site. The design of the six recognition helices used to generate the complete protein was carried out by selecting bifinger units and considering various factors, such as the known binding preferences of zinc finger proteins, the frequency of amino acids at positions -1, 2, 3, and 6 selected in the B2H selection system to bind the desired target bases, avoiding amino acids at positions -1, 2, 3, and 6 selected in B2H to bind multiple different bases, and maintaining background dependence as much as possible by matching flanking bases. The complete ZF sequence is derived from the naturally occurring Zif268 protein, and the selected recognition helices are maintained in the background of the sequences selected in B2H (finger 1-2 or finger 2-3 of Zif268). The bifinger units are linked by a linker TGSQKP (SEQ ID NO: 651), where the 6 bp binding sites are contiguous, and by a linker TGGGGSQKP (SEQ ID NO: 652), where 1 bp separates the 6 bp binding sites. Ultimately, 209 ZFPs targeting 49 different binding sites were selected for initial screening of PCSK9 in HeLa cells.

[0407] Figure 1 The overlap of gRNA and zinc finger proteins mapped to the PCSK9 target region was shown.

[0408] Example 4: Guide RNA screening in HeLa cells

[0409] Primary screening of PCSK9-targeting gRNAs was performed in HeLa cells. The gRNA sequences were ordered as DNA fragments from Twist Biosciences, with the u6 promoter sequence preceding the gRNA coding sequence.

[0410] HeLa cells were transfected with gRNA in DNA form and CRISPR-off. Six 96-well plates (Sigma-Aldrich catalog number M2936) were seeded with 12,000 HeLa cells per well (ATCC catalog number CCL-2) in standard medium supplemented with 10% fetal bovine serum (Thermo Fisher catalog number A4766) v / v, 1x GlutaMAX™ (Thermo Fisher catalog number 35050061), and 1x penicillin-streptomycin (Thermo Fisher catalog number 15140122) in DMEM (Thermo Fisher catalog number 11-965-092). After plating, cells were allowed to grow in an incubator at 37°C and 5% CO2 for 24 hours. 25 ng of each gRNA fragment and 50 ng of CRISPR-off (SEQ ID NO: 658) plasmid were resuspended in DPBS buffer (Thermo Fisher catalog number 14190144) to a concentration of 7.5 ng / μL. Additionally, 10 ng of EF1a:purinomycin anti-plasmid was added to the transfection mixture to achieve a total DNA load of 85 ng. Transfection mixtures were prepared by adding the resuspended DNA components to the Mirus® TransIT®-LT1 transfection reagent (Mirus catalog number MIR2300) according to the manufacturer's instructions. 10 μL of each transfection mixture was added in duplicate to a total of six screening plates. The positive control used was CRISPRi (dCas9-KRAB), which has two gRNA targeting sites near the TSS. The two CRISPRi positive control gRNAs used were gRNA004 and gRNA005, as noted in the gRNA sequence table. These control conditions are referred to as “CRISPRi-1” and “CRISPRi-2” in the primary screening data table, respectively. The negative controls are CRISPR-off without gRNA, CRISPR-off with non-PCSK9 locus (CD151) targeting gRNA, and empty vector (pUC19; NEB catalog number N3041S).

[0411] 24 hours after transfection, puromycin resistance selection was performed. The cell culture medium was completely aspirated, and all wells were washed three times with DPBS buffer (Thermo Fisher catalog number 14190144). 200 μL of 1 μg / μL puromycin was then added to all wells of the selection plate.

[0412] Forty-eight hours post-transfection, cells were passaged. The cell culture medium was completely aspirated, and all wells were washed three times with DPBS buffer (Thermo Fisher catalog number 14190144). Cells were enzymatically stripped for five minutes at 37°C with 25 μL trypsin-EDTA (0.25%) (Thermo Fisher catalog number 25200056). Trypsinized cells were resuspended 1:8 in fresh standard medium and replate at a 1:4 ratio 72 hours after medium replacement.

[0413] To measure the level of secreted PCSK9 protein, the culture medium was harvested 24 hours after replacement, and the relative cell counts of the cell plates were determined using the Promega Cell Titer Glo™ assay (catalog number G7570) according to the manufacturer's recommendations. PCSK9 protein levels were assessed using the BioLegend LEGEND MAX™ Human PCSK9 ELISA Kit (catalog number 443107). The harvested culture medium was plated, and all subsequent steps were performed strictly according to the manufacturer's recommendations. Final plate readouts were performed at 450 nm on a Perkin Elmer® VICTOR® Nivo™ F instrument. The function was fitted to a standard curve using GraphPad Prism software, and interpolation was performed for unknown samples. PCSK9 ELISA results were normalized using the Cell Titer Glo® assay (Promega catalog number G7571) to correct for any well-to-well variability in cell counts.

[0414] More than 200 gRNAs were tested, of which 40 were identified as head sequences. Figure 2 (Head sequence indicated by a darker circle). The sequences and potency of the tested gRNAs are shown in Table 7 (SEQ: SEQ ID NO). Relative PCSK9 secretion (“Control PCSK9%”) represents the mean PCSK9 protein level in the treated sample, expressed as a percentage of the mean of all non-targeting gRNA (CD151) negative control conditions. Robust silencing of PCSK9 (30%–40% of negative control levels) was observed in cells treated with multiple gRNA candidate therapies. The head 40 gRNAs with the best PCSK9 protein knockout were selected as sgRNAs for further follow-up studies.

[0415] Table 7. Target sequences of the tested gRNAs

[0416]

[0417]

[0418]

[0419]

[0420]

[0421]

[0422]

[0423]

[0424]

[0425]

[0426]

[0427]

[0428]

[0429]

[0430]

[0431]

[0432]

[0433] The best-performing gRNA was found to be closely aligned with the transcription start site of the PCSK9 gene. Figure 3 ).

[0434] Following this primary selection, secondary selection was performed using 40 head gRNAs in RNA form. Forty guide sequences were chemically synthesized and co-transfected with in vitro transcribed mRNAs encoding CRISPR-off, CRISPRi, or WT Cas9 constructs. Secreted PCSK9 levels were measured at 7 and 28 days post-transfection.

[0435] To generate in vitro transcribed CRISPR-off, CRISPRi, and WT Cas9 effector mRNAs, plasmid constructs encoding these proteins were linearized using the MfeI restriction endonuclease from NEB® (catalog number R3589S). Following the manufacturer's instructions, an in vitro transcription reaction was established using the CellScript T7 mScript™ Standard mRNA Production System (catalog number C-MSC100625) with 1 μg of linearized template. The resulting RNA had a 5' cap 1 structure and was 3' polyadenylated. The transcribed RNA was purified using the Qiagen RNeasy® Mini Kit (catalog number 74104).

[0436] Terminally modified sgRNAs purified using standard desalting methods were obtained from Integrated DNA Technologies. The three nucleotides at the 5' and 3' ends of each guide sequence were 2'-O-methyl modified. The three nucleotide internucleotides at the 3' and 5' ends were phosphate thioester nucleotide internucleotides (Table 8; SEQ: SEQ ID NO). In the table, mX (i.e., mA, mC, mG, or mU) represents a 2'-O-methyl modified ribonucleotide, rX (i.e., rA, rC, rG, or rU) indicates a native ribonucleotide, and * indicates a phosphate thioester bond. All nucleotide internucleotides that are not phosphate thioester bonds are phosphate ester bonds.

[0437] Table 8. Target sequences of the head 40 gRNAs from secondary HeLa cell screening

[0438]

[0439]

[0440]

[0441]

[0442]

[0443]

[0444]

[0445]

[0446]

[0447]

[0448]

[0449]

[0450]

[0451]

[0452]

[0453]

[0454]

[0455]

[0456]

[0457]

[0458] HeLa cells were reverse transfected in 96-well plates with 25 ng effector and 12.5 ng sgRNA using the TransIT®-X2 transfection reagent from Mirus (catalog number MIR6003). Conditioned medium was harvested weekly for up to four weeks and used to measure secreted PCSK9 levels using the LEGEND MAX™ Human PCSK9 ELISA kit from BioLegend (catalog number 443107). ELISA data were normalized relative to cell number using the CellTiter-Glo® kit from Promega (catalog number G7571). While PCSK9 silencing using CRISPRi (dCas9-KRAB) was transient and returned to baseline by day 28, several sgRNAs co-transfected with the CRISPR-off (DNMT3A-3L-dCas9-KRAB) construct showed robust and durable silencing. Figure 4A Of the 40 guide sequences tested with CRISPR-off, 16 showed greater silencing efficacy than WTCas9 (Table 9).

[0459] Table 9. Relative PCSK9 expression in HeLa cells treated with modified gRNA and CRISPR-off

[0460]

[0461]

[0462] RNA was extracted at two-week intervals using the Rapid RNA 96 Kit (catalog number R1053) from Zymo Research. qPCR was performed using the qScript XLT One-Step RT-qPCR ToughMix (catalog number 95134-500) from Quantabio and TaqMan assays (PCSK9: Hs00545399_m1, PPIA: Hs99999904_m1). PCSK9 levels were normalized using PPIA. Relative quantification was performed using the delta-delta CT method.

[0463] It was found that the inhibition of PCSK9 secretion was associated with mRNA silencing on day 14. Figure 4B ).

[0464] In HeLa cells, modRNA004 and modRNA111 were tested to inhibit PCSK9 secretion for 60 days. Figure 5 WTCas9 was co-transfected with modRNA180 as a positive control. Cells were treated with 25 ng of effector and 12.5 ng of gRNA. This showed that modRNA004 and modRNA111 mediated persistent silencing of PCSK9, comparable to that achieved with WTCas9 via gene editing in HeLa cells.

[0465] Furthermore, the results showed that epigenetic silencing was maintained in HeLa cells treated with simvastatin. Statin treatment is known to increase PCSK9 secretion via a transcriptional mechanism. In epigenetically silenced HeLa cells, statin treatment did not increase PCSK9 secretion. Figure 6 ).

[0466] Example 5: Assay of guide RNA in Huh7 hepatocellular carcinoma cell line

[0467] The Huh7 hepatocellular carcinoma cell line is suitable for high-throughput screening and transfection. Thirteen guide sequences from the HeLa screening head were tested in the Huh7 hepatocellular carcinoma cell line; these 13 guide sequences showed either complete homology or a single mismatch with the cynomolgus monkey PCSK9 gene. Epigenetic silencing using the CRISPR-off construct was shown to be stable within 7 days. Figure 7 ).

[0468] Subsequently, the specificity of fusion proteins 11 and 13 to PCSK9-targeting gRNA 041 was measured in Huh7 cells (see Table 12). Differential gene expression was assessed by RNAseq of treated and untreated cells. Values ​​were visualized in a volcano plot. Figure 25 No significant off-target effects were detected.

[0469] Example 6: Determination of guide RNA in primary human and cynomolgus monkey hepatocytes

[0470] Primary human and cynomolgus monkey HepatoPac® cultures from BioIVT were used to test gRNA efficacy in primary hepatocytes. HepatoPac® cultures were maintained according to the manufacturer's recommendations. Briefly, HepatoPac® maintenance medium was thawed and prepared within 30 minutes of cell arrival. After medium replacement, cells were allowed to acclimatize for two days in a 37°C, 10% CO2 incubator. On the second day after reception, primary cultures were administered multiple concentrations of CRISPR-off+sgRNA, GFP-mRNA, and WTCRISPR Cas9 using a standard procedure. The medium was replaced every other day for up to four weeks to assess the persistence and / or heritability of silencing. PCSK9 silencing was assessed every seven days by ELISA to measure the level of secreted PCSK9 in the medium. PCSK9 concentrations were controlled in total hepatocytes using a human albumin ELISA (Thermo Fisher®). Data were then presented as a percentage of total PCSK9 secretion relative to the GFP-mRNA negative control. Specificity could be assessed by isolating primary human and cynomolgus monkey hepatocytes from mouse fibroblast feeder layers using a magnetic bead-based antibody method (Miltenyi Biotec). After isolating primary hepatocytes from the feeder layer, the cells were processed for RNA-seq evaluation and whole-genome bisulfite sequencing.

[0471] Thirteen head gRNAs in RNA form were selected for testing in primary human hepatocytes (PHH). Head gRNAs were selected based on (i) PCSK9 silencing efficiency and persistence in HeLa cells and (ii) whether they had a perfect match with the human PCSK9 gene and at most one mismatch with the non-human primate PCSK9 gene. Combinations of gRNAs were also tested to determine their efficacy and persistence. Negative controls were CRISPR-off only, gRNA fragment only (modRNA003), and CRISPRi only. The positive control was CRISPRi co-transfected with modRNA004 (Table 10; SEQ: SEQ ID NO; NHP: non-human primates). All tested gRNAs were predicted to bind to both human and non-human primate PCSK9.

[0472] Table 10. gRNAs screened in primary human hepatocytes

[0473]

[0474] Robust PCSK9 silencing was observed. For some gRNAs, a reduction of >70% in secreted PCSK9 was observed on day 7 (depending on transfection efficiency).

[0475] Example 7: ZF assay in HeLa cells

[0476] A total of 209 zinc finger proteins (structures shown in SEQ ID NO: 659) were designed from 49 PCSK9 target sites (selected from chromosome 1 of GRCh38, between 55038548 and 55040548) of the ZF library. No other exact matches of the target sites were found in the human genome (GRCh38).

[0477] HeLa cells were transfected with the ZF-off construct in DNA form. Six 96-well plates (Sigma-Aldrich catalog number M2936) were seeded with 12,000 HeLa cells (ATCC catalog number CCL-2) per well in standard medium supplemented with 10% fetal bovine serum (Thermo Fisher catalog number A4766) v / v, 1x GlutaMAX™ (Thermo Fisher catalog number 35050061), and 1x penicillin-streptomycin (Thermo Fisher catalog number 15140122) in DMEM (Thermo Fisher catalog number 11-965-092). After plating, cells were allowed to grow for 24 hours in a 37°C incubator with 5% CO2. 10 ng of the ZF-off plasmid was resuspended in DPBS buffer (Thermo Fisher catalog number 14190144) to a concentration of 7.5 ng / μL. In addition, 10 ng of EF1a:purinomycin resistance plasmid and 65 ng of empty vector (pUC19) were added to the transfection mixture to achieve a total DNA load of 85 ng. The transfection mixture was created by adding DNA resuspended in serum-free OPTI-MEM medium (Thermo Fisher® catalog number 31985062) and Mirus® TransIT®-LT1 transfection reagent (MIR2300) according to the manufacturer's instructions. Two 10 μL copies of the transfection mixture were added to a total of six screening plates. The positive control was CRISPR-off (SEQ ID NO: 658) with high-performance gRNA (gRNA009). The negative control was ZF-off with a non-PCSK9 locus target (CLTA) and empty vector (pUC19; NEB catalog number N3041S).

[0478] ZF screening yielded hits with activity comparable to CRISPR. Figure 8 Candidates with high silencing efficiency were selected for subsequent experiments. Figure 9 The ZF screening results based on distance from the TSS are shown. A total of 209 ZFs were screened, and their PCSK9 knockout activity relative to the negative control is shown in Table 11 below.

[0479] Table 11. Activity of ZF-off constructs

[0480]

[0481]

[0482] Several ZF-off constructs, when combined with gRNA003, are shown to be more effective than WTCas9 and CRISPR-off in silencing PCSK9. The target sites of the ZFP domain and the ZF sequences (F1 to F6) in these ZF-off constructs are shown in Table 1.

[0483] Example 8: Full-specific screening of constructs in primary human hepatocytes

[0484] The specificity of CRISPR-off and ZF-off constructs for silencing PCSK9 was tested in primary human hepatocytes. Readouts assessing specificity were obtained using RNA-seq, methylation arrays, and whole-genome bisulfite sequencing. Changes in whole-genome expression and methylation after epigenetic editing were analyzed compared to a negative control.

[0485] The specificity of five gRNAs was tested in PXB cells. Fresh human hepatocytes were isolated from PXB mouse models. Figure 26A The long-term stability and functionality of hepatocytes, as well as robust PCSK9 secretion, were confirmed. Cells were treated with the epigenetic repressor PLA2628 (fusion protein 12 in Example 12) or a control. PCSK9 secretion was measured and plotted as the percentage of secreted PCSK9 produced by the negative control (PLA2628 with non-PCSK9-targeting gRNA, “off-target”) (time process shown in Figure 1). Figure 26B The level on day 14 was shown in the middle; Figure 27 (in Chinese). Construct specificity was tested for all gRNAs by RNAseq on day 14 post-delivery. Volcano plots of the two exemplary RNAs evaluated (RNA041 and RNA049, sequences shown in Table 12) are shown below. Figure 26C As shown, no significant off-target effects were observed for any of the five tested gRNAs.

[0486] Example 9: CpG methylation pattern

[0487] The CpG methylation patterns in human hepatocytes (e.g., primary cells or cell lines) treated with CRISPR-off or ZF-off were investigated. Hybridization capture assays were performed on bisulfite-treated DNA to investigate the methylation patterns at CpG sites induced by CRISPR-off or ZF-off in a 1 kb region surrounding the PCSK9 TSS.

[0488] In mice treated with CRISPR-off (gRNA 041 and 049) or ZF-off (ZFP152) constructs, whole-liver hybrid capture assays were performed at d90 after partial hepatectomy and compared with untreated or vector-treated control mice, as described in more detail below. Methylation was investigated in approximately 5 kb of genomic regions, including the PCSK9 promoter region. Low levels of CpG methylation were observed in untreated and vector-treated control mice, while significant levels of CpG methylation were observed in CRISPR-Off and ZF-Off treated mice (data not shown).

[0489] Example 10: Stable PCSK9 silencing via epigenetic editing in mice with wild-type PCSK9

[0490] The ability of CRISPR-off and ZF-off constructs to mediate epigenetic silencing of endogenous PCSK9 was tested in vivo. Constructs were delivered using a single IV administration of mRNA (and, for CRISPR-off silencing, gRNA). Silencing was tested in wild-type mice for periods ranging from 2 to 6 months. Readouts were serum PCSK9 and serum cholesterol levels. A subset of each cohort was selected for liver hematoxylin and eosin (H&E) staining RNAseq and analysis. Robust, stable, and heritable PCSK9 silencing was observed for several constructs.

[0491] Example 11: Stable PCSK9 silencing via epigenetic editing in mice expressing transgenic human PCSK9

[0492] Three different mouse strains expressing transgenic human PCSK9 are available: hPCSK9-Tg (mPCSK9+ / -) heterozygous mice, hPCSK9-Tg (mPCSK9+ / +) homozygous mice, and hPCSK9-Tg (mPCSK9- / -) mice. The available hPCSK9-Tg (mPCSK9- / -) mouse strain is C57BL / 6J-Pcsk9- / -Tg (RP11-55M23-AbsI), which expresses human PCSK9 under the control of its own promoter. Figure 10See, for example, Weider et al., J Biol Chem (2016) 291(32): 16659-71.

[0493] CRISPR-off and ZF-off constructs were tested in hPCSK9-Tg (mPCSK9- / -) mice expressing hPCSK9. The constructs used were: CRISPR / wtCas9 (SEQ ID NO: 2) with gRNA g079; CRISPRi (NLS-NLS-dCas9-NLS-KOX1KRAB-NLS-NLS) with gRNA g041; CRISPR-OFF PLA2628 (fusion protein 12 as provided in Example 12 of this document, SEQ ID NO: 1519) with gRNA g056; CRISPR-OFF PLA2628 with gRNAs g041 and g049; CRISPR-OFF PLA1489 (fusion protein 11 as provided in Example 12 of this document, SEQ ID NO: 1517) with gRNAs g041 and g049; and ZF-OFF ZFP152ADD as shown below:

[0494]

[0495] (SEQ ID NO: 1527)

[0496] Construct delivered via a single IV administration of an epigenetic silencing agent Figure 20B ); 3 mg / kg of the formulation was administered to each mouse. The guide RNAs selected for use are shown in Table 12. Serum ELISA for human PCSK9 was performed 42 days after IV administration. Figure 21A ) and 84 days Figure 21B The efficacy of PCSK9 silencing was measured. Mice treated with constructs that did not show persistent PCSK9 silencing were sacrificed at 6 weeks. Persistent silencing was observed with PLA2628 (used with gRNAs g041 and g049), PLA1489 (used with gRNAs g041 and g049), and ZFP152ADD.

[0497] Table 12: Guide RNAs tested in transgenic mice

[0498]

[0499] In further experiments, g041 and g49 were synthesized using the following chemical modification patterns (where 'm' indicates ribonucleotides modified with 2'-OMe; 'r' indicates unmodified ribonucleotides; and '*' indicates phosphate thioester bonds):

[0500] mA*mC*mU*rGrCrUrGrGrCrUrCrArCrUrCrCrUrCrCrGrUrUrUrUrArGrArGmCmUmAmGmAmAmAmUmAmGmCrArGrUrUrArArArArU rArArGrGrCrUrArGrUrCrCrGrUrUrArUrCrAmAmCmUmUmGmAmAmAmAmGmUmGrGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (mg041, SEQ IDNO: 1528)

[0501] mA*mU*mC*rGrUrCrCrGrArUrGrGrGrGrCrUrCrUrGrGrGrUrUrUrUrArGrArGmCmUmAmGmAmAmUmAmGmCrArArGrUrUrArArArArU rArArGrGrCrUrArGrUrCrCrGrUrUrArUrCrAmAmCmUmUmGmAmAmAmAmGmUmGrGmCmAmCmCmGmAmGmUmCmGmGmUmGmCmU*mU*mU*mU (mg049, SEQ IDNO: 1529)

[0502] The modified guide RNA was tested in hPCSK9-Tg (mPCSK9- / -) mice expressing hPCSK9. The guide RNA was administered alone (mg041 or mg049) or together with mRNA constructs encoding the following epigenetic repressor proteins (mg041 and mg049):

[0503]

[0504] The construct and guide RNA were delivered to each mouse at a 1:1 ratio of 0.375 mg / kg via a single administration of epigenetic repressor mRNA and gRNA. The efficacy of PCSK9 silencing was measured at 7 and 14 days post-administration by serum ELISA for human PCSK9. A guide sequence with corresponding terminal modifications (having three 5' and three 3' nucleotides modified with 2'-OMe and phosphate thioester bonds) was used as a positive control. Durable silencing of PCSK9 was observed for all guide sequences, and silencing was equivalent between the terminally modified control guide RNA and the tested modified guide RNA at both time points. The silencing levels were similar when tested at both time points.

[0505] Example 12: Fusion protein with variant NLS conformation

[0506] Several improved fusion protein constructs with significantly higher epigenetic silencing activity were developed using variant nuclear localization sequence (NLS) conformations.

[0507] Several constructs of NLS domains with variant configurations were constructed. Figure 11A and 11B ), and tested at the PCSK9 locus in HeLa cells ( Figure 12A-12B The construct is in Hepa1-6 (). Figure 13 ) and HuH7 ( Figures 14A-14C and Figure 15 Further testing was conducted at the PCSK9 locus in the fusion protein. The amino acid and DNA sequences of an exemplary fusion protein construct are shown below:

[0508]

[0509]

[0510]

[0511]

[0512]

[0513]

[0514]

[0515]

[0516]

[0517]

[0518]

[0519]

[0520]

[0521]

[0522]

[0523]

[0524]

[0525]

[0526]

[0527]

[0528]

[0529]

[0530]

[0531]

[0532]

[0533]

[0534]

[0535]

[0536]

[0537]

[0538]

[0539]

[0540]

[0541]

[0542]

[0543]

[0544]

[0545]

[0546]

[0547]

[0548]

[0549]

[0550]

[0551]

[0552]

[0553]

[0554]

[0555]

[0556]

[0557]

[0558]

[0559]

[0560]

[0561]

[0562]

[0563]

[0564]

[0565]

[0566]

[0567]

[0568]

[0569]

[0570]

[0571]

[0572]

[0573]

[0574]

[0575]

[0576]

[0577]

[0578]

[0579] Figure 14A The sequence can be found below:

[0580]

[0581] Figure 15 The sequence can be found below:

[0582]

[0583]

[0584]

[0585]

[0586]

[0587]

[0588] Cell culture and transfection

[0589] HeLa (ATCC-CRM-CCL-2), Hepa1-6 (PCSK9-IRES-TdTomato), Huh7 (SekisuiXenoTech, LLC), and HEK293T Griptite (CLTA-GFP) cells were cultured in DMEM containing 10% FBS. All experiments in HeLa and Huh7 cells were performed using chemically synthesized guide RNA and in vitro transcribed effector constructs. HeLa cells were reverse transfected using the TransIT-X2 transfection reagent from Mirus (catalog number MIR6003). Huh7 cells were reverse transfected using the MessengerMAX reagent from Invitrogen (catalog number LMRNA003). Secreted PCSK9 levels were measured at specified time points using the LEGEND MAX™ Human PCSK9 ELISA kit from Biolegend (catalog number 443107). ELISA data were normalized for cell number using the CellTiter-Glo kit from Promega (catalog number G7571).

[0590] HEK293T Griptite cells containing GFP knocked into the CLTA locus were co-transfected with plasmids encoding effector constructs and human CLTA guide RNA using the TransIT-X2 transfection reagent from Mirus (catalog number MIR6003). GFP expression was measured by FACS as a substitute for CLTA expression.

[0591] Hepa1-6 cells were co-transfected with plasmids encoding effector constructs and mouse PCSK9 guide RNA using the SF Cell Line 96-well nuclear transfection kit (catalog number V4SC-2096, program code: CM-138) from the Lonza Amaxa 4D nuclear transfection apparatus. FACS analysis of TdTomato expression was performed at specified time points as a proxy for PCSK9 levels.

[0592] Effector constructs and in vitro transcription of synthesized gRNA

[0593] Following the manufacturer's instructions, an in vitro transcription reaction was established using the CellScript T7 mScript™ standard mRNA production system (catalog number C-MSC100625) with 1 μg of linearized effector template to obtain RNA with a Cap 1 structure at the 5' end and 3' polyadenylated. Terminally modified sgRNA with three 2'O-methyl modified nucleotides at both the 5' and 3' ends was obtained from Integrated DNA Technologies.

[0594] Methylation spectrum

[0595] Genomic DNA was extracted from each well of a 96-well culture plate using the DNAdvance Tissue DNA Extraction Kit (Beckman Coulter). After quantification of the genomic DNA using the High Sensitivity DNA 1X Kit (Quant-IT), each genomic DNA sample was bisulfite-converted using the EZ-96 DNA Methylation Gold MagPrep Kit (Zymo Research) according to the manufacturer's instructions. For hybridization capture experiments, DNA libraries were prepared using the xGen™ Methyl-Seq DNA Library Preparation Kit (IDT), and hybridization capture was performed using the xGen™ DNA Library Hybridization Capture Kit (IDT). For amplicon sequencing experiments, DNA libraries were prepared using the xGen™ Methyl-Seq DNA Library Preparation Kit (IDT), and hybridization capture was performed using the xGen™ DNA Library Hybridization Capture Kit (IDT). The bisulfite-converted DNA from each sample was used to inoculate PCR corresponding to each of the two VIM amplicon using the Platinum Taq Kit (Invitrogen). Before sequencing via commercial service (Azenta), the merged products were cleaned using an AMPure XP kit (Beckman Coulter) and fragment size was assessed via a D1000 screentape on a Tapestation 4200 (Agilent).

[0596] Example 13: Bacterial DNA Methyltransferase

[0597] In this experiment, a group of bacterial proteins were screened for their DNA methyltransferase activity in mammalian cells. The epigenetic silencing activity of these bacterial DNA methyltransferases was tested by fusing the N-terminus of the bacterial DNA methyltransferases (Table 13) with a dCas9 domain using the experimental procedure of Example 1. These constructs were then transfected into reporter cell lines that expressed GFP under the control of the mammalian promoter of CTLA4.

[0598] Table 13: Bacterial DNA Methyltransferases

[0599]

[0600]

[0601] M.SssI DNA methyltransferase can effectively methylate DNA in mammalian cells ( Figure 16 It has a stable silence of up to 30 days. Figure 16 The sequence can be found below:

[0602]

[0603] On day 29, the methylation profiles of these cells were analyzed, confirming 20% ​​methylation of the target gene. Figure 17-18 ).

[0604] Using the experimental procedures of Example 1, the epigenetic silencing activities of three other orthologous DNA methyltransferases (predicted to be closely related to M.SssI) were identified and tested.

[0605] Table 14: Bacterial Methyltransferases

[0606]

[0607] Table 14 shows DNA methyltransferases predicted to have similar or improved functions to M. SssI. Sequences were tested in a CRISPR-off background, replacing mouse DNMT3A / DNMT3L, and their functions were compared with those of M. SssI DNA methyltransferases in terms of silencing the PCSK9 locus in the HeLa TdTomato system to identify novel features and improved functions.

[0608] Example 14: Alternative KRAB Domain

[0609] In this embodiment, a fusion protein was constructed using an alternative KRAB domain (Table 15), and when tested using the experimental procedure of Example 1, it showed improved activity compared to CRISPR-off. Figures 19A-19D ).

[0610] Table 15: Alternative KRAB Domains

[0611]

[0612]

[0613] Example 15: ZIM 3 Fusion Construct

[0614] Novel fusions of ZIM3 and KOX1KRAB were generated. Both ZIM3 and KOX1KRAB are KRAB family proteins with broad homology. Therefore, sequences representing the midpoint between ZIM3 and KOX1KRAB were designed. These KOX1KRAB and ZIM3 constructs encode small regions of KOX1KRAB and ZIM3 concentrated around the zinc finger domain of the protein. While the regions of use for KOX1KRAB and ZIM3 are very similar within the first approximately 75 bp of their sequences, ZIM3 also has a small α-helix region at its C-terminus, which is absent in KOX1KRAB. The KOX1KRAB-FL sequence includes a KOX1KRAB sequence equivalent to this additional fragment, while the ZIM3 truncation removes this additional fragment from the ZIM3 sequence. The ZIM3 / KOX1KRAB chimera is a fusion of the N- and C-terminal fragments of the two proteins. ZIM3-like KOX1KRAB variants were first assembled by BLAST from ZIM3 or KOX1KRAB proteins from non-human species to assemble the 100 closest homologs (“families”) of each gene; second, three members of the KOX1KRAB family most similar to ZIM3 and three members of the ZIM3 family most similar to KOX1KRAB were identified; third, the KOX1KRAB-FL sequence was rationally modified to be similar to each of the three groups (Table 16).

[0615] Table 16: ZIM-KOX1KRAB chimeric protein

[0616]

[0617] Example 16: PCSK9 Silent Dose Response Experiment

[0618] Different doses of CRISPR-off and ZF-off constructs were tested in mice expressing human PCSK9. Constructs exhibiting persistent PCSK9 silencing at 3 mg / kg (see Example 11) were selected for further studies. These included PLA2628 with gRNAs G041 and g049; PLA1489 with gRNAs G041 and g049; and ZFP152ADD, as disclosed in Example 11. Constructs were delivered as described in Example 11. Each construct of interest was tested at 0.2 mg / kg, 0.375 mg / kg, 0.75 mg / kg, and 3 mg / kg. Baseline hPCSK9 levels were determined two weeks prior to construct administration. Figure 22 The efficacy of PCSK9 in silencing was measured over 28 days. Results for PLA1489 and ZFP152ADD were shown in [data missing]. Figure 23A and Figure 23BThe attached figure shows serum human PCSK9 levels as measured by ELISA after IV. Durable silencing of hPCSK9 was observed with each construct at several test doses, with levels at approximately 10% or less of baseline.

[0619] Example 17: PCSK9 silencing is persistent after partial hepatectomy

[0620] Partial hepatectomy is an established method for inducing liver regeneration in mice and has been used to test the persistence of genetic and epigenetic effects in mice. Mice were treated with an epigenetic editor targeting PCSK9 using a single IV administration. Eleven mice in each group were given (1) a sham administration (mediator-negative control only), (2) 1.5 mg / kg of the epigenetic editor, or (3) 3 mg / kg of the epigenetic editor. The mice were then divided into cohort A and cohort B. After 3 months, mice in cohort B underwent surgical resection of 70% of their liver, while mice in cohort A did not undergo surgery. After 2 months, the livers of mice in cohort B regenerated. Serum levels of PCSK9 were measured periodically after injection in both cohorts. Figure 24 The silencing of PCSK9 in mice in both cohort A and cohort B was persistent 140 days after injection.

[0621] sequence

[0622] The SEQ ID NO (SEQ) of the nucleotide (nt) and amino acid (aa) sequences described in this disclosure is listed below.

[0623]

[0624]

[0625]

[0626]

[0627]

[0628]

[0629]

[0630]

[0631]

[0632]

[0633]

[0634]

[0635]

[0636]

[0637]

[0638]

[0639]

[0640]

[0641]

[0642]

[0643]

[0644]

[0645]

[0646]

[0647]

[0648]

[0649]

[0650]

[0651]

[0652]

[0653]

[0654]

[0655]

[0656]

[0657]

[0658]

[0659]

[0660]

[0661]

[0662]

[0663]

[0664]

[0665]

[0666]

[0667]

[0668]

[0669]

[0670]

[0671]

[0672]

[0673]

[0674]

[0675]

[0676]

[0677]

[0678]

[0679]

[0680]

[0681]

[0682]

[0683]

[0684]

[0685]

[0686]

[0687]

[0688]

[0689]

[0690]

[0691]

[0692]

[0693]

[0694]

[0695]

[0696]

[0697]

[0698]

[0699]

[0700]

[0701]

[0702]

[0703]

Claims

1. A system for inhibiting the transcription of the human PCSK9 gene in human cells, optionally human hepatocytes, said system comprising... a) One or more fusion proteins, said fusion proteins collectively comprising DNA methyltransferase (DNMT) domain and / or domain recruiting DNMT, optionally wherein the DNMT domain and / or the recruiting domain comprises a DNMT3A domain and / or a DNMT3L domain, and optionally wherein the recruited DNMT is DNMT3A, and Transcription repressor domain, Each domain is linked to a DNA-binding domain that binds to a target sequence in the human PCSK9 gene, wherein the target sequence comprises the sequence of SEQ ID NO: 687, SEQ ID NO: 1039, SEQ ID NO: 1044, or SEQ ID NO: 1046; or b) One or more nucleic acid molecules that encode the one or more fusion proteins.

2. The system of claim 1, wherein the DNA-binding domain comprises a death CRISPR Cas (dCas) domain, and wherein the system comprises a guide RNA that targets the fusion protein to one or more sequences selected from SEQ ID NO: 1039, 1044, and 1046 in the PCSK9 gene.

3. The system of claim 2, wherein the system comprises (i) one or more guide RNAs, the guide RNAs comprising any one of SEQ ID NO: 1491-1493, or (ii) one or more nucleic acid molecules encoding the one or more guide RNAs of (i).

4. The system of claim 3, wherein the guide RNA comprises the sequence mA*mC*mU*rGrCrCrUrGrGrCrUrCrArCrUrCrCrCrCrCrCrGrUrUrUrUrArGrArGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArGrGrCrUrArGrUrCrUrCrCrCrCrUrUrUrArUrCrAmAmCmUmUmGmAmAmAmAmAmGmUmGrGmCmAmCmCmGmAmGmUmCmGmUmGmGmGmCmAmCmGmAmGmUmGmGmGmGmGmGmGmUmGmCmU*mU*mU*mU (mg041, SEQ ID NO: 1528).

5. The system of claim 3, wherein the guide RNA comprises the sequence mA*mU*mC*rGrUrCrCrGrArUrGrGrGrGrCrUrCrUrGrGrGrUrUrUrUrArGrArGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArGrGrCrUrArGrUrCrUrCrCrCrGrUrUrArUrCrAmAmCmUmUmGmAmAmAmAmAmGmUmGrGmCmAmCmCmGmAmGmUmGmGmGmCmAmCmGmAmGmUmGmGmGmGmGmGmGmGmGmGmCmU*mU*mU*mU (mg049, SEQ ID NO: 1529).

6. The system of any one of claims 2-5, wherein the dCas domain comprises a dCas9 sequence, optionally having at least 90% identity with SEQ ID NO: 12 or 13.

7. The system of claim 1, wherein the DNA binding domain comprises a ZFP domain, the ZFP domain binding to a nucleotide target sequence comprising the sequence of SEQ ID NO:

687.

8. The system of claim 7, wherein the ZFP domain comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

9. The system of any one of claims 1-8, wherein the DNMT3A domain comprises a sequence having at least 90% identity with SEQ ID NO: 574 or 575.

10. The system of any one of claims 1-9, wherein the DNMT3L domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO:578-581.

11. The system of any one of claims 1-9, wherein the DNMT3L domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO:582-603.

12. The system of any one of claims 1-8, wherein the DNMT domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO:601-603.

13. The system of any one of claims 1-12, wherein the transcriptional repressor domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 33-570.

14. The system of any one of claims 1-12, wherein the transcriptional repressor domain comprises a KRAB domain derived from KOX1, ZIM3, ZFP28 or ZN627.

15. The system of claim 14, wherein the KRAB domain comprises a sequence having at least 90% identity with a sequence selected from SEQ ID NO: 89, 116, 245 and 255.

16. The system of any one of claims 1-12, wherein the transcriptional repressor domain comprises a fusion of the N-terminal and C-terminal regions of ZIM3 and KOX1KRAB, and optionally comprises the amino acid sequence of SEQ ID NO: 571 or 572.

17. The system of any one of claims 1-12, wherein the transcriptional repressor domain is derived from KAP1, MECP2, HP1a / CBX5, HP1b, CBX8, CDYL2, TOX, TOX3, TOX4, EED, EZH2, RBBP4, RCOR1, or SCML2.

18. The system of any one of claims 1-17, wherein the system comprises: a) A fusion protein comprising a DNMT3A domain, a DNMT3L domain, a transcriptional repressor domain, and a DNA-binding domain. Optionally, one or both of the DNMT3A and DNMT3L domains are human, and The DNA-binding domain mentioned therein is a death CRISPR-Cas domain or a ZFP domain; or b) A nucleic acid molecule that encodes the fusion protein.

19. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, the DNMT3A domain, the first peptide linker, the DNMT3L domain, the second peptide linker, the DNA-binding domain, the third peptide linker, and the transcriptional repressor domain.

20. The system of claim 19, wherein the fusion protein comprises, from the N-terminus to the C-terminus, the DNMT3A domain, the first peptide linker, the DNMT3L domain, the second peptide linker, the first nuclear localization signal (NLS), the DNA-binding domain, the second NLS, the third peptide linker, and the transcriptional repressor domain.

21. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, a first nuclear localization signal (NLS), the DNMT3A domain, a first peptide linker, the DNMT3L domain, a second peptide linker, the DNA-binding domain, the third peptide linker, the transcriptional repressor domain, and a second NLS.

22. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, first and second nuclear localization signals (NLS), the DNMT3A domain, the first peptide linker, the DNMT3L domain, the second peptide linker, the DNA-binding domain, the third peptide linker, the transcriptional repressor domain, and third and fourth NLS.

23. The system of any one of claims 18-22, wherein the transcriptional repressor domain is a KRAB domain, optionally a KOX1, ZFP28, ZN627 or ZIM3 KRAB domain.

24. The system of any one of claims 19-23, wherein one or both of the second peptide connector and the third peptide connector are XTEN connectors, optionally selected from XTEN80 and XTEN16, and further optionally wherein the second peptide connector is XTEN80 and the third peptide connector is XTEN16.

25. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a dSpCas9 domain, a second NLS, an XTEN16 peptide linker, and a human KOX1 KRAB domain.

26. The system of any one of claims 1-6 or 9-15, wherein the fusion protein comprises SEQ ID NO:1519 or a sequence that is at least 90% identical thereto.

27. The system of any one of claims 1 or 7-15, wherein the fusion protein comprises SEQ ID NO: 688 or at least 90% identical thereto, and comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

28. The system of any one of claims 1 or 7-15, wherein the fusion protein comprises SEQ ID NO: 689 or at least 90% identical thereto, and comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

29. The system of any one of claims 1 or 7-15, wherein the fusion protein comprises SEQ ID NO: 690 or at least 90% identical thereto, and comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

30. The system of any one of claims 1 or 7-15, wherein the fusion protein comprises SEQ ID NO: 1527 or at least 90% identical thereto, and comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

31. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a first NLS, a ZFP domain, a second NLS, an XTEN16 linker, and a human KOX1 KRAB domain.

32. The system of claim 31, wherein the fusion protein comprises the sequence of SEQ ID NO: 659 or at least 90% identical thereto, or SEQ ID NO: 1496 or at least 90% identical thereto, optionally wherein the ZFP comprises the F1-F6 amino acid sequence of ZF034 as shown in Table 1.

33. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a ZFP domain, an XTEN16 peptide linker, a human KOX1 KRAB domain, and third and fourth NLS, optionally wherein the fusion protein comprises SEQ ID NO: 1514 or at least 90% identical to it, or SEQ ID NO: 1525 or at least 90% identical to it.

34. The system of claim 18, wherein the fusion protein comprises, from the N-terminus to the C-terminus, first and second NLS, a human DNMT3A domain, a first peptide linker, a human DNMT3L domain, an XTEN80 peptide linker, a dSpCas9 domain, an XTEN16 peptide linker, a human ZIM3 KRAB domain, and third and fourth NLS.

35. The system of claim 34, wherein the fusion protein comprises SEQ ID NO: 698 or at least 90% identical to it.

36. A human cell or a descendant of said cell, said human cell comprising the system of any one of claims 1-35, optionally said cell being a hepatocyte.

37. A pharmaceutical composition comprising the system and pharmaceutically acceptable excipients of any one of claims 1-35.

38. A method of treating a patient in need of treatment, the method comprising administering to the patient the system of any one of claims 1-35 or the pharmaceutical composition of claim 37, optionally intravenously.

39. The method of claim 38, wherein the patient Suffering from heart disease, Having elevated low-density lipoprotein cholesterol (LDL-C) or hypercholesterolemia, At risk of developing myocardial infarction, stroke, or unstable angina, and / or Patients with primary hyperlipidemia may be either heterozygous familial hypercholesterolemia (HeFH) or homozygous familial hypercholesterolemia (HoFH).

40. The system of any one of claims 1-35 or the pharmaceutical composition of claim 37, for treating a patient in need, optionally in the method of claim 38 or 39.

41. Use of the system of any one of claims 1-35 in the preparation of a medicament for treating a patient in need, optionally in the method of claim 38 or 39.

Citation Information

Patent Citations

  • RNA-guided nucleases and active fragments and variants thereof and methods of use

    US11162114B2

  • Delivery system for functional nucleases

    US20160200779A1

  • Switchable cas9 nucleases and uses thereof

    US20160208288A1

  • Selection of sites for targeting by zinc finger proteins and methods of designing zinc finger proteins to bind to preselected sites

    US6453242B1

  • Regulation of endogenous gene expression in cells using zinc finger proteins

    US6534261B1