DNA methyltransferase-like protein (DNMT3l) or DNA methyltransferase 3a (DNMT3a) repressor systems for epigenetic editing
Fusion proteins with a DNA-binding domain and catalytically inactive DNA methyltransferase effector domain address the issue of undesired effects in targeted transcriptional repression, achieving precise and specific gene repression.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2026-03-26
AI Technical Summary
Existing approaches for targeted transcriptional repression often result in undesired effects, necessitating improved fusion proteins and DNA-targeting systems that can selectively repress gene expression without recruiting heterochromatin-inducing factors or having DNA methyltransferase activity.
Development of fusion proteins comprising a DNA-binding domain for targeting to a specific site, such as PCSK9, and an effector domain with a catalytically inactive DNA methyltransferase domain, devoid of DNA methyltransferase activity, heterochromatin-inducing factor recruitment, and H3K4meO peptides, to achieve targeted gene repression.
The fusion proteins effectively repress target genes with high specificity and reduced off-target effects, providing a controlled and precise mechanism for epigenetic editing.
Smart Images

Figure US2025047433_26032026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 224742003640DNA METHYLTRANSFERASE-LIKE PROTEIN (DNMT3L) OR DNA METHYLTRANSFERASE 3A (DNMT3A) REPRESSOR SYSTEMS FOR EPIGENETIC EDITINGCross-Reference to Related Applications
[0001] This application claims priority from U.S. provisional application No. 63 / 698,024 filed September 23, 2024, the contents of which are incorporated by reference in its entirety.Incorporation by Reference of Sequence Listing
[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a 224742003640SeqList.xml, created September 23, 2025, which is 782,562 bytes in size. The information in the electronic format of the Sequence Listing is herein incorporated by reference in its entirety.Field
[0003] The present disclosure relates in some aspects to fusion proteins for targeted transcriptional repression. Also provided are DNA-targeting systems, such as CRISPR / Cas-based DNA-targeting systems, comprising an effector domain comprising DNA methyltransferase 3L proteins or domains thereof. Also provided are DNA-targeting systems, such as CRISPR / Cas-based DNA-targeting systems, comprising an effector domain comprising variant DNA methyltransferase 3A domains or functionally active portions thereof. In some aspects, the compositions and methods provided herein facilitate targeted transcriptional repression by targeting the effector domain or variants thereof to a target site, such as a target site for a target gene. In some aspects, also provided are methods and uses related to the provided fusion proteins, effectors or DNA-targeting systems or combinations thereof, for example in connection with therapeutic applications.Background
[0004] Targeted epigenetic modification can be used in aspects such as investigating biology and regulation of gene expression. However, existing approaches and systems for targeted transcriptional repression may result in undesired effects in various contexts. Improved fusion proteins and DNA-targeting systems are needed. Provided are embodiments that meet such and other needs.Attorney Docket No. 224742003640Summary
[0005] Provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, and (b) an effector domain that recruits domains with DNA methyltransferase activity to the target site for PCSK9, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides. In some of any embodiments, the effector domain comprises a catalytically inactive DNA methyltransferase domain or portion thereof.
[0006] Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, and (b) an effector domain comprising a catalytically inactive DNA methyltransferase domain or portion thereof, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides.
[0007] Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, and (b) an effector domain comprising a catalytically inactive DNA methyltransferase domain or portion thereof, wherein the length of the fusion protein minus (a) is less than 750 amino acids.
[0008] In some of any embodiments, the length of the fusion protein minus (a) is less than 500 amino acids. In some of any embodiments, the length of the fusion protein minus (a) is less than 370 amino acids. In some of any embodiments, the fusion protein is devoid of any domains with DNA methyltransferase activity and domains capable of recruiting heterochromatin inducing factors. In some of any embodiments, the fusion protein is devoid of any H3K4meO peptides.
[0009] Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA-binding domain for targeting to a target site for PCSK9 and an effector domain that is a single effector domain, wherein the single effector domain comprises a catalytically inactive DNA methyltransferase domain or a functional portion thereof, and wherein the single effector domain is less than 600 amino acids in length. Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein consists essentially of: (a) a DNA-binding domain for targeting to a target site for PCSK9, and (b) an effector domain, wherein the effector domain comprises a catalytically inactive DNA methyltransferase domain or a functional portion thereof, and wherein the effector domain is less than 600 amino acids in length.Attorney Docket No. 224742003640
[0010] In some of any embodiments, the catalytically inactive DNA methyltransferase domain or portion thereof comprises a DNMT3L protein or portion thereof that recruits domains with DNA methyltransferase activity to the target site for PCSK9.
[0011] In some of any embodiments, the length of the effector domain is less than 510 amino acids. In some of any embodiments, the length of the effector domain is less than 400 amino acids.
[0012] In some of any embodiments, the length of the effector domain is less than 300 amino acids. In some of any embodiments, the effector domain is devoid of any domains with DNA methyltransferase activity and domains capable of recruiting heterochromatin inducing factors. In some of any embodiments, the fusion protein is devoid of any H3K4meO peptides.
[0013] In some of any embodiments, the DNMT3L protein or portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemesfrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some of any embodiments, the DNMT3L protein or portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus.
[0014] In some of any embodiments, the portion of the DNMT3L protein is a contiguous portion that is less than a full-length DNMT3L MTase-like domain and comprises at least 10 amino acids from a reference DNMT3L MTase-like domain, wherein the contiguous portion of at least 10 amino acids is involved in a DNMT3A-DNMT3L interface. In some of any embodiments, the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5 (SEQ ID NO: 126), wherein Xi is L or M, X2is L or I, X3is L or R, X4 is K or R, and X5 is P or Q. In some of any embodiments, the contiguous portion comprises the sequence set forth in X1DX2X3X4X5X6RFLX7 (SEQ ID NO: 127), wherein XI is E or D, X2 is L or Q, X3 is D, E, or M, X4 is V or T, X5 is A or T, X6 is S, T, or V, and X7 is E or Q. In some of any embodiments, the contiguous portion comprises the sequence set forth as amino acid residues 257-273 or 292-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:128. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:129. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:130. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:131. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:132. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO:133.
[0015] In some of any embodiments, the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5X6SX7X8PFFWX9FX10DNLX11LX12X13X14DX15X16X17X18X19RFLX20Attorney Docket No. 224742003640(SEQ ID NO: 134), wherein Xi is L or M, X2 is L or I, X3 is L or R, X4 is K or R, X5 is P or Q, Xe is G or E, X7 is P, Q, or absent, Xs is R or Q, X9 is M or I, Xw is V or M, Xu is V or L, X12 is N or T, X13 is K or E, X14 is E or D, X15 is L or Q, Xie is D, E, or M, X17 is V or T, Xis is A or T, X19 is S, T, or V, and X20 is E or Q. In some of any embodiments, the contiguous portion comprises the sequence set forth as amino acid residues 257-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 135. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 136. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 137.
[0016] In some of any embodiments, the contiguous portion comprises the sequence set forth in XiX2VRX3DVEX4WGPFDLX5YGX6TX7PLGX8X9CDRXioPXiiWYXi2FQFHRXi3LQYAXi4PXi5Xi6Xi7SX 18X19PFFWX2OFX21DNLX22LX23X24X25DX26X27X28X29X3ORFLX31 (SEQ ID NO: 138), wherein XI is D or N, X2 is T or V, X3 is K or R, X4 is E or K, X5 is V or L, X6 is A or S, X7 is P or Q, X8 is H or S, X9 is T or S, X10 is P or C, XI 1 is S or G, X12 is L or M, X13 is L or I, X14 is L or R, X15 is K or R, Xie is P or Q, X17 is G or E, Xi8is P, Q, or absent, X19 is R or Q, X20 is M or I, X21 is V or M, X22 is V or L, X23 is N or T, X24 is K or E, X25 is E or D, X26 is L or Q, X27 is D, E, or M, X28is V or T, X29 is A or T, X30 is S, T, or V, and X31 is E or Q. In some of any embodiments, the contiguous portion comprises the sequence set forth as amino acid residues 225-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 139. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 140. In some of any embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 141.
[0017] In some of any embodiments, the contiguous portion comprises at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 75 amino acids.
[0018] In some of any embodiments, the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some of any embodiments, the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17. In some of any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some of any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15. In some ofAttorney Docket No. 224742003640 any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some of any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16. In some of any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some of any embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17.
[0019] In some of any embodiments, the DNMT3L protein or portion thereof is a DNMT3L MTase-like domain. In some of any embodiments, the DNMT3L protein or portion thereof is a portion of a DNMT3L MTase-like domain.
[0020] In some of any embodiments, the DNMT3L protein or portion thereof comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17.
[0021] In some of any embodiments, the effector domain further comprises a DNMT3L ADD domain. In some of any embodiments, the effector domain, from N-terminus to C-terminus, comprises the DNMT3L ADD domain and the DNMT3L MTase-like domain. In some of any embodiments, the effector domain, from N-terminus to C-terminus, comprises the DNMT3L MTase-like domain and the DNMT3L ADD domain. In some of any embodiments, the DNMT3L ADD domain comprises the sequence set forth in any one of SEQ ID NOs: 109, 110, 379, and 380. In some of any embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 379. In some of any embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 380. In some of any embodiments, the DNMT3LAttorney Docket No. 224742003640ADD domain comprises the sequence set forth in SEQ ID NO: 109 or 110. In some of any embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109. In some of any embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 110.
[0022] In some of any embodiments, the effector domain comprises the sequence set forth in any one of SEQ ID NO: 378, 381, or 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in any one of SEQ ID NO: 378, 381, or 382. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 378, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 378. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 381, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 381. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 382.
[0023] In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 107, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 107. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 108, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 108.
[0024] Provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a reference DNMT3A sequence at a position selected from among 686, 707-721, 756, 766, 771, 831-848, 854, 855, 860, 873, 876, 879, and 881-887, corresponding to numbering of positions set forth in SEQ ID NO: 183. In some of any of such embodiments, the one or more amino acid substitutions are within a DNA binding region, wherein the DNA binding region is a catalytic loop, a target recognition domain (TRD), and a homodimer interface, or a combination thereof.Attorney Docket No. 224742003640
[0025] Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a position of a DNA binding region of a reference DNMT3A sequence, wherein the DNA binding region is a catalytic loop, a target recognition domain (TRD), and a homodimer interface, or a combination thereof.
[0026] In some of any of such embodiments, the reference DNMT3A sequence comprises the sequence of amino acids set forth in SEQ ID NO: 183 or a functionally active portion thereof. In some of any of such embodiments, the functionally active portion thereof comprises a contiguous sequence contained within amino acid residues 612-912, with reference to positions set forth in SEQ ID NO: 183. In some of any of such embodiments, the functionally active portion thereof comprises a methyltransferase (MTase) domain. In some of any of such embodiments, the reference DNMT3A sequence comprises the sequence set forth in SEQ ID NO: 113. In some of any of such embodiments, the reference DNMT3A sequence consists of the sequence set forth in SEQ ID NO: 113.
[0027] In some of any of such embodiments, the catalytic loop is set forth as amino acid residues 707- 721 corresponding to numbering of positions set forth in SEQ ID NO: 183. In some of any of such embodiments, the one or more amino acid substitutions are at a position 707-721.
[0028] In some of any of such embodiments, the TRD is set forth as amino acid residues 831-848 corresponding to numbering of positions set forth in SEQ ID NO: 183. In some of any of such embodiments, one or more amino acid substitutions are at a position 831-848.
[0029] In some of any of such embodiments, the homodimer interface is set forth as amino acid residues 873-887 corresponding to numbering of positions set forth in SEQ ID NO: 183. In some of any of such embodiments, one or more amino acid substitutions are at a position 873-887.
[0030] In some of any of such embodiments, one or more amino acid substitutions are at a position selected from among 686, 711, 714, 756, 766, 771, 831, 832, 835, 836, 838, 841, 844, 845, 847, 854, 855, 860, 873, 876, 879, 881, 882, 883, and 887. In some of any of such embodiments, the one or more amino acid substitutions are selected from D686A, N711A, S714A, E756A, K766E, R771Q, R831A, R831E, T832A, T835A, T835E, R836E, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, W860A, H873A, D876G, N879A, S881A, R882H, L883A, R887A, and R887E, or a conservative amino acid substitution thereof.
[0031] In some of any of such embodiments, the one or more amino acid substitutions are selected from D686A, N711A, S714A, K766E, R771Q, R831A, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, D876G, N879A, L883A, and R887A, or any combination thereof.Attorney Docket No. 224742003640
[0032] In some of any of such embodiments, the one or more amino acid substitutions are 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions.
[0033] In some of any of such embodiments, the one or more amino acid substitutions are a single amino acid substitution selected from D686A, N711A, S714A, K766E, R771Q, R831A, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, D876G, N879A, L883A, and R887A. In some of any of such embodiments, the one or more amino acid substitutions are two amino acid substitutions selected from D686A, N711A, S714A, K766E, R771Q, R831A, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, D876G, N879A, L883A, and R887A.
[0034] In some of any of such embodiments, at least one amino acid substitution is selected from K841A, D845K, N711A, R887A, D686A, R831A, K766E, K855E, and N879A. In some of any of such embodiments, at least one further amino acid substitution is selected from D686A, N711A, S714A, E756A, K766E, R771Q, R831A, R831E, T832A, T835A, T835E, R836E, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, W860A, H873A, D876G, N879A, S881A, R882H, L883A, R887A, and R887E, or a conservative amino acid substitution thereof.
[0035] In some of any of such embodiments, the one or more amino acid substitutions are selected from K841A, D845K, N711A, R887A, D686A, R831A, K766E, K855E, and N879A, or a combination thereof. In some of any of such embodiments, one or more amino acid substitutions are two amino acid substitutions selected from K841A, D845K, N711A, R887A, D686A, R831A, K766E, K855E, and N879A.
[0036] In some of any of such embodiments, at least one amino acid substitution is D686A. In some of any of such embodiments, at least one amino acid substitution is N711A. In some of any of such embodiments, at least one amino acid substitution is K766E. In some of any of such embodiments, at least one amino acid substitution is R831A. In some of any of such embodiments, at least one amino acid substitution is K841A. In some of any of such embodiments, at least one amino acid substitution is D845K. In some of any of such embodiments, at least one amino acid substitution is K855E. In some of any of such embodiments, at least one amino acid substitution is N879A. In some of any of such embodiments, at least one amino acid substitution is R887A.
[0037] In some of any of such embodiments, the variant DNMT3A domain or functionally active portion thereof comprises the sequence set forth in any one of SEQ ID NOS: 184-212.
[0038] In some of any of such embodiments, at least one amino acid substitution is selected from S714A, D845K, N711A, E854H, N838A, R771Q, R831A, R887A, L883A, S881A, and H847E. In some of any of such embodiments, at least one further amino acid substitution is selected from E756A, T835E, W860A, R831E, R887E, K844E, R836E, R882H, T832A, K841E, T835A, D876G, N879A, K841A, K766E, K855E, D686A, and H837A.Attorney Docket No. 224742003640
[0039] In some of any of such embodiments, the one or more amino acid substitutions are two amino acid substitutions selected from D686A, K855E, K766E, K841A, N879A, R887A, R831A, R771Q, N838A, N711A, D845K, and S714A.
[0040] Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a reference DNMT3A sequence at positions selected from among 613, 621, 630, 631, 632, 635, 651, 659, 676, 677, 680, 688, 693, 694, 720, 721, 729, 736, 739, 742, 744, 749, 766, 767, 771, 783, 789, 790, 792, 803, 812, 821, 826, 823, 829, 831, 836, 841, 844, 847, 855, 866, 873, 882, 885, 887, 891, 899, 900, and 906, corresponding to numbering of position set forth in SEQ ID NO: 183, wherein the substituted amino acid is alanine (A), aspartic acid (D), glutamic acid (E), glutamine (Q), or asparagine (N).
[0041] In some of any of such embodiments, the effector domain represses transcription of one or more target genes.
[0042] In some of any of such embodiments, the effector domain is a multipartite effector that further comprises one or more domains selected from a DNA methyltransferase domain, a repressor domain capable of recruiting heterochromatin inducing factors, or combinations thereof. In some of any of such embodiments, the repressor domain is selected from a KRAB repressor domain, ERF repressor domain, Mxil repressor domain, SID4X repressor domain, Mad-SID repressor domain, LSD1 repressor domain, EZH2 repressor domain, or variant of any of the foregoing. In some of any of such embodiments, the repressor domain is selected from a KRAB repressor domain, ERF repressor domain, Mxil repressor domain, SID4X repressor domain, Mad-SID repressor domain, LSD1 repressor domain, EZH2 repressor domain, or variant of any of the foregoing comprising the sequence set forth in any one of SEQ ID NOS: 19 and 267-276, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some of any of such embodiments, the repressor domain is selected from a KRAB domain or an EZH2 domain.
[0043] In some of any of such embodiments, the DNA methyltransferase domain has DNA methyltransferase activity or is a catalytically inactive regulatory factor of DNA methyltransferases to inhibit transcription or increase the duration of inhibition of one or more target genes. In some of any of such embodiments, the DNA methyltransferase domain is a DNMT3L domain.
[0044] In some of any of such embodiments, the effector domain is a multipartite effector that comprises the variant DNMT3A domain; and a KRAB domain or an EZH2 domain. In some of any of such embodiments, the effector domain is a multipartite effector that comprises the variant DNMT3A domain; aAttorney Docket No. 224742003640DNMT3L domain; and a KRAB domain or an EZH2 domain. In some of any of such embodiments, the fusion protein comprises from N-terminus to C-terminus: (i) the variant DNMT3A domain, (ii) the DNA- binding domain, and (iii) a KRAB domain or an EZH2 domain. In some of any of such embodiments, the fusion protein comprises from N-terminus to C-terminus: (i) the variant DNMT3A domain, (ii) a DNMT3L domain, (iii) the DNA-binding domain, and (iv) a KRAB domain or an EZH2 domain.
[0045] In some of any of such embodiments, the KRAB domain is selected from: a KRAB domain from KOX1, a KRAB domain from ZIM3, and a KRAB domain from ZNF324. In some of any of such embodiments, the KRAB domain comprises the sequence set forth in any one of SEQ ID NOS: 19 and 268- 270, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0046] In some of any of such embodiments, the EZH2 domain comprises the sequence set forth in SEQ ID NO: 271, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
[0047] In some of any of such embodiments, the DNMT3L domain comprises an DNMT3L MTase-like domain. In some of any of such embodiments, the DNMT3L domain comprises any one of the sequences set forth in SEQ ID NO: 15-17, a portion thereof, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 15. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 16. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 17.
[0048] In some of any of such embodiments, the DNMT3L domain further comprises an DNMT3L ADD domain. In some of any of such embodiments, the DNMT3L domain comprises, from N terminus to C terminus, the DNMT3L ADD domain and the DNMT3L MTase-like domain. In some of any of such embodiments, the DNMT3L domain comprises, from N terminus to C terminus, the DNMT3L MTase-like domain and the DNMT3L ADD domain. In some of any of such embodiments, the ADD domain comprises the sequence set forth in SEQ ID NO: 109, or an amino acid sequence with at least 90%, 91%, 92%, 93%,Attorney Docket No. 22474200364094%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the ADD comprises the sequence set forth in SEQ ID NO: 109.
[0049] In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 116, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 116.
[0050] In some of any of such embodiments, the fusion protein comprises a variant DNMT3A- DNMT3L (DNMT3A / L) fusion protein. In some of any of such embodiments, the variant DNMT3A / L fusion protein comprises a reference DNMT3A / L fusion sequence of amino acids set forth in SEQ ID NO: 264 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the reference DNMT3A / L fiision sequence of amino acids is set forth in SEQ ID NO: 264, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the reference DNMT3A / L fusion sequence of amino acids is set forth in SEQ ID NO: 264. In some of any of such embodiments, the reference DNMT3A / L fusion sequence of amino acids is set forth in SEQ ID NO: 265, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the reference DNMT3A / L fusion sequence of amino acids is set forth in SEQ ID NO: 265.
[0051] In some of any embodiments, the DNA-binding domain is selected from: a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or a variant thereof; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an LScel enzyme or a variant thereof. In some of any of such embodiments, the DNA-binding domain is a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or variant thereof. In some of any of such embodiments, the Cas protein or variant thereof is capable of complexing with a guide RNA (gRNA).
[0052] In some of any embodiments, the target site for PCSK9 is located within 500 bp of the genomic coordinate chrl:55,039,548. In some of any embodiments, the target site for PCSK9 is located within 300 bp of the genomic coordinate chr 1 :55,039,548. In some of any embodiments, the target site for PCSK9 is located within 200 bp of the genomic coordinate chr 1:55, 039, 548. In some of any embodiments, the target site for PCSK9 is located within 100 bp of the genomic coordinate chrl:55,039,548. In some of any embodiments, the target site for PCSK9 is located within 80 bp of the genomic coordinate chrl:55,039,548. In some of any embodiments, the target site for PCSK9 is located within 80 bp of the genomic coordinate chrl:55,039,548. In some of any embodiments, the target site for PCSK9 is within the coordinates chrl:Attorney Docket No. 22474200364055,039,338-55,039,658. In some of any embodiments, the target site for PCSK9 is within the coordinates chrl: 55,039,470-55,039,597. In some of any embodiments, the target site for PCSK9 is or comprises the coordinates chrl: 55,039,538-55,039,557. In some of any embodiments, the target site comprises the sequence set forth in any one of SEQ ID NOs: 39 and 142-144, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing. In some of any embodiments, the target site comprises the sequence set forth in SEQ ID NO: 39, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing. In some of any embodiments, the target site comprises the sequence set forth in SEQ ID NO: 142, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing. In some of any embodiments, the target site comprises the sequence set forth in SEQ ID NO: 143, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing. In some of any embodiments, the target site comprises the sequence set forth in SEQ ID NO: 144, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
[0053] Also provided herein is a fusion protein for targeted gene repression, wherein the fiision protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, wherein the target site for PCSK9 is located within 300 bp of the genomic coordinate chrl:55,039,548, and (b) an effector domain comprising a DNMT3L protein or portion thereof comprising the sequence set forth in any one of SEQ ID NOs: 15-17 or an amino acid sequence that has at least 90% identity to any one of SEQ ID NOs: 15-17, and wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides. Also provided herein is a fusion protein for targeted gene repression, wherein the fusion protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, wherein the target site for PCSK9 is within the coordinates chrl: 55,039,470-55,039,597, and (b) an effector domain comprising a DNMT3L protein or portion thereof comprising the sequence set forth in any one of SEQ ID NOs: 15-17 or an amino acid sequence that has at least 90% identity to any one of SEQ ID NOs: 15-17, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence setAttorney Docket No. 224742003640 forth in SEQ ID NO: 16. In some of any embodiments, the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17.
[0054] In some of any embodiments, the effector domain is independently fused to the N-terminus, the C-terminus, or both the N-terminus and the C-terminus, of the DNA-binding domain. In some of any embodiments, the fusion protein comprises, from N-terminus to C-terminus, the DNMT3L protein or portion thereof and the DNA-binding domain. In some of any embodiments, the fusion protein comprises, from N- terminus to C-terminus, the DNA-binding domain and the DNMT3L protein or portion thereof.
[0055] In some of any embodiments, the DNA-binding domain is a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or variant thereof, and the Cas protein is capable of complexing with a gRNA for targeting the DNA-binding domain to the target site for PCSK9. In some of any embodiments, the Cas protein or variant thereof is a deactivated (dCas) protein. In some of any embodiments, the dCas protein lacks nuclease activity. In some of any embodiments, the dCas protein is a dCasl2 protein. In some of any embodiments, the dCas protein is a dCas9 protein.
[0056] In some of any embodiments, the dCas9 protein is a Staphylococcus aureus dCas9 (dSaCas9) protein. In some of any embodiments, the dSaCas9 comprises at least one amino acid mutation selected from D10A and N580A, with reference to numbering of positions of SEQ ID NO: 41. In some of any embodiments, the dSaCas9 protein comprises the sequence set forth in SEQ ID NO: 40, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the dSaCas9 is set forth in SEQ ID NO: 40.
[0057] In some of any embodiments, the dCas9 protein is a Streptococcus pyogenes dCas9 (dSpCas9) protein. In some of any embodiments, the dSpCas9 protein comprises at least one amino acid mutation selected from D10A and H840A, with reference to numbering of positions of SEQ ID NO: 42. In some of any embodiments, the dSpCas9 comprises the sequence set forth in SEQ ID NO: 18, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the dSpCas9 is set forth in SEQ ID NO: 18.
[0058] In some of any embodiments, the gRNA comprises a gRNA spacer sequence comprising the sequence set forth in SEQ ID NO: 43.
[0059] In some of any embodiments, the DNA-binding domain is a zinc finger protein (ZFP).
[0060] In some of any embodiments, the DNA-binding domain is a zinc finger protein (ZFP), wherein the ZFP binds to a target site in a gene or a regulatory DNA element thereof.Attorney Docket No. 224742003640
[0061] In some of any embodiments, the fusion protein further comprises a nuclear localization signal (NLS). In some of any embodiments, the NLS is a nucleoplasmin NLS or an SV40 NLS. In some of any embodiments, the NLS comprises the sequence set forth in SEQ ID NO: 44 or 34.
[0062] In some of any embodiments, the fusion protein comprises a linker. In some of any embodiments, the linker comprises the sequence set forth in any one of SEQ ID NOS: 33, 45-47, and 120, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the linker comprises the sequence set forth in any one of SEQ ID NOS: 33 and 45-47, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the linker comprises the sequence set forth in SEQ ID NO: 33, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the linker comprises the sequence set forth in SEQ ID NO: 33.
[0063] In some of any embodiments, the linker comprises the sequence set forth in SEQ ID NO: 120, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the linker comprises the sequence set forth in SEQ ID NO: 120.
[0064] In some of any embodiments, the fusion protein comprises the sequence set forth in any of SEQ ID NOs: 1-3, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 1 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 1. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 2 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 2. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 3 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 3.
[0065] In some of any embodiments, the fusion protein comprises the sequence set forth in any one of SEQ ID NOS: 213-241, 243, 245-261, 277, and 278, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in any one of SEQ ID NOS: 213-215, 217-219, 225-232, 235, 236, 239, and 240, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in any one of SEQ ID NOS: 213, 214, 217, 219, 226, 229, 232, 236, and 240, or a sequence havingAttorney Docket No. 224742003640 at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 236, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 236. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: Til, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 277. In some of any of such embodiments, the fusion protein comprises the sequence set forth in any one of SEQ ID NOs: 243, 245-261, and 278, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in any one of SEQ ID NOs: 243, 245, 247, 251, 254, 257, 259, 261, and 278, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 259, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 259. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 278, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some of any of such embodiments, the fusion protein comprises the sequence set forth in SEQ ID NO: 278.
[0066] Provided herein is a DNA-targeting system for gene repression, comprising: a) a fusion protein provided herein; and b) one guide RNA (gRNA) that targets the DNA-binding domain to the target site for PCSK9.
[0067] Provided herein is a DNA-targeting system for gene repression, comprising: a) the fusion protein as disclosed herein; and b) at least one guide RNA (gRNA) that targets the DNA-binding domain to a target site in a gene or a regulatory DNA element thereof.
[0068] Also provided herein is a DNA-targeting system for gene repression, comprising: a) the fusion protein as disclosed herein; and b) two or more different gRNAs, wherein each gRNA targets the DNA- binding domain to a different target site in a gene or a regulatory DNA element thereof. In some of any of such embodiments, any two of the different target sites are in the same gene or regulatory DNA element thereof, or in different genes or regulatory DNA elements thereof.
[0069] Provided herein is a polynucleotide encoding a fusion protein provided herein. Also provided herein is polynucleotide encoding the DNA-targeting system provided herein.Attorney Docket No. 224742003640
[0070] Provided herein is a vector comprising a fusion protein provided herein, a DNA-targeting system provided herein, or a polynucleotide provided herein. In some of any embodiments, the vector is a viral vector. In some of any embodiments, the viral vector is an adeno-associated virus (AAV) vector. In some of any embodiments, the vector is a non-viral vector. In some of any embodiments, the non-viral vector is selected from: a lipid nanoparticle, a liposome, an exosome, or a cell penetrating peptide.
[0071] Provided herein is a lipid nanoparticle comprising a fusion protein provided herein, a DNA- targeting system provided herein, or a polynucleotide provided herein.
[0072] Provided herein is a method of targeted gene repression, comprising introducing into a cell: the fusion protein disclosed herein, the DNA-targeting system disclosed herein, the polynucleotide disclosed herein, the vector disclosed herein, or the lipid nanoparticle disclosed herein. In some of any of such embodiments, the cell is a cell from and / or in a subject. In some of any of such embodiments, the introducing is by transient delivery into the cell. In some of any of such embodiments, the transient delivery comprises electroporation, transfection, or transduction. In some of any of such embodiments, the fusion protein, the DNA-targeting system, and / or the polynucleotide is transiently expressed and / or transiently present in the cell for a period of time after the introducing. In some of any of such embodiments, the gene repression is a reduction in the expression of the gene. In some of any of such embodiments, the gene repression is a reduction in the transcription of the gene. In some of any of such embodiments, the gene repression is sustained. In some of any of such embodiments, the gene repression is sustained until after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any of such embodiments, the gene repression is sustained for at least 1 day, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at least 3 months, at least 6 months, or at least 1 year, or more, after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any of such embodiments, the gene repression is sustained at a fold-change of 1.0, 0.9, 0.8, 0.7, 0.6, 0.5, 0.4, 0.3, or less relative to the use of a DNA-targeting system using the same fusion protein but with a non-targeting gRNA, after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any of such embodiments, the off-target methylation activity is reduced by more than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%, as compared to the use of a fusion protein comprising the sequence of amino acids set forth in SEQ ID NO: 14.
[0073] Provided herein is a method of targeted PCSK9 gene repression, comprising introducing into a cell comprising a domain with DNA methyltransferase activity: a fusion protein provided herein, a DNA- targeting system provided herein, a polynucleotide provided herein, a vector provided herein, or a lipid nanoparticle provided herein. In some of any embodiments, the cell is a cell from and / or in a subject. In someAttorney Docket No. 224742003640 of any embodiments, the introducing is by transient delivery into the cell. In some of any embodiments, the transient delivery comprises electroporation, transfection, or transduction. In some of any embodiments, the fusion protein, the DNA-targeting system, and / or the polynucleotide is transiently expressed and / or transiently present in the cell for a period of time after the introducing. In some of any embodiments, the PCSK9 gene repression is a reduction in the expression of a PCSK9 gene. In some of any embodiments, the PCSK9 gene repression is a reduction in the transcription of a PCSK9 gene. In some of any embodiments, the PCSK9 gene repression is sustained. In some of any embodiments, the PCSK9 gene repression is sustained until after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any embodiments, the PCSK9 gene repression is sustained for at least 1 day, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at least 3 months, at least 6 months, or at least 1 year, or more, after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any embodiments, the PCSK9 gene repression is sustained at a fold-change of 0.2 or less relative to not introducing the fusion protein, DNA-targeting system, polynucleotide, vector, or lipid nanoparticle into the cell, after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell. In some of any embodiments, off-target methylation activity is reduced by more than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%, as compared to the use of a fusion protein comprising the sequence of amino acids set forth in SEQ ID NO: 14.Brief Description of the Drawings
[0074] FIG. 1A shows on-target repression of an exemplary Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, or fold-change in HBV mRNA relative to a TBP housekeeping control and untreated cells, over 42 days in Hep3B cells transfected with a gRNA targeting HBV and 1.3 pg / mL mRNA encoding a reference fusion protein (DNMT3A-mDNMT3L-dSpCas9-KRAB) or a fusion protein designed to recruit endogenous MTases containing: (a) a mouse DNMT3L (mDNMT3L) MTase-like domain, (b) a mDNMT3L MTase-like domain and a KRAB domain, or (c) a mDNMT3L MTase-like domain, a KRAB domain, and a H3K4meO synthetic peptide. Also depicted is results for repression following transfection with a catalytically dead fusion protein (C710A) or when cells were not transfected (untreated).
[0075] FIG. IB shows on-target repression of an exemplary Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, or mean methylation in HBV CpG Island 2 (CGI2) , over day 7 post-transfection in Hep3B cells transfected with a gRNA targeting HBV and mRNA encoding a referenceAttorney Docket No. 224742003640 fusion protein (DNMT3A-mDNMT3L-dSpCas9-KRAB) or a fusion protein designed to recruit endogenous MTases containing: (a) mDNMT3L MTase-like domain and a KRAB domain, or (b) a mDNMT3L MTase- like domain, a KRAB domain, and a H3K4meO synthetic peptide at two different doses. Also depicted is results for repression following transfection with a catalytically dead fusion protein (C710A) or when cells were not transfected (untreated).
[0076] FIG. 2 shows on-target repression of an exemplary Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, or fold-change in HBV mRNA relative to a TBP housekeeping control and untreated cells, over 30 days in Hep3B cells transfected with 0.94 pg / mL mRNA encoding a reference fusion proteins DNMT3A-mDNMT3L-ZFP-KRAB or DNMT3A-ZFP-KRAB, or a fusion protein designed to recruit endogenous MTases (mDNMT3L-ZFP-KRAB), where the ZFP targets the exemplary HBV gene. Also depicted is results for repression when cells were not transfected (untreated).
[0077] FIG. 3A and 3B show results following methylation sequencing in Hep3B cells that demonstrate levels of off-target methylation as mean percent methylation across all regions captured by the off-target custom panel (FIG. 3A) and number of CpGs for each treatment that are differentially methylated compared to the untreated control (FIG. 3B) 7 days post-transfection of two different doses of tested fusion proteins. Tested fusion proteins contained: (a) mDNMT3L MTase-like domain and a KRAB domain or (b) a mDNMT3L MTase-like domain, a KRAB domain, and a H3K4meO synthetic peptide, as compared to a reference fusion protein (DNMT3A-mDNMT3L-dSpCas9-KRAB) and negative controls, including untreated cells and a catalytically dead fusion protein (C710A).
[0078] FIG. 4 shows on-target repression of an exemplary PCSK9 gene, or fold-change in PCSK9 mRNA relative to a TBP housekeeping control and untreated cells, over 21 days in Huh7 cells transfected with a gRNA targeting HBV and 0.94 pg / mL mRNA encoding a reference fusion protein (DNMT3A- mDNMT3L-dSpCas9-KRAB) or one of several tested fusion proteins designed to recruit endogenous MTases. Tested fusion proteins contained a DNMT3L domain from one of three species: human (FIG. 4, left), mouse (FIG. 4, middle), and Apodemus (FIG. 4, right) and are set forth in Table El. Also depicted is results for repression using negative controls, such as untreated cells, a dSpCas9-only construct, or a dSpCas9-KRAB construct.
[0079] FIG. 5 shows on-target repression of an exemplary PCSK9 gene expressed as fold-change in PCSK9 mRNA at day 14 post-transfection (FIG. 5, left) or at day 42 post-transfection (FIG. 5, right) in Hep3B cells transfected with a gRNA targeting PCSK9 and mRNA at a dose of 0.94 pg / mL encoding one of two reference fusion protein containing DNMT3A and mouse DNMT3L (mDNMT3L) (DNMT3A- mDNMT3L-dSpCas9-KRAB or DNMT3A-mDNMT3L-dSpCas9), a reference DNMT3A-mDNMT3L-18aa- dSpCas9-KRAB as set forth in SEQ ID NO: 112, or one of several tested fusion proteins designed to recruitAttorney Docket No. 224742003640 endogenous MTases as set forth in Table El. Also depicted is results for repression using negative controls, such as untreated cells or a dSpCas9-KRAB construct.
[0080] FIG. 6 shows on-target repression of an exemplary HBV gene expressed as fold-change in HBV mRNA at day 14 post-transfection (FIG. 6, left) or at day 42 post-transfection (FIG. 6, right) in Hep3B cells transfected with a gRNA targeting HBV and mRNA at a dose of 0.94 pg / mL encoding one of two reference fusion protein containing DNMT3A and mouse DNMT3L (mDNMT3L) (DNMT3A-mDNMT3L-dSpCas9- KRAB or DNMT3A-mDNMT3L-dSpCas9), a reference DNMT3A-mDNMT3L-18aa-dSpCas9-KRAB as set forth in SEQ ID NO: 112, or one of several tested fusion proteins designed to recruit endogenous MTases as set forth in Table El. Also depicted is results for repression using negative controls, such as untreated cells or a dSpCas9-KRAB construct.
[0081] FIG. 7A shows the fold change in PCSK9 mRNA expression in Huh7 cells relative to the untreated cells at Day 34 post-transfection following lipofection of the indicated fusion proteins from Table El in combination with an exemplary PCSK9-1 gRNA.
[0082] FIG. 7B shows the fold change in PCSK9 mRNA expression in Hep3B cells relative to the untreated cells at Day 34 post-transfection following lipofection of the indicated fusion proteins from Table El in combination with an exemplary PCSK9-1 gRNA.
[0083] FIG. 8 shows the number of differentially methylated loci (CpGs) from the custom off-target panel in Huh7 cells for each fusion protein at Day 7 and Day 45 post-transfection following transfection of either ApDNMT3L-dSpCas9 (As3L) or DNMT3A-mDNMT3L-dSpCas9-KRAB (D3A-m3L-KOXl) positive control fusion proteins in combination with exemplary PCSK9-1 gRNA.
[0084] FIG. 9A shows schematics of three tested dCas-9 effector fusion proteins.
[0085] FIG. 9B shows a dose response curve for fold change in PSCK9 mRNA relative to the lipid only controls for each fusion protein shown in FIG. 9A in combination with an exemplary PCSK9-1 gRNA.
[0086] FIG. 10A shows the positions of the exemplary PCSK9-targeting gRNAs along the CpG island relative to the transcriptional start site (TSS) and Exon 1 of PCSK9.
[0087] FIG. 10B shows the fold change in PCSK9 mRNA expression in Huh7 cells relative to the untreated cells at Day 4 post-transfection following transfection with the indicated fusion proteins in combination with an exemplary PCSK9-targeting gRNA or different negative controls.
[0088] FIG. 10C shows the fold change in PCSK9 mRNA expression in Huh7 cells relative to the nontargeting gRNA (gNT) with the negative control fusion protein at Day 28 post-transfection following transfection with the indicated fusion proteins in combination with an exemplary PCSK9-targeting gRNA or different negative controls.
[0089] FIG. 10D shows the fold change in PCSK9 mRNA expression in Huh7 cells relative to the nonAttorney Docket No. 224742003640 targeting gRNA (gNT) with the negative control fusion protein over time for each indicated fusion protein in combination with guides PCSK9-1 (left), PCSK9-2 (middle), and PCSK9-4 (right).
[0090] FIG. 11 shows durable repression of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, at day 4 (left panel) and day 14 (right panel) in Hep3B cells transfected with a gRNA targeting HBV and mRNA encoding a fusion protein containing one of 30 variant DNMT3A domains fused to a DNA binding protein and a KRAB domain. Also depicted is results for repression following transfection with control constructs not containing a DNMT3A domain or containing an unmodified (wild-type) DNMT3A. Durable repression was measured as fold-change in HBV mRNA.
[0091] FIG. 12 shows results following capture methylation sequencing in HepG2.NTCP cells that demonstrate how a subset of fusion proteins containing a variant DNMT3A domain fused to a DNA binding protein and a KRAB domain show decreased numbers of differentially methylated loci (DML) as compared to a fusion protein containing a wild-type (WT) DNMT3A domain.
[0092] FIG. 13 depicts the level of durable repression of hepatitis B virus (HBV), i.e., on-target methylation activity, in gray and off-target methylation activity in black achieved in different fusion proteins containing one of various variant DNMT3A domains (i.e., DNMT3A mutants) fused to a DNA binding protein and a KRAB domain.
[0093] FIG. 14 depicts the level of durable repression of hepatitis B virus (HBV), i.e., on-target methylation activity at Day 4 post-transfection (left), Day 21 post-transfection (middle), and Day 28 posttransfection (right) in Hep3B cells following lipofectamine-based delivery of HBV-targeting gRNA and 1.3 pg / mL mRNA encoding one of 18 fusion proteins set forth in Table E4. As controls, cells were also left untreated or delivered mRNA encoding one of several indicated control fusion and either HBV-targeting gRNA (gHBV) or non-targeting gRNA (gNT). Durable repression was measured as fold-change in HBV mRNA.
[0094] FIG. 15 depicts the level of durable repression of hepatitis B virus (HBV), i.e., on-target methylation activity, in dark gray (right-hand axis) from Day 21 post-transfection in Hep3B cells and off- target methylation activity in light gray (left-hand axis) from Day 2 post-transfection in HepG2 cells achieved in one of 18 different fusion proteins and a reference, i.e., wild-type (WT), fusion protein set forth in Table E4.
[0095] FIGs. 16A-16H show on-target repression over 35 days (top) in Hep3B cells delivered 1.3 pg / mL mRNA as compared to off-target methylation activity at Day 2 post-transfection (bottom) in Hep3G cells after delivery of one of the 8 lead fusion protein candidates set forth in Table E4: N879A (FIG. 16A), R855E (FIG. 16B), R831A (FIG. 16C), N711A (FIG. 16D), D845K (FIG. 16E), K766E (FIG. 16F),Attorney Docket No. 224742003640R887A (FIG. 16G), and N838A (FIG. 16H). Also depicted is the on- and off-target activity in cells transfected with a reference fusion protein (DNMT3A-3L-dSpCas9-KRAB, or WT; SEQ ID NO: 14) or a catalytically dead fusion protein (C710A; SEQ ID NO: 23). For on-target activity (top), additional controls included untreated cells and transfection with a dSpCas9-KRAB construct.
[0096] FIGs. 17A-17C show methylation data in Hep3B cells left untreated or transfected with 1.3 pg / mL mRNA encoding a fusion protein containing anN879A DNMT3A variant, a catalytically dead DNMT3A variant (C710A), or a reference wild-type DNMT3A variant (WT). Methylation data includes on- target methylation data generated by qPCR over 42 days (FIG. 17A), on-target methylation data from Day 7 post-transfection generated by methyl sequencing (FIG. 17B), and off-target methylation data from Day 7 post-transfection (FIG. 17C). For on-target methylation data (FIG. 17A), cells were also transfected with mRNA encoding a fusion protein containing only dSpCas9 and a KRAB domain as an additional control.
[0097] FIG. 18 shows durable repression of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, at day 4 (left panel) and day 29 (right panel) in Hep3B cells transfected with mRNA encoding a fusion protein containing one of 18 variant DNMT3A domains fused to an HBV-targeting zinc finger protein (ZFP) and a KRAB domain. Also depicted is results for repression following no transfection (untreated) or transfection with control constructs that contain an unmodified (wild-type) DNMT3A, a catalytically dead DNMT3A variant (C710A), and / or a non-targeting ZFP (NT-ZFP). Durable repression was measured as fold-change in HBV mRNA.
[0098] FIGs. 19A-19C show methylation data in Hep3B cells left untreated or transfected with 0.94 pg / mL mRNA encoding an HBV-targeting ZFP fusion protein containing an N879A DNMT3A variant (N879A), a catalytically dead DNMT3A variant (C710A), or a reference wild-type DNMT3A MTase (WT). Other controls included no transfection (untreated), transfection with a fusion protein containing a wild-type DNMT3A MTase, DNMT3L MTase-like domain, a dSpCas9, and a KRAB domain as set forth in SEQ ID NO: 112 (DNMT3A-DNMT3L-18aa-dSpCas9-KRAB, or D3AL-18aa-dSpCas9-KRAB) and HBV-targeting gRNA, or a fusion protein with an N879A DNMT3A variant MTase but no ZFP (N879A_No_ZFP). Methylation data includes on-target methylation data generated by qPCR over 29 days (FIG. 19A), on-target methylation data from Day 7 post-transfection generated by methyl sequencing (FIG. 19B), and off-target methylation data from Day 7 post-transfection (FIG. 19C).
[0099] FIG. 20 shows durable repression, as shown by fold-change in HBV mRNA, on day 21 posttransfection for cells transfected with 1.3 pg / mL of mRNA encoding fusion proteins containing dSpCas9 (x- axis) with a given DNMT3A variant versus on day 22 post-transfection for cells transfected with 0.94 pg / mL of mRNA encoding fusion proteins containing ZFP (y-axis) with a given DNMT3A variant. Each point represents a different DNMT3A variant, whose mutation is as labeled on the graph and corresponds to theAttorney Docket No. 224742003640 variant DNMT3A domains set forth in Table E4 or Table E5.
[0100] FIG. 21 shows on-target repression using qPCR of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription over 42 days in Hep3B cells transfected with a gRNA targeting HBV and 1.3 pg / mL mRNA encoding a fusion protein containing at least one ADD domain (indicated by an asterisk (*) following the domain the ADD domain is linked to) or a reference fusion protein with no ADD domain. Also depicted is results for repression following transfection with a catalytically dead fusion protein (C710A) or when cells were not transfected (untreated).
[0101] FIG. 22 shows on-target repression using qPCR of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription at day 4 (left) and day 42 (right) in Hep3B cells transfected with a gRNA targeting HBV and mRNA encoding a fusion protein containing at least one ADD domain (indicated by an asterisk following the domain the ADD domain is linked to) or a reference fusion protein with no ADD domain at one of two indicated doses. Also depicted is results for repression following transfection with a catalytically dead fusion protein (C710A), a fusion protein with a non-targeting gRNA (gNT), a control dSpCas9-KRAB fusion protein, or when cells were not transfected (untreated).
[0102] FIG. 23 shows on-target repression of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription at day 7 in Hep3B cells transfected with a gRNA targeting HBV and mRNA encoding a fusion protein containing at least one ADD domain (indicated by an asterisk following the domain the ADD domain is linked to) or a reference fusion protein with no ADD domain at one of two indicated doses. Also depicted is results for repression following transfection with a catalytically dead fusion protein (C710A) or when cells were not transfected (untreated).
[0103] FIG. 24A and FIG. 24B show results following methylation sequencing in Hep3B cells that demonstrate levels of off-target methylation as mean percent methylation across all regions captured by the off-target custom panel (FIG. 24A) and number of CpGs for each treatment that are differentially methylated compared to the untreated control (FIG. 24B) 7 days post-transfection of two indicated doses of tested fusion proteins. Tested fusion proteins contained an ADD domain (indicated by an asterisk following the domain the ADD domain is linked to) as compared to a reference fusion protein with no ADD domain (DNMT3A- DNMT3L-dSpCas9-KRAB) and negative controls, including untreated cells and a catalytically dead fusion protein (C710A).
[0104] FIG. 25 shows the on-target repression using qPCR of an exemplary target gene, a Hepatitis B viral (HBV) gene associated with HBV viral replication and transcription, over 30 days in Hep3B cells transfected with mRNA encoding a fusion protein containing a HBV-targeting zinc finger protein (ZFP) and at least one ADD domain (indicated by an asterisk following the domain the ADD domain is linked to) or aAttorney Docket No. 224742003640 reference fusion protein with no ADD domain (DNMT3A-mDNMT3L-ZFP-KRAB) at one of two indicated doses. Also depicted is results for repression following transfection with a control ZFP-KRAB fusion protein or when cells were not transfected (untreated).
[0105] FIG. 26A and FIG. 26B shows a dose response curve for fold change in PCSK9 mRNA relative to lipid only controls for assessing on-target potency using fusion proteins containing a D3A-3L-dSpCas9- KRAB fusion protein, where the 3L domain contains either a mouse DNMT3L MTase-like domain (FIG. 26A) or a human DNMT3L MTase-like domain (FIG. 26B) with an ADD domain (indicated by an asterisk (*) following the domain the ADD domain is linked to) or without an ADD domain in combination with an exemplary PCSK9-l-targeting gRNA.
[0106] FIG. 27 shows the level of durable repression of PCSK9, i.e., on-target methylation activity, in Huh7 cells using qPCR following transfection of mRNA encoding fusion proteins using different DNMT3 engineering strategies: addition of an ADD domain (indicated by an asterisk), use of variant DNMT3A domain (N879A), or use of DNMT3L with no DNMT3A domain (mDNMT3L-dSpCas9-KRAB). mRNA was delivered at two different doses (0.54 pg / mL, left; 0.94 pg / mL, right) over 28 days. Negative controls included leaving cells untreated or delivering cells mRNA encoding dSpCas9-KRAB (as set forth in SEQ ID NO: 22) or DNMT3A(C710A)-DNMT3L-dSpCas9-KRAB (C710A; set forth in SEQ ID NO: 23). Positive controls included delivering cells mRNA encoding the reference fusion protein DNMT3A-DNMT3L- dSpCas9-KRAB (SEQ ID NO: 14) or DNMT3A-DNMT3L-18aa-dSpCas9-KRAB (SEQ ID NO: 112). Durable repression was measured as fold change in PCSK9 mRNA as compared to untreated cells.
[0107] FIG. 28 with partial views FIG. 28A and FIG. 28B depicts an exemplary sequence alignment to depict identification of corresponding residues in a sequence compared to a reference sequence. The symbol “*” between two aligned amino acid indicates that the aligned amino acids are identical. The symbol “-“ indicates a gap in the alignment. Exemplary, non-limiting positions for amino acid substitution described herein are indicated with bold text. Based on the alignment of two similar sequences having identical residues in common, a skilled artisan can identify “corresponding” positions in a sequence by comparison to a reference sequence using conserved and identical amino acid residues as guides. Shown in the figure is an exemplary alignment of a reference mouse DNMT3L protein sequence set forth in SEQ ID NO: 107 (“mouse,” which contains the full-length mouse DNMT3L protein sequence with an ADD domain and a MTase-like domain but no initiating methionine residue) with a human DNMT3L protein sequence set forth in SEQ ID NO: 108 (“human,” which contains the full-length human DNMT3L protein sequence with an ADD domain and a MTase-like domain but no initiating methionine residue); aligning identical residues demonstrates, for example, that amino acid residue S60 in SEQ ID NO: 107 corresponds to residue S26 in SEQ ID NO: 108. It is within the level of a skilled artisan to carry out similar alignments between twoAttorney Docket No. 224742003640 similar protein sequences to identify corresponding residues, including based on the exemplification and description herein. Primary domains are annotated using italics with reference to the ADD domain and bold in reference to the MTase-like domain. A key is provided to correlate specific domains and regions of the DNMT3L protein involved in the recruitment of DNMT3 A to the appropriate annotation. It is understood that to the extent that residues of a domain or region involved in DNMT3A recruitment are present in a reference sequence that the corresponding domain or region involved in DNMT3A recruitment would similarly align between sequences.
[0108] FIG. 29 depicts the strategy of increasing the specificity of a fusion protein comprising a DNMT3A domain to transcriptionally repress gene expression. The fusion protein is composed of a DNA binding domain and a multipartite effector comprising a DNMT3A MTase domain and a repressor domain capable of recruiting heterochromatin inducing factors, e.g., a KRAB domain. In the top panel, the fusion protein containing a wild-type DNMT3A domain may rely more on the MTase for DNA binding, leading to off-target activity. In the variation depicted in the bottom panel, a variant MTase, such as a variant DNMT3A described herein, has reduced DNA binding affinity, so the fusion protein relies more on the DNA binding domain for DNA binding. Specificity is more driven by the DNA binding domain, leading to increased on- target activity.
[0109] FIG. 30 depicts an exemplary sequence alignment to depict identification of corresponding residues in a sequence compared to a reference sequence. The symbol “*” between two aligned amino acid indicates that the aligned amino acids are identical. The symbol indicates a gap in the alignment. Exemplary, non-limiting positions for amino acid substitution described herein are indicated with bold text. Based on the alignment of two similar sequences having identical residues in common, a skilled artisan can identify “corresponding” positions in a sequence by comparison to a reference sequence using conserved and identical amino acid residues as guides. Shown in the figure is an exemplary alignment of a reference DNMT3A domain sequence set forth in SEQ ID NO: 183 (“DNMT3A_Protein,” which contains the full- length DNMT3A protein sequence with three primary domains and an initiating methionine residue) with a functionally active portion of a DNMT3A domain sequence set forth in SEQ ID NO: 113 (“DNMT3A_Domain,” which contains the full-length sequence of only one primary domain, the MTase domain); aligning identical residues demonstrates, for example, that amino acid residue N711 in SEQ ID NO: 183 corresponds to residue N 100 in SEQ ID NO: 113 and amino acid residue S714 in SEQ ID NO: 183 corresponds to residue S 103 in SEQ ID NO: 113. It is within the level of a skilled artisan to carry out similar alignments between two similar protein sequences to identify corresponding residues, including based on the exemplification and description herein. Primary domains are annotated using boxes with reference to the “DNMT3A_Protein” sequence and DNA binding regions are annotated using underlining with reference toAttorney Docket No. 224742003640 the “DNMT3A_Domain” sequence. A key is provided to correlate specific domains and DNA binding regions to the appropriate annotation. It is understood that to the extent that residues of a domain are present in a reference sequence that the corresponding primary domains and DNA binding regions would similarly align between sequences.Detailed Description
[0110] Provided herein are fusion proteins for targeted gene repression. DNA methyltransferase families, such as DNMT3, contain enzymes that are able to methylate DNA and thus lead to gene repression. Included within the DNMT3 family are catalytic proteins, such as DNA methyltransferase 3A (DNMT3A) and DNA methyltransferase B (DNMT3B), and non-catalytic regulatory proteins, such as DNA methyltransferase 3L (DNMT3L). In some aspects, provided herein are fusion proteins utilizing different strategies for engineering DNMT3 proteins, domains, or portions thereof contained within fusion protein to improve targeted gene repression. In some embodiments, the different strategies for engineering DNMT3 improve specificity of targeted gene repression, such as reducing off-target effects. In some embodiments, the fusion protein is for recruiting domains with DNA methyltransferase activity to the target site for PCSK9. For example, in some embodiments, the fusion protein comprises a DNA-binding domain for targeting to a target site for PCSK9 and an effector domain that recruits DNA methyltransferase activity to the target site for PCSK9, such as any fusion protein described herein in Section I.A. In other embodiments, the fusion protein comprises a variant DNMT3A domain or functionally active portion thereof, such as any fusion protein described in Section I.B. For example, in some embodiments, the fusion protein comprises a DNA- binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof described herein in Section IV.[oni] Provided herein are fusion proteins comprising: (a) a DNA-binding domain, e.g., Cas, such as a dead Cas (dCas), for targeting to a target site for PCSK9, and (b) an effector domain that recruits domains with DNA methyltransferase activity to the target site for PCSK9. In some embodiments, the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail. In some embodiments, the length of the fusion protein minus the DNA-binding domain is less than 750 amino acids. In some embodiments, the effector domain comprises a DNA methyltransferase 3L (DNMT3L) protein or portion thereof. In some embodiments, the effector domain is less than 600 amino acids in length.
[0112] In some embodiments, the fusion proteins are for targeted transcriptional repression of PCSK9. In some aspects, also provided herein are DNA-target systems comprising any of the provided fusionAttorney Docket No. 224742003640 proteins, that are capable of inducing targeted transcriptional repression of PCSK9. Also provided are polynucleotides, vectors, pluralities and combinations thereof, that encode the fusion proteins, DNA- targeting systems, or components thereof. Also provided are cells, such as engineered cells, that encode the fusion proteins, DNA-targeting systems, or components thereof. Also provided are methods and uses related to the provided compositions, for example in repressing transcription of a target gene or modifying a phenotype of a cell, including in the treatment of diseases or conditions related to PCSK9 expression, such as Familial Hypercholesterolemia (FH) and / or cardiovascular disease, such as atherosclerotic cardiovascular disease (ASCVD).
[0113] T argeted epigenetic modulation is an approach for investigating biology and therapeutic applications. Sequence-specific DNA-targeting systems, such as zinc finger proteins, transcription-activator- like effectors, and CRISPR / Cas systems can be programmed by a user to target sequences of interest. These DNA-targeting systems can be used to recruit effector proteins such as transcriptional and epigenetic modulators to endogenous genomic loci, for example to activate or repress transcription of a target gene.
[0114] Despite the potential for targeted transcriptional repression as a therapeutic or investigative tool, fusion proteins used for transcriptional repression effector may target other sites besides the particular target site, leading to a loss of specificity or non-specific binding. Specifically, DNA methyltransferases, such as DNA methyltransferase 3A (DNMT3A) or 3B (DNMT3B) bind DNA through sequence-dependent and independent interactions to recognize a CpG and transfer a methyl group to the C5 position of the cytosine. Traditional fusion proteins for transcriptional repression directly fuse the catalytic domain of a methyltransferase domain (e.g., DNMT3A or DNMT3B MTase domain) to a programmable DNA binding domain, such as a dCas9 or zinc finger protein (ZFP), that redirects the localization of the fusion protein in the genome. However, the inherent DNA binding ability of the methyltransferase domain can result in off- target methylation. Further, attempts to modify the fusion protein to increase specificity may lead to different levels of decreased transcription or may induced decreased transcription for different amounts of time for a given target site for a gene of interest. For example, a fusion protein may only transiently decrease transcription or may induce durably (e.g., heritably) decreased transcription.
[0115] The specificity and degree of transcriptional repression can affect the therapeutic potential of the transcriptional repression. For example, targeted transcription repression of the target gene PCSK9 can be used for treatment of diseases or conditions, such as treatment of FH and / or cardiovascular disease, such as ASCVD. For example, rare gain-of-function mutations in proprotein convertase subtilisin / kexin type 9 (PCSK9) have been associated with FH. Mutations in PCSK9 can also alter expression of LDLR, and the upregulation of LDLR is associated with protection from cardiovascular disease. PCSK9 negatively regulates cell surface expression of LDLR by binding to the epidermal growth factor-like repeat A (EGF-A) domain ofAttorney Docket No. 224742003640LDLR and targeting LDLR for lysosomal degradation. Mutations resulting in increased PCSK9 activity thus decrease the ability of cells to express LDLR at the cell surface, resulting in elevated circulating LDL. Conversely, loss-of-function PCSK9 mutations are associated with lowered LDL levels and protection from cardiovascular disease (Lagace, T.A. Curr. Opin. Lipidol. 25(5):387-393 (2014); Peterson, A. S. et al. J. Lipid Res. 49(6): 1152-1156 (2008); Bouhairie, V. E. et al. Cardiol. Clin. 33(2): 169-179 (2015)). Thus, targeted transcriptional repression of PCSK9 is a promising treatment for diseases such as FH and cardiovascular disease, such as ASCVD.
[0116] However, non-specific binding of a fusion protein for targeted transcriptional repression of PCSK9 may result in off-target activity, e.g., off-target repression, which may result in loss of gene function at unintended targets that interfere with therapeutic effects of PCSK9 transcriptional repression or other undesired phenotypes.
[0117] Provided embodiments herein address the need for a fusion protein for targeted repression of PCSK9 with minimal off-target activity by not directly fusing the domain with DNA methyltransferase activity to the DNA binding domain. Rather, because many cells (e.g., somatic cells) already express endogenous DNMT3A and DNMT3B, the improved fusion protein comprises a catalytically inactive cofactor DNA methyltransferase 3L (DNMT3L) protein or portion thereof to recruit the endogenous DNA methylation machinery to the target site for PCSK9 without the direct fusion of DNMT3A or DNMT3B to the fusion protein. It is shown herein that such a strategy reduces off-target binding as compared to a fusion protein that includes a DNMT3A domain fused to the DNA-binding domain.
[0118] However, endogenous DNMT3A and DNMT3B expressed in cells contain N-terminal regulatory domains, such as the PWWP domain and the ADD domain, that may interfere with the efficiency of on- target methylation relative to known fusion proteins for targeted transcriptional repression that comprise the catalytic domain of a DNA methyltransferase (for known fusion proteins for targeted transcription repression comprising the catalytic domain of a DNA methyltransferase, see, e.g., Siddique et al., J. Molecular Biol., 2013). Specifically, the ADD domain prevents methylation of regions marked by tri-methylation of the histone tail H3K4 (H3K4me3), a chromatin mark associated with the promoters of actively transcribed genes that may otherwise be a target for de novo methylation by DNMT3A or DNMT3B. The ADD domain sterically interferes with the catalytic portion of methyltransferase domains (e.g., DNMT3A) and the DNMT3A-DNMT3L binding interface on DNMT3L in the presence of H3K4me3. However, when the H3K4 tail is not methylated (H3K4meO), the ADD domain can swing out of the inhibitory confirmation to bind to it, thus enabling DNA methylation.
[0119] Previous methods have attempted to overcome the inhibition caused by endogenous DNMT3A or DNMT3B or the ADD domain by additionally adding domains capable of recruiting heterochromatinAttorney Docket No. 224742003640 inducing factors, such as KRAB domains (see, e.g., WO 2016 / 063264), or by supplying a synthetic H3K4meO histone tail (see, e.g., WO 2024 / 173896 and US 2024 / 0279623). It was thought either of these strategies was necessary to enable robust on-target methylation by fusion proteins for targeted transcriptional repression that recruit endogenous methyltransferase domains.
[0120] However, as shown in the Examples herein, neither a KRAB domain nor synthetic H3K4meO histone tail is necessary for robust on-target methylation when targeting PCSK9. Surprisingly, it is found herein that a fusion protein comprising a DNA-binding domain for targeting to a target site for PCSK9 and an effector domain comprising a DNMT3L protein or portion thereof, and wherein the fusion protein is not fused to domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail, is sufficient for robust on-target methylation when targeting PCSK9.
[0121] Together, the fusion proteins for targeted transcription repression provided herein enable on- target transcriptional repression of PCSK9 with increased specificity by minimizing off-target activity. In particular, the provided fusion protein is sufficient to repress PCSK9 using only a DNA-binding domain targeting a target site for PCSK9 and a DNMT3L protein or portion thereof, as compared to other methods that utilize additional components. The elimination of such components thus enables a fusion protein that is smaller and easier to deliver into cells for therapeutic applications, such as the treatment of FH and / or cardiovascular disease.
[0122] Provided herein are fusion proteins of a DNA-binding domain, e.g., Cas, such as a dead Cas (dCas), and an effector domain comprising a variant DNA methyltransferase 3A (DNMT3A) domain or a functionally active portion thereof. In some embodiments, the fusion proteins are for targeted transcriptional repression of gene expression. In some aspects, also provided herein are DNA-targeting systems comprising any of the provided fusion proteins, that are capable of inducing targeted transcriptional repression of target genes, for example, when recruited to a target site at the target gene. In some aspects, the DNA-binding domain of a provided DNA-targeting system is an RNA-guided endonuclease (e.g., inactivated endonuclease, such as dCas) that includes one or more gRNAs. In some aspects, the fusion protein and DNA-targeting systems lead to decreased transcription of an endogenous gene, when recruited to a target site at the endogenous gene. Also provided are polynucleotides, vectors, pluralities and combinations thereof, that encode the effector domains, fusion proteins, DNA-targeting systems, gRNAs or components thereof. Also provided are cells, such as engineered cells, that encode the effector domains, fusion proteins, DNA- targeting systems, gRNAs or components thereof. Also provided are methods and uses related to the provided compositions, for example in repressing transcription of a target gene or modifying a phenotype of a cell, including in connection with therapeutic applications.Attorney Docket No. 224742003640
[0123] In some embodiments, the effector domain of a provided fusion protein is a multipartite effector that further comprises one or more domains selected from a DNA methyltransferase domain, a repressor domain capable of recruiting heterochromatin inducing factors, or combinations thereof. In some embodiments, the DNA methyltransferase domain is a catalytically inactive regulatory factor of DNA methyltransferases. In some embodiments, the DNA methyltransferase domain is a DNMT3L domain. In some embodiments, the repressor domain comprises a KRAB repressor domain, ERF repressor domain, Mxil repressor domain, SID4X repressor domain, Mad-SID repressor domain, LSD1 repressor domain, EZH2 repressor domain, or variant of any of the foregoing.
[0124] Targeted epigenetic modulation is an approach for investigating biology and therapeutic applications. Sequence-specific DNA-targeting systems, such as zinc finger proteins, franscription-activator- like effectors, and CRISPR / Cas systems can be programmed by a user to target sequences of interest. These DNA-targeting systems can be used to recruit effector proteins such as transcriptional and epigenetic modulators to endogenous genomic loci, for example to activate or repress transcription of a target gene.
[0125] Despite the potential for targeted transcriptional repression as a therapeutic or investigative tool, the ability to repress transcription can be unpredictable or unreliable and is dependent on the specific effectors that are recruited to particular target sites. Only a handful of transcriptional repression effector domains have been frequently used for targeted transcriptional repression. For a given target site for a gene of interest, some transcriptional repressor domains may lead to decreased transcription of the gene, and others may not. In addition, different transcriptional repression effector domains may lead to different levels of decreased transcription or may induce decreased transcription for different amounts of time. For example, a transcriptional repression effector domain may only transiently decrease transcription or may induce durably (e.g., heritably) decreased transcription. The effect of a given transcriptional repression effector domain being recruited to a particular target site on the transcription of a gene is also generally unpredictable. Often, a transcriptional repression effector domain must be tested at several target sites to identify a suitable target site for transcriptional repression of the gene of interest. Additionally, a transcriptional repression effector domain may target other sites besides the particular target site, leading to a loss of specificity or non-specific binding.
[0126] The predictability, degree of transcriptional repression, and specificity of binding can affect the therapeutic potential of the transcriptional repression. For example, weak transcriptional repression may not decrease transcription of the target gene sufficiently to result in a therapeutic effect. In some cases, strong or durable transcriptional repression may lead to a therapeutic effect. Therapeutic potential for human subjects of some transcriptional repression effector domains may also be limited by the immunogenicity of the domain, for example if the domain is from a non-human organism. Further, certain transcriptional repressionAttorney Docket No. 224742003640 effector domains (e.g., effector domains comprising DNA methyltransferases, e.g., DNMT3A), bind DNA through sequence-dependent and sequence-independent interactions that may redirect the localization of a fusion protein such that it binds non-specifically. FIG. 29 (top panel) depicts an example of a fusion protein with a DNA binding domain and multipartite effector comprising a DNMT3A MTase domain (MTase) and repressor domain (KRAB) where the DNA binding location of the fusion protein can, in some cases, be driven by the DNA binding affinity of the MTase rather than the DNA-binding domain, resulting in nonspecific binding. Non-specific binding of a transcriptional repression effector domain may result in off-target activity, e.g., off-target repression, which may result in loss of gene function at unintended targets that interfere with therapeutic effects or other undesired phenotypes. These challenges raise the need for an expanded and improved transcriptional repression effector domains, fusion proteins, and DNA-targeting systems containing transcriptional repression effector domains comprising DNA methyltransferases for use in targeted transcriptional repression.
[0127] In some embodiments, the provided embodiments address these needs. In some embodiments, the provided variant DNMT3A domains or functionally active portions thereof when fused with a DNA- binding domain (e.g., dCas in combination with a gRNA) allow for improved on-target specificity to target the transcriptional repression effector domains to specific target sites. In some aspects, improved on-target specificity may be measured by maintaining sufficient on-target activity. In some aspects, improved on-target specificity may be measured by minimizing off-target activity.
[0128] In some aspects, the variant DNMT3A domains or functionally active portions thereof described herein have substitutions that may reduce DNA binding affinity, which may allow for a DNA binding domain rather than the DNMT3A domain to direct DNA binding, leading to on-target specificity (see FIG. 29, bottom panel). In some aspects, the variant DNMT3A domains or functionally active portion thereof described herein have substitutions that may increase catalytic activity, which may allow for sufficient on- target activity. In some aspects, the variant DNMT3A domains or functionally active portion thereof described herein have substitutions that may reduce DNA binding affinity and / or increase catalytic activity, allowing for improved on-target activity while minimizing off-target activity. In some aspects, the transcriptional repressor domains are derived from human genes, thereby reducing potential for immunogenicity in human subjects. In some aspects, the expanded set of fusion proteins may allow for increased control of the degree of transcriptional repression at a given locus. In some aspects, a transcriptional repression effector domain or fusion protein provided herein may provide an increased degree of transcriptional repression, or increased durability of transcriptional repression, when targeted to a target site. In some aspects, the increased specificity, degree, or durability of transcriptional repression may increase the therapeutic effect of the targeted transcriptional repression, e.g., by reducing the need forAttorney Docket No. 224742003640 repeated administration and / or by increasing the effect of administration or lowering risks due to off-target editing.
[0129] All publications, including patent documents, scientific articles and databases, referred to in this application are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication were individually incorporated by reference. If a definition set forth herein is contrary to or otherwise inconsistent with a definition set forth in the patents, applications, published applications and other publications that are herein incorporated by reference, the definition set forth herein prevails over the definition that is incorporated herein by reference.
[0130] The section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.I. FUSION PROTEINS
[0131] In some aspects, provided herein are fusion proteins utilizing a strategy for engineering DNMT3.
[0132] In some embodiments, the fusion protein is for recruiting domains with DNA methyltransferase activity to the target site for PCSK9. In some embodiments, the fusion protein is any fusion protein described herein in Section I.A. In some embodiments, the fusion protein comprises: (a) a DNA-binding domain, such as any described herein in Section I.A. 1, for targeting to a target site for PCSK9 (e.g., any target site for PCSK9 described herein in Section II. A) and (b) an effector domain that recruits domains with DNA methyltransferase activity to the target site for PCSK9, such as any effector domain described herein in Section I.A.2.
[0133] In some embodiments, the fusion protein comprises a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof. In some embodiments, the fusion protein is any fusion protein described herein in Section I.B. In some embodiments, the fusion protein targets a target site and / or target gene, such as any described herein in Section II. In some embodiments, the fusion protein comprises (a) a DNA-binding domain or a component thereof, such as any described herein in Section I.B. 1, and (b) an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, such as any described here in Section IV. In some embodiments, the effector domain can include the variant DNMT3A fused to a DNMT3L as described in Section I.B.2.a. In some embodiments, the effector domain can include the variant DNMT3A as a multipartite fusion as described in Section I.B.2.b.
[0134] In some embodiments, any two components of the fusion protein may be fused directly (i.e., without an intervening amino acid sequence). In some embodiments, any two components of the fusionAttorney Docket No. 224742003640 protein may be fused indirectly, e.g., via an intervening amino acid sequence, such as a linker or nuclear localization signal (NLS), such as any linker or NLS described herein.
[0135] In some embodiments, the fusion protein comprises one or more linkers. In some embodiments, the one or more linkers connect any two components of the fusion protein. A linker may be included anywhere in the polypeptide sequence of the fusion protein, for example, between a transcriptional activation domain and the DNA-binding domain or a component thereof. A linker may be of any length and designed to promote or restrict the mobility of components in the fusion protein. In some embodiments, inclusion of a linker in the fusion protein enhances the function of the fusion protein. For example, inclusion of the linker in the fusion protein may lead to enhanced activation of the target gene in comparison to a comparable fusion protein without the linker.
[0136] A linker may comprise any amino acid sequence of about 2 to about 100, about 5 to about 80, about 10 to about 60, or about 20 to about 50 amino acids. A linker may comprise an amino acid sequence of at least about 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 or 85 amino acids. A linker may comprise an amino acid sequence of less than about 100, 90, 80, 70, 60, 50, or 40 amino acids. A linker may include sequential or tandem repeats of an amino acid sequence that is 2 to 20 amino acids in length. Linkers may be rich in amino acids glycine (G), serine (S), and / or alanine (A). Linkers may include, for example, a GS linker. An exemplary GS linker is represented by the sequence GGGGS (SEQ ID NO: 82). A linker may comprise repeats of a sequence, for example as represented by the formula (GGGGS)n, wherein n is an integer that represents the number of times the GGGGS sequence is repeated (e.g., between 1 and 10 times). The number of times a linker sequence is repeated can be adjusted to optimize the linker length and achieve appropriate separation of the functional domains. Other examples of linkers may include, for example, GGGGG (SEQ ID NO: 83), GGAGG (SEQ ID NO: 84), GGGGSSS (SEQ ID NO: 85), GGGGAAA (SEQ ID NO: 86), GGSGG (SEQ ID NO: 47), or GSGSG (SEQ ID NO: 87).
[0137] In some embodiments, the linker is an XTEN linker. In some aspects, an XTEN linker is a recombinant polypeptide (e.g., an unstructured recombinant peptide) lacking hydrophobic amino acid residues. Exemplary XTEN linkers are described in, for example, Schellenberger et al., Nature Biotechnology 27, 1186-1190 (2009) or WO 2021 / 247570. In some embodiments, the linker comprises a linker described in WO 2021 / 247570. In some aspects, the linker is or comprises the sequence set forth in SEQ ID NO: 33, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 33. In some aspects, the linker comprises the sequence set forth in SEQ ID NO: 33, or a contiguous portion of SEQ ID NO: 33 of at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70 or 75 amino acids. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 33. In some embodiments, the linker consists of the sequence set forth inAttorney Docket No. 224742003640SEQ ID NO: 33. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 45 or SEQ ID NO: 46, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to the foregoing. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 45, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 45. In some aspects, the linker comprises the sequence set forth in SEQ ID NO:45, or a contiguous portion of SEQ ID NO: 45 of at least 5, 10, or 15 amino acids. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 45. In some embodiments, the linker consists of the sequence set forth in SEQ ID NO: 45. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 46, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 46. In some aspects, the linker comprises the sequence set forth in SEQ ID NO: 46, or a contiguous portion of SEQ ID NO: 46 of at least 5, 10, 15, 20, or 25 amino acids. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 46. In some embodiments, the linker consists of the sequence set forth in SEQ ID NO: 46. In some embodiments, the linker comprises any one of the sequences set forth in SEQ ID NOS: 38 and 82-92, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the linker comprises any one of the sequences set forth in SEQ ID NOS: 38 and 82-92. In some embodiments, the linker consists of any one of the sequences set forth in SEQ ID NOS: 38 and 82-92. In some embodiments, the linker comprises any one of the sequences set forth in SEQ ID NOS: 82-92, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the linker comprises any one of the sequences set forth in SEQ ID NOS: 82-92. In some embodiments, the linker consists of any one of the sequences set forth in SEQ ID NOS: 82-92.
[0138] In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 384, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 384. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 384 or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 384. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 384. In some embodiments, the linker consists of the sequence set forth in SEQ ID NO: 384
[0139] In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 120, a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 120. In some embodiments, the linker comprises the sequence setAttorney Docket No. 224742003640 forth in SEQ ID NO: 120 or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 120. In some embodiments, the linker comprises the sequence set forth in SEQ ID NO: 120. In some embodiments, the linker consists of the sequence set forth in SEQ ID NO: 120. Appropriate linkers may be selected or designed based rational criteria known in the art, for example as described in Chen et al. Adv. Drug Deliv. Rev. 65(10): 1357-1369 (2013).
[0140] In some embodiments, the fusion protein comprises one or more nuclear localization signal (NLS). In some embodiments, a fusion protein described herein comprises one or more nuclear localization sequences (NLSs), such as about or more than about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more NLSs. When more than one NLS is present, each may be selected independently of the others, such that a single NLS may be present in more than one copy and / or in combination with one or more other NLSs present in one or more copies. Non-limiting examples of NLSs include an NLS sequence derived from: the NLS of the SV40 virus large T-antigen, having the amino acid sequence PKKKRKV (SEQ ID NO: 34); the NLS from nucleoplasmin (e.g. the nucleoplasmin bipartite NLS with the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 44)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 93) or RQRRNELKRSP (SEQ ID NO: 94); the hRNPAl M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 95); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 96) of the IBB domain from importin-alpha; the sequences VSRKRPRP (SEQ ID NO: 97) and PPKKARED (SEQ ID NO: 98) of the myoma T protein; the sequence PQPKKKPL (SEQ ID NO: 99) of human p53; the sequence SALIKKKKKMAP (SEQ ID NO: 100) of mouse c-abl IV; the sequences DRLRR (SEQ ID NO: 101) and PKQKKRK (SEQ ID NO: 102) of the influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 103) of the Hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 104) of the mouse Mxl protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 105) of the human poly(ADP -ribose) polymerase; and the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 106) of the steroid hormone receptors (human) glucocorticoid. The NLS may comprise a portion of any of the foregoing. In some embodiments, the one or more NLSs are of sufficient strength to drive localization and / or accumulation of the fusion protein in the nucleus of a eukaryotic cell, such as in a detectable amount. In some embodiments, the strength of nuclear localization activity may derive from the number of NLSs in the fusion protein, the particular NLS(s) used, or a combination of these factors. Detection of accumulation in the nucleus may be performed by any suitable technique. For example, a detectable marker may be fused to the fusion protein, such that location within a cell may be visualized, such as in combination with a means for detecting the location of the nucleus (e.g., a stain specific for the nucleus such as DAPI). Cell nuclei may also be isolated from cells, the contents of which may then be analyzed by any suitable process for detecting protein, such asAttorney Docket No. 224742003640 immunohistochemistry, Western blot, or enzyme activity assay. Accumulation in the nucleus may also be determined indirectly, such as by an assay for the effect of the fusion protein (e.g., an assay for altered gene expression activity in a cell transformed with the fusion protein), as compared to a control condition (e.g., an untransformed cell).
[0141] In some aspects, the fusion protein comprises one or more tags, linkers and / or NLS sequences. In some embodiments, exemplary tags, linkers and / or NLS sequences can be any described herein.
[0142] In some cases, sequences provided herein, including amino acid sequences for the DNA- targeting systems or fusion proteins provided herein, contain sequences of one or more tags, linkers and / or NLS sequences. In some aspects, it is understood that the exemplary tags, linkers and / or NLS sequences are not required or are not the sole or exclusive tags, linkers and / or NLS sequences that can be employed in the DNA-targeting systems or fusion proteins. In some aspects, sequences containing tags, linkers and / or NLS sequences are exemplary, and are not limited to the specific tags, linkers and / or NLS sequences contained in the described sequences. In some aspects, alternative tags, linkers and / or NLS sequences can be can be employed in the DNA-targeting systems or fusion proteins, or the DNA-targeting system or fusion protein in some cases does not contain or lacks a tag, linker and / or NLS. In some aspects, alternative tags, linkers and / or NLS sequences include other known tags, linkers and / or NLS sequences that have similar function or serve similar purposes.
[0143] It is understood that “consists essentially of’ permits the presence of one or more tags, linkers, and / or NLS sequences. The term “consists essentially of’ (which can be used interchangeably with “consisting essentially of’ or other grammatical variations) thus means that the polynucleotide sequence encoding the fusion protein or the amino sequence of the fusion protein includes those DNA-binding domains and effector domains identified and excluding any other DNA-binding domains and effector domains not so identified, but in which the polynucleotide sequence or amino acid sequence may contain other sequence components. In some embodiments, such other sequence components do not substantially alter the DNA-binding activity or the transcriptional repression activity of the fusion protein.A. Fusion Proteins for Recruiting Domains with DNA Methyltransferase Activity to the Target Site for PCSK9
[0144] In some aspects, provided herein are fusion proteins. In some embodiments, the fusion protein comprises: (a) a DNA-binding domain for targeting to a target site for PCSK9, such as any DNA-binding domain described herein in Section I.A. 1, and (b) an effector domain that recruits domains with DNAAttorney Docket No. 224742003640 methyltransferase activity to the target site for PCSK9, such as any effector domain described herein in Section I.A.2.
[0145] In some embodiments, the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). For example, the fusion protein is not fused to a DNMT3A or DNMT3B domain. Without wishing to be bound by theory, it is theorized that the presence of the effector domain, such as a DNMT3L protein or portion thereof, is alone sufficient to recruit methyltransferase machinery (e.g., DNMT3A or DNMT3B) in cells.
[0146] In some embodiments, the fusion protein is devoid of domains capable of recruiting heterochromatin inducing factors or peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptide). In some embodiments, the fusion protein is devoid of domains capable of recruiting heterochromatin inducing factors, wherein domains capable of recruiting heterochromatin inducing factors include a KRAB domain, ERF repressor domain, MXI1 domain, SID4X domain, MAD-SID domain, LSD1, EZH2, a partially or fully functional fragment or domain of any of the foregoing, or a combination of any of the foregoing.
[0147] In some embodiments, the fusion protein is devoid of domains capable of recruiting heterochromatin inducing factors, wherein the domains capable of recruiting heterochromatin inducing factors comprise a KRAB domain. The Kriippel associated box (KRAB) domain is a domain present in many zinc finger protein-based transcription factors. The KRAB domain comprises charged amino acids and can be divided into sub-domains A and B. The KRAB domain recruits corepressors KAP1 (KRAB-associated protein- 1), epigenetic readers such as heterochromatin protein 1 (HP1), and other chromatin modulators to induce transcriptional repression through heterochromatin formation. KRAB-mediated gene repression is associated with loss of histone H3 -acetylation and an increase in H3 lysine 9 trimethylation (H3K9me3) at the repressed gene promoters. KRAB domains, including in dCas fusion proteins, have been described, for example, in WO 2017 / 180915, WO 2014 / 197748, US 2019 / 0127713, WO 2013 / 176772, Urrutia R. et al. Genome Biol. 4, 231 (2003), Groner A. C. et al. PLoS Genet. 6, el000869 (2010).
[0148] In some embodiments, the fusion protein minus the DNA-binding domain is less than 1000 amino acids, less than 950 amino acids, less than 900 amino acids, less than 850 amino acids, less than 800 amino acids, less than 750 amino acids, less than 700 amino acids, less than 650 amino acids, less than 600 amino acids, less than 550 amino acids, less than 500 amino acids, less than 450 amino acids, less than 400 amino acids, less than 350 amino acids, less than 300 amino acids, less than 250 amino acids, less than 200 amino acids, less than 150 amino acids, or less than 100 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 500 amino acids, less than 490 amino acids, less thanAttorney Docket No. 224742003640480 amino acids, less than 470 amino acids, less than 460 amino acids, less than 450 amino acids, less than440 amino acids, less than 430 amino acids, less than 420 amino acids, less than 410 amino acids, less than400 amino acids, less than 390 amino acids, less than 380 amino acids, less than 370 aminos, less than 360 amino acids, less than 350 amino acids, less than 340 amino acids, less than 330 amino acids, less than 320 amino acids, less than 310 amino acids, or less than 300 amino acids.
[0149] In some embodiments, the fusion protein minus the DNA-binding domain is less than 750 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 500 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 400 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 370 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 350 amino acids. In some embodiments, the fusion protein minus the DNA-binding domain is less than 300 amino acids.
[0150] In some embodiments, the fusion protein is capable of targeting the effector domain to a target site for PCSK9. In some embodiments, the fusion protein is capable of being targeted to a target site for PCSK9, by virtue of the DNA-binding domain or component thereof. Without wishing to be bound by theory, it is believed the effector domain can recruit DNA methyltransferase domains with methyltransferase activity to the target site for PCSK9. In some aspects, targeting of the effector or the fusion protein decreases transcription of PCSK9. In some aspects, the decreased transcription of PCSK9 is sustained, i.e., durable. In some aspects, the targeting is specific, i.e., low off-target activity.
[0151] In some embodiments, any two or more domains of the fusion protein are heterologous, i.e., the domains are from different species, or at least one of the domains is not found in nature. In some aspects, the fusion protein is an engineered fusion protein, i.e., the fusion protein is not found in nature.
[0152] In some embodiments, the fusion protein comprises its constituent components (e.g., domains) in any suitable arrangement, orientation, or order. For example, an effector domain (e.g., a DNMT3L protein or portion thereof) may be fused to the N-terminus or C-terminus of the DNA-binding domain of the fusion protein. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, the effector domain (e.g., the DNMT3L protein or portion thereof) and the DNA-binding domain. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, the DNA-binding domain and the effector domain (e.g., the DNMT3L protein or portion thereof).
[0153] In some embodiments, the effector domain (e.g., a DNMT3L protein or portion thereof) is the only transcriptional repressor domain fused to the DNA-binding domain. In some embodiments, the fusion protein consists essentially of a DNA-binding domain and a DNMT3L domain or portion thereof, whereinAttorney Docket No. 224742003640DNMT3L is the only repressor domain present.1. DNA-binding domains
[0154] In some embodiments, provided are DNA-binding domains. In some aspects, a DNA-binding domain is capable of specifically targeting (e.g., binding or hybridizing to) a target site, such as any target site in the PCSK9 locus or described herein in Section ILA. In some aspects, a DNA-binding domain targets a specific sequence of nucleotides, such as a DNA sequence. In some aspects, a DNA-binding domain can be engineered (e.g. designed or programmed) to target a specific target site. In some aspects, a DNA-binding domain recruits an effector domain (e.g., a DNMT3L protein or portion thereof) described herein in Section I. A.2 to the target site, wherein the effector domain can recruit DNA methyltransferases and thereby induce targeted gene repression.
[0155] In some embodiments, the DNA-binding domain comprises a CRISPR associated (Cas) protein, zinc finger protein (ZFP), transcription activator-like effectors (TALE), meganuclease, homing endonuclease, LScel enzyme, or variants thereof. In some embodiments, the DNA-binding domain comprises a catalytically inactive (e.g., nuclease-inactive or nuclease-inactivated) variant of any of the foregoing. In some embodiments, the DNA-binding domain comprises a deactivated Cas9 (dCas9) protein or variant thereof that is a catalytically inactivated so that it is inactive for nuclease activity and is not able to cleave DNA.
[0156] In some embodiments, the DNA-binding domain is a component of a DNA-targeting system, such as any described herein in Section III, comprising a Cas-gRNA combination, comprising a Cas protein or variant thereof and at least one guide RNA (gRNA). In some embodiments, the gRNA binds to the target site. In some embodiments, the gRNA comprises a spacer sequence that is capable of targeting and / or hybridizing to the target site. In some embodiments, the gRNA is capable of complexing with the Cas protein or variant thereof, e.g., via a scaffold sequence of the gRNA. In some aspects, the gRNA directs or recruits the Cas protein or variant thereof to the target site.
[0157] Exemplary components and features of the DNA-binding domains, including for CRISPR / Cas- based, ZFN-based, and TALE-based DNA-binding domains are provided below. a. CRISPR / Cas-based DNA-binding domains
[0158] Provided herein are DNA-binding domains based on CRISPR / Cas systems, i.e., CRISPR / Cas- based DNA-binding domains, that are able to bind to a target site or a combination of target sites. In some embodiments, the CRISPR / Cas-based DNA-binding domain is nuclease inactive, deactivated or nuclease- dead, such as a dCas (e.g., dCas9) so that the system binds to the target site without mediating nucleic acidAttorney Docket No. 224742003640 cleavage. In some embodiments, the CRISPR / Cas-based DNA-binding domain can include any known Cas protein or variant thereof, and generally a nuclease-inactive or dCas.
[0159] The CRISPR system (also known as CRISPR / Cas system, or CRISPR-Cas system) refers to a conserved microbial nuclease system, found in the genomes of bacteria and archaea, that provides a form of acquired immunity against invading phages and plasmids. Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR), refers to loci containing multiple repeating DNA elements that are separated by non-repeating DNA sequences called spacers. Spacers are short sequences of foreign DNA that are incorporated into the genome between CRISPR repeats, serving as a 'memory' of past exposures. Spacers encode the DNA-targeting portion of RNA molecules that confer specificity for nucleic acid cleavage by the CRISPR system. CRISPR loci contain or are adjacent to one or more CRISPR-associated (Cas) genes, which can act as RNA-guided nucleases for mediating the cleavage, as well as non-protein coding DNA elements that encode RNA molecules capable of programming the specificity of the CRISPR-mediated nucleic acid cleavage.
[0160] In Type II CRISPR / Cas systems with the Cas protein Cas9, two RNA molecules and the Cas9 protein form a ribonucleoprotein (RNP) complex to direct Cas9 nuclease activity. The CRISPR RNA (crRNA) contains a spacer sequence that is complementary to a target nucleic acid sequence (target site), and that encodes the sequence specificity of the complex. The trans-activating crRNA (tracrRNA) base-pairs to a portion of the crRNA and forms a structure that complexes with the Cas9 protein, forming a Cas / RNA RNP complex.
[0161] Naturally occurring CRISPR / Cas systems, such as those with Cas9, have been engineered to allow efficient programming of Cas / RNA RNPs to target desired sequences in cells of interest, both for geneediting and modulation of gene expression. The tracrRNA and crRNA have been engineered to form a single chimeric guide RNA molecule, commonly referred to as a guide RNA (gRNA), for example as described in WO 2013 / 176772, WO 2014 / 093661, WO 2014 / 093655, Jinek, M. et al. Science 337(6096): 816-21 (2012), or Cong, L. et al. Science 339(6121): 819-23 (2013), and as described herein, for example, in Section I. A. l.a.2. The spacer sequence of the gRNA can be chosen by a user to target the Cas / gRNA RNP complex to a desired locus, e.g., a desired target site in the target gene.
[0162] Cas proteins have also been engineered to be catalytically inactivated or nuclease inactive to allow targeting of Cas / gRNA RNPs without inducing cleavage at the target site. Mutations in Cas proteins can reduce or abolish nuclease activity of the Cas protein, rendering the Cas protein catalytically inactive. Cas proteins with reduced or abolished nuclease activity are referred to as deactivated Cas or dead Cas (dCas), or nuclease-inactive Cas (iCas) proteins, as referred to interchangeably herein. In some aspects, the dCas or iCas can still bind to target site in the DNA in a site- and / or sequence-specific manner, as long as itAttorney Docket No. 224742003640 retains the ability to interact with the guide RNA (gRNA) which directs the Cas-gRNA combination to the target site.
[0163] dCas-fusion proteins with transcriptional and / or epigenetic regulators have been used as a versatile platform for ectopically regulating gene expression in target cells. These include fusion of a Cas with an effector domain, such as a transcriptional activator or transcriptional repressor. For example, fusing dCas9 with a transcriptional activator such as VP64 (a polypeptide composed of four tandem copies of VP 16, a 16 amino acid transactivation domain of the Herpes simplex virus) can result in increased expression of a targeted gene. Alternatively, fusing dCas9 with a transcriptional repressor such as KRAB (Kriippel associated box) can result in reduced expression of a targeted gene. A variety of dCas-fusion proteins with transcriptional regulators have been engineered, for example as described in WO 2014 / 197748, WO 2016 / 130600, WO 2017 / 180915, WO 2021 / 226555, WO 2021 / 226077, WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, WO 2021 / 247570, Adli, M. Nat. Commun. 9, 1911 (2018), Perez-Pinera, P. et al. Nat. Methods 10, 973-976 (2013), Mali, P. et al. Nat. Biotechnol. 31, 833-838 (2013), Maeder, M. L. et al. Nat. Methods 10, 977-979 (2013), Gilbert, L. A. et al. Cell 154(2):442-451 (2013), and Nunez, J.K. et al. Cell 184(9):2503-2519 (2021).1) Cas proteins
[0164] In some aspects, the DNA-binding domain comprises a CRISPR-associated (Cas) protein. In some embodiments, the Cas protein is a variant Cas protein, such as a Cas protein derived from or based on a naturally occurring Cas protein or portion thereof. In some embodiments, the variant Cas protein comprises one or more modifications, mutations, or amino acid substitutions in comparison to the naturally occurring Cas protein. In particular embodiments provided herein, the Cas protein is nuclease-inactive (i.e., is a dCas protein).
[0165] In some embodiments, the Cas protein is derived from a Class 1 CRISPR system (i.e., multiple Cas protein system), such as a Type I, Type III, or Type IV CRISPR system. In some embodiments, the Cas protein is derived from a Class 2 CRISPR system (i.e., single Cas protein system), such as a Type II, Type V, or Type VI CRISPR system. In some embodiments, the Cas protein is derived from a Type V CRISPR system.
[0166] CRISPR / Cas systems may be multi-protein systems or single effector protein systems. Multiprotein, or Class 1, CRISPR systems include Type I, Type III, and Type IV systems. In some aspects, Class 2 systems include a single effector molecule and include Type II, Type V, and Type VI. In some embodiments, the DNA targeting system comprises components of CRISPR / Cas systems, such as a Type I, Type II, Type III, Type IV, Type V, or Type VI CRISPR system. In some embodiments, the Cas protein isAttorney Docket No. 224742003640 from a Class 1 CRISPR system (i.e., multiple Cas protein system), such as a Type I, Type III, or Type IV CRISPR system. In some embodiments, the Cas protein is from a Class 2 CRISPR system (i.e., single Cas protein system), such as a Type II, Type V, or Type VI CRISPR system.
[0167] Various CRISPR / Cas systems and associated Cas proteins for use in gene editing and regulation have been described, for example in Moon et al. Exp. Mol. Med. 51, 1-11 (2019), Zhang, F. Q. Rev. Biophys. 52, E6 (2019), and Makarova et al. Methods Mol. Biol. 1311:47-75 (2015).
[0168] Type I CRISPR / Cas systems employ a large multisubunit ribonucleoprotein (RNP) complex called Cascade that recognizes double-stranded DNA (dsDNA) targets. After target recognition and verification, Cascade recruits the signature protein Cas3, a fused helicase-nuclease, to degrade DNA.
[0169] In some embodiments, the Cas protein is from a Type II CRISPR system. Exemplary Cas proteins of a Type II CRISPR system include Cas9. In some embodiments, the Cas protein is from a Cas9 protein or variant thereof, for example as described in WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, WO 2014 / 093655, Jinek. et al. Science 337(6096):816-21 (2012), Mali et al. Science 339(6121): 823-6 (2013), Cong et al. Science 339(6121): 819-23 (2013), Perez-Pinera et al. Nat. Methods 10, 973-976 (2013), or Mali et al. Nat. Biotechnol. 31, 833-838 (2013). In Type II CRISPR / Cas systems with the Cas protein Cas9, two RNA molecules and the Cas9 protein form a ribonucleoprotein (RNP) complex to direct Cas9 nuclease activity. The CRISPR RNA (crRNA) contains a spacer sequence that is complementary to a target nucleic acid sequence (target site), and that encodes the sequence specificity of the complex. The trans-activating crRNA (tracrRNA) base-pairs to a portion of the crRNA and forms a structure that complexes with the Cas9 protein, forming a Cas / RNA RNP complex. Cas9 mediates cleavage of target DNA if a correct protospacer-adjacent motif (PAM) is also present at the 3' end of the protospacer. For protospacer targeting, the sequence must be immediately followed by the protospacer-adjacent motif (PAM), a short sequence recognized by the Cas9 nuclease that is required for DNA cleavage.
[0170] Different Type II systems have differing PAM requirements. The S. pyogenes CRISPR system may have the PAM sequence for this Cas9 (SpCas9) as 5'-NRG-3', where R is either A or G, and characterized the specificity of this system in human cells. A unique capability of the CRISPR / Cas9 system is the straightforward ability to simultaneously target multiple distinct genomic loci by co-expressing a single Cas9 protein with two or more sgRNAs. For example, the Streptococcus pyogenes Type II system typically prefers to use an “NGG” (SEQ ID NO: 71) sequence, where “N” can be any nucleotide, but also accepts other PAM sequences, such as “NAG” in engineered systems (Hsu et al., Nature Biotechnology (2013) doi: 10. 1038 / nbt.2647). Similarly, the Cas9 derived from Neisseria meningitidis (NmCas9) normally has a native PAM of NNNNGATT (SEQ ID NO: 73), but has activity across a variety of PAMs, including a highly degenerate NNNNGNNN PAM (Esvelt et al. Nature Methods (2013) doi: 10. 1038 / nmeth.2681). In anotherAttorney Docket No. 224742003640 example, the Cas9 derived from Campylobacter jejuni typically uses 5'-NNNNACAC-3' or 5'-NNNNRYAC- 3' (SEQ ID NO: 74) PAM sequences, where “N” can be any nucleotide, “R” can be either guanine (G) or adenine (A), and “Y” can be either cytosine (C) or thymine (T). In some aspects, the PAM sequences for spacer targeting depends on the type, ortholog, variant or species of the Cas protein.
[0171] In some embodiments, the Cas protein is derived from a Cas9 protein or variant thereof, for example as described in WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, WO 2014 / 093655, Jinek, M. et al. Science 337(6096): 816-21 (2012), Mali, P. et al. Science 339(6121):823-6 (2013), Cong, L. et al. Science 339(6121): 819-23 (2013), Perez-Pinera, P. et al. Nat. Methods 10, 973-976 (2013), or Mali, P. et al. Nat. Biotechnol. 31, 833-838 (2013). Various CRISPR / Cas systems and associated Cas proteins for use in gene editing and regulation have been described, for example in Moon, S.B. et al. Exp. Mol. Med. 51, 1-11 (2019), Zhang, F. Q. Rev. Biophys. 52, E6 (2019), and Makarova K.S. et al. Methods Mol. Biol. 1311:47-75 (2015).
[0172] In some embodiments, the Cas9 protein comprises a sequence from a Cas9 molecule of S. aureus. In some embodiments, the Cas9 protein comprises a sequence set forth in SEQ ID NO: 41 or SEQ ID NO: 48, or a variant thereof, such as an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 41 or SEQ ID NO: 48. In some embodiments, the Cas9 protein comprises a sequence from a Cas9 molecule of S. pyogenes. In some embodiments, the Cas9 protein comprises a sequence set forth in SEQ ID NO: 42 or SEQ ID NO: 49, or a variant thereof, such as an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 42 or SEQ ID NO: 49.
[0173] In Type III systems, the RNP complex is multimeric with a helicoid structure similar to Cascade. In contrast to Type I CRISPR / Cas systems, the Type III RNP complex recognizes complementary RNA sequences instead of dsDNA. RNA recognition stimulates a nonspecific DNA cleavage activity of the exemplary Type III Cas 10 nuclease that is part of the RNP complex, such that DNA cleavage is achieved cotranscriptionally.
[0174] In some embodiments, the Cas protein is from a Type V CRISPR system. Exemplary Cas proteins of a Type V CRISPR system include Cas 12a (also known as Cpfl), Cas 12b (also known as C2cl), Casl2e (also known as CasX), Cas 12k (also known as C2c5), Cas 14a, and Cas 14b. In some embodiments, the Cas protein is from a Casl2 protein (i.e., Cpfl) or variant thereof, for example as described in WO 2017 / 189308, WO2019 / 232069 and Zetsche et al. Cell. 163(3):759-71 (2015).
[0175] Exemplary Type V systems include those based on a Casl2 effector, and the C-terminus with only one RuvC endonuclease domain is the defining characteristic of the Type V systems. The RuvC nuclease domain cleaves dsDNA adjacent to protospacer adjacent motif (PAM) sequences and singleAttorney Docket No. 224742003640 stranded DNA (ssDNA) nonspecifically. The Type V systems can be further divided into subtypes, each characterized by different signature proteins, PAM sequences, and properties. Non-limiting exemplary Cas proteins derived from Type V CRISPR systems include Casl2a (Cpfl), UnlCasl2fl, Casl2j (CasPhi, such as CasPhi-2), Casl2k, and CasMini. For example, Type V-A includes, for example, Casl2a, which uses “TTTV” (SEQ ID NO: 77) PAM sequence, where “V” is adenine (A), cytosine (C), or guanine (G). Type V- F is includes, for example, Casl2f, which can use “TTTR” ,where “R” is G or A, or “TTTN” , where “N” is any nucleotide. Type V-K is includes, for example, Casl2k, which uses “GGTT” PAM sequence.
[0176] In some embodiments, the Cas 12a protein comprises a sequence from a Cas 12a molecule of Acidaminococcus sp, such as an AsCasl2a set forth in SEQ ID NO: 50 or SEQ ID NO: 51, or a variant thereof, such as an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 50 or SEQ ID NO: 51.
[0177] Non-limiting examples of Cas9 orthologs from other bacterial strains include but are not limited to: Cas proteins identified in Acaryochloris marina MBIC11017; Acetohalobium arabaticum DSM 5501; Acidithiobacillus caldus; Acidithiobacillus ferrooxidans ATCC 23270; Alicyclobacillus acidocaldarius LAA1; Alicyclobacillus acidocaldarius subsp. acidocaldarius DSM 446; Allochromatium vinosum DSM 180; Ammonifex degensii KC4; Anabaena variabilis ATCC 29413; Arthrospira maxima CS-328; Arthrospira platensis str. Paraca; Arthrospira sp. PCC 8005; Bacillus pseudomycoides DSM 12442; Bacillus selenitireducens MLS 10; Burkholderiales bacterium 1 1 47; Caldicelulosiruptor becscii DSM 6725; Campylobacter jejuni; Candidatus Desulforudis audaxviator MP104C; Caldicellulosiruptor hydrothermalis 108; Clostridium phage c-st; Clostridium botulinum A3 str. Loch Maree; Clostridium botulinum Ba4 str. 657; Clostridium difficile QCD-63q42; Crocosphaera watsonii WH 8501; Cyanothece sp. ATCC 51142; Cyanothece sp. CCY0110; Cyanothece sp. PCC 7424; Cyanothece sp. PCC 7822; Exiguobacterium sibiricum 255-15; Finegoldia magna ATCC 29328; Ktedonobacter racemifer DSM 44963; Lactobacillus delbrueckii subsp. bulgaricus PB2003 / 044-T3-4; Lactobacillus salivarius ATCC 11741; Listeria innocua; Lyngbya sp. PCC 8106; Marinobacter sp. ELB17; Methanohalobium evestigatum Z-7303; Microcystis phage Ma-LMMOl; Microcystis aeruginosa NIES-843; Microscilla marina ATCC 23134; Microcoleus chthonoplastes PCC 7420; Neisseria meningitidis; Nitrosococcus halophilus Nc4; Nocardiopsis dassonvillei subsp. dassonvillei DSM 43111; Nodularia spumigena CCY9414; Nostoc sp. PCC 7120; Oscillatoria sp. PCC 6506; Pelotomaculum thermopropionicum SI; Pefrotoga mobilis SJ95; Polaromonas naphthalenivorans CJ2; Polaromonas sp. JS666; Pseudoalteromonas haloplanktis TAC125; Streptomyces pristinaespiralis ATCC 25486; Streptomyces pristinaespiralis ATCC 25486; Streptococcus thermophilus; Streptomyces viridochromogenes DSM 40736; Streptosporangium roseum DSM 43021; Synechococcus sp. PCC 7335; and Thermosipho africanus TCF52B (Chylinski et al., RNA Biol., 2013; 10(5): 726-737).Attorney Docket No. 224742003640
[0178] In some embodiments, the DNA-targeting systems or fusion proteins comprise a Cas protein, such as a Cas protein set forth in any one of SEQ ID NOS: 41, 42, and 48-59, or a variant thereof, such as an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any one of SEQ ID NOS: 41, 42, and 48-59. In some aspects, the Cas protein lacks an initial methionine residue. In some aspects, the Cas protein comprises an initial methionine residue.
[0179] In some aspects, the Cas protein is a variant that lacks nuclease activity (i.e., is a dCas protein). In some embodiments, the Cas protein is mutated so that nuclease activity is reduced or eliminated. Such Cas proteins are referred to as deactivated Cas or dead Cas (dCas) or nuclease-inactive Cas (iCas) proteins, as referred to interchangeably herein. In some embodiments, the Cas protein is a variant Cas9 protein that lacks nuclease activity or that is a deactivated Cas9 (dCas9, or iCas9) protein.
[0180] In some aspects, in the provided fusion proteins, the DNA-binding domain, e.g., Cas, is a deactivated Cas (dCas), or a nuclease-inactive Cas (iCas). In some embodiments, the component of the DNA-binding domain, such as a protein component, comprises a Cas9 variant such as a deactivated Cas9 or inactivated Cas9. In some embodiments, the component of the DNA-binding domain, such as a protein component, comprises a Casl2a variant such as a deactivated Casl2a (Cpfl) or inactivated Casl2a (Cpfl). In some aspects, the Cas9 protein may be mutated so that the nuclease activity is deactivated or inactivated (also referred to as dCas9 or iCas9). In some aspects, the Cas protein is a variant that lacks nuclease activity (i.e., is a dCas protein). In some embodiments, the Cas protein is mutated so that nuclease activity is reduced or eliminated. Such Cas proteins are referred to as deactivated Cas or dead Cas (dCas) or nuclease-inactive Cas (iCas) proteins, as referred to interchangeably herein. In some embodiments, the variant Cas protein is a variant Cas9 protein that lacks nuclease activity or that is a deactivated Cas9 (dCas9, or iCas9) protein. In some embodiments, the variant Cas protein is a variant Cpfl protein that lacks nuclease activity or that is a deactivated Casl2a (dCasl2a, or iCasl2a) protein.
[0181] In some embodiments, Cas proteins are engineered to be catalytically inactivated or nuclease inactive to allow targeting of Cas / gRNA RNPs without inducing cleavage at the target site. Mutations in Cas proteins can reduce or abolish nuclease activity of the Cas protein, rendering the Cas protein catalytically inactive. Cas proteins with reduced or abolished nuclease activity are referred to as deactivated Cas (dCas), or nuclease-inactive Cas (iCas) proteins, as referred to interchangeably herein. In some aspects, the dCas or iCas can still bind to target site in the DNA in a site- and / or sequence-specific manner, as long as it retains the ability to interact with the guide RNA (gRNA) which directs the Cas-gRNA combination to the target site.
[0182] In some aspects, the dCas or iCas exhibits reduced or no endodeoxyribonuclease activity. For example, an exemplary dCas or iCas, for example dCas9 or iCas9, exhibits less than about 20%, less thanAttorney Docket No. 224742003640 about 15%, less than about 10%, less than about 5%, less than about 1%, or less than about 0. 1%, of the endodeoxyribonuclease activity of a wild-type Cas protein, e.g., a wild-type Cas9 protein. In some embodiments, the dCas or iCas, for example dCas9 or iCas9, exhibits substantially no detectable endodeoxyribonuclease activity. In some embodiments, an exemplary dCas or iCas, for example dCas9 or iCas9, comprises one or more amino acid mutations, substitutions, deletions or insertions at a position corresponding to a position selected from DIO, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or a A987, with reference to a wild-type Streptococcus pyogenes Cas9 (SpCas9), for example, with reference to numbering of positions of a SpCas9 sequence set forth in SEQ ID NO: 42. In some aspects, the dCas9 or iCas9 comprises one or more amino acid mutations, substitutions, deletions or insertions corresponding to D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A, with reference to a wild-type Streptococcus pyogenes Cas9 (SpCas9), for example, with reference to numbering of positions of a SpCas9 sequence set forth in SEQ ID NO: 42. Corresponding positions for mutations can be determined based on sequence alignments and determination of sequence conservation, for example, as described in WO 2013 / 171772 for Cas9 proteins from various species. In some aspects, the dCas protein lacks an initial methionine residue. In some aspects, the dCas protein comprises an initial methionine residue.
[0183] In some embodiments, the dCas9 protein comprises a sequence derived from a naturally occurring Cas9 molecule, or variant thereof. In some embodiments, the dCas9 protein can comprise a sequence derived from a naturally occurring Cas9 molecule of S. pyogenes, S. thermophilus, S. aureus, C. jejuni, N. meningitidis, F. novicida, S. canis, S. auricularis, or variant thereof. In some embodiments, the dCas9 protein comprises a sequence derived from a naturally occurring Cas9 molecule of S. aureus. In some embodiments, the dCas9 protein comprises a sequence derived from a naturally occurring Cas9 molecule of S. pyogenes. In some embodiments, the dCas9 protein comprises a sequence from a Cas9 molecule of C. jejuni.
[0184] Exemplary deactivated Cas9 (dCas9) derived from S. pyogenes contains silencing mutations of the RuvC and HNH nuclease domains (D10A and H840A), for example as described in WO 2013 / 176772, WO 2014 / 093661, Jinek et al. Science 337(6096): 816-21 (2012), and Qi et al. Cell 152(5): 1173-83 (2013). Exemplary dCas variants derived from the Cas 12 system (i.e. Cpfl) are described, for example in WO 2017 / 189308 and Zetsche et al. Cell 163(3):759-71 (2015). Conserved domains that mediate nucleic acid cleavage, such as RuvC and HNH endonuclease domains, are readily identifiable in Cas orthologues, and can be mutated to produce inactive variants, for example as described in Zetsche et al. Cell 163(3):759-71 (2015). Other exemplary Cas orthologs or variants include engineered variants based on a Casl2f (also known as Casl4), including those described in Xu et al., Mol. Cell 81 (20):4333-4345 (2021).Attorney Docket No. 224742003640
[0185] In some embodiments, the DNA-binding domain comprises a Cas-gRNA combination that includes (a) a Cas protein or a variant thereof and (b) at least one gRNA. In some embodiments, the variant Cas protein lacks nuclease activity or is a deactivated Cas (dCas) protein. In some embodiments, the gRNA is capable of complexing with the Cas protein or variant thereof. In some embodiments, the gRNA comprises a gRNA spacer sequence that is capable of hybridizing to the target site or is complementary to the target site at a target gene.
[0186] In some embodiments, the Cas protein or a variant thereof is a Cas9 protein or a variant thereof. In some embodiments, the variant Cas protein is a variant Cas9 protein that lacks nuclease activity or that is a deactivated Cas9 (dCas9) protein. In some embodiments, the Cas9 protein or a variant thereof is a Staphylococcus aureus Cas9 (SaCas9) protein or a variant thereof. In some embodiments, the variant Cas9 protein is a Staphylococcus aureus dCas9 protein (dSaCas9) that comprises at least one amino acid mutation selected from D10A and N580A, with reference to numbering of positions of SEQ ID NO: 48. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 40 or SEQ ID NO: 60, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 40, which lacks an initial methionine residue. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 60, which includes an initial methionine residue.
[0187] In some embodiments, the Cas9 protein or variant thereof is a Streptococcus pyogenes Cas9 (SpCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Streptococcus pyogenes dCas9 (dSpCas9) protein that comprises at least one amino acid mutation selected from D10A and H840A, with reference to numbering of positions of SEQ ID NO: 42. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 18 or SEQ ID NO: 61, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 18, which lacks an initial methionine residue. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 61, which includes an initial methionine residue.
[0188] In some embodiments, the Cas9 protein or variant thereof is a Campylobacter jejuni Cas9 (CjCas9) protein or a variant thereof. In some embodiments, the variant Cas9 comprises at least one amino acid mutation compared to the sequence set forth in SEQ ID NO: 56 or 57. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 56, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 57, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In someAttorney Docket No. 224742003640 embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 57, which lacks an initial methionine residue. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 56, which includes an initial methionine residue.
[0189] In some embodiments, the Cas protein or a variant thereof is a Casl2a protein or a variant thereof. In some embodiments, the variant Cas protein is a variant Cas 12a protein that lacks nuclease activity or that is a deactivated Casl2a (dCasl2a) protein. In some embodiments, the Casl2a protein or variant thereof is a Acidaminococcus sp. Cas 12a (AsCasl2a) protein or a variant thereof. In some embodiments, the variant Casl2a is a Acidaminococcus sp. dCasl2a (dAsCasl2a) protein that comprises at least one amino acid mutation compared to the sequence set forth in SEQ ID NO: 50 or 51. In some embodiments, the variant Cas 12a protein comprises the sequence set forth in SEQ ID NO: 63, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Cas 12a protein comprises the sequence set forth in SEQ ID NO: 64, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Casl2a protein comprises the sequence set forth in SEQ ID NO: 64, which lacks an initial methionine residue. In some embodiments, the variant Cas 12a protein comprises the sequence set forth in SEQ ID NO: 63, which includes an initial methionine residue.
[0190] In some embodiments, the Cas protein or a variant thereof is a CasPhi-2 protein or a variant thereof. In some embodiments, the variant Cas protein is a variant CasPhi-2 protein that lacks nuclease activity or that is a deactivated CasPhi-2 (dCasPhi-2) protein. In some embodiments, the variant CasPhi-2 comprises at least one amino acid mutation compared to the sequence set forth in SEQ ID NO: 52 or 53. In some embodiments, the variant CasPhi-2 protein comprises the sequence set forth in SEQ ID NO: 62, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant CasPhi-2 protein comprises the sequence set forth in SEQ ID NO: 65, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant CasPhi-2 protein comprises the sequence set forth in SEQ ID NO: 65, which lacks an initial methionine residue. In some embodiments, the variant CasPhi-2 protein comprises the sequence set forth in SEQ ID NO: 62, which includes an initial methionine residue.
[0191] In some embodiments, the Cas protein or a variant thereof is a UnlCasl2fl protein or a variant thereof. In some embodiments, the variant Cas protein is a variant UnlCasl2fl protein that lacks nuclease activity or that is a deactivated UnlCasl2fl (dUnlCasl2fl) protein. In some embodiments, the variant UnlCasl2fl comprises at least one amino acid mutation compared to the sequence set forth in SEQ ID NO: 54 or 55. In some embodiments, the variant UnlCasl2fl protein comprises the sequence set forth in SEQ IDAttorney Docket No. 224742003640NO: 54, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant UnlCasl2fl protein comprises the sequence set forth in SEQ ID NO: 55, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant UnlCasl2fl protein comprises the sequence set forth in SEQ ID NO: 55, which lacks an initial methionine residue. In some embodiments, the variant UnlCasl2fl protein comprises the sequence set forth in SEQ ID NO: 54, which includes an initial methionine residue.
[0192] In some embodiments, the Cas protein or a variant thereof is a Casl2k protein or a variant thereof. In some embodiments, the Casl2k protein comprises the sequence set forth in SEQ ID NO: 58, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the Casl2k protein comprises the sequence set forth in SEQ ID NO: 59, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the Cas 12k protein comprises the sequence set forth in SEQ ID NO: 59, which lacks an initial methionine residue. In some embodiments, the Casl2k protein comprises the sequence set forth in SEQ ID NO: 58, which includes an initial methionine residue.
[0193] In some embodiments, the Cas protein or a variant thereof is a CasMini protein or a variant thereof, such as an engineered Cas protein or variant based on a Casl2f (also known as Casl4), including those described in Xu et al., Mol. Cell 81(20):4333-4345 (2021) or set forth in SEQ ID NO: 66. In some embodiments, the variant Cas protein is a variant CasMini protein that lacks nuclease activity or that is a deactivated CasMini (dCasMini) protein. In some embodiments, the variant CasMini comprises at least one amino acid mutation compared to the sequence set forth in SEQ ID NO: 66. In some embodiments, the variant CasMini protein comprises the sequence set forth in SEQ ID NO: 66, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the CasMini protein comprises the sequence set forth in SEQ ID NO: 66. In some embodiments, the variant CasMini protein comprises the sequence set forth in SEQ ID NO: 67 or 68, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the CasMini protein comprises the sequence set forth in SEQ ID NO: 68, which lacks an initial methionine residue. In some embodiments, the CasMini protein comprises the sequence set forth in SEQ ID NO: 67, which includes an initial methionine residue.
[0194] In some embodiments, the Cas protein is a split protein, i.e., comprises two or more separate polypeptide domains that interact or self-assemble to form a functional fusion protein. In some embodiments, the split Cas protein is assembled from separate polypeptide domains comprising trans-splicing inteins. Inteins are internal protein elements that self-excise from their host protein and catalyze ligation of flankingAttorney Docket No. 224742003640 sequences with a peptide bond. In some embodiments, the split Cas protein is assembled from a first polypeptide comprising an N-terminal intein and a second polypeptide comprising a C-terminal intein. In some embodiments, the N terminal intein is the N terminal Npu Intein set forth in SEQ ID NO: 69. In some embodiments, the C terminal intein is the C terminal Npu intein set forth in SEQ ID NO: 70.
[0195] Also provided are Cas proteins comprising a first polypeptide of a split variant Cas protein comprising an N-terminal fragment of a Cas protein and an N-terminal Intein, and any of the multipartite effector domains provided herein. Also provided are fusion proteins comprising a first polypeptide of a split variant Cas protein comprising an N-terminal fragment of a Cas protein and an N-terminal Intein, and any of the multipartite effector domains provided herein, wherein the multipartite effector domain increases transcription of a target locus. In some aspects, the first polypeptide of the split variant Cas protein, and a second polypeptide of the split variant Cas protein comprising a C-terminal fragment of the variant Cas protein and a C-terminal Intein, are present in proximity or present in the same cell, the N-terminal Intein and C-terminal Intein self-excise and ligate the N-terminal fragment and the C-terminal fragment of the variant Cas9 to form a full-length variant Cas9 protein.
[0196] Also provided are fusion proteins comprising a second polypeptide of a split variant Cas protein comprising a C-terminal fragment of a Cas protein and a C-terminal Intein and any of the multipartite effector domains provided herein. Also provided are fusion proteins comprising a second polypeptide of a split variant Cas protein comprising a C-terminal fragment of a Cas protein and a C-terminal Intein and any of the multipartite effector domains provided herein, wherein the multipartite effector domain decreases transcription of a target locus. In some aspects, the second polypeptide of the split variant Cas protein, and a first polypeptide of the split variant Cas protein comprising an N-terminal fragment of the variant Cas protein and an N-terminal Intein, are present in proximity or present in the same cell, the N-terminal Intein and C- terminal Intein self-excise and ligate the N-terminal fragment and the C-terminal fragment of the variant Cas9 to form a full-length variant Cas9 protein.
[0197] In some embodiments, the split Cas protein comprises a split dCas9, such as a split dSpCas9. In an exemplary embodiment, a first polypeptide comprises an N-terminal fragment of dSpCas9, followed by an N terminal Npu Intein, and a second polypeptide comprises a C terminal Npu Intein, followed by a C- terminal fragment of dSpCas9. In some embodiments, the N- and C-terminal fragments of the fusion protein are split at position 573Glu of the SpCas9 molecule, with reference to SEQ ID NO: 42.
[0198] In some aspects, the N-terminal Npu Intein and C-terminal Npu Intein may self-excise and ligate the two fragments, thereby forming the full-length dSpCas9 fusion protein when expressed in a cell.
[0199] In some embodiments, the polypeptides of a split Cas protein may interact non-covalently to form a complex that recapitulates the activity of the non-split protein. For example, two domains of a CasAttorney Docket No. 224742003640 enzyme expressed as separate polypeptides may be recruited by a gRNA to form a ternary complex that recapitulates the activity of the full-length Cas enzyme in complex with the gRNA, for example as described in Wright et al. PNAS 112(10):2984-2989 (2015). In some embodiments, assembly of the split protein is inducible (e.g., light inducible, chemically inducible, small-molecule inducible).
[0200] In some aspects, the two polypeptides of a split Cas protein may be delivered and / or expressed from separate vectors, such as any of the vectors described herein. In some embodiments, the two polypeptides of a split fusion protein may be delivered to a cell and / or expressed from two separate AAV vectors, i.e., using a split AAV-based approach, for example as described in WO 2017 / 197238.
[0201] Approaches for the rationale design of split Cas proteins and their delivery are described, for example, in WO 2016 / 114972, WO 2017 / 197238, Zetsche, et al. Nat. Biotechnol. 33(2): 139-42 (2015), Wright et al. PNAS 112(10):2984-2989 (2015), Truong, et al. Nucleic Acids Res. 43, 6450-6458 (2015), and Fine et al. Sci. Rep. 5, 10777 (2015).
[0202] DNA-targeting systems, in some cases comprising a fusion protein, such as dCas-fusion proteins include fusion of the Cas with an effector domain, such as a transcription activation domain. Any of a variety of effector domains, for example those that decrease transcription from the target locus, including any described herein, for example, in Section I.A.2 and I.B.2, can be used.
[0203] In some aspects, provided is a DNA-targeting system comprising a fusion protein comprising a DNA-binding domain comprising a nuclease-inactive Cas protein or variant thereof, and an effector domain for decreasing or repressing transcription (i.e., a transcriptional repressor) when targeted to a target site in a gene or regulatory element thereof. In some aspects, the DNA-targeting system also includes one or more gRNA, provided in combination or as a complex with the dCas protein or variant thereof, for targeting of the DNA-targeting system to the target site. In some embodiments, the fusion protein is guided to a specific target site sequence of the target gene by the guide RNA, wherein the effector domain mediates targeted epigenetic modification to increase or promote transcription of the target gene.2) Guide RNAs
[0204] In some embodiments, the Cas protein (e.g., dCas9) is provided in combination or as a complex with one or more guide RNA (gRNA). In some aspects, the gRNA is a nucleic acid that promotes the specific targeting or homing of the gRNA / Cas ribonucleoprotein (RNP) complex to the target site of the target gene, such as any described above. In some embodiments, a target site of a gRNA may be referred to as a protospacer.
[0205] Provided herein are gRNAs, such as gRNAs that target or bind to a target site, such as any described herein. In some embodiments, the gRNA is capable of complexing with the Cas protein or variantAttorney Docket No. 224742003640 thereof. In some embodiments, the gRNA comprises a gRNA spacer sequence (i.e., a spacer sequence or a guide sequence) that is capable of hybridizing to the target site, or that is complementary to the target site. In some embodiments, the gRNA comprises a scaffold sequence that complexes with or binds to the Cas protein.
[0206] In some embodiments, the gRNAs provided herein are chimeric gRNAs. In general, gRNAs can be unimolecular (i.e., composed of a single RNA molecule), or modular (comprising more than one, and typically two, separate RNA molecules). Modular gRNAs can be engineered to be unimolecular, wherein sequences from the separate modular RNA molecules are comprised in a single gRNA molecule, sometimes referred to as a chimeric gRNA, synthetic gRNA, or single gRNA. In some embodiments, the chimeric gRNA is a fusion of two non-coding RNA sequences: a crRNA sequence and a tracrRNA sequence, for example as described in WO 2013 / 176772, or Jinek, M. et al. Science 337(6096):816-21 (2012). In some embodiments, the chimeric gRNA mimics the naturally occurring crRNA:tracrRNA duplex involved in the Type II Effector system, wherein the naturally occurring crRNA:tracrRNA duplex acts as a guide for the Cas9 protein.
[0207] In some aspects, the gRNA comprises a scaffold sequence, which complexes with the Cas protein. Scaffold sequences are Cas protein-specific, such that different scaffold sequences complex (i.e. are compatible) with different Cas molecules. gRNAs are readily designed with scaffold sequences that are compatible with the desired Cas protein.
[0208] In some aspects, the spacer sequence of a gRNA is a polynucleotide sequence comprising at least a portion that has sufficient complementarity with the target site to hybridize with the target site and direct sequence-specific binding of a CRISPR complex to the sequence of the target site. Full complementarity is not necessarily required, provided there is sufficient complementarity to cause hybridization and promote formation of a CRISPR complex. In some embodiments, the gRNA comprises a spacer sequence that is complementary, e.g., at least 80%, 85%, 90%, 95%, 98%, 99%, or 100% (e.g., fully complementary), to the target site. A spacer sequence may be selected to reduce the degree of secondary structure within the spacer sequence. Secondary structure may be determined by any suitable polynucleotide folding algorithm.
[0209] In some embodiments, the gRNA spacer sequence is between about 14 nucleotides (nt) and about 26 nt, or between 16 nt and 22 nt in length. In some embodiments, the gRNA spacer sequence is 14 nt, 15 nt, 16 nt, 17 nt, 18 nt, 19 nt, 20 nt, 21 nt or 22 nt, 23 nt, 24 nt, 25 nt, or 26 nt in length. In some embodiments, the gRNA spacer sequence is 18 nt, 19 nt, 20 nt, 21 nt or 22 nt in length.
[0210] A target site of a gRNA may be referred to as a protospacer. In some aspects, the spacer is designed to target a protospacer (i.e. target site) with a specific protospacer-adjacent motif (PAM), i.e. a sequence immediately adjacent to the protospacer that contributes to and / or is required for Cas bindingAttorney Docket No. 224742003640 specificity. Different CRISPR / Cas systems have different PAM requirements for targeting. For example, S. pyogenes Cas9 targets sequences having the PAM 5’-NGG-3’ (SEQ ID NO: 71); S. aureus Cas9 uses the PAM 5’- NNGRRT-3’ (SEQ ID NO: 72); N. meningitidis Cas9 targets sequences having the PAM 5'- NNNNGATT -3’ (SEQ ID NO: 73); C. jejuni Cas9 targets sequences having the PAM 5'-NNNNRYAC-3', (SEQ ID NO: 74); S. thermophilus targets sequences having the PAM 5’-NNAGAAW-3’ (SEQ ID NO: 75); F. Novicida Cas9 targets sequences having the PAM 5’-NGG-3’ (SEQ ID NO: 71); T. denticola Cas9 targets sequences having the PAM 5’-NAAAAC-3’ (SEQ ID NO: 76); Casl2a (also known as Cpfl) targets sequences having the PAM 5’-TTTV-3’ (SEQ ID NO: 77). Cas proteins may target or be engineered to target sequences having different PAMs from those listed above. For example, variant SpCas9 proteins may target sequences having a PAM selected from: 5’-NGG-3’ (SEQ ID NO: 71), 5’-NGAN-3’ (SEQ ID NO: 78), 5’- NGNG-3’(SEQ ID NO: 79), 5’-NGAG-3’(SEQ ID NO: 80), or 5’-NGCG-3’(SEQ ID NO: 81).
[0211] In some embodiments, the gRNA (including the guide sequence) will comprise the base uracil (U), whereas DNA encoding the gRNA molecule will comprise the base thymine (T). While not wishing to be bound by theory, in some embodiments, it is believed that the complementarity of the guide sequence with the target sequence contributes to specificity of the interaction of the gRNA molecule / Cas molecule complex with a target nucleic acid. It is understood that in a guide sequence and target sequence pair, the uracil bases in the guide sequence will pair with the adenine bases in the target sequence. A gRNA spacer sequence herein may be defined by the DNA sequence encoding the gRNA spacer, and / or the RNA sequence of the spacer.
[0212] In some embodiments, one, more than one, or all of the nucleotides of a gRNA can have a modification, e.g., to render the gRNA less susceptible to degradation and / or improve bio-compatibility. By way of non-limiting example, the backbone of the gRNA can be modified with a phosphorothioate, or other modification s). In some cases, a nucleotide of the gRNA can comprise a 2’ modification, e.g., a 2- acetylation, e.g., a 2’ methylation, or other modification(s).
[0213] Methods for designing gRNAs and exemplary binding domains can include those described in, e.g., International PCT Pub. Nos. WO 2014 / 197748, WO 2016 / 130600, WO 2017 / 180915, WO 2021 / 226555, WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, WO 2014 / 093655, WO 2015 / 089427, WO 2016 / 049258, WO 2016 / 123578, WO 2021 / 076744, WO 2014 / 191128, WO 2015 / 161276, WO 2017 / 193107, and WO 2017 / 093969.
[0214] In some embodiments, a gRNA provided herein targets a target site, such as any target site described herein in Section ILA or in the PCSK9 locus. In some embodiments, the gRNA targets a target site for the target gene PCSK9, such as for any suitable target gene (e.g., target site and target genes in Section ILA). In some embodiments the gRNA hybridizes to the sequence complementary to the sequence defined asAttorney Docket No. 224742003640 the target site. The strand of the target nucleic acid comprising the target site sequence may be referred to as the “complementary strand” of the target nucleic acid. gRNAs that target a target site in the PCSK9 locus are known. International published PCT Appl. No. WO 2023 / 250511, WO 2023 / 173110, WO 2023 / 093862, WO 2023 / 215711, WO 2023 / 240076, and US published Appl. No. US20240052328, the disclosures of which are incorporated by reference in their entireties.
[0215] In some embodiments, the gRNA targets a target site that comprises a sequence selected from any one of SEQ ID NOS: 43, 155-157, and 162-167, a contiguous portion thereof of at least 14 nucleotides, a complementary sequence of any of the foregoing, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any of the foregoing. In some embodiments, the target site is a contiguous portion of any one of SEQ ID NOS: 43, 155- 157, and 162-167 that is 14, 15, 16, 17, 18 or 19 nucleotides in length. In some embodiments, the target site is set forth in any one of SEQ ID NOS: 43, 155-157, and 162-167.
[0216] In some embodiments, the gRNA targets a target site that comprises a sequence selected from any one of SEQ ID NOS: 43 and 155-157, a contiguous portion thereof of at least 14 nucleotides, a complementary sequence of any of the foregoing, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any of the foregoing. In some embodiments, the target site is a contiguous portion of any one of SEQ ID NOS: 43 and 155-157 that is 14, 15, 16, 17, 18 or 19 nucleotides in length. In some embodiments, the target site is set forth in any one of SEQ ID NOS: 43 and 155-157.
[0217] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO:43 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO:43.
[0218] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 155 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 155.
[0219] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 156 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 156.
[0220] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 157 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 157.
[0221] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 162 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets aAttorney Docket No. 224742003640 target site that includes or is set forth in SEQ ID NO: 162.
[0222] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 163 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 163.
[0223] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 164 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 164.
[0224] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 165 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 165.
[0225] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 166 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 166.
[0226] In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 167 or a contiguous portion thereof of at least 14 nucleotides. In some embodiments, the gRNA targets a target site that includes or is set forth in SEQ ID NO: 167.
[0227] In some embodiments, the gRNA further comprises a scaffold sequence. In some embodiments, the scaffold sequence comprises any one of the sequences set forth in SEQ ID NOs: 182 and 389-391, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to all or a portion thereof. In some embodiments, the scaffold sequence comprises any one of the sequences set forth in SEQ ID NOs: 182 and 389-391.
[0228] In some embodiments, an exemplary scaffold sequence for S. pyogenes Cas9 comprises a sequence set forth in SEQ ID NO: 182 or SEQ ID NO: 389, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity thereof. In some embodiments, the scaffold sequence comprises the sequence set forth in SEQ ID NO: 182 (GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACU UGAAAAAGUGGCACCGAGUCGGUGC), or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to all or a portion thereof. In some embodiments, the scaffold sequence is set forth in SEQ ID NO: 182. In some embodiments, the scaffold sequence comprises the sequence set forth in SEQ ID NO: 389 (GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCU), or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to all or a portion thereof. In someAttorney Docket No. 224742003640 embodiments, the scaffold sequence is set forth in SEQ ID NO: 389.
[0229] In some embodiments, an exemplary scaffold sequence comprises a 3’ polyU sequence. In some embodiments, an exemplary scaffold sequence for S. pyogenes Cas9 comprises a sequence set forth in SEQ ID NO: 390 or SEQ ID NO: 391. In some embodiments, the scaffold sequence comprises the sequence set forth in SEQ ID NO: 390 (GUUUUAGAGCUAGAAAUAGCAAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCUUUU), or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to all or a portion thereof. In some embodiments, the scaffold sequence is set forth in SEQ ID NO: 390. In some embodiments, the scaffold sequence comprises the sequence set forth in SEQ ID NO: 391 (GUUUAAGAGCUAUGCUGGAAACAGCAUAGCAAGUUUAAAUAAGGCUAGUCCGUUAUCAACU UGAAAAAGUGGCACCGAGUCGGUGCUUU), or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to all or a portion thereof. In some embodiments, the scaffold sequence is set forth in SEQ ID NO: 391.
[0230] In some embodiments, a gRNA provided herein comprises a spacer sequence selected from any one of SEQ ID NOS: 43, 155-157, and 162-167. In some embodiments, a gRNA provided herein comprises a spacer sequence selected from any one of SEQ ID NOS: 43 and 155-157. In some embodiments, the gRNA further comprises a scaffold sequence set forth in any one of SEQ ID NOs: 182 and 389-391, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity thereto. In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 182, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to SEQ ID NO: 182. In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 182. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 168-171 and 176- 181, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 168-171 and 176-181. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 168-171 and 176-181. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 168-171, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 168-171. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 168-171.
[0231] In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 391, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%,Attorney Docket No. 22474200364099.5%, 99.9%, or 100% sequence identity to SEQ ID NO: 391. In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 391. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 392-395 and 400-405, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 392-395 and 400-405. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 392-395 and 400-405. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 392-395, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 392-395. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 392-395.
[0232] In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO:389, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to SEQ ID NO: 389. In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 389. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 406-409 and 414-419, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 406-409 and 414-419. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 406-409 and 414-419. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 406-409, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 406-409. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 406-409.
[0233] In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO:390, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to SEQ ID NO: 390. In some embodiments, the gRNA further comprises a scaffold sequence set forth in SEQ ID NO: 390. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 420-423 and 428-432, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 420-423 and 428-432. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 420-423 and 428-432. In some embodiments, the gRNA comprises the sequence selected from any one of SEQ ID NOS: 420-423, or a sequence having at or at least 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100% sequence identity to any one of SEQ ID NO: 420-423. In some embodiments, the gRNA is set forth in any one of SEQ ID NOS: 420-423.Attorney Docket No. 224742003640
[0234] In some embodiments, any one of the provided gRNA sequences is complexed with or is provided in combination with a Cas9. In some embodiments, the Cas9 is a dCas9. In some embodiments, the dCas9 is a dSpCas9, such as a dSpCas9 set forth in SEQ ID NO: 18, or a variant and / or fusion thereof. b. Other DNA-binding domains
[0235] In some of any of the provided embodiments, the DNA-binding domain comprises a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an LScel enzyme or a variant thereof. In some embodiments, the DNA-binding domain binds to the target site, e.g., at the endogenous locus and / or for the target gene. In some embodiments, the DNA-binding domain comprises a catalytically inactive variant of any of the foregoing.
[0236] In some embodiments, the DNA-binding domain comprises a zinc finger protein (ZFP), i.e., is a ZFP-based DNA-binding domain. In some embodiments, a zinc finger protein (ZFP), a zinc finger DNA binding protein, or zinc finger DNA binding domain, is a protein, or a domain within a larger protein, that binds DNA in a sequence-specific manner through one or more zinc fingers, which are regions of amino acid sequence within the binding domain, having a structure that is stabilized through coordination of a zinc ion. The term zinc finger DNA binding protein is often abbreviated as zinc finger protein or ZFP. Among the ZFPs are artificial, or engineered, ZFPs, comprising ZFP domains targeting specific DNA sequences, typically 9-18 nucleotides long, generated by assembly of individual fingers. ZFPs include those in which a single finger domain is approximately 30 amino acids in length and contains an alpha helix containing two invariant histidine residues coordinated through zinc with two cysteines of a single beta turn, and having two, three, four, five, or six fingers. Generally, sequence-specificity of a ZFP may be altered by making amino acid substitutions at the four helix positions (-1, 2, 3, and 6) on a zinc finger recognition helix. Thus, for example, the ZFP or ZFP-containing molecule is non-naturally occurring, e.g., is engineered to bind to a target site of choice.
[0237] In some cases, the DNA-targeting system is or comprises a zinc-finger DNA binding domain fused to an effector domain. In some embodiments, zinc fingers are custom-designed (i.e., designed by the user), or obtained from a commercial source. Various methods for designing zinc finger proteins are available. For example, methods for designing zinc finger proteins to bind to a target DNA sequence of interest are described, for example in Liu, Q. et al., PNAS, 94(11):5525-30 (1997); Wright, D.A. et al., Nat. Protoc., 1(3): 1637-52 (2006); Gersbach, C.A. et al., Acc. Chem. Res., 47(8):2309-18 (2014); Bhakta M.S. et al., Methods Mol. Biol., 649:3-30 (2010); and Gaj et al., Trends Biotechnol, 31(7):397-405 (2013). In addition, various web-based tools for designing zinc finger proteins to bind to a DNA target sequence of interest are publicly available. See, for example, the Zinc Finger Tools design web site from ScrippsAttorney Docket No. 224742003640 available on the world wide web at scripps.edu / barbas / zfdesign / zfdesignhome.php. Various commercial services for designing zinc finger proteins to bind to a DNA target sequence of interest are also available. See, for example, the commercially available services or kits offered by Creative Biolabs (world wide web at creative-biolabs.com / Design-and-Synthesis-of-Artificial-Zinc-Finger-Proteins.html), the Zinc Finger Consortium Modular Assembly Kit available from Addgene (world wide web at addgene.org / kits / zfc- modular-assembly / ), or the CompoZr Custom ZFN Service from Sigma Aldrich (world wide web at sigmaaldrich.com / life-science / zinc-fmger-nuclease-technology / custom-zfn.html). For example, platforms for zinc-finger construction are available that provide specifically targeted zinc fingers for thousands of targets. See, e.g., Gaj et al., Trends in Biotechnology, 2013, 31(7), 397-405. Some gene-specific engineered zinc fingers are available commercially. In some cases, commercially available zinc fingers are used or are custom designed.
[0238] In some embodiments, the ZFP binds to, or is capable of binding to (i.e., targets), a target site described herein, such as any target site described in Section II. A or in the PCSK9 locus. In some embodiments, the ZFP facilitates target-specific binding of a fusion protein comprising the ZFP.
[0239] ZFPs that bind to a target site in the PCSK9 locus are known. See, e.g., WO 2023 / 215711, the disclosure of which is incorporated by reference in its entirety.
[0240] In some embodiments, the DNA-binding domain is based on transcription activator-like effectors (TALEs), i.e., is a TALE-based DNA-binding domain. TALEs are proteins naturally found in Xanthomonas bacteria. TALEs comprise a plurality of repeated amino acid sequences, each repeat having binding specificity for one base in a target sequence. Each repeat comprises a pair of variable residues in position 12 and 13 (repeat variable diresidue; RVD) that determine the nucleotide specificity of the repeat. In some embodiments, RVDs associated with recognition of the different nucleotides are HD for recognizing C, NG for recognizing T, NI for recognizing A, NN for recognizing G or A, NS for recognizing A, C, G or T, HG for recognizing T, IG for recognizing T, NK for recognizing G, HA for recognizing C, ND for recognizing C, HI for recognizing C, HN for recognizing G, NA for recognizing G, SN for recognizing G or A and YG for recognizing T, TL for recognizing A, VT for recognizing A or G and SW for recognizing A. In some embodiments, RVDs can be mutated towards other amino acid residues in order to modulate their specificity towards nucleotides A, T, C and G and in particular to enhance this specificity. Binding domains with similar modular base-per-base nucleic acid binding properties can also be derived from different bacterial species. These alternative modular proteins may exhibit more sequence variability than TALE repeats.
[0241] In some embodiments, a “TALE DNA binding domain” or “TALE” is a polypeptide comprising one or more TALE repeat domains / units. The repeat domains, each comprising a repeat variable diresidue (RVD), are involved in binding of the TALE to its cognate target DNA sequence. A single “repeat unit” (alsoAttorney Docket No. 224742003640 referred to as a “repeat”) is typically 33-35 amino acids in length and exhibits at least some sequence homology with other TALE repeat sequences within a naturally occurring TALE protein. TALE proteins may be designed to bind to a target site using canonical or non-canonical RVDs within the repeat units. See, e.g., U.S. Pat. Nos. 8,586,526 and 9,458,205.
[0242] In some embodiments, a TALE is a fusion protein comprising a nucleic acid binding domain derived from a TALE and an effector domain. In some embodiments, one or more sites in an endogenous locus can be targeted by engineered TALEs.
[0243] ZFP and TALE-based DNA-binding domains can be engineered to bind to a predetermined nucleotide sequence, for example via engineering (altering one or more amino acids) of the recognition helix region of a naturally occurring zinc finger protein, by engineering of the amino acids in a TALE repeat involved in DNA binding (the repeat variable diresidue or RVD region), or by systematic ordering of modular DNA-binding domains, such as TALE repeats or ZFP domains. Therefore, engineered ZFP or TALE proteins are proteins that are non-naturally occurring. Non-limiting examples of methods for engineering ZFPs and TALEs are design and selection. A designed protein is a protein not occurring in nature whose design / composition results principally from rational criteria. Rational criteria for design include application of substitution rules and computerized algorithms for processing information in a database storing information of existing ZFP or TALE designs (canonical and non-canonical RVDs) and binding data. See, for example, U.S. Pat. Nos. 9,458,205; 8,586,526; 6,140,081; 6,453,242; and 6,534,261; see also WO 98 / 53058; WO 98 / 53059; WO 98 / 53060; WO 02 / 016536 and WO 03 / 016496.2. Effector Domains for Transcriptional Repression
[0244] In some aspects, provided herein are effector domains for transcriptional repression. Also provided are fusion proteins comprising an effector domain, such as any described herein in Section I. In some aspects, provided herein are effector domains comprising a DNMT3L protein or portion thereof. In some aspects, the effector domain is capable of recruiting domains with DNA methyltransferase activity, such as DNMT3A or DNMT3B, expressed in a cell to the target site for PCSK9, such as any target site described herein in Section ILA. In some embodiments, the effector domain recruits domains with DNA methyltransferase activity to the target site for PCSK9 and is thus capable of decreasing transcription for PCSK9 when present at a target site at the PCSK9 locus, for example decreasing transcription of PCSK9 when recruited to a target site for PCSK9. In some aspects, the effector domains are targeted to the target site via a DNA-binding domain, such as a CRISPR / Cas-based, ZFN -based, or TALE-based DNA-binding domain, including any of the DNA-binding domains described herein, for example, in Section LA. 1.
[0245] Repression of gene expression of endogenous genes, such as human genes, can be achieved byAttorney Docket No. 224742003640 targeting (e.g., via a CRISPR-based, ZFN-based, or TALE-based DNA-binding domain) the effector domains to a target site for PCSK9.
[0246] In some embodiments, an effector domain provided herein comprises a catalytically inactive DNA methyltransferase domain or portion thereof. In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. DNMT3L (DNA (cytosine-5)-methyltransferase 3-like) is a catalytically inactive regulatory factor of DNA methyltransferases that can either promote or inhibit DNA methylation depending on the context. DNMT3L interacts with DNMT3A and significantly enhances its catalytic activity. For instance, DNMT3L interacts with the catalytic domain of DNMT3A to form a heterodimer, demonstrating that DNMT3L has dual functions of binding an unmethylated histone tail and activating DNA methyltransferase. Without wishing to be bound by theory, it is also believed that a DNMT3L protein or portion thereof is sufficient to recruit and bind to domains with DNA methyltransferase activity, such as DNMT3A or DNMT3B, without being fused to the domains with DNA methyltransferase activity.
[0247] DNMT3L is characterized by a regulatory ATRX-DNMT3-DNMT3L (ADD) domain and an MTase-like domain. An exemplary full-length amino acid sequence of a wild-type (also called “unmodified”) DNMT3L polypeptide derived from a mouse comprising both an ADD domain and an MTase-like domain is set forth in SEQ ID NO: 107. The exemplary wild-type DNMT3L polypeptide set forth in SEQ ID NO: 107 does not include a methionine. An exemplary full-length amino acid sequence of a wild-type DNMT3L polypeptide derived from a mouse comprising both an ADD domain and an MTase-like domain and including an initial methionine is set forth in SEQ ID NO: 266. With reference to the exemplary wild-type DNMT3L polypeptide set forth in SEQ ID NO: 107, the ADD domain is the contiguous sequence set forth as amino acid residues 67-206 or amino acid residues 74-206. With reference to the exemplary wildtype DNMT3L polypeptide set forth in SEQ ID NO: 266, the ADD domain is the contiguous sequence set forth as amino acid residues 68-207 or amino acid residues 75-207. With reference to the exemplary wildtype DNMT3L set forth in SEQ ID NO: 107, the MTase-like domain is the contiguous sequence set forth as amino acid residues 207-420. With reference to the exemplary wild-type DNMT3L set forth in SEQ ID NO: 266, the MTase-like domain is the contiguous sequence set forth as amino acid residues 208-421. It is within the level of a skilled artisan to identify domains in an DNMT3L protein, including portion thereof, e.g., an MTase-like domain, such as by alignment of a reference sequence (e.g. SEQ ID NO: 107) with other DNMT3L sequences, e.g., human DNMT3L protein sequence set forth in SEQ ID NO: 108. An exemplary alignment identifying domains is exemplified in FIG. 28, which shows residues in SEQ ID NO: 108 (“human”) that correspond to the numbering of positions in SEQ ID NO: 107 (“mouse”).
[0248] A DNMT3L ADD domain is characterized by controlling DNA methyltransferase activity byAttorney Docket No. 224742003640 preventing methylation of regions marked by tri-methylation of the histone tail H3K4 (H3K4me3). H3K4me3 is a chromatin mark associated with the promoters of actively transcribed genes, including transcribed CpG island-containing genes with hypomethylated promoters that may otherwise be prime targets for de novo methylation by domains and / or proteins with methyltransferase activity, such as DNMT3A or DNMT3B. The ADD domain sterically interferes with the DNMT3A-DNMT3L binding interface on the DNMT3L. When the H3K4 is tri-methylated, the ADD domain cannot bind to the histone tail, and so the DNMT3A or DNMT3B protein cannot methylate. However, if a gene is no longer actively transcribed and the H3K4 tail is not methylated, the ADD domain can bind H3K4meO, and swing out of the inhibitory confirmation enabling DNA methylation. Further, the ADD domain has also been demonstrated to interact with the common suite of heterochromatin proteins, analogously to a KRAB domain.
[0249] A DNMT3L MTase-like domain is characterized by associating with proteins with methyltransferase activity (e.g., DNMT3A or DNMT3B) to both stimulate the DNA methylation activity and also to enhance the recruitment of proteins such as DNMT3A or DNMT3B to genomic sites to be methylated. Without wishing to be bound by theory, it is believed that the DNMT3L MTase-like domain alone is sufficient to recruit domains and / or proteins with methyltransferase activity, such as DNMT3A or DNMT3B. Specifically, regions of the DNMT3L MTase-like domain are known to be involved in the binding or recruitment of DNMT3A through their role as part of the DNMT3A-DNMT3L binding interface on the DNMT3L protein. Exemplary regions of the DNMT3L MTase-like domain involved in the DNMT3A-DNMT3L binding interface are known in the literature, see, e.g., Jia et al., Nature, 2007 and Jurkowska et al., Nucleic Acids Res. , 2008. The exemplary wild-type DNMT3L polypeptide set forth in SEQ ID NO: 108 does not include a methionine. An exemplary full-length amino acid sequence of a wild-type human DNMT3L polypeptide and including an initial methionine is set forth in SEQ ID NO: 279. With reference to the exemplary wild-type human DNMT3L polypeptide set forth in SEQ ID NO: 108, the region of the DNMT3L MTase-like domain involved in the DNMT3A-DNMT3L binding interface includes the contiguous sequence set forth as amino acid residues 225-233, 257-273, and / or 291-302. With reference to the exemplary wild-type human DNMT3L polypeptide set forth in SEQ ID NO: 279, the region of the DNMT3L MTase-like domain involved in the DNMT3A-DNMT3L binding interface includes the contiguous sequence set forth as amino acid residues 226-234, 258-274 and / or 292-303. Without wishing to be bound by theory, it is believed that a pair of phenylalanine residues (F) play a role in forming the DNMT3A-DNMT3L interface. An exemplary pair of phenylalanine residues include F296 and F336 with reference to the exemplary wild-type mouse DNMT3L set forth in SEQ ID NO: 107. An exemplary pair of phenylalanine residues include F297 and F337 with reference to the exemplary wild-type mouse DNMT3L set forth in SEQ ID NO: 279. It is within the level of a skilled artisan to identify regions involved in theAttorney Docket No. 224742003640DNMT3A-DNMT3L binding interface in an DNMT3L protein, e.g., a pair of phenyalanine residues, such as by alignment of a reference sequence (e.g. SEQ ID NO: 107 or SEQ ID NO: 108) with other DNMT3L sequences.
[0250] In some embodiments, the effector domain is 50, 100, 250, 300, 350, 400, 450, 50, 550, or 600 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, or 421 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is at least 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids in length. In some embodiments, the effector domain is 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is 22, 37, 42, 47, 49, 57, 61, 62, 70, 72, 76, or 80 amino acids in length, or within a range defined by any of the foregoing.
[0251] In some embodiments, the effector domain is less than 800, less than 700, less than 600, less than 550, less than 540, less than 530, less than 520, less than 510, less than 500, less than 490, less than 480, less than 470, less than 460, less than 450, less than 440, less than 430, less than 420, less than 410, less than400, less than 390, less than 380, less than 370, less than 360, less than 350, less than 340, less than 330, less than 320, less than 310, less than 300, less than 290, less than 280, less than 270, less than 260, less than 250, less than 200, less than 100, less than 50, less than 40, less than 30, less than 25, less than 20, or less than 15 amino acids in length. In some embodiments, the effector domain is less than 800 amino acids in length. In some embodiments, the effector domain is less than 600 amino acids in length. In some embodiments, the effector domain is less than 520 amino acids in length. In some embodiments, the effector domain is less than 510 amino acids in length. In some embodiments, the effector domain is less than 400 amino acids in length. In some embodiments, the effector domain is less than 300 amino acids in length. In some embodiments, the effector domain is less than 250 amino acids in length.
[0252] In some embodiments, the effector domain is between 10 and 800, between 10 and 700, between 10 and 600, between 10 and 550, between 10 and 510, between 10 and 500, between 10 and 400, between 10 and 300, between 10 and 250, between 10 and 200, between 10 and 100, between 10 and 50, 15 and 800, between 15 and 700, between 15 and 600, between 15 and 550, between 15 and 510, between 15 and 500,Attorney Docket No. 224742003640 between 15 and 400, between 15 and 300, between 15 and 250, between 15 and 200, between 15 and 100, between 15 and 50, 40 and 800, between 40 and 700, between 40 and 600, between 40 and 550, between 40 and 510, between 40 and 500, between 40 and 400, between 40 and 300, between 40 and 250, between 40 and 200, between 40 and 100, between 40 and 50, 100 and 800, between 100 and 700, between 100 and 600, between 100 and 550, between 100 and 510, between 100 and 500, between 100 and 400, between 100 and 300, between 100 and 250, between 100 and 200, between 250 and 800, between 250 and 700, between 250 and 600, between 250 and 550, between 250 and 510, between 250 and 500, between 250 and 400, between 250 and 300, between 300 and 800, between 300 and 700, between 300 and 600, between 300 and 550, between 300 and 510, between 300 and 500, between 300 and 400, between 350 and 800, between 350 and 700, between 350 and 600, between 350 and 550, between 350 and 510, between 350 and 500, between 350 and 400, between 400 and 800, between 400 and 700, between 400 and 600, between 400 and 550, between 400 and 510, between 400 and 500, between 500 and 800, between 500 and 700, between 500 and 600, between 500 and 550, between 500 and 510, between 510 and 800, between 510 and 700, between 510 and 600, or between 510 and 550 amino acids in length. In some embodiments, the effector domain is between 250 and 510 amino acids in length. In some embodiments, the effector domain is between 250 and 400 amino acids in length.
[0253] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the effector domain is a DNMT3L domain. In some embodiments, the DNMT3L domain comprises a DNMT3L ADD domain, a DNMT3L MTase-like domain, or a DNMT3L ADD domain and a DNMT3L MTase-like domain. In some embodiments, the effector domain is a DNMT3L MTase-like domain. In some embodiments, the DNMT3L protein or portion thereof is selected from one of the following species: Acomys russatus, Ailuropoda melanoleuca, Apodemus sylvaticus, Arvicanthis niloticus, Bos indicusm, Callithrix jacchus, Camelus bactrianus, Capricomis sumatraensis, Carlito syrichta, Castor canadensis, Cavia porcellus, Chinchilla lanigera, Choloepus didactylus, Chrysochloris asiatica, Cricetulus griseus, Cynocephalus volans, Dasypus novemcinctus, Desmodus rotund s, Diceros bicomis minor, Dipodomys spectabilis, Echinops telfairi, Elephantulus edwardii, Enhydra lutris kenyoni, Eptesicus fuscus, Equus caballus, Erinaceus europaeus, Eschrichtius robustus, Eubalaena glacialis, Eulemur rufifrons, Fukomys damarensis, Globicephala melas, Heterocephalus glaber, Hippopotamus amphibius kiboko, Hipposideros armiger, Homo sapiens, Hyaena hyaena, Jaculus jaculus, Lipotes vexillifer, Loxodonta Africana, Macaca fascicularis, Manis pentadactyla, Marmota monax, Mastomys coucha, Meriones unguiculatus, Mesoplodon densirostris, Microtus ochrogaster, Miniopterus natalensis, Molossus molossus, Monodelphis domestica, Monodon monoceros, Muntiacus muntjak, Mus musculus, Mustela nigripes, Myodes glareolus, Myotis lucifugus, Myotis yumanensis, Nannospalax galili, Neofelis nebulosa, NeogaleAttorney Docket No. 224742003640 vison, Neotoma lepida, Notamacropus eugenii, Nyctereutes procyonoides, Nycticebus coucang, Ochotona curzoniae, Ochotona princeps, Octodon degus, Odobenus rosmarus divergens, Onychomys torridus, Orcinus orca, Orycteropus afer afer, Ovis aries, Perognathus longimembris pacificus, Phacochoerus africanus, Phoca vitulina, Phocoena Phocoena, Phocoena sinus, Phyllostomus hastatus, Physeter macrocephalus, Pipistrellus kuhlii, Propithecus coquereli, Pteronotus mesoamericanus, Pteropus vampyrus, Rattus norvegicus, Rhinolophus ferrumequinum, Rousettus aegyptiacus, Saccopteryx leptura, Sagmatias obliquidens, Saimiri boliviensis boliviensis, Sarcophilus harrisii, Sigmodon hispidus, Smutsia gigantea, Stumira hondurensis, Sus scrofa, Trichechus manatus latirostris, Tupaia chinensis, Urocitellus parryii, Ursus arctos, Vicugna pacos, Vombatus ursinus, and Vulpes vulpes.
[0254] Table t lists the SEQ ID NOs for full-length DNMT3L protein sequences from the species listed above as well the corresponding start and stop amino acid positions for the ADD and MTase-like domains for each sequence. For example, the protein sequence for the full-length DNMT3L protein from Apodemus sylvaticus (Wood mouse) is set forth in SEQ ID NO: 282. With reference to SEQ ID NO: 282, the DNMT3L ADD domain from Apodemus sylvaticus starts at amino acid position 75 and ends at amino acid position 207 while the DNMT3L MTase domain from Apodemus sylvaticus starts at amino acid position 208 and ends at amino acid position 420.Table 1. DMNT3L orthologsAtorney Docket No. 224742003640Atorney Docket No. 224742003640Atorney Docket No. 224742003640Attorney Docket No. 224742003640
[0255] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the DNMT3L protein comprises any one of the sequences set forth in SEQ ID NOs: 266, 279, and 280-377, as shown in Table 1. In some embodiments, the effector domain comprises the DNMT3L ADD domain of any one of SEQ ID NOs: 266, 279, and 280-377, corresponding to the sequence encompassed by the ADD domain start and stop amino acid positions set forth in Table 1. In some embodiments, the effector domain comprises the DNMT3L MTase-like domain of any one of SEQ ID NOs: 266, 279, and 280-377, corresponding to the sequence encompassed by MTase-like domain start and stop amino acid positions set forth in Table 1. In some embodiments, the effector domain comprises the DNMT3L ADD domain and the amino acid sequence of the DNMT3L MTase-like domain of any one of SEQ ID NOs: 266, 279, and 280-377, corresponding to the sequence encompassed by the ADD domain start and stop amino acid positions and the sequence encompassed by MTase-like domain start and stop amino acid positions set forth in Table 1.
[0256] In some embodiments, the DNMT3L protein or portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemestrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some embodiments, the DNMT3L protein of portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the DNMT3L protein of portion thereof is from Homo sapiens. In some embodiments, the DNMT3L protein of portion thereof is from Mus musculus. In some embodiments, the DNMT3L protein or portion thereof is from Apodemus sylvaticus.Attorney Docket No. 224742003640
[0257] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the effector domain is a contiguous portion of a DNTM3L protein that comprises at least 10 amino acids from a reference DNMT3L MTase-like domain. In some embodiments, the effector domain comprises less than a full-length DNMT3L MTase-like domain. In some embodiments, the effector domain comprises a DNMT3L MTase-like domain. In some embodiments, the effector domain is a DNMT3L MTase-like domain. In some embodiments, the effector domain comprises a contiguous portion of the DNMT3L protein that is greater than a full-length DNMT3L MTase-like domain, for example the effector domain may comprise a DNMT3L MTase-like domain and fiirther comprise a DNMT3L ADD domain or portion thereof. In some embodiments, the effector domain comprises a DNMT3L protein. In some embodiments, the effector domain is a DNMT3L protein.
[0258] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the DNMT3L protein or portion thereof comprises a portion of the DNMT3L protein that is a contiguous portion that is less than a full-length DNMT3L MTase-like domain and comprises at least 10 amino acids from a reference DNMT3L MTase-like domain. In some embodiments, the contiguous portion is involved in a DNMT3A-DNMT3L interface. Without wishing to be bound by theory, it is believed that a contiguous portion of a DNMT3L protein that is involved in a DNMT3A-DNMT3L interface is able to recruit domains with DNA methyltransferase activity (e.g., DNMT3A) to the target site for PCSK9, such as any target site described in Section II. Provided herein are portions of the DNMT3L protein that are a contiguous portion that comprise at least 10 amino acids from a reference DNMT3L MTase-like domain, such as any reference DNMT3L MTase-like domain described herein, and are further involved in a DNMT3A-DNMT3L interface.
[0259] In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 257-273 or 292-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 258-274 or 293-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 257-273 from the reference DNMT3L MTase- like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 258-274 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 292-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residuesAttorney Docket No. 224742003640293-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 257-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 258-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 225-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 226-303 from the reference DNMT3L MTase- like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279.
[0260] In some embodiments, the reference DNMT3L MTase-like domain is a DNMT3L MTase-like domain selected from one of the following species: Acomys russatus, Ailuropoda melanoleuca, Apodemus sylvaticus, Arvicanthis niloticus, Bos indicusm, Callithrix jacchus, Camelus bactrianus, Capricomis sumatraensis, Carlito syrichta, Castor canadensis, Cavia porcellus, Chinchilla lanigera, Choloepus didactylus, Chrysochloris asiatica, Cricetulus griseus, Cynocephalus volans, Dasypus novemcinctus, Desmodus rotundus, Diceros bicomis minor, Dipodomys spectabilis, Echinops telfairi, Elephantulus edwardii, Enhydra lufris kenyoni, Eptesicus fuscus, Equus caballus, Erinaceus europaeus, Eschrichtius robustus, Eubalaena glacialis, Eulemur rufifrons, Fukomys damarensis, Globicephala melas, Heterocephalus glaber, Hippopotamus amphibius kiboko, Hipposideros armiger, Homo sapiens, Hyaena hyaena, Jaculus jaculus, Lipotes vexillifer, Loxodonta Africana, Macaca fascicularis, Manis pentadactyla, Marmota monax, Mastomys coucha, Meriones unguiculatus, Mesoplodon densirosfris, Microtus ochrogaster, Miniopterus natalensis, Molossus molossus, Monodelphis domestica, Monodon monoceros, Muntiacus muntjak, Mus musculus, Mustela nigripes, Myodes glareolus, Myotis lucifugus, Myotis yumanensis, Nannospalax galili, Neofelis nebulosa, Neogale vison, Neotoma lepida, Notamacropus eugenii, Nyctereutes procyonoides, Nycticebus coucang, Ochotona curzoniae, Ochotona princeps, Octodon degus, Odobenus rosmarus divergens, Onychomys torridus, Orcinus orca, Orycteropus afer afer, Ovis aries, Perognathus longimembris pacificus, Phacochoerus africanus, Phoca vitulina, Phocoena Phocoena, Phocoena sinus, Phyllostomus hastatus, Physeter macrocephalus, Pipisfrellus kuhlii, Propithecus coquereli, Pteronotus mesoamericanus, Pteropus vampyrus, Rattus norvegicus, Rhinolophus ferrumequinum, Rousettus aegyptiacus, Saccopteryx leptura, Sagmatias obliquidens, Saimiri boliviensis boliviensis, Sarcophilus harrisii, Sigmodon hispidus, Smutsia gigantea, Stumira hondurensis, Sus scrofa, Trichechus manatus latirosfris, Tupaia chinensis, Urocitellus parryii, Ursus arctos, Vicugna pacos, Vombatus ursinus, and Vulpes vulpes.
[0261] In some embodiments, the reference DNMT3L MTase-like domain is a DNMT3L MTase-likeAttorney Docket No. 224742003640 domain selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemestrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some embodiments, the reference DNMT3L MTase- like domain is a DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus.
[0262] In some embodiments, the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some embodiments, the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the reference DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17.
[0263] In some embodiments, the reference DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Mus musculus. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth as amino residues 208-421 corresponding to the numbering of positions in SEQ ID NO: 266. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 15. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15. In some embodiments, the reference DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 15.
[0264] In some embodiments, the reference DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth as amino residues 174-386 corresponding to the numbering of positions in SEQ ID NO: 279. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 16. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16. In some embodiments, the reference DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 16.
[0265] In some embodiments, the reference DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth as amino residues 208-420 corresponding to the numbering of positions in SEQ ID NO: 282. In some embodiments, the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 17. In some embodiments, theAttorney Docket No. 224742003640 reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17. In some embodiments, the reference DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 17.
[0266] In some embodiments, the portion of the DNMT3L protein is a contiguous portion that is less than a full-length DNMT3L MTase-like domain and comprises at least 10 amino acids from a reference DNMT3L MTase-like domain, wherein the contiguous portion of at least 10 amino acids is involved in a DNMT3A-DNMT3L interface. In some embodiments, the contiguous portion comprises any one of the sequences set forth in SEQ ID NOs: 126, 127, 134, and 138. In some embodiments, the contiguous portion is any one of the sequences set forth in SEQ ID NOs: 126, 127, 134, and 138.
[0267] In some embodiments, the contiguous portion comprises any one of the sequences set forth in SEQ ID NOs: 128-133, 135-137, and 139-141, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some embodiments, the contiguous portion comprises any one of the sequences set forth in SEQ ID NOs: 128-133, 135-137, and 139-141. In some embodiments, the contiguous portion is any one of the sequences set forth in SEQ ID NOs: 128-133, 135-137, and 139-141.
[0268] In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 257-273 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 258-274 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5 (SEQ ID NO: 126), wherein Xi is L or M, X2is L or I, X3 is L or R, X4 is K or R, and X5 is P or Q.
[0269] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Mus musculus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 128, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 128.
[0270] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the reference DNMT3L MTase-like domain is SEQ IDAttorney Docket No. 224742003640NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 129, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 129.
[0271] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 130, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 130.
[0272] In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 292-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 293-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the contiguous portion comprises the sequence set forth in X1DX2X3X4X5X6RFLX7 (SEQ ID NO: 127), wherein XI is E or D, X2 is L or Q, X3 is D, E, or M, X4 is V or T, X5 is A or T, X6 is S, T, or V, and X7 is E or Q.
[0273] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Mus musculus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 131, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 131.
[0274] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16, or an aminoAttorney Docket No. 224742003640 acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 132, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 132.
[0275] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 133, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 133.
[0276] In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 257-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 258-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5X6SX7X8PFFWX9FX10DNLX11LX12X13X14DX15X16X17X18X19RFLX20 (SEQ ID NO: 134), wherein Xi is L or M, X2 is L or I, X3 is L or R, X4 is K or R, X5 is P or Q, Xe is G or E, X7 is P, Q, or absent, Xs is R or Q, X9 is M or I, Xw is V or M, Xu is V or L, X12 is N or T, X13 is K or E, X14 is E or D, X15 is L or Q, Xie is D, E, or M, X17 is V or T, Xis is A or T, X19 is S, T, or V, and X20 is E or Q.
[0277] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Mus musculus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 135, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 135.
[0278] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-likeAttorney Docket No. 224742003640 domain from Homo sapiens. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 136, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 136.
[0279] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 137, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 137.
[0280] In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 225-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108. In some embodiments, the contiguous portion comprises the sequence set forth as amino residues 226-303 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 279. In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the contiguous portion comprises the sequence set forth in XiX2VRX3DVEX4WGPFDLX5YGX6TX7PLGX8X9CDRXioPXiiWYXi2FQFHRXi3LQYAXi4PXi5Xi6Xi7SX 18X19PFFWX2OFX21DNLX22LX23X24X25DX26X27X28X29X3ORFLX31 (SEQ ID NO: 138), wherein XI is D or N, X2 is T or V, X3 is K or R, X4 is E or K, X5 is V or L, X6 is A or S, X7 is P or Q, X8 is H or S, X9 is T or S, X10 is P or C, XI 1 is S or G, X12 is L or M, X13 is L or I, X14 is L or R, X15 is K or R, Xie is P or Q, X17 is G or E, Xi8is P, Q, or absent, X19 is R or Q, X20 is M or I, X21 is V or M, X22 is V or L, X23 is N or T, X24 is K or E, X25 is E or D, X26 is L or Q, X27 is D, E, or M, X28is V or T, X29 is A or T, X30 is S, T, or V, and X31 is E or Q.
[0281] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Mus musculus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, orAttorney Docket No. 22474200364099% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO:15. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 139, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 139.
[0282] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO:16. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 140, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 140.
[0283] In some embodiments, the reference DNMT3L MTase-like domain is the DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the reference DNMT3L MTase-like domain is SEQ ID NO: 17. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 141, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity. In some embodiments, the contiguous portion comprises the sequence set forth in SEQ ID NO: 141.
[0284] In some embodiments, the contiguous portion comprises at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 amino acids.
[0285] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the DNMT3L protein or portion thereof comprises a DNMT3L MTase-like domain. Provided herein are DNMT3L MTase-like domains.
[0286] In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain selected from one of the following species: Acomys russatus, Ailuropoda melanoleuca, Apodemus sylvaticus, Arvicanthis niloticus, Bos indicusm, Callithrix jacchus, Camelus bactrianus, Capricomis sumatraensis, Carlito syrichta, Castor canadensis, Cavia porcellus, Chinchilla lanigera, Choloepus didactylus, Chrysochloris asiatica, Cricetulus griseus, Cynocephalus volans, Dasypus novemcinctus, Desmodus rotundus, Diceros bicomis minor, Dipodomys spectabilis, Echinops telfairi, Elephantulus edwardii, Enhydra lufris kenyoni, Eptesicus fuscus, Equus caballus, Erinaceus europaeus, Eschrichtius robustus, Eubalaena glacialis, Eulemur rufifrons, Fukomys damarensis, Globicephala melas, Heterocephalus glaber,Attorney Docket No. 224742003640Hippopotamus amphibius kiboko, Hipposideros armiger, Homo sapiens, Hyaena hyaena, Jaculus jaculus, Lipotes vexillifer, Loxodonta Africana, Macaca fascicularis, Manis pentadactyla, Marmota monax, Mastomys concha, Meriones unguiculatus, Mesoplodon densirostris, Microtus ochrogaster, Miniopterus natalensis, Molossus molossus, Monodelphis domestica, Monodon monoceros, Muntiacus muntjak, Mus musculus, Mustela nigripes, Myodes glareolus, Myotis lucifugus, Myotis yumanensis, Nannospalax galili, Neofelis nebulosa, Neogale vison, Neotoma lepida, Notamacropus eugenii, Nyctereutes procyonoides, Nycticebus coucang, Ochotona curzoniae, Ochotona princeps, Octodon degus, Odobenus rosmarus divergens, Onychomys torridus, Orcinus orca, Orycteropus afer afer, Ovis aries, Perognathus longimembris pacificus, Phacochoerus africanus, Phoca vitulina, Phocoena Phocoena, Phocoena sinus, Phyllostomus hastatus, Physeter macrocephalus, Pipistrellus kuhlii, Propithecus coquereli, Pteronotus mesoamericanus, Pteropus vampyrus, Rattus norvegicus, Rhinolophus ferrumequinum, Rousettus aegyptiacus, Saccopteryx leptura, Sagmatias obliquidens, Saimiri boliviensis boliviensis, Sarcophilus harrisii, Sigmodon hispidus, Smutsia gigantea, Stumira hondurensis, Sus scrofa, Trichechus manatus latirostris, Tupaia chinensis, Urocitellus parryii, Ursus arctos, Vicugna pacos, Vombatus ursinus, and Vulpes vulpes.
[0287] In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemesfrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain from Mus musculus. In some embodiments, the DNMT3L MTase-like domain is a DNMT3L MTase-like domain from Apodemus sylvaticus.
[0288] In some embodiments, the DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing. In some embodiments, the DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17. In some embodiments, the DNMT3L MTase-like domain is any one of the sequences set forth in SEQ ID NOs: 15-17.
[0289] In some embodiments, the DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Mus musculus. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth as amino residues 208-421 corresponding to the numbering of positions in SEQ ID NO: 266. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, orAttorney Docket No. 22474200364099% sequence identity to SEQ ID NO: 15. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15. In some embodiments, the DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 15.
[0290] In some embodiments, the DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Homo sapiens. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth as amino residues 174-386 corresponding to the numbering of positions in SEQ ID NO: 297. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 16. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16. In some embodiments, the DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 16.
[0291] In some embodiments, the DNMT3L MTase-like domain comprises a DNMT3L MTase-like domain from Apodemus sylvaticus. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth as amino residues 208-420 corresponding to the numbering of positions in SEQ ID NO: 282. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 17. In some embodiments, the DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17. In some embodiments, the DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 17.
[0292] In some embodiments, the effector domain further comprises a DNMT3L ADD domain. For example, the effector domain comprises a DNMT3L MTase-like domain, such as any provided DNMT3L MTase-like domain, or a portion thereof and further comprises a DNMT3L ADD domain. Exemplary effector domain sequences comprising a mouse DNMT3L ADD domain and a mouse DNMT3L MTase-like domain are set forth in SEQ ID NO: 116 and SEQ ID NO: 378. An exemplary effector domain comprising a mouse DNMT3L ADD domain and a mouse DNMT3L MTase-like domain is the sequence set forth in SEQ ID NO: 116.
[0293] In some embodiments, the DNMT3L ADD domain is a DNMT3L ADD domain selected from one of the following species: Acomys russatus, Ailuropoda melanoleuca, Apodemus sylvaticus, Arvicanthis niloticus, Bos indicusm, Callithrix jacchus, Camelus bactrianus, Capricomis sumatraensis, Carlito syrichta, Castor canadensis, Cavia porcellus, Chinchilla lanigera, Choloepus didactylus, Chrysochloris asiatica, Cricetulus griseus, Cynocephalus volans, Dasypus novemcinctus, Desmodus rotundus, Diceros bicomis minor, Dipodomys spectabilis, Echinops telfairi, Elephantulus edwardii, Enhydra lutris kenyoni, Eptesicus fuscus, Equus caballus, Erinaceus europaeus, Eschrichtius robustus, Eubalaena glacialis, Eulemur rufifrons,Attorney Docket No. 224742003640Fukomys damarensis, Globicephala melas, Heterocephalus glaber, Hippopotamus amphibius kiboko, Hipposideros armiger, Homo sapiens, Hyaena hyaena, Jaculus jaculus, Lipotes vexillifer, Loxodonta Africana, Macaca fascicularis, Manis pentadactyla, Marmota monax, Mastomys concha, Meriones unguiculatus, Mesoplodon densirostris, Microtus ochrogaster, Miniopterus natalensis, Molossus molossus, Monodelphis domestica, Monodon monoceros, Muntiacus muntjak, Mus musculus, Mustela nigripes, Myodes glareolus, Myotis lucifugus, Myotis yumanensis, Nannospalax galili, Neofelis nebulosa, Neogale vison, Neotoma lepida, Notamacropus eugenii, Nyctereutes procyonoides, Nycticebus coucang, Ochotona curzoniae, Ochotona princeps, Octodon degus, Odobenus rosmarus divergens, Onychomys torridus, Orcinus orca, Orycteropus afer afer, Ovis aries, Perognathus longimembris pacificus, Phacochoerus africanus, Phoca vitulina, Phocoena Phocoena, Phocoena sinus, Phyllostomus hastatus, Physeter macrocephalus, Pipistrellus kuhlii, Propithecus coquereli, Pteronotus mesoamericanus, Pteropus vampyrus, Rattus norvegicus, Rhinolophus ferrumequinum, Rousettus aegyptiacus, Saccopteryx leptura, Sagmatias obliquidens, Saimiri boliviensis boliviensis, Sarcophilus harrisii, Sigmodon hispidus, Smutsia gigantea, Stumira hondurensis, Sus scrofa, Trichechus manatus latirostris, Tupaia chinensis, Urocitellus parryii, Ursus arctos, Vicugna pacos, Vombatus ursinus, and Vulpes vulpes.
[0294] In some embodiments, the DNMT3L ADD domain is a DNMT3L ADD domain selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemestrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some embodiments, the DNMT3L ADD domain is a DNMT3L ADD domain selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the DNMT3L ADD domain is from Homo sapiens or Mus musculus. In some embodiments, the DNMT3L ADD domain is a DNMT3L ADD domain from Homo sapiens. In some embodiments the DNMT3L ADD domain is a DNMT3L ADD domain from Mus musculus. In some embodiments, the DNMT3L ADD domain is a DNMT3L ADD domain from Apodemus sylvaticus.
[0295] In some embodiments, the DNMT3L ADD domain comprises any one of the sequences set forth in SEQ ID NOs: 109, 110, 379, and 380, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the DNMT3L ADD domain comprises any one of the sequences set forth in SEQ ID NOs: 109, 110, 379, and 380. In some embodiments, the DNMT3L ADD domain is any one of the sequences set forth in SEQ ID NOs: 109, 110, 379, and 380.
[0296] In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109 or 110, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the DNMT3L ADD domainAttorney Docket No. 224742003640 comprises the sequence set forth in SEQ ID NO: 109 or 110. In some embodiments, the DNMT3L ADD domain is the sequence set forth in SEQ ID NO: 109 or 110.
[0297] In some embodiments, the DNMT3L ADD domain is from Mus musculus. In some embodiments, the reference DNMT3L ADD domain comprises the sequence set forth as amino residues 68- 207 or amino acid residues 75-207 corresponding to the numbering of positions in SEQ ID NO: 266. In some embodiments, the reference DNMT3L ADD domain comprises the sequence set forth as amino acid residues 75-207 corresponding to the numbering of positions in SEQ ID NO: 266. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 109. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 379, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 379. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 379. In some embodiments, the DNMT3L ADD domain is the sequence set forth in SEQ ID NO: 109. In some embodiments, the DNMT3L ADD domain is the sequence set forth in SEQ ID NO: 379.
[0298] In some embodiments, the DNMT3L ADD domain is from Homo sapiens. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth as amino residues 41-173 corresponding to the numbering of positions in SEQ ID NO: 279. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 110, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 110. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 110. In some embodiments, the DNMT3L ADD domain is the sequence set forth in SEQ ID NO: 110.
[0299] In some embodiments, the DNMT3L ADD domain is from Apodemus sylvaticus. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth as amino residues 75-207 corresponding to the numbering of positions in SEQ ID NO: 282. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 380, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 380. In some embodiments, the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 380. In some embodiments, the DNMT3L ADD domain is the sequence set forth in SEQ ID NO: 380.
[0300] In some embodiments, the effector domain comprises, from N-terminus to C-terminus, the DNMT3L ADD domain and DNMT3L MTase-like domain or portion thereof. In some embodiments, the effector domain comprises, from N-terminus to C-terminus, the DNMT3L MTase-like domain or portionAttorney Docket No. 224742003640 thereof and the DNMT3L ADD domain. In some embodiments, the effector domain comprises, from N- terminus to C-terminus, the DNMT3L ADD domain and DNMT3L MTase-like domain. In some embodiments, the effector domain comprises, from N-terminus to C-terminus, the DNMT3L MTase-like domain and the DNMT3L ADD domain.
[0301] In some embodiments, the effector domain comprises any one of the sequences set forth in SEQ ID NO: 378, 381, and 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the effector domain comprises any one of the sequences set forth in SEQ ID NO: 378, 381, and 382. In some embodiments, the effector domain is any one of the sequences set forth in SEQ ID NO: 378, 381, and 382.
[0302] In some embodiments, the effector domain is from Mus musculus. In some embodiments, the effector domain comprises the sequence set forth as amino residues 75-207 and 208-421 corresponding to the numbering of positions in SEQ ID NO: 266. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 378, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 378. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 378. In some embodiments, the effector domain is the sequence set forth in SEQ ID NO: 378.
[0303] In some embodiments, the effector domain is from Homo sapiens. In some embodiments, the effector domain comprises the sequence set forth as amino residues 41-173 and 174-386 corresponding to the numbering of positions in SEQ ID NO: 279. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 381, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 381. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 381. In some embodiments, the effector domain is the sequence set forth in SEQ ID NO: 381.
[0304] In some embodiments, the effector domain is from Apodemus sylvaticus. In some embodiments, the effector domain comprises the sequence set forth as amino residues 75-207 and 208-420 corresponding to the numbering of positions in SEQ ID NO: 282. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 382. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 382. In some embodiments, the effector domain is the sequence set forth in SEQ ID NO: 382.
[0305] In some embodiments, the effector domain comprises a DNMT3L protein or portion thereof. In some embodiments, the DNMT3L protein or portion thereof comprises a DNMT3L protein. In some embodiments, the DNMT3L protein or portion thereof is a DNMT3L protein. Provided herein are DNMT3LAttorney Docket No. 224742003640 proteins.
[0306] In some embodiments, the DNMT3L protein is selected from one of the following species: Acomys russatus, Ailuropoda melanoleuca, Apodemus sylvaticus, Arvicanthis niloticus, Bos indicusm, Callithrix jacchus, Camelus bactrianus, Capricomis sumatraensis, Carlito syrichta, Castor canadensis, Cavia porcellus, Chinchilla lanigera, Choloepus didactylus, Chrysochloris asiatica, Cricetulus griseus, Cynocephalus volans, Dasypus novemcinctus, Desmodus rotund s, Diceros bicomis minor, Dipodomys spectabilis, Echinops telfairi, Elephantulus edwardii, Enhydra lutris kenyoni, Eptesicus fuscus, Equus caballus, Erinaceus europaeus, Eschrichtius robustus, Eubalaena glacialis, Eulemur rufifrons, Fukomys damarensis, Globicephala melas, Heterocephalus glaber, Hippopotamus amphibius kiboko, Hipposideros armiger, Homo sapiens, Hyaena hyaena, Jaculus jaculus, Lipotes vexillifer, Loxodonta Africana, Macaca fascicularis, Manis pentadactyla, Marmota monax, Mastomys coucha, Meriones unguiculatus, Mesoplodon densirostris, Microtus ochrogaster, Miniopterus natalensis, Molossus molossus, Monodelphis domestica, Monodon monoceros, Muntiacus muntjak, Mus musculus, Mustela nigripes, Myodes glareolus, Myotis lucifugus, Myotis yumanensis, Nannospalax galili, Neofelis nebulosa, Neogale vison, Neotoma lepida, Notamacropus eugenii, Nyctereutes procyonoides, Nycticebus coucang, Ochotona curzoniae, Ochotona princeps, Octodon degus, Odobenus rosmarus divergens, Onychomys torridus, Orcinus orca, Orycteropus afer afer, Ovis aries, Perognathus longimembris pacificus, Phacochoerus africanus, Phoca vitulina, Phocoena Phocoena, Phocoena sinus, Phyllostomus hastatus, Physeter macrocephalus, Pipistrellus kuhlii, Propithecus coquereli, Pteronotus mesoamericanus, Pteropus vampyrus, Rattus norvegicus, Rhinolophus ferrumequinum, Rousettus aegyptiacus, Saccopteryx leptura, Sagmatias obliquidens, Saimiri boliviensis boliviensis, Sarcophilus harrisii, Sigmodon hispidus, Smutsia gigantea, Sturnira hondurensis, Sus scrofa, Trichechus manatus latirostris, Tupaia chinensis, Urocitellus parryii, Ursus arctos, Vicugna pacos, Vombatus ursinus, and Vulpes vulpes.
[0307] In some embodiments, the DNMT3L protein is selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemestrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus. In some embodiments, the DNMT3L protein is selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus. In some embodiments, the DNMT3L protein is from Homo sapiens or Mus musculus. In some embodiments, the DNMT3L protein is from Homo sapiens. In some embodiments, the DNMT3L protein is from Mus musculus. In some embodiments, the DNMT3L protein is from Apodemus sylvaticus.
[0308] In some embodiments, the DNMT3L protein comprises any one of the sequences set forth in SEQ ID NOs: 266, 269, and 280-377, or an amino acid sequence that has at least 90%, 91%, 92%, 93%,Attorney Docket No. 22474200364094%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the DNMT3L protein comprises any one of the sequences set forth in SEQ ID NOs: 266, 269, and 280-377. In some embodiments, the DNMT3L protein is any one of the sequences set forth in SEQ ID NOs: 266, 269, and 280-377.
[0309] In some embodiments, the DNMT3L protein does not contain an N-terminal methionine. In some embodiments, the DNMT3L protein comprises any one of the sequences set forth in SEQ ID NOs: 107, 108, and 383 or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the DNMT3L protein comprises any one of the sequences set forth in SEQ ID NOs: 107, 108, and 383. In some embodiments, the DNMT3L protein is any one of the sequences set forth in SEQ ID NOs: 107, 108, and 383. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 107 or 108, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 107 or 108. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 107 or 108.
[0310] In some embodiments, the DNMT3L protein is from Mus musculus. In some embodiments, the DNMT3L protein contains an N-terminal methionine, such as set forth in SEQ ID NO: 266. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 266, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 266. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 266. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 266. In some embodiments, the DNMT3L protein does not contain an N-terminal methionine, such as set forth in SEQ ID NO 107. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 107, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 107. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 107. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 107.
[0311] In some embodiments, the DNMT3L protein is from Homo sapiens. In some embodiments, the DNMT3L protein contains an N-terminal methionine, such as set forth in SEQ ID NO: 279. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 279, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 279. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 279. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 279. In some embodiments, the DNMT3L protein does not contain an N-terminal methionine, such as set forth in SEQ IDAttorney Docket No. 224742003640NO 108. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 108, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 108. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 108. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 108.
[0312] In some embodiments, the DNMT3L protein is from Apodemus sylvaticus. In some embodiments, the DNMT3L protein contains an N-terminal methionine, such as set forth in SEQ ID NO: 282. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 282, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 282. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 282. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 282. In some embodiments, the DNMT3L protein does not contain an N-terminal methionine, such as set forth in SEQ ID NO 383. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 383, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to SEQ ID NO: 383. In some embodiments, the DNMT3L protein comprises the sequence set forth in SEQ ID NO: 383. In some embodiments, the DNMT3L protein is the sequence set forth in SEQ ID NO: 383.
[0313] In some embodiments, an effector domain provided herein is a single transcriptional repressor domain. In some embodiments, the single transcriptional effector domain is a DNMT3L domain or portion thereof. In some embodiments, the DNMT3L domain or portion thereof is the only transcriptional repressor domain. In some embodiments, the effector domain consists of a DNMT3L domain or portion thereof as the sole transcriptional repressor domain. In some embodiments, the single transcriptional repressor domain is exclusively a DNMT3L domain. In some embodiments, the effector domain does not include any other transcriptional repressor domains, such as DNMT3A, KRAB, SID, or other repressor motifs. In some embodiments, the effector domain is free of other transcriptional effector repressor domains, including but not limited to DNMT3A, KRAB domains, SID domains, or other repressor motifs. In some embodiments, transcriptional repression is mediated exclusively by the DNMT3L domain or portion thereof.
[0314] In some embodiments, an effector domain is or comprises a DNMT3L protein or portion thereof that is able to recruit domains or proteins with DNA methyltransferase activity that then decrease the transcription from a gene by at least 5%, 10%, 20%, 30%, 40% or 50%, 60%, 70%, 80%, 85%, 90%, or 100% or more, such as 2-fold, 5-fold, 10-fold, 20-fold, 30-fold, 40-fold, 50-fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold or more, compared to the absence of the effector domain. In some embodiments, an effector is or comprises a DNMT3L protein or portion thereofAttorney Docket No. 224742003640 that decrease the transcription from a gene by at least 5%, 10%, 20%, 30%, 40% or 50%, 60%, 70%, 80%, 85%, 90%, or 100% or more, such as 2-fold, 5-fold, 10-fold, 20-fold, 30-fokl, 40-fold, 50-fold, 60-fold, 70- fold, 80-fold, 90-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold or more, compared to the absence of the effector domain.3. Exemplary Fusion Proteins
[0315] In some aspects, provided are fusion proteins comprising any of the DNA-binding domains described herein (e.g., in Section I.A. 1) and any of the effector domains comprising a DNMT3L domain or functional portion thereof (e.g., in Section I.A.2).
[0316] In some embodiments, the fusion protein comprises a DNA-binding domain comprises a CRISPR-associated (Cas) protein. In some embodiments, the fusion protein comprises a dCas9 protein. In some embodiments, the fusion protein comprising a dCas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1-3, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-3. In some embodiments, the fusion protein comprising a dCas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1-3. In some embodiments, the fusion protein comprising a dCas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1-3, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to any one of SEQ ID NOs: 1-3, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein comprising a dCas9 protein comprises the amino acid sequence set forth in any one of SEQ ID NOs: 1-3, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein comprising a dCas9 protein is the amino acid sequence set forth in any one of SEQ ID NOs: 1-3.
[0317] In some embodiments, the fusion protein comprises an effector domain comprising a DNMT3L protein or portion thereof. In some embodiments, the DNMT3L protein or portion thereof thereof is from a Mus musculus. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 1, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 1. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 1, or a sequence having at least about 80%, 85%, 90%, 91%,Attorney Docket No. 22474200364092%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 1, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 1, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a nonmethylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein is the amino acid sequence set forth in SEQ ID NO: 1.
[0318] In some embodiments, the fusion protein comprises an effector domain comprising a DNMT3L protein or portion thereof, wherein the DNMT3L protein or portion thereof is from Homo sapiens. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 2, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 2. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 2, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 2, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 2, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non- methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein is the amino acid sequence set forth in SEQ ID NO: 2.
[0319] In some embodiments, the fusion protein comprises an effector domain comprising a DNMT3L domain or portion thereof, wherein the DNMT3L domain or portion thereof is from Apodemus sylvaticus. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 3, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 3. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 3, or a sequence having at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to SEQ ID NO: 3, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meOAttorney Docket No. 224742003640 peptides). In some embodiments, the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 3, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a nonmethylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein is the amino acid sequence set forth in SEQ ID NO: 3.
[0320] In some embodiments, the fusion protein comprises a DNA-binding domain comprising a CRISPR-associated (Cas) protein. In some embodiments, the fusion protein comprises a dCas9 protein. In some emboidments, the fusion protein comprises an effector domain comprising a DNMT3L domain or portion thereof. In some embodiments, the DNMT3L domain or portion thereof is the only transcriptional repressor domain fused to dCas9. In some embodiments, the fusion protein consists essentially of a dCas9 protein and a DNMT3L domain or portion thereof, wherein DNMT3L is the only repressor domain present. In some embodiments, the DNMT3L domain is a DNMT3L MTase-like domain. In some embodiments, the fusion protein comprises a dCas9 protein and a DNMT3L domain. In some embodiments, the fusion protein comprises a dCas9 protein and a DNMT3L MTase-like domain. In some embdoiments, the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and peptides comprising a portion of a non-methylated histone tail (e.g., H3K4meO peptides). In some embodiments, the fusion protein consists essentially of a dCas9 protein and a DNMT3L MTase-like domain. In some embodiments, the fusion protein comprises from N-terminus to C-terminus a dCas9 protein, a linker, and a DNMT3L domain or portion thereof. In some embodiments, the fusion protein comprises from N-terminus to C-terminus a DNMT3L domain or portion thereof, a linker, and a dCas9 protein.B. Fusion Proteins Comprising a Variant DNMT3A Domain or Functionally Active Portion Thereof
[0321] In some aspects, provided herein are fusion proteins. In some embodiments, the fusion protein targets a target site and / or target gene, such as any described herein in Section II. In some embodiments, the fusion protein comprises (a) a DNA-binding domain or a component thereof, such as any described herein in Section I.B. 1, and (b) an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, such as any described here in Section IV. In some embodiments, the effector domain can include the variant DNMT3A fused to a DNMT3L as described in Section I.B.2.a. In some embodiments, the effector domain can include the variant DNMT3A as a multipartite fusion as described in Section I.B.2.b. In some embodiments, the effector domain is a multipartite effector that further comprises one or more domains selected from a DNA methyltransferase domain, a repressor domain capable of recruiting heterochromatinAttorney Docket No. 224742003640 inducing factors, or combinations thereof. In some embodiments, the fusion protein is capable of targeting the effector domain that comprises a variant DNMT3A domain or functionally active portion thereof, including a multipartite effector, to a target site for a target gene. In some embodiments, the fusion protein is capable of being targeted to a target site for a target gene, by virtue of the DNA-binding domain or component thereof. In some aspects, targeting of the effector, such as the multipartite effector, or the fusion protein decreases transcription of the target gene. In some aspects, the decreased transcription of the target gene is sustained, i.e., durable. In some aspects, the targeting is specific, i.e., low off-target activity.
[0322] In some embodiments, any two or more domains of the fusion protein are heterologous, i.e., the domains are from different species, or at least one of the domains is not found in nature. In some aspects, the fusion protein is an engineered fiision protein, i.e., the fiision protein is not found in nature.
[0323] In some embodiments, the fusion protein comprises its constituent components (e.g., domains) in any suitable arrangement, orientation, or order. For example, a variant DNMT3A domain or functionally active portion thereof may be fused to the N-terminus or C-terminus of the DNA-binding domain of the fusion protein and / or to the N-terminus or C-terminus of another DNA methyltransferase domain and / or repressor domain.1. DNA-binding domains
[0324] In some embodiments of the fusion protein comprising an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, provided are DNA-binding domains. In some aspects, a DNA-binding domain is capable of specifically targeting (e.g., binding or hybridizing to) a target site, such as any target site described herein in Section II. In some aspects, a DNA-binding domain targets a specific sequence of nucleotides, such as a DNA sequence. In some aspects, a DNA-binding domain can be engineered (e.g. designed or programmed) to target a specific target site. In some aspects, a DNA-binding domain recruits an effector domain, including a multipartite effector, comprising a variant DNMT3A domain or functionally active portion thereof described herein in Section IV to the target site, thereby inducing targeted gene repression.
[0325] In some embodiments, the DNA-binding domain comprises a CRISPR associated (Cas) protein, zinc finger protein (ZFP), transcription activator-like effectors (TALE), meganuclease, homing endonuclease, LScel enzyme, or variants thereof. In some embodiments, the DNA-binding domain comprises a catalytically inactive (e.g., nuclease-inactive or nuclease-inactivated) variant of any of the foregoing. In some embodiments, the DNA-binding domain comprises a deactivated Cas9 (dCas9) protein or variant thereof that is a catalytically inactivated so that it is inactive for nuclease activity and is not able to cleave DNA.Attorney Docket No. 224742003640
[0326] In some embodiments, the DNA-binding domain is a component of a DNA-targeting system, such as any described in Section III, which comprises a Cas-gRNA combination, comprising a Cas protein or variant thereof and at least one guide RNA (gRNA). In some embodiments, the gRNA binds to the target site. In some embodiments, the gRNA comprises a spacer sequence that is capable of targeting and / or hybridizing to the target site. In some embodiments, the gRNA is capable of complexing with the Cas protein or variant thereof, e.g., via a scaffold sequence of the gRNA. In some aspects, the gRNA directs or recruits the Cas protein or variant thereof to the target site.
[0327] Exemplary components and features of the DNA-binding domains, including for CRISPR / Cas- based, ZFN-based, and TALE-based DNA-binding domains are provided below. a. CRISPR / Cas-based DNA-binding domains
[0328] In some embodiments of the fusion protein comprising a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, the DNA-binding domain is based on a CRISPR / Cas system, i.e., CRISPR / Cas-based DNA-binding domains, that are able to bind to a target site or a combination of target sites. In some embodiments, the DNA-binding domain is any CRISPR / Cas-based DNA-binding domain provided herein in Section I.A. l.a.1) Cas proteins
[0329] In some embodiments of the fusion protein comprising a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, the DNA-binding domain comprises a CRISPR-associated (Cas) protein. In some embodiments, the DNA-binding domain comprises any Cas protein provided herein in Section I.A. l.a. 1.
[0330] In some embodiments, the Cas protein is a Cas9 protein or variant thereof. In some embodiments, the Cas9 protein or variant thereof is a Streptococcus pyogenes Cas9 (SpCas9) protein or a variant thereof. In some embodiments, the variant Cas9 is a Streptococcus pyogenes dCas9 (dSpCas9) protein that comprises at least one amino acid mutation selected from D10A and H840A, with reference to numbering of positions of SEQ ID NO: 42. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 18 or SEQ ID NO: 61, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 18, which lacks an initial methionine residue. In some embodiments, the variant Cas9 protein comprises the sequence set forth in SEQ ID NO: 61, which includes an initial methionine residue.Attorney Docket No. 2247420036402) Guide RNAs
[0331] In some embodiments of the fusion protein comprising a DNA-binding domain that is a Cas protein and an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, the Cas protein (e.g., dCas9) is provided in combination or as a complex with one or more guide RNA (gRNA). In some aspects, the gRNA is a nucleic acid that promotes the specific targeting or homing of the gRNA / Cas ribonucleoprotein (RNP) complex to the target site of the target gene, such as any described in Section II. In some embodiments, a target site of a gRNA may be referred to as a protospacer.
[0332] In some embodiments, the gRNA is any gRNA described herein in Section II.A. l.a.2.
[0333] In some embodiments, the gRNA provided herein targets a target site, such as any target site of interest for targeted transcriptional repression. In some embodiments, the gRNA targets a target site for a target gene, such as for any suitable target gene (e.g., target site and target genes in Section II). In some embodiments the gRNA hybridizes to the sequence complementary to the sequence defined as the target site. The strand of the target nucleic acid comprising the target site sequence may be referred to as the “complementary strand” of the target nucleic acid. gRNAs that target a target site in a gene or a regulatory DNA element thereof are known, for example, in, e.g., PCT Appl. No. WO 2023 / 250511, WO 2023 / 173110, WO 2023 / 093862, WO 2023 / 215711, WO 2023 / 240076, W02024 / 040254, W02024 / 064910, US published Appl. No. 20240052328, US published Appl. No. 20240067968, and US published Appl. No. 20240067969, the disclosures of which are incorporated by reference in their entireties. b. Other DNA-binding domains
[0334] In some embodiments of the fusion protein comprising a DNA-binding domain and an effector domain comprising a variant DNMT3A domain or any functionally active portion thereof, the DNA-binding domain comprises a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an LScel enzyme or a variant thereof. In some embodiments, the fusion protein comprises any DNA-binding domain provided herein in Section LA. Lb.
[0335] In some embodiments, the DNA-binding domain comprises a zinc finger protein (ZFP), i.e., is a ZFP-based DNA-binding domain. In some embodiments, the ZFP DNA-binding domain is any ZFP provided herein in Section I. A. 1.b.
[0336] In some embodiments, the ZFP binds to, or is capable of binding to (i.e., targets), a target site described herein, such as any target site described in Section II. In some embodiments, the ZFP binds to a target site in a gene or a regulatory DNA element thereof. In some embodiments, the ZFP facilitates targetspecific binding of a fusion protein comprising the ZFP.Attorney Docket No. 224742003640
[0337] ZFPs that bind to a target site in a gene or a regulatory DNA element thereof are known, for example, in, e.g., WO 2023 / 215711, WO 2024 / 064910, WO 2024 / 040254, US published Appl. No. 20240067969, and US published Appl. No. 20240067968, the disclosures of which are incorporated by reference in their entireties.
[0338] In some aspects, provided herein is a ZFP, such as a HBV-targeting ZFP as described herein. In some embodiments, the ZFP targets a target site comprising the nucleotide sequence set forth in SEQ ID NO: 26, a contiguous portion thereof of at least 12 nt, or a complementary sequence of any of the foregoing. In some embodiments, the fusion protein targets a target site comprising the nucleotide sequence set forth in SEQ ID NO: 26. In some embodiments, the target site is double-stranded DNA. In some embodiments, the ZFP of the fusion protein comprises six zinc fingers denoted Fl through F6 in order from N-terminus to C- terminus, each comprising a corresponding zinc finger recognition region Fl through F6, and the amino acid sequence of each zinc finger recognition region is as follows: Fl: QSAHRKN (SEQ ID NO: 27); F2: TSSNRKT (SEQ ID NO: 28); F3: RSDNLSA (SEQ ID NO: 29); F4: RNNDRKT (SEQ ID NO: 30); F5: TSGSLSR (SEQ ID NO: 31); F6: QAGHLAK (SEQ ID NO: 32). In some embodiments, the ZFP of the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 21, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the ZFP of the fusion protein comprises the amino acid sequence set forth in SEQ ID NO: 21.2. Effector Domains for Transcriptional Repression
[0339] In some aspects, provided herein are effector domains for transcriptional repression. Also provided are fusion proteins comprising transcriptional repression effector domains (also referred to as effector domains) comprising a variant DNMT3A domain or functionally active portion thereof, such as any described herein in Section IV. In some aspects, provided herein are effector domains comprising a variant DNMT3A domain or functionally active portion thereof fused to a DNMT3L domain, such as any described here in Section I.B.2.a. Also provided herein are multipartite effectors for transcriptional repression, e.g., an effector domain comprising a variant DNMT3A domain or functionally active portion thereof and further comprising DNA methyltransferase domains, repressor domains capable of recruiting heterochromatin inducing factors, or combinations thereof, such as any described herein. In some aspects, the DNA methyltransferase and repressor domains decrease, or are capable of decreasing, transcription of an endogenous locus when recruited to a target site at the endogenous locus, for example decreasing transcription of a gene when recruited to a target site for the gene. In some aspects, the DNA methyltransferase and repressor domains are provided as part of a fusion protein, such as any describedAttorney Docket No. 224742003640 herein. In some aspects, the repressor domains are targeted to one or more target sites for a gene (or multiple genes) to induce, catalyze, or lead to decreased transcription of the gene. In some aspects, the effector domains are targeted to the target site via a DNA-binding domain, such as a CRISPR / Cas-based, ZFN -based, or TALE-based DNA-binding domain, including any of the DNA-binding domains described herein, for example, in Section I.B. 1.
[0340] In some embodiments, the effector domain, such as a multipartite effector, induces, catalyzes, or leads to transcription repression, transcription co-repression, histone modification, histone acetylation, histone deacetylation, nucleosome remodeling, chromatin remodeling, heterochromatin formation, proteolysis, ubiquitination, deubiquitination, phosphorylation, dephosphorylation, splicing, DNA methylation, DNA demethylation, histone methylation, histone demethylation, or DNA base oxidation. In some embodiments, the effector domain induces, catalyzes, or leads to transcription repression or transcription co-repression. In some embodiments, the effector domain induces transcription repression. In some embodiments, the effector domain has one of the aforementioned activities itself (i.e., acts directly). In some embodiments, the effector domain recruits and / or interacts with a protein or polypeptide domain that has one of the aforementioned activities (i.e., acts indirectly).
[0341] Repression of gene expression of endogenous genes, such as human genes, can be achieved by targeting (e.g., via a CRISPR-based, ZFN-based, or TALE-based DNA-binding domain) the effector domains to a target site for the genes, such as regulatory DNA elements thereof (e.g., a promoter or enhancer).
[0342] In some embodiments, an effector domain provided herein comprises a variant DNMT3A domain or functionally active portion thereof. In some embodiment, an effector domain is a multipartite effector that further comprises a domain from a human protein. In some embodiments, an effector domain further comprises any portion of the protein that is capable of acting as an effector domain as described herein. In some embodiments, a multipartite effector is or comprises a variant DNMT3A domain or functionally active portion thereof and a portion, fragment, domain or variant of a human protein, such as a portion, fragment, domain or variant of a human protein, that exhibits transcriptional repression, is capable of inducing or repressing transcription from a gene), is a functional effector domain, and / or has a function of transcription repression. In some embodiments, a multipartite effector is or comprises a variant DNMT3A domain or functionally active portion thereof and a partially or fully functional portion, a partially or fully functional fragment, a partially or fully functional domain or a partially or fully functional variant of a human protein, that decreases the transcription from a gene by at least 5%, 10%, 20%, 30%, 40% or 50%, 60%, 70%, 80%, 85%, 90%, or 100% or more, such as 2-fold, 5-fold, I O-fold, 20-fold, 30-fold, 40-fold, 50- fold, 60-fold, 70-fold, 80-fold, 90-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 1000-fold or more.Attorney Docket No. 224742003640 compared to the absence of the effector domain, including a multipartite effector.
[0343] In some embodiments, the effector domain is 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is at least 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids in length. In some embodiments, the effector domain is 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is 22, 37, 42, 47, 49, 57, 61, 62, 70, 72, 76, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the effector domain is at least 22, 37, 42, 47, 49, 57, 61, 62, 70, 72, 76, or 80 amino acids in length. In some embodiments, the effector domain is between 10 and 80, 20 and 70, 30 and 80, 30 and 70, 30 and 60, 40 and 80, 40 and 70, 40 and 60, 40 and 50, 50 and 80, 50 and 70, 50 and 60 amino acids in length.
[0344] Gene expression of endogenous mammalian genes, such as human genes, can be achieved by targeting a fusion protein comprising a DNA-binding domain, such as a dCas9, and an effector domain or multipartite effector comprising a variant DNMT3A domain or functionally active portion thereof to mammalian genes or regulatory DNA elements thereof (e.g., a promoter or enhancer) via one or more gRNAs. Any of a variety of effector domains comprising a variant DNMT3A domain or functionally active portion thereof for transcriptional repression are known and can be used in accord with the provided embodiments. Effector domains, as well as transcriptional repression of target genes using Cas fusion proteins with the effector domains, are described, for example, in WO 2014 / 197748, WO 2017 / 180915, WO 2021 / 226077, WO 2013 / 176772, WO 2014 / 152432, WO 2014 / 093661, Adli, M. Nat. Commun. 9, 1911 (2018), and Gilbert, L. A. et al. Cell 154(2): 442-451 (2013). a. DNMT3A-DNMT3L Fusions
[0345] In some aspects, provided herein is an effector domain that comprises a fusion of a variant DNMT3A domain or functionally active portion thereof, such as any described herein in Section IV, and DNMT3L (DNMT3A / L, i.e., D3AL). DNMT3L is described herein in Section I.B.2.b. l.
[0346] In some embodiments, the effector domain comprises a fusion of a variant DNMT3A domain or functionally active portion thereof, such as any described herein in Section IV, and a DNMT3L domain (DNMT3A / L, i.e., D3AL), such as any DNMT3L domain described herein in Section I.B.2.b. 1. In some embodiments, the DNMT3L domain comprises a DNMT3L MTase-like domain. In some embodiments, the DNMT3L domain comprises a DNMT3L MTase-like domain and further comprises a DNMT3L ADDAttorney Docket No. 224742003640 domain. In some embodiments, the DNMT3L domain consists of a DNMT3L MTase-like domain and a DNMT3L ADD domain.
[0347] In some embodiments, the DNMT3L domain comprises, from N to C terminus, the DNMT3L ADD domain and DNMT3L MTase-like domain. In some embodiments, the DNMT3L domain comprises, form N to C terminus, the DNMT3L MTase-like domain and the DNMT3L ADD domain.
[0348] In some embodiments, the effector domain comprises a variant DNMT3A / L fusion containing the one or more amino acid modifications as described in Section IV that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the reference (e.g., unmodified) DNMT3A / L fusion domain or fragment thereof, such as with the amino acid sequence of SEQ ID NO: 264 or SEQ ID NO: 265. In some embodiments, the effector domain comprises a variant DNMT3A / L fusion containing the one of more amino acid modifications as described herein in Section I that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the reference (e.g., unmodified) DNMT3A / L fusion domain set forth in SEQ ID NO: 264 and retains the amino acid modification(s), e.g. substitution(s) therein not present in the reference (e.g., unmodified) DNMT3A / L fusion. In some embodiments, the effector domain comprises a variant DNMT3A / L fusion containing the one of more amino acid modifications as described herein in Section IV that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the reference (e.g., unmodified) DNMT3A / L fusion domain set forth in SEQ ID NO: 265 and retains the amino acid modification s), e.g. substitution(s) therein not present in the reference (e.g., unmodified) DNMT3A / L fiision. b. Multipartite Effectors for Transcriptional Repression
[0349] In some aspects, provided herein are multipartite effectors comprising a variant DNMT3A domain or a functionally active portion thereof and further comprising one or more domains selected from a DNA methyltransferase domain, a repressor domain capable of recruiting heterochromatin inducing factors (referred to as a repressor domain), or combinations thereof. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and further comprises a DNA methylfransferase domain and a repressor domain. In some embodiments, the multipartite effector is a variant DNMT3A domain or functionally active portion thereof as described herein in Section IV, a DNA methyltransferase domain herein provided in Section I.B.2.b. I, and a repressor domain herein provided in Section I.B.2.b.2.
[0350] In some aspects, the multipartite effector decreases transcription of an endogenous locus when recruited to a target site at the endogenous locus. For example, the multipartite effector decreases transcription of a gene when recruited (e.g., targeted to) a target site for the gene, such as a regulatory DNAAttorney Docket No. 224742003640 element (e.g., a promoter or enhancer). In some embodiments, the multipartite effector represses, catalyzes, or leads to decreased transcription of a gene when ectopically recruited to the gene or a DNA regulatory element thereof. In some embodiments, a multipartite effector induces, catalyzes, or leads to transcription repression, transcription co-repression, histone modification, histone acetylation, histone deacetylation, nucleosome remodeling, chromatin remodeling, heterochromatin formation, proteolysis, ubiquitination, deubiquitination, phosphorylation, dephosphorylation, splicing, DNA methylation, DNA demethylation, histone methylation, histone demethylation, or DNA base oxidation. In some embodiments, the multipartite effector induces, catalyzes, or leads to transcription repression or transcription co-repression. In some embodiments, the multipartite effector induces transcription repression. In some embodiments, the multipartite effector has one of the aforementioned activities itself (i.e., acts directly). In some embodiments, the multipartite effector domain recruits and / or interacts with a protein or polypeptide domain that has one of the aforementioned activities (i.e., acts indirectly).
[0351] In some embodiments, a fusion protein comprises a variant DNMT3A domain or functionally active portion thereof, a DNA methyltransferase domain, DNA binding domain, and / or a variable repressor domain (e.g., KRAB, EZH2, or variants thereof), in any sequential order or arrangement. In some embodiments, the fusion protein comprises a variant DNMT3A domain or functionally active portion thereof, a DNMT3L domain, DNA binding domain, and / or a repressor domain (e.g. KRAB, EZH2, or variants thereof), in any sequential order or arrangement. In some embodiments, the variant DNMT3A domain or functionally active portion thereof and DNMT3L domain is a fusion (i.e., variant DNMT3A / L fusion).
[0352] In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C-terminus: (i) variant DNMT3A or DNMT3A / L fusion, (ii) DNA binding domain, and (iii) repressor domain. In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C- terminus: (i) variant DNMT3A or DNMT3A / L fusion, (ii) repressor domain, and (iii) DNA binding domain. In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C-terminus: (i) DNA binding domain, (ii) variant DNMT3A or DNMT3A / L fusion, and (iii) repressor domain. In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C-terminus: (i) DNA binding domain, (ii) repressor domain, and (iii) variant DNMT3A or DNMT3A / L fusion. In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C-terminus: (i) repressor domain, (ii) variant DNMT3A or DNMT3A / L fusion, and (iii) DNA binding domain. In some embodiments, the sequential order of the fusion protein comprises from N-terminus to C-terminus: (i) repressor domain, (ii) DNA binding domain, and (iii) variant DNMT3A or DNMT3A / L fusion.Attorney Docket No. 2247420036401) DNA Methyltransferase Domains
[0353] In some aspects, provided herein are DNA methyltransferase (DNMT) domains. In some embodiments, an effector domain comprises a variant DNMT3A domain or functionally active portion thereof, as described herein in Section IV, and further comprises a DNMT domain (e.g., DNMT3 domain) or a variant thereof. DNMT3, including in dCas fusion proteins, have been described, for example, in US20190127713, Liu, X. S. et al. Cell 167, 233-247.el7 (2016), Lei, Y. et al. Nat. Commun. 8, 16026 (2017). In some embodiments, the effector domain is a multipartite effector that comprises a variant DNMT3A domain or functionally active portion thereof and further comprises at least one DNMT3 domain or a variant thereof. In some embodiments, the effector domain is a multipartite effector that comprises at least one DNMT3L domain, or a variant thereof. In some embodiments, the effector domain is a multipartite effector that comprises at least one DNMT3L domain or a variant thereof. An exemplary DNMT3L domain is set forth in SEQ ID NO: 15. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or a functionally active portion thereof and further comprises the sequence set forth in SEQ ID NO: 15, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0354] In some embodiments, the effector domain comprises a fusion of a variant DNMT3A domain or functionally active portion thereof and DNMT3L (DNMT3A / L, i.e., D3AL), as described herein in Section I.B.2.a.2) Repressor Domains
[0355] In some aspects, provided herein are repressor domains. In some embodiments, a multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof, as described herein in Section IV, and further comprises a repressor domain such as any described herein, a portion thereof, a partially or fully functional fragment or domain thereof, or a combination of any of the foregoing. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and further comprises a repressor domain capable of recruiting heterochromatin inducing factors. In some embodiments, the variable repressor domain a KRAB domain, ERF repressor domain, MXI 1 domain, SID4X domain, MAD-SID domain, LSD1, EZH2, a partially or fully functional fragment or domain of any of the foregoing, or a combination of any of the foregoing. In some embodiments, the effector domain is a multipartite effector that comprises a variant DNMT3A domain or functionally active portion thereof and further comprises a KRAB or a variant thereof. In some embodiments, one or more transcriptional repression domain comprises a domain of a protein selected from among a KRAB domain from KOX1, a KRAB domain from ZIM3, and a KRAB domain from ZNF324. In some embodiments, the repressor domain isAttorney Docket No. 224742003640EZH2 or a variant thereof. In some embodiments, the repressor domain comprises a transcriptional repressor domain described in WO 2021 / 226077. In some aspects, the fusion protein comprising a variant DNMT3A domain or functionally active portion thereof and any of the aforementioned repressor domains is targeted to a target site for a gene and leads to decreased transcription of the gene.
[0356] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and a KRAB domain, or a variant thereof. The KRAB-containing zinc finger proteins make up the largest family of transcriptional repressors in mammals. The Kriippel associated box (KRAB) domain is a repressor domain present in many zinc finger protein-based transcription factors. The KRAB domain comprises charged amino acids and can be divided into sub-domains A and B. The KRAB domain recruits corepressors KAP1 (KRAB-associated protein- 1), epigenetic readers such as heterochromatin protein 1 (HP1), and other chromatin modulators to induce transcriptional repression through heterochromatin formation. KRAB-mediated gene repression is associated with loss of histone H3- acetylation and an increase in H3 lysine 9 trimethylation (H3K9me3) at the repressed gene promoters. KRAB domains, including in dCas fusion proteins, have been described, for example, in WO 2017 / 180915, WO 2014 / 197748, US 2019 / 0127713, WO 2013 / 176772, Urrutia R. et al. Genome Biol. 4, 231 (2003), Groner A. C. et al. pLoS Genet. 6, el000869 (2010). In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one KRAB domain or a variant thereof. An exemplary KRAB domain is set forth in SEQ ID NO: 267. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 267, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one KRAB domain variant. An exemplary KRAB domain variant is a KRAB domain from a human KOX1 (also called zinc finger protein 10 or ZNF10), a member of the KRAB C2H2 zinc finger family protein, that confers strong transcriptional repressor activities even to remote promoter positions. In some embodiments, the KRAB domain from a human KOX1 is set forth in SEQ ID NOs: 19 or 268. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 19, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 268, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, an exemplary KRAB domain variant is aAttorney Docket No. 224742003640KRAB domain from zinc finger imprinted 3 (ZIM3), a transcriptional regulator found in zinc finger proteins. In some embodiments, the KRAB domain from ZIM3 is set forth in SEQ ID NO: 269. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 269, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing. In some embodiments, an exemplary KRAB domain variant is a KRAB domain from zinc finger protein 324 (ZNF324), a protein involved in transcriptional regulation. In some embodiments, the KRAB domain from ZNF324 is set forth in SEQ ID NO: 270. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 270, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0357] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one ERF repressor domain, or a variant thereof. ERF (ETS2 repressor factor) is a strong transcriptional repressor that comprises a conserved DNA binding domain and represses transcription via a distinct domain at the carboxyl-terminus of the protein. ERF repressor domains, including in dCas fusion proteins, have been described, for example, in W02017180915, WO2014197748, WO2013176772, Mavrothalassitis, G., Ghysdael, J. Proteins of the ETS family with transcriptional repressor activity. Oncogene 19, 6524-6532 (2000). In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one ERF repressor domain or a variant thereof. An exemplary ERF repressor domain is set forth in SEQ ID NO: 272. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 272, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0358] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one MXI1 domain, or a variant thereof. The MXI1 domain functions by antagonizing the myc transcriptional activity by competing for binding to myc-associated factor x (MAX). MXI1 domains, including in dCas fusion proteins, have been described, for example, in W02017180915, WO2014197748, US20190127713. In some embodiments, the multipartite effector comprises a variant DNMT3 A domain or functionally active portion thereof and at least one MXI 1 domain or a variant thereof. An exemplary MXI1 domain is set forth in SEQ ID NO: 273. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 273, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.Attorney Docket No. 224742003640
[0359] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one SID4X domain, or a variant thereof. The mSin3 interacting domain (SID) is present on different transcription repressor proteins. It interacts with the paired amphipathic alpha-helix 2 (PAH2) domain of mSin3, a transcriptional repressor domain that is attached to transcription repressor proteins such as the mSin3 A corepressor. A dCas9 molecule can be fused to four concatenated mSin3 interaction domains (SID4X). SID domains, including in dCas fusion proteins, have been described, for example, in W02017180915, WO2014197748, WO2014093655. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one SID domain or a variant thereof. An exemplary SID domain is set forth in SEQ ID NO: 274. In some embodiments, the effector domain comprises the sequence set forth in SEQ ID NO: 274, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0360] In some embodiments, the multipartite effector that comprises a variant DNMT3A domain or functionally active portion thereof and at least one MAD domain, or a variant thereof. The MAD family proteins, Madl, Mxil, Mad3, and Mad4, belong to the basic helix-loop-helix-zipper class and contain a conserved N terminal region (termed Sin3 interaction domain (SID)) necessary for repressional activity. MAD-SID domains, including in dCas fusion proteins, have been described, for example, in W02017180915, WO2014197748, WO2013176772. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and at least one MAD-SID domain or a variant thereof. An exemplary MAD-SID domain is set forth in SEQ ID NO: 275. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 275, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0361] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and a LSD1 domain. LSD1 (also known as Lysine-specific histone demethylase 1A) is a histone demethylase that can demethylate lysine residues of histone H3, thereby acting as a coactivator or a corepressor, depending on the context. LSD1, including in dCas fusion proteins, has been described, for example, in WO 2013 / 176772, WO 2014 / 152432, and Kearns, N. A. et al. Nat. Methods. 12(5):401-403 (2015). An exemplary LSD1 polypeptide is set forth in SEQ ID NO: 276. In some embodiments, the multipartite effector that comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 276, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity toAttorney Docket No. 224742003640 any of the foregoing.
[0362] In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and an EZH2 domain or a variant thereof. EZH2 (also known as Histonelysine N -methyltransferase EZH2) is a Catalytic subunit of the PRC2 / EED-EZH2 complex, which methylates 'Lys-9' (H3K9me) and 'Lys-27' (H3K27me) of histone H3, in some aspects leading to transcriptional repression of the affected target gene. EZH2, including in dCas fusion proteins, has been described, for example, in O’Geen, H. et al., Epigenetics Chromatin. 12(1):26 (2019). An exemplary EZH2 polypeptide is set forth in SEQ ID NO: 271. In some embodiments, the multipartite effector comprises a variant DNMT3A domain or functionally active portion thereof and the sequence set forth in SEQ ID NO: 271, or a portion thereof, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
[0363] In some embodiments, the repressor domain comprises or is selected from a repressor domain shown in Table 2, or a domain or a portion thereof, such as a contiguous portion thereof of at least 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids, such as at least 20 amino acids, or a variant thereof or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of repressor domain shown in Table 2, or a domain or a portion thereof, such as a contiguous portion thereof of at least 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids, such as at least 20 amino acids, or a variant thereof. In some embodiments, the repressor domain comprises or is selected from a repressor domain shown in Table 2, or a domain or a portion thereof, such as a contiguous portion thereof of at least 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids, such as at least 20 amino acids, or a variant thereof.
[0364] In some embodiments, the repressor domain comprises or is selected from a repressor domain shown in Table 2, or a domain or a portion thereof, such as a contiguous portion thereof of 10, 15, 20, 22, 25, 30, 35, 37, 40, 42, 45, 47, 49, 50, 55, 57, 60, 61, 62, 65, 70, 72, 75, 76, or 80 amino acids in length, or within a range defined by any of the foregoing. In some embodiments, the repressor domain comprises or is selected fro...
Claims
Attorney Docket No. 224742003640CLAIMS1. A fusion protein for targeted gene repression, wherein the fusion protein comprises:(a) a DNA-binding domain for targeting to a target site for PCSK9, and(b) an effector domain that recruits domains with DNA methyltransferase activity to the target site for PCSK9, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides.
2. The fusion protein of claim 1, wherein the effector domain comprises a catalytically inactive DNA methyltransferase domain or portion thereof.
3. A fusion protein for targeted gene repression, wherein the fusion protein comprises:(a) a DNA-binding domain for targeting to a target site for PCSK9, and(b) an effector domain comprising a catalytically inactive DNA methyltransferase domain or portion thereof, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides.
4. A fusion protein for targeted gene repression, wherein the fusion protein comprises:(a) a DNA-binding domain for targeting to a target site for PCSK9, and(b) an effector domain comprising a catalytically inactive DNA methyltransferase domain or portion thereof, wherein the length of the fusion protein minus (a) is less than 750 amino acids.
5. The fusion protein of any one of claims 1-4, wherein the length of the fusion protein minus (a) is less than 500 amino acids.
6. The fusion protein of any one of claims 1-5, wherein the length of the fusion protein minus (a) is less than 370 amino acids.
7. The fusion protein of any one of claims 4-6, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity and domains capable of recruiting heterochromatin inducing factors.Attorney Docket No. 2247420036408. The fusion protein of any one of claims 4-7, wherein the fusion protein is devoid of any H3K4meO peptides.
9. A fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA- binding domain for targeting to a target site for PCSK9 and an effector domain that is a single effector domain, wherein the single effector domain comprises a catalytically inactive DNA methyltransferase domain or a functional portion thereof, and wherein the single effector domain is less than 600 amino acids in length.
10. A fusion protein for targeted gene repression, wherein the fusion protein consists essentially of:(a) a DNA-binding domain for targeting to a target site for PCSK9, and(b) an effector domain, wherein the effector domain comprises a catalytically inactive DNA methyltransferase domain or a functional portion thereof, and wherein the effector domain is less than 600 amino acids in length.
11. The fusion protein of claim 10, wherein the catalytically inactive DNA methyltransferase domain or portion thereof comprises a DNMT3L protein or portion thereof that recruits domains with DNA methyltransferase activity to the target site for PCSK9.
12. The fusion protein of any one of claims 1-11, wherein the length of the effector domain is less than 510 amino acids.
13. The fusion protein of any one of claims 1-12, wherein the length of the effector domain is less than 400 amino acids.
14. The fusion protein of any one of claims 1-13, wherein the length of the effector domain is less than 300 amino acids.Attorney Docket No. 22474200364015. The fusion protein of any one of claims 9-14, wherein the effector domain is devoid of any domains with DNA methyltransferase activity and domains capable of recruiting heterochromatin inducing factors.
16. The fusion protein of any one of claims 9-15, wherein the fusion protein is devoid of any H3K4meO peptides.
17. The fusion protein of any one of claims 11-16, wherein the DNMT3L protein or portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, Apodemus sylvaticus, Rattus norvegicus, Bos taurus, Papio Anubis, Cebus imitator, Macaca nemestrina, Pongo abelii, Lexodonta Africana, Pan troglodytes, and Chlorocebus sabaeus.
18. The fusion protein of any one of claims 11-17, wherein the DNMT3L protein or portion thereof is selected from one of the following organisms: Homo sapiens, Mus musculus, and Apodemus sylvaticus.
19. The fusion protein of any one of claims 11-18, wherein the portion of the DNMT3L protein is a contiguous portion that is less than a full-length DNMT3L MTase-like domain and comprises at least 10 amino acids from a reference DNMT3L MTase-like domain, wherein the contiguous portion of at least 10 amino acids is involved in a DNMT3A-DNMT3L interface.
20. The fusion protein of claim 19, wherein the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5 (SEQ ID NO: 126), wherein Xi is L or M, X2is L or I, X3is L or R, X4 is K or R, and X5 is P or Q.
21. The fusion protein of claim 19, wherein the contiguous portion comprises the sequence set forth in X1DX2X3X4X5X6RFLX7 (SEQ ID NO: 127), wherein XI is E or D, X2 is L or Q, X3 is D, E, or M, X4 is V or T, X5 is A or T, X6 is S, T, or V, and X7 is E or Q.
22. The fusion protein of any one of claims 19-21, wherein the contiguous portion comprises the sequence set forth as amino acid residues 257-273 or 292-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108.Attorney Docket No. 22474200364023. The fusion protein of any one of claims 19, 20, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 128.
24. The fusion protein of any one of claims 19, 20, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 129.
25. The fusion protein of any one of claims 19, 20, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 130.
26. The fusion protein of any one of claims 19, 21, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 131.
27. The fusion protein of any one of claims 19, 21, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 132.
28. The fusion protein of any one of claims 19, 21, and 22, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 133.
29. The fusion protein of any one of claims 19-22, wherein the contiguous portion comprises the sequence set forth in WYX1FQFHRX2LQYAX3PX4X5X6SX7X8PFFWX9FX10DNLX11LX12X13X14DX15X16X17X18X19RFLX20 (SEQ ID NO: 134), wherein Xi is L or M, X2 is L or I, X3 is L or R, X4 is K or R, X5 is P or Q, Xe is G or E, X7 is P, Q, or absent, Xs is R or Q, X9 is M or I, Xw is V or M, Xu is V or L, X12 is N or T, X13 is K or E, X14 is E or D, X15 is L or Q, Xie is D, E, or M, X17 is V or T, Xis is A or T, X19 is S, T, or V, and X20 is E or Q.
30. The fusion protein of any one of claims 19-22 and 29, wherein the contiguous portion comprises the sequence set forth as amino acid residues 257-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108.
31. The fusion protein of any one of claims 19-22, 29, and 30, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 135.Attorney Docket No. 22474200364032. The fusion protein of any one of claims 19-22, 29, and 30, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 136.
33. The fusion protein of any one of claims 19-22, 29, and 30, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 137.
34. The fusion protein of any one of claims 19-22, 29, and 30, wherein the contiguous portion comprises the sequence set forth in XiX2VRX3DVEX4WGPFDLX5YGX6TX7PLGX8X9CDRXioPXiiWYXi2FQFHRXi3LQYAXi4PXi5Xi6Xi7SX 18X19PFFWX2OFX21DNLX22LX23X24X25DX26X27X28X29X3ORFLX31 (SEQ ID NO: 138), wherein XI is D or N, X2 is T or V, X3 is K or R, X4 is E or K, X5 is V or L, X6 is A or S, X7 is P or Q, X8 is H or S, X9 is T or S, X10 is P or C, XI 1 is S or G, X12 is L or M, X13 is L or I, X14 is L or R, X15 is K or R, Xie is P or Q, X17 is G or E, Xi8is P, Q, or absent, X19 is R or Q, X20 is M or I, X21 is V or M, X22 is V or L, X23 is N or T, X24 is K or E, X25 is E or D, X26 is L or Q, X27 is D, E, or M, X28is V or T, X29 is A or T, X30 is S, T, or V, and X31 is E or Q.
35. The fusion protein of any one of claims 19-22, 29, 30, and 34, wherein the contiguous portion comprises the sequence set forth as amino acid residues 225-302 from the reference DNMT3L MTase-like domain, corresponding to numbering of positions set forth in SEQ ID NO: 108.
36. The fusion protein of any one of claims 19-22, 29, 30, 34, and 35, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 139.
37. The fusion protein of any one of claims 19-22, 29, 30, 34, and 35, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 140.
38. The fusion protein of any one of claims 19-22, 29, 30, 34, and 35, wherein the contiguous portion comprises the sequence set forth in SEQ ID NO: 141.
39. The fusion protein of any one of claims 19-38, wherein the contiguous portion comprises at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or at least 75 amino acids.Attorney Docket No. 22474200364040. The fusion protein of any one of claims 19-39, wherein the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing.
41. The fusion protein of any one of claims 19-40, wherein the reference DNMT3L MTase-like domain comprises any one of the sequences set forth in SEQ ID NOs: 15-17.
42. The fusion protein of any one of claims 19-23, 26, 29-31, 34-36, and 39-41, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing.
43. The fusion protein of any one of claims 19-23, 26, 29-31, 34-36, and 39-42, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 15.
44. The fusion protein of any one of claims 19-22, 24, l, 29, 30, 32, 34, 35, 37, and 39-41, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing.
45. The fusion protein of any one of claims 19-22, 24, l, 29, 30, 32, 34, 35, 37, 39-41, and 44, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 16.
46. The fusion protein of any one of claims 19-22, 25, 28-30, 33-35, and 38-41, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing.
47. The fusion protein of any one of claims 19-22, 25, 28-30, 33-35, 38-41, and 46, wherein the reference DNMT3L MTase-like domain comprises the sequence set forth in SEQ ID NO: 17.Attorney Docket No. 22474200364048. The fusion protein of any one of claims 11, 17, and 18, wherein the DNMT3L protein or portion thereof is a DNMT3L MTase-like domain.
49. The fusion protein of any one of claims 11 and 17-47, wherein the DNMT3L protein or portion thereof is a portion of a DNMT3L MTase-like domain.
50. The fusion protein of any one of claims 11, 17, 18, and 48, wherein the DNMT3L protein or portion thereof comprises any one of the sequences set forth in SEQ ID NOs: 15-17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity of any of the foregoing.
51. The fusion protein of any one of claims 11, 17, 18, 48, and 50, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
52. The fusion protein of any one of claims 11, 17, 18, 48, 50, and 51, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15.
53. The fusion protein of any one of claims 11, 17, 18, 48, and 50, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
54. The fusion protein of any one of claims 11, 17, 18, 48, 50, and 53, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16.
55. The fusion protein of any one of claims 11, 17, 18, 48, and 50, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
56. The fusion protein of any one of claims 11, 17, 18, 48, 50, and 55, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17.Attorney Docket No. 22474200364057. The fusion protein of any one of claims 48-56, wherein the effector domain further comprises a DNMT3L ADD domain.
58. The fusion protein of claim 57, wherein the effector domain, from N-terminus to C-terminus, comprises the DNMT3L ADD domain and the DNMT3L MTase-like domain.
59. The fusion protein of claim 57, wherein the effector domain, from N-terminus to C-terminus, comprises the DNMT3L MTase-like domain and the DNMT3L ADD domain.
60. The fusion protein of any one of claims 57-59, wherein the DNMT3L ADD domain comprises the sequence set forth in any one of SEQ ID NOs: 109, 110, 379, and 380.
61. The fusion protein of any one of claims 57-60, wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 379.
62. The fusion protein of any one of claims 57-60, wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 380.
63. The fusion protein of any one of claims 57-60, wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109 or 110.
64. The fusion protein of any one of claims 57-60 and 63, wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109.
65. The fusion protein of any one of claims 57-60 and 63, wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 110.
66. The fusion protein of any one of claims 1-65, wherein the effector domain comprises the sequence set forth in any one of SEQ ID NO: 378, 381, or 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
67. The fusion protein of any one of claims 1-66, wherein the effector domain comprises the sequence set forth in any one of SEQ ID NO: 378, 381, or 382.Attorney Docket No. 22474200364068. The fusion protein of any one of claims 1-60 and 66, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 378, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
69. The fusion protein of any one of claims 1-60, 66, and 68, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 378.
70. The fusion protein of any one of claims 1-60 and 66, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 381, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
71. The fusion protein of any one of claims 1-60, 66, and 70, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 381.
72. The fusion protein of any one of claims 1-60 and 66, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 382, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
73. The fusion protein of any one of claims 1-60, 66, and 72, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 382.
74. The fusion protein of any one of claims 1-11, 17, 18, 48, 50-52, 57, 58, 60, 63, and 64, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 107, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
75. The fusion protein of any one of claims 1-11, 17, 18, 48, 50-52, 57, 58, 60, 63, 64, and 74, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 107.
76. The fusion protein of any one of claims 1-11, 17, 18, 48, 50, 51, 55-58, 60, 63, and 65, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 108, or an amino acid sequenceAttorney Docket No. 224742003640 that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
77. The fusion protein of any one of claims 1-11, 17, 18, 48, 50, 51, 55-58, 60, 63, 65, and 76, wherein the effector domain comprises the sequence set forth in SEQ ID NO: 108.
78. The fusion protein of any one of claims 1-77, wherein the DNA-binding domain is selected from: a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or a variant thereof; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an LScel enzyme or a variant thereof.
79. The fusion protein of any one of claims 1-78, wherein the target site for PCSK9 is located within 500 bp of the genomic coordinate chrl:55,039,548.
80. The fusion protein of any one of claims 1-79, wherein the target site for PCSK9 is located within 300 bp of the genomic coordinate chrl:55,039,548.
81. The fusion protein of any one of claims 1-80, wherein the target site for PCSK9 is located within 200 bp of the genomic coordinate chrl:55,039,548.
82. The fusion protein of any one of claims 1-81, wherein the target site for PCSK9 is located within 100 bp of the genomic coordinate chrl:55,039,548.
83. The fusion protein of any one of claims 1-82, wherein the target site for PCSK9 is located within 80 bp of the genomic coordinate chrl:55,039,548.
84. The fusion protein of any one of claims 1-81, wherein the target site for PCSK9 is within the coordinates chrl: 55,039,338-55,039,658.
85. The fusion protein of any one of claims 1-84, wherein the target site for PCSK9 is within the coordinates chrl: 55,039,470-55,039,597.Attorney Docket No. 22474200364086. The fusion protein of any one of claims 1-85, wherein the target site for PCSK9 is or comprises the coordinates chrl: 55,039,538-55,039,557.
87. The fusion protein of any one of claims 1-85, wherein the target site comprises the sequence set forth in any one of SEQ ID NOs: 39 and 142-144, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
88. The fusion protein of any one of claims 1-87, wherein the target site comprises the sequence set forth in SEQ ID NO: 39, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
89. The fusion protein of any one of claims 1-87, wherein the target site comprises the sequence set forth in SEQ ID NO: 142, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
90. The fusion protein of any one of claims 1-87, wherein the target site comprises the sequence set forth in SEQ ID NO: 143, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
91. The fusion protein of any one of claims 1-87, wherein the target site comprises the sequence set forth in SEQ ID NO: 144, a contiguous portion thereof of at least 14 nucleotides (nt), or a complementary sequence of any of the foregoing.
92. A fusion protein for targeted gene repression, wherein the fusion protein comprises:(a) a DNA-binding domain for targeting to a target site for PCSK9, wherein the target site for PCSK9 is located within 300 bp of the genomic coordinate chrl:55,039,548, and(b) an effector domain comprising a DNMT3L protein or portion thereof comprising the sequence set forth in any one of SEQ ID NOs: 15-17 or an amino acid sequence that has at least 90% identity to any one of SEQ ID NOs: 15-17, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides.
93. A fusion protein for targeted gene repression, wherein the fusion protein comprises:Attorney Docket No. 224742003640(a) a DNA-binding domain for targeting to a target site for PCSK9, wherein the target site for PCSK9 is within the coordinates chrl: 55,039,470-55,039,597, and(b) an effector domain comprising a DNMT3L protein or portion thereof comprising the sequence set forth in any one of SEQ ID NOs: 15-17 or an amino acid sequence that has at least 90% identity to any one of SEQ ID NOs: 15-17, wherein the fusion protein is devoid of any domains with DNA methyltransferase activity, domains capable of recruiting heterochromatin inducing factors, and H3K4meO peptides.
94. The fusion protein of claim 92 or claim 93, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
95. The fusion protein of any one of claims 92-94, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 15.
96. The fusion protein of claim 92 or claim 93, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
97. The fusion protein of any one of claims 92, 93, and 96, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 16.
98. The fusion protein of claim 92 or claim 93, wherein the DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
99. The fusion protein of any one of claims 92, 93, and 98, wherein DNMT3L protein or portion thereof comprises the sequence set forth in SEQ ID NO: 17.
100. The fusion protein of any one of claims 1-99, wherein the effector domain is independently fused to the N-terminus, the C-terminus, or both the N-terminus and the C-terminus, of the DNA-binding domain.Attorney Docket No. 224742003640101. The fusion protein of any one of claims 11-100, wherein the fusion protein comprises, from N-terminus to C-terminus, the DNMT3L protein or portion thereof and the DNA-binding domain.
102. The fusion protein of any one of claims 11-100, wherein the fusion protein comprises, from N-terminus to C-terminus, the DNA-binding domain and the DNMT3L protein or portion thereof.
103. The fusion protein of any one of claims 1-102, wherein the DNA-binding domain is a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or variant thereof, and the Cas protein is capable of complexing with a gRNA for targeting the DNA-binding domain to the target site for PCSK9.
104. The fusion protein of claim 103, wherein the Cas protein or variant thereof is a deactivated (dCas) protein.
105. The fusion protein of claim 104, wherein the dCas protein lacks nuclease activity.
106. The fusion protein of claim 104 or claim 105, wherein the dCas protein is a dCasl2 protein.
107. The fusion protein of claim 104 or claim 105, wherein the dCas protein is a dCas9 protein.
108. The fusion protein of claim 107, wherein the dCas9 protein is a Staphylococcus aureus dCas9 (dSaCas9) protein.
109. The fusion protein of claim 108, wherein the dSaCas9 comprises at least one amino acid mutation selected from D10A and N580A, with reference to numbering of positions of SEQ ID NO: 41.
110. The fusion protein of claim 108 or claim 109, wherein the dSaCas9 protein comprises the sequence set forth in SEQ ID NO: 40, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
111. The fusion protein of any one of claims 108-110, wherein the dSaCas9 is set forth in SEQ ID NO: 40.Attorney Docket No. 224742003640112. The fusion protein of claim 107, wherein the dCas9 protein is a Streptococcus pyogenes dCas9 (dSpCas9) protein.
113. The fusion protein of claim 112, wherein the dSpCas9 protein comprises at least one amino acid mutation selected from D10A and H840A, with reference to numbering of positions of SEQ ID NO: 41.
114. The fusion protein of claim 112 or claim 113, wherein the dSpCas9 comprises the sequence set forth in SEQ ID NO: 18, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
115. The fusion protein of any one of claims 112-114, wherein the dSpCas9 is set forth in SEQ ID NO: 18.
116. The fusion protein of any one of claims 103-115, wherein the gRNA comprises a gRNA spacer sequence comprising the sequence set forth in any one of SEQ ID NOs: 43 and 155-157.
117. The fusion protein of any one of claims 103-116, wherein the gRNA comprises a gRNA spacer sequence comprising the sequence set forth in SEQ ID NO: 43.
118. The fusion protein of any one of claims 1-102, wherein the DNA-binding domain is a zinc finger protein (ZFP).
119. The fusion protein of any one of claims 1-118, wherein the fusion protein further comprises a nuclear localization signal (NLS).
120. The fusion protein of claim 119, wherein the NLS is a nucleoplasmin NLS or an SV40 NLS.
121. The fusion protein of claim 119 or 120, wherein the NLS comprises the sequence set forth in SEQ ID NO: 44 or 34.
122. The fusion protein of any one of claims 1-121, wherein the fusion protein comprises a linker.Attorney Docket No. 224742003640123. The fusion protein of claim 122, wherein the linker comprises the sequence set forth in any one of SEQ ID NOS: 33 and 45-47, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
124. The fusion protein of claim 122 or claim 123, wherein the linker comprises the sequence set forth in SEQ ID NO: 33, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
125. The fusion protein of any one of claims 122-124, wherein the linker comprises the sequence set forth in SEQ ID NO: 33.
126. The fusion protein of any one of claims 1-18, 48, 50-56, 78-101, 103-105, 107, 112-117, and 119-125, wherein the fusion protein comprises the sequence set forth in any of SEQ ID NOs: 1-3, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
127. The fusion protein of any one of claims 1-18, 48, 50-52, 78-95, 100, 101, 103-105, 107, 112- 117, and 119-126, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 1 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
128. The fusion protein of any one of claims 1-18, 48, 50-52, 78-95, 100, 101, 103-105, 107, 112- 117, and 119-127, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 1.
129. The fusion protein of any one of claims 1-18, 48, 50, 53, 54, 78-93, 96, 97, 100, 101, 103- 105, 107, 112-117, and 119-126, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 2 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
130. The fusion protein of any one of claims 1-18, 48, 50, 53, 54, 78-93, 96, 97, 100, 101, 103- 105, 107, 112-117, 119-126, and 129, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 2.Attorney Docket No. 224742003640131. The fusion protein of any one of claims 1-18, 48, 50, 55, 56, 78-93, 98-101, 103-105, 107, 112-117, and 119-126, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 3 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
132. The fusion protein of any one of claims 1-18, 48, 50, 55, 56, 78-93, 98-101, 103-105, 107, 112-117, 119-126, and 131, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 3.
133. A fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA- binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a reference DNMT3A sequence at a position selected from among 686, 707-721, 756, 766, 771, 831-848, 854, 855, 860, 873, 876, 879, and 881-887, corresponding to numbering of positions set forth in SEQ ID NO: 183.
134. A fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA- binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a position of a DNA binding region of a reference DNMT3A sequence, wherein the DNA binding region is a catalytic loop, a target recognition domain (TRD), and a homodimer interface, or a combination thereof.
135. The fusion protein of claim 133 or claim 134, wherein the reference DNMT3A sequence comprises the sequence of amino acids set forth in SEQ ID NO: 183 or a functionally active portion thereof.
136. The fusion protein of any one of claims 133-135, wherein the functionally active portion thereof comprises a contiguous sequence contained within amino acid residues 612-912, with reference to positions set forth in SEQ ID NO: 183.
137. The fusion protein of any one of claims 133-136, wherein the functionally active portion thereof comprises a methyltransferase (MTase) domain.Attorney Docket No. 224742003640138. The fusion protein of any one of claims 133-137, wherein the reference DNMT3A sequence comprises the sequence set forth in SEQ ID NO: 113.
139. The fusion protein of any one of claims 133-138, wherein one or more amino acid substitutions are at a position selected from among 686, 711, 714, 756, 766, 771, 831, 832, 835, 836, 838, 841, 844, 845, 847, 854, 855, 860, 873, 876, 879, 881, 882, 883, and 887.
140. The fusion protein of any one of claims 133-139, wherein the one or more amino acid substitutions are selected from D686A, N711A, S714A, E756A, K766E, R771Q, R831A, R831E, T832A, T835A, T835E, R836E, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, W860A, H873A, D876G, N879A, S881A, R882H, L883A, R887A, and R887E, or a conservative amino acid substitution thereof.
141. The fusion protein of any one of claims 133-140, wherein the one or more amino acid substitutions are selected from D686A, N711A, S714A, K766E, R771Q, R831A, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, D876G, N879A, L883A, and R887A, or any combination thereof.
142. The fusion protein of any one of claims 133-141, wherein the one or more amino acid substitutions are a single amino acid substitution selected from D686A, N711A, S714A, K766E, R771Q, R831A, N838A, K841A, K841E, K844E, D845K, H847E, E854H, K855E, D876G, N879A, L883A, and R887A.
143. A fusion protein for targeted gene repression, wherein the fusion protein comprises a DNA- binding domain and an effector domain comprising a variant DNMT3A domain or functionally active portion thereof, wherein the variant DNMT3A domain or functionally active portion thereof comprises one or more amino acid substitutions in a reference DNMT3A sequence at positions selected from among 613, 621, 630, 631, 632, 635, 651, 659, 676, 677, 680, 688, 693, 694, 720, 721, 729, 736, 739, 742, 744, 749, 766, 767, 771, 783, 789, 790, 792, 803, 812, 821, 826, 823, 829, 831, 836, 841, 844, 847, 855, 866, 873, 882, 885, 887, 891, 899, 900, and 906, corresponding to numbering of position set forth in SEQ ID NO: 183, wherein the substituted amino acid is alanine (A), aspartic acid (D), glutamic acid (E), glutamine (Q), or asparagine (N).Attorney Docket No. 224742003640144. The fusion protein of any one of claims 133-143, wherein the effector domain represses transcription of one or more target genes.
145. The fusion protein of any one of claims 133-144, wherein the effector domain is a multipartite effector that further comprises one or more domains selected from a DNA methyltransferase domain, a repressor domain capable of recruiting heterochromatin inducing factors, or combinations thereof.
146. The fusion protein of claim 145, wherein the repressor domain is selected from a KRAB domain, ERF domain, Mxil domain, SID4X domain, Mad-SID domain, LSD1 domain, EZH2 domain, or variant of any of the foregoing.
147. The fusion protein of claim 145 or claim 146, wherein the repressor domain comprises the sequence set forth in any one of SEQ ID NOS: 19 and 267-276, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity to any of the foregoing.
148. The fusion protein of any one of claims 145-147, wherein the repressor domain is a KRAB domain.
149. The fusion protein of any one of claims 145-148, wherein the DNA methyltransferase domain is a DNMT3L domain.
150. The fusion protein of any one of claims 133-149, wherein the effector domain is a multipartite effector that comprises the variant DNMT3A domain; a DNMT3L domain; and a KRAB domain.
151. The fusion protein of any one of claims 133-150, wherein the fusion protein comprises from N-terminus to C-terminus: (i) the variant DNMT3A domain, (ii) a DNMT3L domain, (iii) the DNA-binding domain, and (iv) a KRAB domain.
152. The fusion protein of any one of claims 146-151, wherein the KRAB domain is selected from: a KRAB domain from KOX1, a KRAB domain from ZIM3, and a KRAB domain from ZNF324.Attorney Docket No. 224742003640153. The fusion protein of any one of claims 146-152, wherein the KRAB domain comprises the sequence set forth in any one of SEQ ID NOS: 19 and 267-270, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
154. The fusion protein of any one of claims 149-153, wherein the DNMT3L domain comprises an DNMT3L MTase-like domain.
155. The fusion protein of any one of claims 149-154, wherein the DNMT3L domain comprises any one of the sequences set forth in SEQ ID NO: 15-17, a portion thereof, or an amino acid sequence with at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto, optionally wherein the DNMT3L domain comprises the sequence set forth in SEQ ID NO: 15.
156. The fusion protein of any one of claims 149-155, wherein the DNMT3L domain further comprises an DNMT3L ADD domain, optionally wherein the DNMT3L ADD domain comprises the sequence set forth in SEQ ID NO: 109.
157. The fusion protein of any one of claims 133-156, wherein the fusion protein comprises a variant DNMT3A-DNMT3L (DNMT3A / L) fusion protein.
158. The fusion protein of any one of claims 133-157, wherein the DNA-binding domain is selected from: a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or a variant thereof; a zinc finger protein (ZFP); a transcription activator-like effector (TALE); a meganuclease; a homing endonuclease; or an LScel enzyme or a variant thereof, optionally wherein the DNA-binding domain comprises a catalytically inactive variant of any of the foregoing.
159. The fusion protein of any one of claims 133-158, wherein the DNA-binding domain is a Clustered Regularly Interspaced Short Palindromic Repeats associated (Cas) protein or variant thereof, optionally wherein the Cas protein or variant thereof is a deactivated (dCas) protein, more optionally wherein the dCas protein is a dCas9 protein.
160. The fusion protein of claim 159, wherein the dCas9 protein is a Streptococcus pyogenes dCas9 (dSpCas9) protein, optinally wherein the dSpCas9 comprises the sequence set forth in SEQ ID NO:Attorney Docket No. 22474200364018, or an amino acid sequence that has at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
161. The fusion protein of any one of claims 133-158, wherein the DNA-binding domain is a zinc finger protein (ZFP), wherein the ZFP binds to a target site in a gene or a regulatory DNA element thereof.
162. The fusion protein of any one of claims 133-161, wherein the fusion protein comprises a nuclear localization signal (NLS), optionally wherein the NLS comprises the sequence set forth in SEQ ID NO: 34 or 44.
163. The fusion protein of any one of claims 133-162, wherein the fusion protein comprises a linker.
164. The fusion protein of claim 163, wherein the linker comprises the sequence set forth in any one of SEQ ID NOS: 33, 45, 46, 47, and 120, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
165. The fusion protein of any one of claims 133-164, wherein the fusion protein comprises the sequence set forth in any one of SEQ ID NOS: 213-241, 243, 245-261, 277, and 278 or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
166. The fusion protein of any one of claims 133-160 and 162-165, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 277, or a sequence having at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.
167. The fusion protein of any one of claims 133-160 and 162-166, wherein the fusion protein comprises the sequence set forth in SEQ ID NO: 277.
168. A DNA-targeting system for gene repression, comprising the fusion protein of any one of claims 1-167.
169. A DNA-targeting system for gene repression, comprising: a) the fusion protein of any one of claims 1-117, 119-161, and 163-167; andAttorney Docket No. 224742003640 b) one guide RNA (gRNA) that targets the DNA-binding domain to the target site for PCSK9.
170. A polynucleotide encoding the fusion protein of any one of claims 1-167.
171. A polynucleotide encoding the DNA-targeting system of claim 168 or claim 169.
172. A plurality of polynucleotides comprising: a) a polynucleotide encoding the fusion protein of any one of claims 1-117, 119-161, and 163-167; and b) at least one guide RNA (gRNA) that targets the DNA-binding domain of the fusion protein of any one of claims 1-117, 119-161, and 163-167 to the target site.
173. A vector comprising the fusion protein of any one of claims 1-167, the DNA-targeting system of claim 168 or claim 169, the polynucleotide of claim 170 or claim 171, or the plurality of polynucleotides of claim 172.
174. The vector of claim 173, wherein the vector is a viral vector.
175. The vector of claim 174, wherein the viral vector is an adeno-associated virus (AAV) vector.
176. The vector of claim 175, wherein the vector is a non-viral vector.
177. The vector of claim 176, wherein the non-viral vector is selected from: a lipid nanoparticle, a liposome, an exosome, or a cell penetrating peptide.
178. A lipid nanoparticle comprising the fusion protein of any one of claims 1-167, the DNA- targeting system of claim 168 or claim 169, the polynucleotide of claim 170 or claim 171, or the plurality of polynucleotides of claim 172.
179. A method of targeted gene repression, comprising introducing into a cell: the fusion protein of any one of claims 133-167, the DNA-targeting system of claim 168 or claim 169, the polynucleotide of claim 170 or claim 171, the plurality of polynucleotides of claim 172, the vector of any one of claims 173- 177, or the lipid nanoparticle of claim 178.Attorney Docket No. 224742003640180. A method of targeted PCSK9 gene repression, comprising introducing into a cell comprising a domain with DNA methyltransferase activity: the fusion protein of any one of claims 1-167, the DNA- targeting system of claim 168 or claim 169, the polynucleotide of claim 170 or claim 171, the plurality of polynucleotides of claim 172, the vector of any one of claims 173-177, or the lipid nanoparticle of claim 178.
181. The method of claim 179 or claim 180, wherein the cell is a cell from and / or in a subject.
182. The method of any one of claims 179-181, wherein the introducing is by transient delivery into the cell.
183. The method of claim 182, wherein the transient delivery comprises electroporation, transfection, or transduction.
184. The method of any one of claims 179-183, wherein the fusion protein, the DNA-targeting system, the polynucleotide and / or the plurality of polynucleotides is transiently expressed and / or transiently present in the cell for a period of time after the introducing.
185. The method of any one of claims 180-184, wherein the PCSK9 gene repression is a reduction in the expression of a PCSK9 gene.
186. The method of any one of claims 180-185, wherein the PCSK9 gene repression is a reduction in the transcription of a PCSK9 gene.
187. The method of any one of claims 180-186, wherein the PCSK9 gene repression is sustained.
188. The method of any one of claims 180-187, wherein the PCSK9 gene repression is sustained until after the fusion protein, the DNA-targeting system, the polynucleotide, and / or the plurality of polynucleotides is no longer expressed and / or present in the cell.
189. The method of any one of claims 180-188, wherein the PCSK9 gene repression is sustained for at least 1 day, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 1 month, at leastAttorney Docket No. 2247420036403 months, at least 6 months, or at least 1 year, or more, after the fusion protein, the DNA-targeting system, the polynucleotide, and / or the plurality of polynucleotides is no longer expressed and / or present in the cell.
190. The method of any one of claims 180-189, wherein the PCSK9 gene repression is sustained at a fold-change of 0.2 or less relative to not introducing the fusion protein, DNA-targeting system, polynucleotide, vector, or lipid nanoparticle into the cell, after the fusion protein, the DNA-targeting system, and / or the polynucleotide is no longer expressed and / or present in the cell.
191. The method of any one of claims 179-190, wherein off-target methylation activity is reduced by more than 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%, as compared to the use of a fusion protein comprising the sequence of amino acids set forth in SEQ ID NO: 14.
Citation Information
Patent Citations
Lipids and lipid nanoparticle formulations for delivery of nucleic acids
US10723692B2
Method for gene editing
US10941395B2
Liposomal apparatus and manufacturing methods
US20040142025A1
Systems and methods for manufacturing liposomes
US20070042031A1
Crispr / CAS9-based repressors for silencing gene targets in vivo and methods of use
US20190127713A1