Engineered epigenetic effectors

Ultracompact DNA methyltransferase and TET domains, combined with compact dCas proteins, address the delivery and universality issues of current epigenetic editors, enabling durable and potent gene activation through synergistic transcriptional regulation for in vivo epigenome engineering.

WO2025254779A1PCT designated stage Publication Date: 2025-12-11EPICRISPR BIOTECHNOLOGIES INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/028952
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-27
Filing Date
2025-05-12
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Current epigenetic editors for targeted DNA methylation and demethylation are too large for effective tissue-specific delivery in vivo, and their modulation of gene expression may not universally apply across different chromatin states or levels of gene expression.

Method used

Development of ultracompact DNA methyltransferase and ten eleven translocation (TET) protein domains, combined with compact dCas proteins, to create a versatile platform for modulating gene expression, utilizing synergistic relationships between transcriptional regulation mechanisms to achieve durable and potent gene activation.

Benefits of technology

The compact epigenetic editors enable durable and potent gene expression modulation, allowing for single AAV packaging for in vivo epigenome engineering and achieving synergistic and long-lasting activation of epigenetically silenced loci.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028952_11122025_PF_FP_ABST
    Figure US2025028952_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are engineered epigenetic effectors, such as DNA methyltransferases and methylcytosine dioxygenases that are smaller than their naturally occurring counterparts. The engineered epigenetic effectors find use in regulating expression level of target genes in a cell, e.g., when targeted to a genetic locus of interest by a targeting moiety, such as a Cas protein, coupled thereto.
Need to check novelty before this filing date? Find Prior Art

Description

ENGINEERED EPIGENETIC EFFECTORSREFERENCE TO PRIORITY APPLICATIONS

[0001] This application claims priority to U.S. provisional application numbers: 63 / 764368, filed February 27, 2025; 63 / 701357, filed September 30, 2024; and 63 / 657085, filed June 6, 2024. The content of each of these aforementioned applications is expressly incorporated herein by reference in its entirety.REFERENCE TO SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing in electronic format The Sequence Listing is provided as aa file entitled EPICR035WOSequenceListing.xml, created May 12, 2025, which is 1,717,862 bytes in size. The information in the electronic format of the Sequence Listing is incorporated by reference in its entirety.BACKGROUND

[0003] Various effectors (e.g., transcriptional regulators and / or epigenetic efifectors / editors) can be utilized to regulate expression or activity of a target gene in the cell. For example, a heterologous gene effector can be introduced (e.g., delivered, expressed, etc.) to the cell, and the heterologous gene effector, either alone or along with an additional agent, can effect such regulation of the target gene. In some examples, the additional agent can comprise a clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR- associated protein (Cas) for specifically binding to the target gene (e.g., a target deoxyribonucleic acid (DNA) sequence or ribonucleic acid (RNA) sequence (e.g., foreign DNA sequence or RNA sequence) of the target gene), while the heterologous gene effector can regulate expression or activity level of the target gene. Such gene effectors can be utilized, e.g., as gene therapy to treat or ameliorate a condition (e.g., a disease) of a subjectSUMMARY

[0004] Provided herein is an engineered gene effector comprising a variant methylcytosine dioxygenase that is (1) 400 to 700 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain; or (2) 750 to 1 ,000 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain.

[0005] Also provided is an engineered gene effector comprising a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase comprises the amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023 or 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto, optionally wherein the variant methylcytosine dioxygenase is 400 to 700 amino acids in length.

[0006] Provided herein is an engineered gene effector comprising a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase is a variant murine methylcytosine dioxygenase that is 400 to 1700 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain.

[0007] Further provided is an engineered gene effector comprising a variant DNA methyltransferase, wherein the variant DNA methyltransferase is 180 to 360 amino acids in length and comprises a catalytic domain (CD)-like domain (e.g., a C-terminal CD-like domain).

[0008] Also provided is an engineered gene effector comprising a variant DNA methyltransferase, wherein the variant DNA methyltransferase comprises any one of SEQ ID NOs: 23-29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto, optionally wherein the variant DNA methyltransferase is 180 to 360 amino acids in length.

[0009] Provided herein is a fusion protein comprising: any one of the engineered gene effectors of the present disclosure; and a heterologous endonuclease (e.g., a Cas protein). Also provided is a fusion protein comprising: any one of the engineered gene effectors of the present disclosure; a heterologous polypeptide; and a heterologous endonuclease, wherein the variant methylcytosine dioxygenase is fused to a N-terminus of the heterologous endonuclease, and wherein the heterologous polypeptide is fused to a C-terminus of the heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein or a nucleasedeactivated variant thereof, optionally wherein the heterologous polypeptide is a transcriptional activator.

[0010] Also provided is a fusion protein comprising an amino acid sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to, or is about 100% identical to the amino acid sequence of any one of SEQ ID NOs: 432-438, 783, 795-803, 813- 865, 919-924, 967-993 or 994.

[0011] Further provided are polynucleotides encoding any one of the engineered gene effectors or fusion proteins of the present disclosure, vectors that include any one of the polynucleotides, and cells that include any one of the polynucleotides or vectors of the present disclosure.

[0012] Also provided is a system that includes: any one of the engineered gene effectors of the present disclosure; a heterologous endonuclease coupled to the polypeptide of the engineered gene effector, optionally wherein the heterologous endonuclease is a Cas protein; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein. Further provided is a system that includes: any one of the engineered gene effectors of the present disclosure; a heterologous polypeptide; a heterologous endonuclease, wherein the variant methylcytosine dioxygenase is fused to a N-terminus of the heterologous endonuclease, and wherein the heterologous polypeptide is fused to a C-terminus of the heterologous endonuclease; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein, optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof, optionally wherein the heterologous polypeptide is a transcriptional activator. Also provided is a system that includes any one of the fusion proteins of the present disclosure; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein.

[0013] Provided herein is a polynucleotide or a combination of polynucleotides encoding any one of the systems of the present disclosure, wherein the combination of polynucleotides is configured to express the heterologous endonuclease coupled to the engineered gene effector and the guide nucleic acid in a cell. Also provided is a polynucleotide comprising a nucleotide sequence that is at least 85%, 90%, 95%, 97%, 98%, or 99% identical to the nucleotide sequence of any one of SEQ ID NOs: 442-451, 453-459, 804-812, 86-918, 925-930, or 995-1021, or 1022.

[0014] Also provided is a kit that includes any one of the engineered gene effectors, fusion proteins combinations, systems, polynucleotides, vectors, and / or cells of the present disclosure.

[0015] Provided herein is a method of epigenetic regulation of a target gene in a cell, comprising contacting a cell with any one of the systems or combinations of the present disclosure.

[0016] Also provided is a computer-implemented method of generating a linker sequence, comprising: (a)selecting a plurality of human protein-derived peptide sequences having a length of 25 to 150 amino acids based on a structural confidence metric, wherein at least two of the peptide sequences of the plurality of human protein-derived peptide sequences have different lengths, wherein each peptide sequence of the plurality of human protein- derived peptide sequences is associated with one of a plurality of length categories based on its amino acid sequence length, optionally wherein at least one peptide sequence of the plurality of human protein-derived peptide sequences comprises a nuclear localization signal (NLS); (b)generating a flexibility score for each peptide sequence of the plurality of human protein- derived peptide sequences by sequence modeling, optionally wherein the flexibility score is a normalized B-factor ranking; and (c) for each of the plurality of length categories, selecting one or more candidate linker sequences from the peptide sequences associated with the length category based on the flexibility score of each of the peptide sequences, optionally selecting the one or more candidate linker sequences from the peptide sequences associated with the length category based on one or more functional features associated with each of the peptide sequences, thereby generating one or more candidate human protein-derived linker sequences.

[0017] Further provided is a system comprising a computing device comprising at least one processor and instructions executable by the at least one processor to perform the method (e.g., the computer-implemented method of generating a linker sequence) of the present disclosure. Also provided is a non-transitory computer-readable medium having stored thereon computer-readable instructions that, when executed by a processor, cause the processor to execute the method (e.g., the computer-implemented method of generating a linker sequence) of the present disclosure. Also provided is an engineered human protein-derived linker generated by the method (e.g., the computer-implemented method of generating a linker sequence) of the present disclosure.

[0018] Provided is an engineered human protein-derived linker comprising the amino acid sequence of any one of SEQ ID NO: 463-466, 468-471, 473-476, 478-481, 483-486, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

[0019] Further provided is a polynucleotide comprising: a heterologous polynucleotide sequence comprising (i) a CRISPR target sequence and (ii) a CRISPR protospacer adjacent motif (PAM) sequence and an additional CRISPR PAM sequence that are different, wherein said CRISPR target sequence is flanked by said CRISPR PAM sequence and said additional CRISPR PAM sequence; a promoter region polynucleotide sequence from a target gene; and a reporter gene polynucleotide sequence, optionally wherein the reporter gene is a green fluorescent protein (GFP). Also provided are vectors and cells comprising the polynucleotide of the present disclosure.

[0020] Provided herein is a method of reactivating a target gene in a cell, comprising contacting a cell with any one of the engineered gene effectors (e.g., the engineered gene effector comprising a variant methylcytosine dioxygenase) of the present disclosure, the system that includes the any one of the engineered gene effectors (e.g., the engineered gene effector comprising a variant methylcytosine dioxygenase) of the present disclosure, or the combination of polynucleotides that encodes the system that includes the any one of the engineered gene effectors (e.g., the engineered gene effector comprising a variant methylcytosine dioxygenase) of the present disclosure, wherein the cell comprises a reporter polynucleotide comprising any one of the polynucleotides that includes a heterologous polynucleotide sequence, a promoter region polynucleotide sequence, and a reporter gene polynucleotide sequence.BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Non-limiting aspects of the present disclosure are set forth with particularity in the appended claims. A better understanding of the non-limiting features of the present disclosure will be obtained by reference to the following detailed description that sets forth non-limiting, illustrative embodiments and the accompanying drawings (also “Figure” and “FIG.” herein) of which:

[0022] FIG. 1 shows a TET (Ten-Eleven Translocation) 1 (human) catalytic domain (structure generated by AlphaFold) superimposed with Naegleria Tet-like proteins (ngTETl, 4LT5), revealing that the low-complexity domain (LCD) is largely disoriented.

[0023] FIG. 2 shows a schematic representation of various reporters designed to evaluate the activity of the demethylases.

[0024] FIG. 3 shows reactivation of a silenced EFlα-eGFP synthetic reporter in HEK293T cells via demethylases.

[0025] FIG. 4 shows reactivation of silenced CD81 in HEK293T cells via demethylases.

[0026] FIG. 5 shows a schematic representation of the domain structure of DNMT3A / 3L and the truncation.

[0027] FIGs. 6A-6B show silencing of a EFlα-eGFP synthetic reporter in HEK293T cells via persistent suppressors.

[0028] FIGS. 7A-7F show non-limiting examples of amino acid sequences of ten- eleven translocation (TET) methylcytosine dioxygenases. FIG. 7A shows a non-limiting example of an amino acid sequence of human methylcytosine dioxygenase TET1 (SEQ ID NO: 1). FIG. 7B shows a non-limiting example of an amino acid sequence of human methylcytosine dioxygenase TET2 (SEQ ID NO: 2). FIG. 7C shows a non-limiting example of an amino acid sequence of human methylcytosine dioxygenase TET3 (SEQ ID NO: 3). FIG. 7D shows a non-limiting example of an amino acid sequence of murine methylcytosine dioxygenase TET1 (SEQ ID NO: 4). FIG. 7E shows a non-limiting example of an amino acid sequence of murine methylcytosine dioxygenase TET2 (SEQ ID NO: 5). FIG. 7F shows a non-limiting example of an amino acid sequence of murine methylcytosine dioxygenase TET3 (SEQ ID NO: 6).

[0029] FIG. 8 depicts a computer system that is programmed or otherwise configured to implement methods provided herein.

[0030] FIG. 9 is a schematic diagram showing non-limiting embodiments of a computer-implemented method of the present disclosure.

[0031] FIGs. 10A-10B show expression suppression via persistent suppressors and the compactness of persistent suppressors. FIG. 10A shows suppression of reporter GFP expression in HEK293T cells via persistent suppressors. FIG. 10B shows a schematic representation of the persistent suppressor DNA construct size.

[0032] FIG. 11 shows sustained reactivation (up to 100 days) of silenced CD81 in HEK293T cells via demethylases.

[0033] FIGs. 12A-12E show durable and synergistic gene reactivation via fusion protein variants including a nuclease-deactivated Cas protein (dCas), an epigenetic modulator (EM), e.g., a demethylase, and a transcriptional activator. FIG. 12A shows a schematic representation of the recruitment of a fusion protein including a nuclease dead Cas protein fused with an epigenetic modulator (EM) and a transcriptional activator to a target gene. FIG. 12B depicts a predicted target gene expression outcome for reactivating epigenetically silenced target genes via fusion protein variants. FIG. 12C shows a schematic representation of a synthetic reporter designed and genetically integrated into HEK293T cells to evaluate the durable and synergistic gene reactivation potential. Vertical lines indicate individual CpG residue within the promoter fragment FIG. 12D shows reactivation of a silenced EFlα-eGFP synthetic reporter in HEK293T cells via fusion protein variants. FIG. 12E shows a schematic representation of the fusion protein variant DNA construct size.

[0034] FIG. 12F shows an embodiment of a nucleotide sequence of a fusion protein of the present disclosure.

[0035] FIG. 12G shows an embodiment of a nucleotide sequence of a transcriptional activator of the present disclosure.

[0036] FIGs. 13A-13B show schematic representations of two reporter systems utilized to evaluate DNA demethylation / gene reactivation. FIG. 13A shows a schematic representation of a synthetic reporter system where eGFP expression was silenced by DNA methylases. FIG. 13B shows a schematic representation of an endogenous gene reporter where CD81 expression was silenced by DNA methylases.

[0037] FIG. 14 shows the long-term gene reactivation activity of demethylase variants vl, v2, v3, v4, v5, v6, v8, v9, vlO, and vl 1 fused with dCas9.

[0038] FIG. 15 shows the gene reactivation activity of the additional demethylase variants vl2, vl3, vl4, and vl5 fused with dCas9.

[0039] FIG. 16 shows synergistic reactivation activity of dCas9 - demethylase v6 - transcriptional activator (dCas9-EDV6-EBAct) fusion protein variants at the EFlα-eGFP reporter gene in HEK293T cells.

[0040] FIG. 17 shows synergistic reactivation activity of dCas9-EDV6-EBAct fusion protein variants compared with the synergistic reactivation activity of dCas9 -demethylase vl3 - transcriptional activator (dCas9-EDV13-EBAct) fusion protein variants at the EFlα-eGFP reporter gene in HEK293T cells.

[0041] FIG. 18 shows gene reactivation activity by dCas-cA2-demethylase v6 (dCas-cA2-EDV6) fusion protein variants with different human protein-derived linker sequences at the silenced EFlα-eGFP synthetic reporter gene in HEK293T cells.

[0042] FIG. 19 shows gene reactivation activity by dCas-cA2-EDV6 fusion protein variants compared with the gene reactivation activity by dCas-cA2-demethylase vl3 (dCas-cA2-EDV13) fusion protein variants at the silenced EFlα-eGFP synthetic reporter gene in HEK293T cells.

[0043] FIG.20 shows synergistic activity of dCas-cA2-demethylase-transciptional activator (dCas-cA2-EDV-EBAct) fusion protein variants in reactivating the silenced EFlα- eGFP synthetic reporter gene in HEK293T cells.

[0044] FIG. 21 shows the identification of additional truncation regions for demethylase variants v6 (EDV6) and vl3 (EDV13) via semi-rational design (structure generated by AlphaFold).

[0045] FIG. 22 shows gene reactivation activity by further truncated demethylase (dCas-cA2-EDV_del) fusion protein variants at the silenced EFlα-eGFP synthetic reporter gene in HEK293T cells.

[0046] FIG. 23 is a graph showing percentage methylation of CpG sites in around the TSS of CD81 in silenced cells treated with dCas9-EDV6.DETAILED DESCRIPTIONOVERVIEW

[0047] Provided herein are engineered gene effectors that include an epigenetic effector or editor, for example, a DNA methylase or methylcytosine dioxygenase, that have been engineered to be smaller. In some embodiments, provided herein are engineered ultracompact epigenetic editors for DNA and histone modifications that enable durable epigenetic gene activation and suppression.

[0048] The discovery and characterization of epigenetic editors are accelerating the development of novel therapeutic approaches for targeting human genetic diseases. These advances are exemplified by the significant and lasting silencing of PCSK9, both in laboratory settings and living organisms. CRISPR-mediated epigenetic editing is a promising newtechnology with the potential to greatly enhance the growing arsenal of gene therapy modalities. The advantages are highlighted by its ability to achieve versatile gene expression modulation that could durably last through cell divisions without editing the DNA sequences. However, two obstacles hinder the full realization of this potential: firstly, current epigenetic editors for targeted DNA methylation and demethylation may be too large for effective tissuespecific delivery in vivo; secondly, the modulation of gene expression by DNA methyltransferases and demethylases (e.g., methylcytosine dioxygenases) may not universally apply to all genes across different chromatin states or levels of gene expression.

[0049] To overcome these challenges, in some embodiments, ultracompact yet highly effective DNA methyltransferase L (DNMT3L) domains were developed, along with compact and efficient ten eleven translocation (TET) protein catalytic domains, as well as compact epigenetic gene activators. Combined with compact (<500aa) dCas proteins, this new toolkit offers a versatile and potent platform for modulating gene expression. In some embodiments, DNMT3L variants that are nearly 50% more compact than the full-length (e.g., native, or wild-type) protein were engineered.

[0050] Similarly, in some embodiments, smaller variants of TET catalytic domains (human TET1 and TET3, mouse TET1 and TET3) that are at least 33% smaller than the typical functional domains used for DNA demethylation were developed. These domains were functional in human cells and at endogenous human genes. Additionally, transcriptional activation domains that interact with CBP / p300 were developed, leading to long-lasting activation of target genes in human cells in vitro and in vivo. To address the risk of inconsistent modulation of gene expression across different chromatin states by a single editing mechanism (e.g., DNA or histone modification) alone, the synergistic relationship of transcriptional regulation in eukaryotes, where the transcriptional output driven by two or more mechanisms is often greater than the sum of the output driven by each mechanism individually, was utilized to design engineered gene effector fusion proteins. Additionally, in some embodiments, combining a demethylase with transcriptional activator domains (such as, without limitation, XV1.48 (SEQ ID NO: 20), XV1.2 (SEQ ID NO: 21), VPR (SEQ ID NO: 780), Leutx-Tox4 (SEQ ID NO: 781), Leutx-ZFX, or hvTR-Q2-ZNF) achieves a synergistic effect that further enhances the potency of epigenetic activation. In some embodiments, combining the engineered demethylases with the ultracompact transcriptional activator domains, upontransient delivery in human cells, both synergistic and durable reactivation of an epigenetically silenced loci was achieved, where the fold activation from such combination was greater than the predicted additive effect by a factor of two or more.

[0051] Engineered compact epigenetic gene effectors can provide a versatile and potent platform for gene expression modulation, and can be utilized in epigenetic editing therapeutic payloads. In some embodiments, the size of these compact epigenetic activators or suppressors (e.g., demethylases and DNA methyltransferases) when combined with deactivated Cas proteins as CRISPR-mediated epigenetic editors can enable single AAV packaging for in vivo epigenome engineering.TERMS

[0052] All terms can have their customary and ordinary meaning to one of ordinary skill in the art, in view of the present disclosure, unless otherwise indicated.

[0053] “Methylation” as used herein has its customary and ordinary meaning as understood by one of ordinary skill in the art in view of the present Application, denotes a base modification of a nucleic acid residue that involves addition of a methyl group to the N5 position of cytosine. A methylated nucleic acid can include one or more 5-methylcytosine (5mC) residues in place of a corresponding cytosine residue. A “CpG island” as used herein has its customary and ordinary meaning as understood by one of ordinary skill in the art in view of the present Application, and denotes a region of a nucleic acid (e.g., a genomic DNA, a reported construct) of >~500 bp, having a GC content of >50-55%, and an expected / observed CpG ratio of >0.6-0.65.

[0054] The term “heterologous,” when used herein with reference to a polypeptide sequence or a nucleic acid sequence, indicates that the polypeptide sequence or the nucleic acid sequence is (1) disposed (e.g., in an environment, such as a cell, a virus, or a fusion polypeptide molecule or a fusion polynucleotide molecule) where it is not normally found (e.g., not normally found in nature); or (2) comprises two or more subsequences that are not found in the same relationship to each other as normally found in nature. For example, a polypeptide can comprise a first polypeptide sequence and a second polypeptide sequence that are not found together in a single polypeptide in nature, and thus the first polypeptide sequence and the second polypeptide sequence can be heterologous to each other. In another example, apolynucleotide can comprise a first polynucleotide sequence and a second polynucleotide sequence that are not found together in a single polynucleotide in nature, and thus the first polynucleotide sequence and the second polynucleotide sequence can be heterologous to each other.

[0055] The term “cell” has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a biological cell. A cell can be the basic structural, functional and / or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include: a prokaryotic cell, eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a protozoa cell, a cell from a plant (e.g. cells from plant crops, fruits, vegetables, grains, soy bean, com, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, fems, clubmosses, homworts, liverworts, mosses), an algal cell, (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, C.Agardh, and the like), seaweeds (e g. kelp), a fungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate animal (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal (e.g., a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, a human, etc.), or etcetera. Sometimes a cell is not originating from a natural organism (e.g. a cell can be a synthetically made, sometimes termed an artificial cell).

[0056] The term “nucleotide,” as used herein has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a base-sugar-phosphate combination. A nucleotide can comprise a synthetic nucleotide. A nucleotide can comprise a synthetic nucleotide analog. Nucleotides can be monomeric units of a nucleic acid sequence (e.g. deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, diTP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives can include, for example, [aS]dATP, 7- deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance onthe nucleic acid molecule containing them. The term nucleotide as used herein can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. A nucleotide may be unlabeled or delectably labeled by well- known techniques. Labeling can also be carried out with quantum dots. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels and enzyme labels. Fluorescent labels of nucleotides may include but are not limited fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6- carboxyfiuorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6- carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4- (4'dimethylaminophenylazo) benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyanine and 5-(2'-aminoethyl)aminonaphthalene-l -sulfonic acid (EDANS). Specific examples of fluorescently labeled nucleotides can include [R6G]dUTP, [TAMRA]dUTP, [RllOJdCTP, [R6G] dCTP, [TAMRA] dCTP, [JOE] ddATP, [R6G] ddATP, [FAM] ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROXJddTTP, [dR6G]ddATP, [dRUO]ddCTP, [dTAMRA]ddGTP, and [dROXJddTTP available from Perkin Elmer, Foster City, Calif. FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein- 15-d ATP, Fluorescein- 12-dUTP, Tetramethyl-rodamine- 6-dUTP, IR770-9-dATP, Fluorescein- 12-ddUTP, Fl uorescein- 12-UTP, and Fluorescein- 15-2'- dATP available from Boehringer Mannheim, Indianapolis, Ind.; and Chromosome Labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, B0DIPY-TMR-14-UTP, BODIPY- TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein- 12-dUTP, Oregon Green 488-5- dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12- dUTP available from Molecular Probes, Eugene, Oreg. Nucleotides can also be labeled or marked by chemical modification. A chemically-modified single nucleotide can be biotin- dNTP. Some non-limiting examples of biotinylated dNTPs can include, biotin-dATP (e.g., bio- N6-ddATP, biotin- 14-dATP), biotin-dCTP (e.g., biotin- 11 -dCTP, biotin- 14-dCTP), and biotin-dUTP (e.g. biotin- 11-dUTP, biotin- 16-dUTP, biotin-20-dUTP).

[0057] The term “polynucleotide,” “oligonucleotide,” or “nucleic acid,” as used interchangeably herein each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multi-stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three dimensional structure, and can perform any function, known or unknown. A polynucleotide can comprise one or more analogs (e.g. altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer. Some nonlimiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, florophores (e.g. rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queuosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides can be interrupted by non-nucleotide components.

[0058] The term “sequence identity” generally refers to an exact nucleotide-to- nucleotide or amino acid-to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Typically, techniques for determining sequence identity include determining the nucleotide sequence of a polynucleotide and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Two or more sequences (polynucleotide or amino acid) can be compared bydetermining their “percent identity." The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the longer sequence and multiplied by 100. Percent identity may also be determined, for example, by comparing sequence information using the advanced BLAST computer program, including version 2.2.9, available from the National Institutes of Health. The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87:2264-2268 (1990) and as discussed in Altschul, et al., J. Mol. Biol., 215:403-410 (1990); Karlin And Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993); and Altschul et al, Nucleic Acids Res., 25:3389-3402 (1997). The program may be used to determine percent identity over the entire length of the proteins being compared. Default parameters are provided to optimize searches with short query sequences in, for example, with the blastp program. The program also allows use of an SEG filter to mask-off segments of the query sequences as determined by the SEG program of Wootton andFederhen, Computers and Chemistry 17:149-163 (1993). Ranges of desired degrees of sequence identity are approximately 50% to 100% and integer values therebetween. In general, this disclosure encompasses sequences with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity with any sequence provided herein.

[0059] The term “gene” has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a nucleic acid (e.g., DNA such as genomic DNA and cDNA) and its corresponding nucleotide sequence that is involved in encoding an RNA transcript. The term as used herein with reference to genomic DNA includes intervening, non-coding regions as well as regulatory regions and can include 5' and 3' ends. In some uses, the term encompasses the transcribed sequences, including 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons and introns. In some genes, the transcribed region will contain “open reading frames” that encode polypeptides. In some uses of the term, a “gene" comprises only the coding sequences (e.g., an “open reading frame” or “coding region”) necessary for encoding a polypeptide. In some cases, genes do not encode a polypeptide, for example, ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term “gene” includes not only the transcribed sequences, but in addition, also includes non-transcribed regions including upstream and downstream regulatoryregions, enhancers and promoters. A gene can refer to an “endogenous gene” or a native gene in its natural location in the genome of an organism. A gene can refer to an “exogenous gene” or a non-native gene. A non-native gene can refer to a gene not normally found in the host organism, but which is introduced into the host organism by gene transfer. A non-native gene can also refer to a gene not in its natural location in the genome of an organism. A non-native gene can also refer to a naturally occurring nucleic acid or polypeptide sequence that comprises mutations, insertions and / or deletions (e.g., non-native sequence).

[0060] The term “deletion” generally refers to the removal (or loss) of one or more (or a specified number of) amino acids (e.g., contiguous or non-contiguous amino acids) from a polypeptide sequence, or the removal (or loss) one or more (or a specified number of) nucleic acid bases (e.g., contiguous or non-contiguous nucleic acid bases) from a polynucleotide sequence (e.g., that encodes the polypeptide sequence. The term “internal deletion” generally refers to a deletion that does not include the N- or C-terminus of a polypeptide or the 5' or 3' end of a polynucleotide. A deletion (e.g., an internal deletion) can be identified by comparing to a reference sequence, e.g., by specifying the start and end positions of the deletion relative to the reference sequence. A deletion (e.g., an internal deletion) is different and distinct from a substitution. For example, deletion of at least one amino acid is not followed by an insertion of at least one different amino acid at the same position as the at least one amino acid as compared to a reference polypeptide sequence, such that the size (e.g., a number of the amino acid residue(s)) of a modified (or engineered) polypeptide sequence comprising the deletion of the at least one amino acid is smaller than the reference polypeptide sequence by the size of the at least one amino acid that has been deleted.

[0061] The term “substitution” generally refers to the removal (or loss) of one or more (or a specified number of) amino acids (e.g., contiguous or non-contiguous amino acids) from a polypeptide sequence followed by an insertion of at least one different amino acid at the same position as the at least one amino acid as compared to a reference polypeptide sequence, such that the size (e.g., a number of the amino acid residue(s)) of a modified (or engineered) polypeptide sequence comprising the substitution of the at least one amino acid is the same as the reference polypeptide sequence; or the removal (or loss) of one or more (or a specified number of) nucleic acid bases (e.g., contiguous or non-contiguous nucleic acid bases) from a polynucleotide sequence followed by an insertion of at least one different nucleic acidbase at the same position as the at least one nucleic acid base as compared to a reference polynucleotide sequence, such that the size (e.g., a number of the nucleic acid base(s)) of a modified (or engineered) polynucleotide sequence comprising the substitution of the at least one nucleic acid base is the same as the reference polynucleotide sequence. Substitutions can be conservative, non-conservative, or a combination thereof. The term “conservative amino acid substitution" generally refers to a substitution of one amino acid for another amino acid of similar biochemical properties (e.g., charge, size, and / or hydrophobicity). The term “nonconservative amino acid substitution” generally refers to a substitution of one amino acid for another amino acid with different biochemical properties (e.g., charge, size, and / or hydrophobicity). A conservative amino acid change can be, for example, a substitution that has minimal effect on the secondary or tertiary structure of a polypeptide. In some embodiments, conservative substitutions are shown in Table 0.1 under the heading of “preferred substitutions.” More substantial changes are provided in Table 0.1 under the heading of “exemplary substitutions,” and as further described below in reference to amino acid side chain classes.TABLE 0.1%Amino acids may be grouped into the following numbered classes according to common sidechain properties:(1) hydrophobic: Norleucine, Met, Ala, Vai, Leu, Ile;(2) neutral hydrophilic: Cys, Ser, Thr, Asn, Gin;(3) acidic: Asp, Glu;(4) basic: His, Lys, Arg;(5) residues that influence chain orientation: Gly, Pro;(6) aromatic: Trp, Tyr, Phe.In some embodiments, a conservative substitution includes a substitution among amino acids within one of the classes. Non-conservative substitutions will entail exchanging a member of one of these classes for a member of another class.

[0062] The term “expression” has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to one or more processes by which a polynucleotide is transcribed from a DNA template (such as into an mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene product.” If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell. “Up-regulated,” with reference to expression, generally refers to an increased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression level in a wild-type state while “down-regulated” generally refers to a decreased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequencerelative to its expression in a wild-type state. Expression of a transfected gene can occur transiently or stably in a cell. During “transient expression” the transfected gene is not transferred to the daughter cell during cell division. Since its expression is restricted to the transfected cell, expression of the gene is lost over time. In contrast, stable expression of a transfected gene can occur when the gene is co-transfected with another gene that confers a selection advantage to the transfected cell. Such a selection advantage may be a resistance towards a certain toxin that is presented to the cell.

[0063] The term “expression profile” has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to quantitative (e.g., abundance) and qualitative expression of one or more genes in a sample (e.g., a cell). The one or more genes can be expressed and ascertained in the form of a nucleic acid molecule (e.g., an rriRNA or other RNA transcript). Alternatively or in addition to, the one or more genes can be expressed and ascertained in the form of a polypeptide (e.g., a protein measured via Western blot). An expression profile of a gene may be defined as a shape of an expression level of the gene over a time period (e.g., at least or up to about 1 hour, at least or up to about 2 hours, at least or up to about 3 hours, at least or up to about 4 hours, at least or up to about 5 hours, at least or up to about 6 hours, at least or up to about 7 hours, at least or up to about 8 hours, at least or up to about 9 hours, at least or up to about 10 hours, at least or up to about 11 hours, at least or up to about 12 hours, at least or up to about 16 hours, at least or up to about 18 hours, at least or up to about 24 hours, at least or up to about 36 hours, at least or up to about 48 hours, at least up to about 3 days, at least up to about 4 days, at least up to about 5 days, at least up to about 6 days, at least up to about 7 days, at least up to about 8 days, at least up to about 9 days, at least up to about 10 days, at least up to about 11 days, at least up to about 12 days, at least up to about 13 days, at least up to about 14 days, etc.). Alternatively, an expression profile of a gene may be defined as an expression level of the gene at a time point of interest (e.g. , the expression level of the gene measured at least or up to about 1 hour, at least or up to about 2 hours, at least or up to about 3 hours, at least or up to about 4 hours, at least or up to about 5 hours, at least or up to about 6 hours, at least or up to about 7 hours, at least or up to about 8 hours, at least or up to about 9 hours, at least or up to about 10 hours, at least or up to about 11 hours, at least or up to about 12 hours, at least or up to about 16 hours, at least or up to about 18 hours, at least or up to about 24 hours, at least or up to about36 hours, at least or up to about 48 hours, at least up to about 3 days, at least up to about 4 days, at least up to about 5 days, at least up to about 6 days, at least up to about 7 days, at least up to about 8 days, at least up to about 9 days, at least up to about 10 days, at least up to about 11 days, at least up to about 12 days, at least up to about 13 days, or at least up to about 14 days after treating a cell to induce such expression level.)

[0064] The term “peptide,” “polypeptide," or “protein,” as used interchangeably herein, each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer can be interrupted by non-amino acids. The terms include amino acid chains of any length, including full length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogues. Modified amino acids can include natural amino acids and non-natural amino acids, which have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. Amino acid analogues can refer to amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0065] The term “derivative,” “variant,” or “fragment,” as used herein with reference to a polypeptide, has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a polypeptide related to a wild type or naturally occurring polypeptide, for example either by amino acid sequence, structure (e.g., secondary and / or tertiary), activity (e.g., enzymatic activity) and / or function. Derivatives, variants and fragments of a polypeptide can comprise one or moreamino acid variations (e.g., mutations, insertions, and deletions), truncations, modifications, or combinations thereof compared to a wild type polypeptide.

[0066] The term “engineered,” “chimeric,” or “recombinant,” as used herein with respect to a polypeptide molecule (e.g., a protein), has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a polypeptide molecule having a heterologous amino acid sequence or an altered amino acid sequence as a result of the application of genetic engineering techniques to nucleic acids which encode the polypeptide molecule, as well as cells or organisms which express the polypeptide molecule. The term “engineered” or “recombinant,” as used herein with respect to a polynucleotide molecule (e.g., a DNA or RNA molecule), generally refers to a polynucleotide molecule having a heterologous nucleic acid sequence or an altered nucleic acid sequence as a result of the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning technologies; transfection, transformation and other gene transfer technologies; homologous recombination; site-directed mutagenesis; and gene fusion. In some cases, an engineered or recombinant polynucleotide (e.g., a genomic DNA sequence) can be modified or altered by a gene editing moiety. For example, an heterologous endonuclease (e.g., an engineered Cas protein) as disclosed herein is not a naturally occurring nuclease (e.g., not a naturally occurring Cas protein). In another example, an engineered gene effector as disclosed herein is not a naturally occurring gene effector.

[0067] For example, an engineered nuclease (e.g., an engineered Cas protein) as disclosed herein is not a naturally occurring nuclease (e.g., not a naturally occurring Cas protein). The terms “engineered nuclease” and “engineered nuclease variant” may be used interchangeable herein.

[0068] The terms “engineered” and “modified” are used interchangeably herein.The terms “engineering” and “modifying” are used interchangeably herein. The terms “engineered cell” or “modified cell” are used interchangeably herein. The terms “engineered characteristic” and “modified characteristic” are used interchangeably herein.

[0069] The term “enhanced expression,” “increased expression,” or “upregulated expression” each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to production of a moietyof interest (e.g., a polynucleotide or a polypeptide) to a level that is above a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression can be substantially zero (or null) or higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. The moiety of interest can comprise a heterologous gene or polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced expression of the polypeptide of interest in the host strain.

[0070] TThhee tteerrmm “enhanced activity,” “increased activity,” or “upregulated activity” each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is above a normal level of activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity can be substantially zero (or null) or higher than zero. The moiety of interest can comprise a polypeptide construct of the host strain. The moiety of interest can comprise a heterologous polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced activity of the polypeptide of interest in the host strain.

[0071] The term “reduced expression,” “decreased expression,” or “downregulated expression” each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to a production of a moiety of interest (e.g., a polynucleotide or a polypeptide) to a level that is below a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked- out or knocked-down in the host strain. In some examples, reduced expression of the moiety of interest can include a complete inhibition of such expression in the host strain.

[0072] The term “reduced activity,” “decreased activity,” or “downregulated activity” each has its ordinary and customary meaning as understood by one of ordinary skill in the art in view of the present disclosure, and generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is below a normal levelof activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked- out or knocked-down in the host strain. In some examples, reduced activity of the moiety of interest can include a complete inhibition of such activity in the host strain.

[0073] The term “subject,” “individual,” or “patient,” as used interchangeably herein, generally refers to a vertebrate, preferably a mammal such as a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.

[0074] The term “treatment” or “treating” generally refers to an approach for obtaining beneficial or desired results including but not limited to a therapeutic benefit and / or a prophylactic benefit For example, a treatment can comprise administering a system or cell population disclosed herein. By therapeutic benefit is meant any therapeutically relevant improvement in or effect on one or more diseases, conditions, or symptoms under treatment. For prophylactic benefit, a composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more of the physiological symptoms of a disease, even though the disease, condition, or symptom may not have yet been manifested.

[0075] The term “effective amount” or “therapeutically effective amount” generally refers to the quantity of a composition, for example a composition comprising heterologous polypeptides, heterologous polynucleotides, and / or modified cells (e.g., modified stem cells), that is sufficient to result in a desired activity upon administration to a subject in need thereof. Within the context of the present disclosure, the term “therapeutically effective” generally refers to that quantity of a composition that is sufficient to delay the manifestation, arrest the progression, relieve or alleviate at least one symptom of a disorder treated by the methods of the present disclosure.

[0076] Whenever the term “at least,” “greater than," or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values inthat series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0077] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0078] The term “about” or “approximately” generally mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0079] The use of the alternative (e.g., “or”) should be understood to mean either one, both, or any combination thereof of the alternatives. The term “and / or” should be understood to mean either one, or both of the alternatives.ENGINEERED GENE EFFECTORS

[0080] In some aspects, the present disclosure provides compositions, combinations, systems, and methods that utilize engineered gene effectors (e.g., engineered epigenetic effectors). Engineered gene effectors of the present disclosure can include an engineered epigenetic gene effector (e.g., a DNA methylase or a DNA demethylase). In some embodiments, the engineered gene effectors includes a variant methylcytosine dioxygenase or a variant DNA methyltransferase, as described herein. Variant Methvlcvtosine Dioxygenases

[0081] An engineered gene effector that includes a variant methylcytosine dioxygenase that is 400 to 1,200 amino acids in length and includes a methylcytosine dioxygenase catalytic domain is provided. Also provided are an engineered gene effector thatincludes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length and includes a methylcytosine dioxygenase catalytic domain. Further provided is an engineered gene effector that includes a variant methylcytosine dioxygenase that is 750 to 1,000 amino acids in length and includes a methylcytosine dioxygenase catalytic domain. In general, the variant methylcytosine dioxygenase is a functional variant of a native methylcytosine dioxygenase, as described herein. As such, the variant methylcytosine dioxygenase generally includes a methylcytosine dioxygenase catalytic domain (e.g., the C-terminal catalytic domain of a native methylcytosine dioxygenase, or a functional portion thereof, for example, a portion without a low-complexity domain (LCD), as described herein). In some embodiments, the variant methylcytosine dioxygenase is no more than 1,200 amino acids in length. In some embodiments, the variant methylcytosine dioxygenase is, is at most, or is about 400, 450, 500, 550, 600, 700, 750, 800, 900, 1,000, 1,100, or 1,200 amino acids in length, or optionally, it has a length in a range defined by any two of the preceding values (e.g., 400-1,200 amino acids, 450-1,000 amino acids, 800-1,200 amino acids, 400-700 amino acids, 750-1,000 amino acids, etc.). In some embodiments, the variant methylcytosine dioxygenase is no more than 700 amino acids in length. In some embodiments, the variant methylcytosine dioxygenase is, is at most, or is about 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700 amino acids in length, or optionally, it has a length in a range defined by any two of the preceding values (e.g. , 400-700 amino acids, 450-680 amino acids, 480-600 amino acids, 450-550 amino acids, etc.). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400, 450, 500, 550, 600, 650 or 700 amino acids in length and includes a methylcytosine dioxygenase catalytic domain, or optionally the variant methylcytosine dioxygenase is a length within a range defined by any two of the aforementioned lengths (e.g., 450-600, 500-700, 500-600, 400-650, etc.). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is no more than 500 amino acids in length and includes a methylcytosine dioxygenase catalytic domain. In some embodiments, the variant methylcytosine dioxygenase is no more than 1,000 amino acids in length. In some embodiments, the variant methylcytosine dioxygenase is, is at most, or is about 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1,000 amino acids in length, or optionally, it has a length in a range defined byany two of the preceding values (e.g., 750-1,000 amino acids, 750-980 amino acids, 780-900 amino acids, 750-850 amino acids, 800-900 amino acids, etc.). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 750, 800, 850, 900, 950 or 1,000 amino acids in length and includes a methylcytosine dioxygenase catalytic domain, or optionally the variant methylcytosine dioxygenase is a length within a range defined by any two of the aforementioned lengths (e.g., 750-1,000, 750-900, 800-1000, 800- 900, 700-950, etc.). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is no more than 900 amino acids in length and includes a methylcytosine dioxygenase catalytic domain.

[0082] In some embodiments, “demethylase” denotes an enzyme that can mediate demethylation of a methylated gene or a methylated locus (e.g., a methylated genomic locus, a methylated CpG island, etc.). In some embodiments, the demethylase is a methylcytosine dioxygenase, including, without limitation, any one of the variant methylcytosine dioxygenases described herein.

[0083] In some embodiments, the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase (e.g., relative to the sequence of the native methylcytosine dioxygenase. In some embodiments, the native methylcytosine dioxygenase is a ten-eleven translocation (TET) methylcytosine dioxygenase. As used herein, a “native” protein denotes a protein having the structural features (e.g., amino acid sequence) of a naturally occurring counterpart. In some embodiments, a TET methylcytosine dioxygenase includes a C-terminal catalytic domain. In some embodiments, the C-terminal catalytic domain includes one or more of a double-stranded 0-helix (DSBH) domain, a cysteine-rich domain, and cofactor binding sites. In some embodiments, the C- terminal catalytic domain includes a low-complexity domain (LCD). In some embodiments, the LCD is in the DSBH domain. In some embodiments, a TET methylcytosine dioxygenase includes an N-terminal CXXC zinc finger domain. The native methylcytosine dioxygenase can be any suitable TET methylcytosine dioxygenase. In some embodiments, the native methylcytosine dioxygenase is a human, non-human primate, mouse, rat, rabbit, cat, dog, horse, cow, or pig TET methylcytosine dioxygenase. In some embodiments, the native methylcytosine dioxygenase is a human TET methylcytosine dioxygenase. In some embodiments, the native methylcytosine dioxygenase includes: human TET1; human TET2;or human TET3. In some embodiments, human TET1 is associated with Gene ID 80312. In some embodiments, human TET1 has or includes the amino acid sequence set forth in FIG. 7 A. In some embodiments, human TET2 is associated with Gene ID 54790. In some embodiments, human TET2 has or includes the amino acid sequence set forth in FIG. 7B. In some embodiments, human TET3 is associated with Gene ID 200424. In some embodiments, human TET3 has or includes the amino acid sequence set forth in FIG. 7C. In some embodiments, the native methylcytosine dioxygenase is a murine TET methylcytosine dioxygenase. In some embodiments, the native methylcytosine dioxygenase includes: murine TET1; murine TET2; or murine TET3. In some embodiments, murine TET1 is associated with Gene ID 52463. In some embodiments, murine TET1 has or includes the amino acid sequence set forth in FIG. 7D. In some embodiments, murine TET2 is associated with Gene ID 214133. In some embodiments, murine TET2 has or includes the amino acid sequence set forth in FIG. 7E. In some embodiments, murine TET3 is associated with Gene ID 194388. In some embodiments, murine TET3 has or includes the amino acid sequence set forth in FIG. 7F.

[0084] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 1-6. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 2. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 3. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 4. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 5. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 6.

[0085] The one or more mutations relative to a native methylcytosine dioxygenase can be a substitution, deletion, insertion, etc., of one or more amino acids in the amino acid sequence of the native methylcytosine dioxygenase. In some embodiments, the one or more mutations does not include addition of amino acids at the N-terminal or C-terminal end of the methylcytosine dioxygenase. In some embodiments, the one or more mutations includes a deletion of one or more amino acid residues relative to the native methylcytosine dioxygenase. In some embodiments, the one or more mutations includes a deletion outside a catalytic domain(e.g., C-terminal catalytic domain) of the native methylcytosine dioxygenase. As used herein, “outside" relative to a domain of an amino acid sequence denotes a portion of the amino acid sequence that is either more N-terminal relative to the N -terminus of the domain or more C- terminal relative to the C-terminus of the domain. In some embodiments, the one or more mutations includes a deletion of an N-terminal portion of the native methylcytosine dioxygenase (e.g., while retaining the C-terminal catalytic domain or a functional portion thereof). In some embodiments, the one or more mutations includes a deletion of an N-terminal portion that includes a CXXC zinc finger domain of the native methylcytosine dioxygenase (e.g., while retaining the C-terminal catalytic domain or a functional portion thereof). In some embodiments, the one or more mutations includes a deletion of at least a portion of the LCD.

[0086] In some embodiments, the one or more mutations includes a deletion of one or more sets of consecutive amino acid residues of the native methylcytosine dioxygenase. For example, the variant methylcytosine dioxygenase can include a deletion relative to the native methylcytosine dioxygenase of a set of consecutive amino acid residues from the N-terminal portion of the native methylcytosine dioxygenase and a deletion of another set of amino acid residues from the C-terminal portion of the native methylcytosine dioxygenase. In some embodiments, the one or more mutations includes a deletion of an N-terminal portion that includes a CXXC zinc finger domain of the native methylcytosine dioxygenase, and a deletion of at least a portion of the LCD of the C-terminal catalytic domain of the native methylcytosine dioxygenase. In some embodiments, the one or more mutations includes a deletion of an N- terminal portion of the native methylcytosine dioxygenase, and a deletion of at least a portion of the LCD of the C-terminal catalytic domain of the native methylcytosine dioxygenase.

[0087] In some embodiments, the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion of the low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase. In some embodiments, the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low- complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase.

[0088] In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800 amino acid residues, or more, relative to the native methylcytosine dioxygenase. In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the native methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 400-1,800 amino acid residues, 450-1,700 amino acid residues, 1,000-1,400 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 1,450, 1,500, 1,550, 1,600, 1650, 1,700, or 1,750 amino acid residues, or more, relative to a human TET1 methylcytosine dioxygenase (e.g., as set forth in FIG. 7 A). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the human TET1 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 1450-1,750 amino acid residues, 1,500-1,700 amino acid residues, 1,550-1,700 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of least 1,050, 1,100, 1,150, 1,200, 1,250, 1,300, 1,350, 1,400, 1,450, 1,500, 1,550, 1,600, 1,650, 1,700, 1,750, or 1,800 amino acid residues relative to a human TET2 methylcytosine dioxygenase (e.g., as set forth in FIG. 7B). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the human TET2 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 1,050-1,800 amino acid residues, 1,200-1,800 amino acid residues, 1,100-1,500 amino acid residues, 1,400-1,600 amino acid residues, 1,500-1,700 amino acid residues, 1,100- 1,300 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 700, 750, 800, 850, 900, 950, 1000, 1,100, 1,150, 1,200, 1,250, 1,300, 1,350, or 1,400 amino acid residues, or more, relative to a human TET3 methylcytosine dioxygenase (e.g., as set forth in FIG. 7C). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the humanTET3 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 700-1,400 amino acid residues, 800-1,350 amino acid residues, 1,200-1,350 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 1,200, 1,250, 1,300, 1,350, 1,400, 1,450, 1,500, 1,550, or 1,600 amino acid residues, or more, relative to a murine TET1 methylcytosine dioxygenase (e.g., as set forth inFIG. 7D). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the murine TET1 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 1,200-1 ,600 amino acid residues, 1,250-1 ,550 amino acid residues, 1,300-1,550 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 950, 1,000, 1,050, 1,100, 1,150, 1,200, 1,250, 1,300, 1,350, 1,400, 1,450, 1,500, 1,550, 1,600, 1,650, or 1,700 amino acid residues relative to a murine TET2 methylcytosine dioxygenase (e.g., as set forth in FIG. 7E). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the murine TET2 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 950-1,700 amino acid residues, 1,200-1,700 amino acid residues, 1,300-1,600 amino acid residues, 1,000-1,300 amino acid residues, 1,400-1,550 amino acid residues, 1,000-1,150 amino acid residues, etc.). In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 300, 350, 400, 450, 500, 550, 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1,100, 1,150, 1,200, or 1,250 amino acid residues, or more, relative to a murine TET3 methylcytosine dioxygenase (e.g., as set forth in FIG. 7F). In some embodiments, the one or more mutations includes a deletion of a number of amino acid residues relative to the murine TET3 methylcytosine dioxygenase in a range defined by any two of the preceding values (e.g., 300-1,200 amino acid residues, 350-600 amino acid residues, 1,000-1,250 amino acid residues, 1,100-1,200 amino acid residues, etc.).

[0089] In some embodiments, the one or more mutations includes a deletion of, of about, or of at least 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80% of the amino acid residues, or more, relative to the native methylcytosine dioxygenase sequence. In some embodiments, the one or more mutations includes a deletion of a percentage of amino acid residues in a range defined by any two of the preceding values (e.g., 10-80%, 20-75%, 30- 75%, etc.) relative to the native methylcytosine dioxygenase sequence.

[0090] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-500 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or mmoorree mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET1 (e.g., as set forth in FIG. 7 A). In some embodiments, the engineered gene effectorincludes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-500 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET1 (e.g., as set forth in FIG. 7 A). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-500 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET1 (e.g., as set forth in FIG. 7 A). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of human TET1).

[0091] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g. , 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET2 (e.g. , as set forth in FIG. 7B). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET2 (e.g., as set forth in FIG. 7B). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosinedioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET2 (e.g., as set forth in FIG. 7B). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of human TET2).

[0092] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 1,000 amino acids in length and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET3 (e.g., as set forth in FIG. 7C). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET3 (e.g. , as set forth in FIG. 7C). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET3 (e.g., as set forth in FIG. 7C). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, andwherein the native methylcytosine dioxygenase includes human TET3 (e.g., as set forth in FIG. 7C). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of human TET3).

[0093] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-700 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET1 (e.g., as set forth in FIG. 7D). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-700 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET1 (e.g., as set forth in FIG. 7D). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-700 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET1 (e.g., as set forth in FIG. 7D). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of murine TET1).

[0094] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET2 (e.g., as set forth in FIG. 7E). In some embodiments, the engineered gene effectorincludes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET2 (e.g., as set forth in FIG. 7E). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 700 amino acids in length (e.g., 450-550 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low-complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET2 (e.g., as set forth in FIG. 7E). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of murine TET2).

[0095] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 1,200 amino acids in length (e.g., 450-700 or 1 ,000- 1,200 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET3 (e.g., as set forth in FIG. 7F). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 1,200 amino acids in length (e.g., 450-700 or 1,000-1,200 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET3 (e.g., as set forth in FIG. 7F). In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 400 to 1,200 amino acids in length (e.g., 450-700 or 1 ,000- 1,200 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain,wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase and a deletion of a low- complexity domain (LCD), or a portion thereof, of the C-terminal catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET3 (e.g., as set forth in FIG. 7F). In some embodiments, the variant methylcytosine dioxygenase retains the LCD of the native methylcytosine dioxygenase (e.g., retains the LCD of murine TET3).

[0096] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the variant methylcytosine dioxygenase comprises SEQ ID NO: 13 or 16, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical thereto. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1, and the variant methylcytosine dioxygenase includes at least amino acid residues 1419-1753 of SEQ ID NO: 1. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1, and the variant methylcytosine dioxygenase includes at least amino acid residues 1991-2136 of SEQ ID NO: 1. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1, and the variant methylcytosine dioxygenase includes at least amino acid residues 1987-2136 of SEQ ID NO: 1. In some embodiments, the variant methylcytosine dioxygenase includes at least amino acid residues 1419-1753 and 1991-2136 ofSEQ ID NO: 1. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1 , and the variant methylcytosine dioxygenase includes at least amino acid residues 1987- 2136 of SEQ ID NO: 1. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 1, and the variant methylcytosine dioxygenase includes at least amino acid residues 1419-1753 and 1987-2136 of SEQ ID NO: 1. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1754-1990 of SEQ ID NO: 1. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 1419-1753 and 1987- 2136 of SEQ ID NO: 1.

[0097] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 2. In some embodiments, the variant methylcytosine dioxygenase comprises any one of SEQ ID NOs:17, 585, 940-948, or 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase comprises at least: (i) amino acid residues 1130-1155 of SEQ ID NO: 2, at the N-terminus of the variant methylcytosine dioxygenase; (ii) amino acid residues 1182-1194 of SEQ ID NO: 2; and (iii) amino acid residues 1912-1935 of SEQ ID NO: 2. In some embodiments, the variant methylcytosine dioxygenase comprises at least: (i) amino acid residues 1130-1155 of SEQ ID NO: 2, at the N- terminus of the variant methylcytosine dioxygenase; (ii) amino acid residues 1182-1194 of SEQ ID NO: 2; and (iii) amino acid residues 1912-1935 of SEQ ID NO: 2, where the variant methylcytosine dioxygenase is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to any one of SEQ ID NOs:941 or 944-947. In some embodiments, the variant methylcytosine comprises at least: (i) amino acid residues 1130-1460 and 1848-2002 of SEQ ID NO: 2; (ii) amino acid residues 1130-1463 and 1844-1968 of SEQ ID NO: 2; (iii) amino acid residues 1130-1463 and 1844-1953 of SEQ ID NO: 2; (iv) amino acid residues 1130-1463 and 1844-1935 of SEQ ID NO: 2; (v) amino acid residues 1130-1456 and 1852-1911 of SEQ ID NO: 2; or (vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 2. In some embodiments, the variant methylcytosine comprises at least: (i) amino acid residues 1130-1460 and 1848-2002 of SEQ ID NO: 2; (ii) amino acid residues 1130-1463 and 1844-1968 of SEQ ID NO: 2; (iii) amino acid residues 1130-1463 and 1844-1953 of SEQ ID NO: 2; (iv) amino acid residues 1130-1463 and 1844-1935 of SEQ ID NO: 2; (v) amino acid residues 1130-1456 and 1852-1911 of SEQ ID NO: 2; or (vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 2, where the variant methylcytosine dioxygenase is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to any one of SEQ ID NOs:941 or 945-947. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 2, and the variant methylcytosine dioxygenase includes at least amino acid residues 1130-1463 of SEQ ID NO: 2. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 2, and the variant methylcytosine dioxygenase includes at least amino acid residues 1844-2002 of SEQ ID NO: 2. In some embodiments, the native methylcytosine dioxygenase includes the amino acidsequence set forth in SEQ ID NO: 2, and the variant methylcytosine dioxygenase includes at least amino acid residues 1130-1463 and 1844-2002 of SEQ ID NO: 2. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1464-1843 of SEQ ID NO: 2. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 1130-1463 and 1844-2002 of SEQ ID NO: 2.

[0098] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the variant methylcytosine dioxygenase comprises SEQ ID NO: 14 or 15, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical thereto. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 3, and the variant methylcytosine dioxygenase includes at least amino acid residues 825-1158 of SEQ ID NO: 3. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 3, and the variant methylcytosine dioxygenase includes at least amino acid residues 1636-1795 of SEQ ID NO: 3. In some embodiments, the variant methylcytosine dioxygenase includes at least amino acid residues 825-1158 and 1636- 1795 of SEQ ID NO: 3. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 825-1158 and 1636-1795 of SEQ ID NO: 3. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1159-1635 of SEQ ID NO: 3.

[0099] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase comprises any one of SEQ ID NOs:4, 11, 12, or 931-939, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase comprises at least amino acid residues 1368-1383 of SEQ ID NO: 4, at the N-terminus of the variant methylcytosine dioxygenase. In some embodiments, the variant methylcytosine dioxygenase comprises at least amino acid residues 1368-1383 of SEQ ID NO: 4, at the N-terminus of the variant methylcytosine dioxygenase, where the variant methylcytosine dioxygenase is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to any one of SEQ ID NOs:932-939. In some embodiments, the variant methylcytosine dioxygenase comprises at least: amino acid residues 1368-1383 of SEQ ID NO: 4, at the N-terminus of the variant methylcytosinedioxygenase; and amino acid residues 1564-1582 of SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase comprises at least: amino acid residues 1368-1383 of SEQ ID NO: 4, at the N-terminus of the variant methylcytosine dioxygenase; and amino acid residues 1564-1582 of SEQ ID NO: 4, where the variant methylcytosine dioxygenase is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to any one of SEQ ID NOs:932, or 934-939. In some embodiments, the variant methylcytosine dioxygenase comprises at least: (i) amino acid residues 1368-1723 and 1848-2002 of SEQ ID NO: 4; (ii) amino acid residues 1368-1723 and 1844-1968 of SEQ ID NO: 4; (iii) amino acid residues 1368-1723 and 1844- 1953 of SEQ ID NO: 4; (iv) amino acid residues 1368-1723 and 1844-1935 of SEQ ID NO: 4; (v) amino acid residues 1368-1723 and 1852-1911 of SEQ ID NO: 4; (vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 4; (vii) amino acid residues 1368-1718 and 1909- 1972 of SEQ ID NO: 4; (viii) amino acid residues 1368-1718 and 1904-1972 of SEQ ID NO: 4; or (ix) amino acid residues 1368-1723 and 1902-1972 of SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase comprises at least: (i) amino acid residues 1368-1723 and 1848-2002 of SEQ ID NO: 4; (ii) amino acid residues 1368-1723 and 1844-1968 of SEQ ID NO: 4; (iii) amino acid residues 1368-1723 and 1844-1953 of SEQ ID NO: 4; (iv) amino acid residues 1368-1723 and 1844-1935 of SEQ ID NO: 4; (v) amino acid residues 1368-1723 and 1852-1911 of SEQ ID NO: 4; (vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 4; (vii) amino add residues 1368-1718 and 1909-1972 of SEQ ID NO: 4; (viii) amino acid residues 1368-1718 and 1904-1972 of SEQ ID NO: 4; or (ix) amino acid residues 1368-1723 and 1902-1972 of SEQ ID NO: 4, where the variant methylcytosine dioxygenase is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical to any one of SEQ ID NOs:932, or 934-939. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 4, and the variant methylcytosine dioxygenase includes at least amino acid residues 1368-1732 of SEQ ID NO: 4. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 4, and the variant methylcytosine dioxygenase includes at least amino acid residues 1902-2039 of SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase includes at least amino acid residues 1368-1732 and 1902-2039 of SEQ ID NO: 4. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 4, and the variant methylcytosine dioxygenase includes atleast amino acid residues 1368-2039 of SEQ ID NO: 4. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 4, and the variant methylcytosine dioxygenase includes at least amino acid residues 1368-1732 and 1902-2039 of SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 1368-1732 and 1902-2039 of SEQ ID NO: 4. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1733-1901 of SEQ ID NO: 4.

[0100] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 5, and the variant methylcytosine dioxygenase includes at least amino acid residues 1046-1377 of SEQ ID NO: 5. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 5, and the variant methylcytosine dioxygenase includes at least amino acid residues 1758-1912 of SEQ ID NO: 5. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 5, and the variant methylcytosine dioxygenase includes at least amino acid residues 1046-1377 and 1758-1912 of SEQ ID NO: 5. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 1046-1377 and 1758- 1912 of SEQ ID NO: 5. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1378-1757 of SEQ ID NO: 5.

[0101] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase comprises SEQ ID NO: 9 or 10, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% identical thereto. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 6, and the variant methylcytosine dioxygenase includes at least amino acid residues 698- 1031 of SEQ ID NO: 6. In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 6, and the variant methylcytosine dioxygenase includes at least amino acid residues 1509-1668 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase includes at least amino acid residues 698-1031 and 1509-1668 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase consists of, orconsists essentially of, amino acid residues 698-1031 and 1509-1668 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1032-1508 of SEQ ID NO: 6.

[0102] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 6, and the variant methylcytosine dioxygenase includes at least amino acid residues 2-1031 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase includes at least amino acid residues 2-1031 and 1509- 1668 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 2-1031 and 1509-1668 of SEQ ID NO: 6. In some embodiments, the variant methylcytosine dioxygenase does not include amino acid residues 1032-1508 of SEQ ID NO: 6.

[0103] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 750 to 1,000 amino acids in length (e.g., 800-900 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET1 (e.g., as set forth in FIG. 7 A). In some embodiments, the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes human TET2 (e.g., as set forth in FIG. 7B).

[0104] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase that is 750 to 1 ,000 amino acids in length (e.g., 800-900 amino acids in length) and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET2 (e.g., as set forth in FIG. 7E). In some embodiments, the variant methylcytosine dioxygenase includes one or more mutations relative to a native methylcytosine dioxygenase, wherein the one or more mutations includes a deletion outside a catalytic domain of the native methylcytosine dioxygenase, and wherein the native methylcytosine dioxygenase includes murine TET2 (e.g., as set forth in FIG. 7E).

[0105] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 2, and the variant methylcytosine dioxygenase includes at least amino acid residues 1130-2002 of SEQ ID NO: 2. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residues 1130-2002 of SEQ ID NO: 2.

[0106] In some embodiments, the native methylcytosine dioxygenase includes the amino acid sequence set forth in SEQ ID NO: 5, and the variant methylcytosine dioxygenase includes at least amino acid residues 1046-1912 of SEQ ID NO: 5. In some embodiments, the variant methylcytosine dioxygenase consists of, or consists essentially of, amino acid residuesl046-1912 of SEQ ID NO: 5.

[0107] In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence of any one of the sequences set forth in Table 0.2, Table 19.2 or Table 20.1, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence of any one of SEQ ID NOs: 7-18 or 585-586 (e.g., as set forth in Table 0.2), 931-948 (e.g., as set forth in Table 19.2), 1023, or 1024 (e.g., as set forth in Table 20.1), or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024. In someembodiments, an engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024, with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, an engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023, or 1024, with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0108] In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence of any one of SEQ ID NOs: 10-13, 15-18, 585, 586 (as set forth in Table 0.2), 932-939, 941, or 943-948 (as set forth in Table 19.2), or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence of any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to any one of SEQ ID NOs. 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to any one of SEQ ID NOs: 10-13, ISIS, 585, 586, 932-939, 941, or 943-948.Table 0.2: Compact DNA demethylases (protein sequence) >

[0109] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 7, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 7. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 7.

[0110] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 8, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least85% identical to SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 8. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 8.

[0111] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 9, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 9. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 9.

[0112] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 10, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 10. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 10. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 10. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 10. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 10. Insome embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 10. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 10.

[0113] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 11, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 11. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 11.

[0114] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 12, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 12. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 12.

[0115] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 13, or a sequence that is, is about, or is at least 85%, 90%,95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 13. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 13.

[0116] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 14, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 14. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 14.

[0117] In some embodiments, the variant methylcytosine dioxygenase incl udes the amino acid sequence of SEQ ID NO: 15, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 15. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 15. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 15. Insome embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 15. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 15. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 15. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 15.

[0118] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 16, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 16. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 16.

[0119] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 17, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequenceat least 98% identical to SEQ ID NO: 17. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 17.

[0120] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 18, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 18. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 18.

[0121] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 585, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 585. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 585.

[0122] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 586, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variantmethylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 586. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 586.

[0123] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 931, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 931. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 931.

[0124] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 932, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequenceat least 95% identical to SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 932. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 932.

[0125] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 933, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 933. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 933.

[0126] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 934, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 934. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 934.

[0127] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 935, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 935. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 935.

[0128] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 936, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 936. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 936.

[0129] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 937, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least85% identical to SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 937. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 937.

[0130] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 938, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 938. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 938.

[0131] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 939, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 939. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 939. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 939. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 939. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 939. Insome embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 939. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 939.

[0132] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 940, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 940. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 940.

[0133] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 941, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 941. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 941.

[0134] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 942, or a sequence that is, is about, or is at least 85%,90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 942. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 942.

[0135] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 943, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 943. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 943.

[0136] In some embodiments, the variant methylcytosine dioxygenase incl udes the amino acid sequence of SEQ ID NO: 944, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 944. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 944. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 944. Insome embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 944. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 944. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 944. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 944.

[0137] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 945, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 945. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 945.

[0138] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 946, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequenceat least 98% identical to SEQ ID NO: 946. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 946.

[0139] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 947, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 947. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 946.

[0140] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 948, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 948. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 948.

[0141] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1023, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variantmethylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 1023. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 1023.

[0142] In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 85% identical to SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 90% identical to SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 95% identical to SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 97% identical to SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 98% identical to SEQ ID NO: 1024. In some embodiments, the variant methylcytosine dioxygenase includes an amino acid sequence at least 99% identical to SEQ ID NO: 1024.

[0143] Also provided herein is an engineered gene effector that includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 7-18, 585-586, 931-948, 1023 or 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 7-18, 585-586, 931-948, 1023 or 1024, with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes avariant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of any one of SEQ ID NOs: 7-18, 585-586, 931-948, 1023 or 1024, with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 7-18,585-586, 931- 948, 1023 or 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 7-18, 585- 586, 931-948, 1023 or 1024, with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 7-18, 585-586, 931-948, 1023 or 1024, with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0144] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 7, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 7 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 7 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 7, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 7 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 7 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0145] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 8, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 8 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 8 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 8, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 8 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 8 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0146] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 9, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98% or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 9 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 9 with 0, 1 , 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 9, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 9 with 0, 1 , 2, or 3 aminoacid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 9 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0147] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 10, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 10 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 10 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 10, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 10 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 10 with 0, 1 , 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0148] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 11 , or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 11 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 11 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, thevariant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 11 , or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 11 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 11 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0149] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 12, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 12 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 12 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 12, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 12 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 12 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0150] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 13, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosinedioxygenase includes the amino acid sequence of SEQ ID NO: 13 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 13 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 13, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 13 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 13 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0151] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 14, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 14 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 14 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 14, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 14 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 14 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0152] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 15, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 15 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 15 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 15, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 15 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 15 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0153] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 16, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 16 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 16 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 16, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenaseconsists of or consists essentially of the amino acid sequence of SEQ ID NO: 16 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 16 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0154] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 17, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 17 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 17 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 17, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 17 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 17 with 0, 1 , 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0155] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 18, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 18 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes theamino acid sequence of SEQ ID NO: 18 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 18, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 18 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 18 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0156] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 585, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 585 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 585, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 585 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0157] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 586, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 586 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of orconsists essentially of the amino acid sequence of SEQ ID NO: 586, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 586 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0158] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 930, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 930 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 930, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 930 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0159] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 931, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 931 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 931, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 931 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0160] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 932, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 932 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 932, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 932 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0161] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 933, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 933 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 933, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 933 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0162] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 934, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosinedioxygenase includes the amino acid sequence of SEQ ID NO: 934 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 934, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 934 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0163] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 935, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 935 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 935, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 935 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0164] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 936, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 936 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 936, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In someembodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 936 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0165] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 937, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 937 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 937, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 937 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0166] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 938, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 938 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 938, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 938 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0167] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes theamino acid sequence of SEQ ID NO: 939, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 939 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 939, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 939 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0168] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 940, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 940 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 940, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 940 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0169] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 941, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 941 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservativesubstitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 941, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 941 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0170] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 942, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 942 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 942, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 942 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0171] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 943, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 943 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 943, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of theamino acid sequence of SEQ ID NO: 943 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0172] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 944, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 944 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 944, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 944 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0173] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 945, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 945 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 945, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 945 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0174] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 946, or a sequence that is, is about, or is at least 85%,90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 946 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 946, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 946 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0175] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 947, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 947 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 947, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 947 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0176] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 948, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 948 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of orconsists essentially of the amino acid sequence of SEQ ID NO: 948, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 948 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0177] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1023, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1023 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 1023, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 1023 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0178] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase includes the amino acid sequence of SEQ ID NO: 1024 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 1024, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant methylcytosine dioxygenase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 1024 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, any mutations thereof are conservative substitutions.

[0179] In some embodiments, the engineered gene effector includes a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase is a variant murine methylcytosine dioxygenase that is 400 to 1700 amino acids in length and includes a methylcytosine dioxygenase catalytic domain. In some embodiments, the engineered gene effector comprises a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase is a variant murine methylcytosine dioxygenase that is 400 to 1700 amino acids in length and includes a methylcytosine dioxygenase catalytic domain, wherein the variant methylcytosine dioxygenase comprises an amino acid sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical to the amino acid sequence of any one of SEQ ID NOs: 7-12, 18, 586, 931-939 or 1023.

[0180] In some embodiments, the engineered gene effector of the present disclosure having a variant methylcytosine dioxygenase demethylates a target gene and activates (e.g., reactivates) the target gene in a cell upon expressing the engineered gene effector to be targeted to a locus of the target gene. As used herein, “activate” denotes increasing the expression level or activity of the target gene relative to a suitable reference expression level or reference level of activity. In some embodiments, the engineered gene effector of the present disclosure having a variant methylcytosine dioxygenase demethylates a target gene and at least partially reactivates the target gene in a cell upon expressing the engineered gene effector to be targeted to a locus of the target gene. As used herein, “reactivate” denotes restoring or increasing expression of the target gene from a transcriptionally suppressed or silenced state. The engineered gene effector can be targeted to the target gene or locus using any suitable option. In some embodiments, the engineered gene effector is coupled to a targeting moiety, such as a heterologous endonuclease that is configured to bind to the target gene or locus, as described herein. The engineered gene effector can be targeted to any suitable region of the target gene or locus. In some embodiments, the engineered gene effector is targeted to a locus of the target gene that includes a regulatory sequence (e.g., a promoter, enhancer, etc.) of the target gene. In some embodiments, the engineered gene effector is targeted to a locus of the target gene that includes a promoter region of the target gene. In some embodiments, the target gene is endogenous to the cell. In some embodiments, the target gene is a silenced gene. In some embodiments, a gene is silenced when the gene product (e.g., mRNA or protein encoded by the gene) is notexpressed above a suitable reference level (e.g., an expression level in a suitable negative control). In some embodiments, a gene is silenced when the gene product (e.g., mRNA or protein encoded by the gene) is expressed above a suitable reference level (e.g., an expression level in a suitable negative control) in, in about, or in at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50% of the silenced cells. In some embodiments, the gene that is silenced in a silenced cell is expressed above a suitable reference level (e.g., an expression level in a suitable negative control) in, in about, or in at least 80, 85, 90, 95, 96, 97, 98, 99% or about 100% of unsilenced cells. In some embodiments, a gene is silenced when the gene product (e.g., mRNA or protein encoded by the gene) is not detectably expressed in the silence cells. In some embodiments, a gene is silenced when the gene product (e.g., mRNA or protein encoded by the gene) is detectably expressed in, in about, or in at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50% of the silenced cells. In some embodiments, the gene that is silenced in a silenced cell is detectably expressed in, in about, or in at least 80, 85, 90, 95, 96, 97, 98, 99% or about 100% of unsilenced cells. In some embodiments, the target gene is a methylated gene, such as a hypermethylated gene. In some embodiments, the target gene or locus includes a CpG island that is at least partially methylated. In some embodiments, the target gene or locus includes a CpG island that is methylated. In some embodiments, the target gene or locus includes a CpG island that is hypermethylated. In some embodiments, the target gene includes a regulatory region (e.g., promoter region, enhancer region, etc.) that includes the CpG island that is at least partially methylated. In some embodiments, the target gene includes a regulatory region (e.g., promoter region, enhancer region, etc.) that includes the CpG island that is hypermethylated. In some embodiments, the CpG island is, is about, or is at least 10, 20, 30, 40, 50, 60, 70, 80, 85, 90, 95, 96, 97, 98, 99% or more methylated. In some embodiments, the CpG island is substantially completely methylated.

[0181] An engineered gene effector of the present disclosure having a variant methylcytosine dioxygenase, in some embodiments, can be coupled to one or more heterologous polypeptides, e.g., to add one or more additional gene regulatory functions. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase. The variant methylcytosine dioxygenase can be coupled to the heterologous polypeptide in any suitable manner. In some embodiments, the heterologous polypeptide is fused to an N-terminus or C-terminus of the variantmethylcytosine dioxygenase. In some embodiments, the heterologous polypeptide is fused directly to the variant methylcytosine dioxygenase. In some embodiments, the heterologous polypeptide is fused indirectly to the variant methylcytosine dioxygenase, e.g., via a peptide peptide spacer or linker. In some embodiments, the heterologous polypeptide is fused indirectly to the variant methylcytosine dioxygenase via a heterologous endonuclease. In some embodiments, the variant methylcytosine dioxygenase is fused (e.g., via a linker or peptide spacer) to a N-terminus of a heterologous endonuclease. In some embodiments, the heterologous polypeptide is fused (e.g., via a linker or peptide spacer) to a C-terminus of the heterologous endonuclease. In some embodiments, the variant methylcytosine dioxygenase is fused to a N-terminus of a heterologous endonuclease, and the heterologous polypeptide is fused to a C-terminus of the heterologous endonuclease. In some embodiments, the heterologous polypeptide is or includes a transcriptional activator. In some embodiments, the heterologous polypeptide is or includes one or more (e.g., 1, 2, 3, 4, 5 or more) transcriptional activators.

[0182] The heterologous polypeptide can include any suitable transcriptional activator. In some embodiments, the transcriptional activator is XVI .48. In some embodiments, “XVI.48” denotes a transcriptional activator domain having the following amino acid sequence:GGSDALDDFDLDMLGGSLDDCLPMVDHIEGCLLDLLSDVGQELPD LGDLGGSDALDDFDLDMLGGSLDDCLPMVDHIEGCLLDLLSDVGQELPDLGDL (SEQ ID NO: 20).

[0183] In some embodiments, the transcriptional activator is XVI.2. In some embodiments, “XVI.2” denotes a transcriptional activator domain having the following amino acid sequence:PPAGQSQTPFSPEGPVPSHVSGLDDCLPMVDHIEGCLLDLLSDVGQELPDLGDLGELLCETASPQGPMQSEGGEEGSTESVSVLP (SEQ IDNO: 21).

[0184] In some embodiments, the transcriptional activator is VPR. In some embodiments, “VPR” denotes a transcriptional activator domain having the following amino acid sequence:DALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSDALDDFDLDMLGSGGSGSQYLPDTDDRHRIEEKRKRTYETFKSIMKKSPF SGPTDPRPPPRRIAVPSRSSASVPKPAPQPYPFTSSLSHNYDEFPTM VFPSGQISQASALAPAPPQVLPQAPAPAPAPAMVSALAQAPAPVPV LAPGPPQAVAPPAPKPTQAGEGTLSEALLQLQFDDEDLGALLGNS TDPAVFTDLASVDNSEFQQLLNQGIPVAPHTTEPMLMEYPEAITRLVTGAQRPPDPAPAPLGAPGLPNGLLSGDEDFSSIADMDFSALLSQIS SGSGSGSRDSREGMFLPKPEAGSAISDVFEGREVCQPKRIRPFHPPG SPWANRPLPASLAPTPTGPVHEPVGSLTPAPVPQPLDPAPAVTPEA SHLLEDPDEETSQAVKALREMADTVIPQKEEAAICGQMDLSHPPPR GHLDELTTTLESMTEDLNLDSPLTPELNEILDTFLNDECLLHAMHISTGLSIFDTSLF (SEQ ID NO: 780)

[0185] In some embodiments, the transcriptional activator is Leutx-Tox4. In some embodiments, “Leutx-Tox4” denotes a transcriptional activator domain having the following amino acid sequence:DLREPSGIKNPGGASASARVSSWDSQSYDIEQICLGASNPPWASTL' FEIDEFVKIYDLPGEDDTSSLNQYLFPVCLEYDQLQSSVGSGGSGG SGGSGEFPGGNDNYLTITGPSHPFLSGAETFHTPSLGDEEFEIPPISLDSDPSLAVSDWGHFDDLADPSSSQDGSFSAQYGVQTLDMPV(SEQ ID NO: 781)

[0186] In some embodiments, the heterologous polypeptide is fused indirectly (e.g., via a linker or peptide spacer) to the variant methylcytosine dioxygenase. Any suitable linker or peptide spacer can be used. In some embodiments, the linker or peptide spacer includes the amino acid sequence of any one of SEQ ID NOs: 396-431, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. As used herein, “linker” and “peptide spacer” are used interchangeably.

[0187] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase (e.g., any one of the variant methylcytosine dioxygenase as described herein), wherein the heterologous polypeptide includes the amino acid sequence of any one or more of SEQ ID NOs: 20-21, and 780-781, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.In some embodiments, the heterologous polypeptide includes one or more amino acid sequences of SEQ ID NOs: 20-21, 684, 688, 735, 768, 780, and / or 781, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide includes any two of the amino acid sequences of SEQ ID NOs: 20-21, 684, 688, 735, 768, 780, and / or 781, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide is fused to an N-terminus or C -terminus of the variant methylcytosine dioxygenase. In some embodiments, the heterologous polypeptide is a transcriptional activator, and includes the amino acid sequence of any one or more of SEQ ID NOs: 20-21, and 780-781, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase, wherein the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 20, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase, wherein the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 21, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase, wherein the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 780, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 684, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 688, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 735, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 768, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologouspolypeptide coupled to the variant methylcytosine dioxygenase, wherein the heterologous polypeptide includes the amino acid sequence of SEQ ID NO: 781, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous polypeptide is a transcriptional activator that includes SEQ ID NOs:735 and 684. In some embodiments, the heterologous polypeptide is a transcriptional activator that includes, in relative order from N- to C-terminus, SEQ ID NOs:735 and 684. In some embodiments, the heterologous polypeptide is a transcriptional activator that includes, in relative order from N- to C-terminus, SEQ ID NOs: 735 and 684, with a peptide spacer therebetween. In some embodiments, the heterologous polypeptide is a transcriptional activator that includes SEQ ID NOs:688 and 768. In some embodiments, the heterologous polypeptide is a transcriptional activator that includes, in relative order from N- to C-terminus, SEQ ID NOs: 688 and 768. In some embodiments, the heterologous polypeptide is or includes a transcriptional activator that includes, in relative order from N- to C-terminus, SEQ ID NOs: 688 and 768, with a peptide spacer therebetween. Any suitable peptide spacer can be used, for example and without limitation, any one of the peptide spacers set forth in Table 0.7. In some embodiments, the peptide spacer is or includes SEQ ID NO: 396. In some embodiments, the heterologous polypeptide is or includes a transcriptional activator that includes, from N- to C- terminus, SEQ ID NO:735, SEQ ID NO:396, and SEQ ID NO: 684 (“hvTR-Q2-ZNF’)). In some embodiments, the heterologous polypeptide is or includes a transcriptional activator that includes, from N- to C-terminus, SEQ ID NO: 688, SEQ ID NO:396, and SEQ ID NO: 768 (“Leutx-ZFX”). In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase (e.g., any one of the variant methylcytosine dioxygenaseas described herein), wherein the heterologous polypeptide includes any one or more of the amino acid sequences set forth in Table 0.3 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosine dioxygenase (e.g., any one of the variant methylcytosine dioxygenaseas described herein), wherein the heterologous polypeptide includes any one of the amino acid sequences set forth in SEQ ID NOs: 772-781 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant methylcytosinedioxygenase (e.g., any one of the variant methylcytosine dioxygenaseas described herein), wherein the heterologous polypeptide includes any one of the amino acid sequences set forth in Table 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any one or more of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any one or more of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence having no more than 5, 4, 3, 2, 1 substitutions thereto). In some embodiments, the heterologous polypeptide includes any two or more of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any two or more of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence having no more than 5, 4, 3, 2, 1 substitutions thereto). In some embodiments, the heterologous polypeptide includes any two of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any two of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence having no more than 5, 4, 3, 2, 1 substitutions thereto). In some embodiments, the heterologous polypeptide includes any three of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any four of the amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the heterologous polypeptide includes any two or more different amino acid sequences set forth in Tables 0.3 and / or 0.4 (or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto). In some embodiments, the two or more different amino acid sequences are fused to each other via a peptide spacer. Any suitable peptide spacer can be used, for example and without limitation, any one of the peptide spacers set forth in Table 0.7.Table 0.3: Transcriptional activator amino acid sequencesTable 0.4: Individual peptide amino acid sequencesVariant DNA Methvltransferases

[0188] An engineered gene effector that includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase is 180 to 360 amino acids in length and includes a catalytic domain-(CD-)like domain is provided. In general, the variant DNA methyltransferase is a functional variant of a native DNA methyltransferase, as described herein. In some embodiments, the variant DNA methyltransferase is 180, 190, 200, 210, 220, 230, 240, 250,260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360 amino acids in length. In someembodiments, the variant DNA methyltransferase has a length in a range defined by any two of the preceding values (e.g., 180-360 amino acids, 190-350 amino acids, 190-230 amino acids, 200-300 amino acids, etc.). In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase is at most 300 amino acids in length and includes a C -terminal catalytic domain-(CD-)like domain. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase is at most 230 amino acids in length and includes a C- terminal catalytic domain-(CD-)like domain.

[0189] In some embodiments, the variant DNA methyltransferase includes one or more mutations in the amino acid sequence of a native DNA methyltransferase. In some embodiments, the native DNA methyltransferase is a DNA methyltransferase 3 like (DNMT3L). In some embodiments, the variant DNA methyltransferase includes one or more mutations in the amino acid sequence of a DNA methyltransferase 3 like (DNMT3L). In some embodiments, the DNMT3L includes a C-terminal catalytic domain-like domain. In some embodiments, the DNMT3L includes an N-terminal ADD domain. The DNMT3L can be any suitable DNA methyltransferase. In some embodiments, the DNMT3L is human, non-human primate, mouse, rat, rabbit, cat, dog, horse, cow, or pig DNMT3L. In some embodiments, the DNMT3L is human DNMT3L (Gene ID: 29947). In some embodiments, the DNMT3L includes the amino acid sequence set forth in SEQ ID NO: 22, below.SEQ ID NO: 22 (DNMT3L)1 MAAIPALDPE AEPSMDVILV GSSELSSSVS PGTGRDLIAY EVKANQRNIE 51 DICICCGSLQ VHTQHPLFEG GICAPCKDKF LDALFLYDDD GYQSYCSICC 101 SGETLLICGN PDCTRCYCFE CVDSLVGPGT SGKVHAMSNW VCYLCLPSSR 151 SGLLQRRRKW RSQLKAFYDR ESENPLEMFE TVPVWRRQPV RVLSLFEDIK 201 KELTSLGFLE SGSDPGQLKH WDVTDTVRK DVEEWGPFDL VYGATPPLGH 251 TCDRPPSWYL FQFHRLLQYA RPKPGSPRPF FWMFVDNLVL NKEDLDVASR 301 FEMEPVTIPD VHGGSLQNAV RVWSNIPAIR SRHWALVSEL EELSLLAQNK 351 QSSKLAAKWP TKLVKNCFLP LREYFKYFST ELTSSL

[0190] The one or more mutations can be a substitution, deletion, insertion, etc. In some embodiments, the one or more mutations in the amino acid sequence of DNMT3L includes a deletion of one or more amino acid residues relative to the native sequence of DNMT3L. In some embodiments, the variant DNA methyltransferase includes one or more mutations in the amino acid sequence of a DNA methyltransferase 3 like (DNMT3L), and wherein the one or more mutations in the amino acid sequence of the DNMT3L is or includesa deletion of an N-terminal domain of the DNMT3L. In some embodiments, the variant DNA methyltransferase includes an N-terminal truncation of the DNMT3L. In some embodiments, the variant DNA methyltransferase includes a C-terminal truncation of the DNMT3L. In some embodiments, the variant DNA methyltransferase includes a truncation of the N- and C- terminus of the DNMT3L.

[0191] In some embodiments, the variant DNA methyltransferase comprises at least amino acid residues 178-380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 164-380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 153- 380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 86-380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 68-380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 57-380 of SEQ ID NO: 22. In some embodiments, the variant DNA methyltransferase includes at least amino acid residues 34-380 of SEQ ID NO: 22.

[0192] In some embodiments, the variant DNA methyltransferase includes any one of the sequences set forth in Table 0.5, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase includes any one of SEQ ID NOs: 23-29 (e.g., as set forth in Table 0.5), or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence set forth in any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to any one of SEQ ID NOs: 23-29. In some embodiments, the variant DNAmethyltransferase includes an amino acid sequence at least 99% identical to any one of SEQ ID NOs: 23-29.

[0193] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 23 or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 23. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 23.

[0194] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 24, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 24. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 24.

[0195] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 25, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNAmethyltransferase includes the amino acid sequence of SEQ ID NO: 25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 25. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 25.

[0196] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 26, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 26. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 26.

[0197] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 27, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 27. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 27.

[0198] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 28, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 28. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 28.

[0199] In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, 99% or more identical thereto. In some embodiments, the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 85% identical to SEQ ID NO: 29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 90% identical to SEQ ID NO: 29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 95% identical to SEQ ID NO:29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 97% identical to SEQ ID NO: 29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 98% identical to SEQ ID NO: 29. In some embodiments, the variant DNA methyltransferase includes an amino acid sequence at least 99% identical to SEQ ID NO: 29.Table 0.5: Variant DNA methyltransferase amino acid sequences

[0200] Also provided herein is an engineered gene effector comprising a variantDNA methyltransferase, wherein the variant DNA methyltransferase comprises any one ofSEQ ID NOs: 23-29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase is 180 to360 amino acids in length. In some embodiments, the variant DNA methyltransferase includes any one of SEQ ID NOs: 23-29 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase includes any one of SEQ ID NOs: 23-29with 0, 1, 2, or 3 amino acid residue mutations wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 23-29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 23-29 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 23-29 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0201] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 23, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 23 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 23 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 23, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 23 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 23 with 0, 1 , 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0202] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 24, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes theamino acid sequence of SEQ ID NO: 24 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:24 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 24, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 24 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 24 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0203] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 25, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 25 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:25 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 25, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 25 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 25 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0204] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acidsequence of SEQ ID NO: 26, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 26 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:26 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 26, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 26 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 26 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0205] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 27, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 27 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:27 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 27, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 27 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially ofthe amino acid sequence of SEQ ID NO: 27 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0206] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 28, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 28 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:28 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 28, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 28 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 28 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0207] In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO: 29 with 0, 1 , 2, or 3 amino acid residue mutations. In some embodiments, the engineered gene effector includes a variant DNA methyltransferase, wherein the variant DNA methyltransferase includes the amino acid sequence of SEQ ID NO:29 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 29, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In someembodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 29 with 0, 1, 2, or 3 amino acid residue mutations. In some embodiments, the variant DNA methyltransferase consists of or consists essentially of the amino acid sequence of SEQ ID NO: 29 with 0, 1, 2, or 3 amino acid residue mutations, wherein any mutations thereof are conservative substitutions.

[0208] An engineered gene effector of the present disclosure having a variant DNA methyltransferase, in some embodiments, can be coupled to one or more heterologous polypeptides, e.g., to add one or more additional gene regulatory functions. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase. In some embodiments, the heterologous polypeptide is fused to an N-terminus or C -terminus of the variant DNA methyltransferase. In some embodiments, the heterologous polypeptide is fused to the N-terminus of the variant DNA methyltransferase. In some embodiments, the heterologous polypeptide is fused to the C-terminus of the variant DNA methyltransferase. In some embodiments, the heterologous polypeptide is fused directly to the N-terminus or C-terminus of the variant DNA methyltransferase.

[0209] In some embodiments, the heterologous polypeptide is fused indirectly (e.g., via a linker or peptide spacer) to the N-terminus or C-terminus of the variant DNA methyltransferase. Any suitable linker or peptide spacer can be used. In some embodiments, the linker or peptide spacer is or includes GKESGSVGGSGGSSEQLAQFRSLDG (SEQ ID NO: 32). In some embodiments, the engineered gene effector includes, from N-terminus to C- terminus, any one of the heterologous polypeptides as described herein, the linker or peptide spacer of SEQ ID NO: 32, and any one of the variant DNA methyltransferases as set forth herein.

[0210] In some embodiments, the heterologous polypeptide is a transcriptional repressor. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N-terminus or C-terminus of the variant DNA methyltransferase, and wherein the heterologous polypeptide is a transcriptional repressor.

[0211] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419,hvTR_Q9WT06, cds_NZ WFIYO 1000004. l_cds_WP_l 55864260.1_2251 , VGLL4.1,VGLL4.2, VGLL4.3, and a variant thereof.

[0212] In some embodiments, the heterologous polypeptide is or includes or comprises a Krueppel-associated box (KRAB). KRAB is a domain (e.g. having about 75 amino acid residues or less) that can be found in eukaryotic Krueppel-type C2H2 zinc finger proteins (ZFPs). The KRAB is a benchmark gene effector capable of repressing a target gene in a cell. However, in some cases, the KRAB may not be optimal or sufficient for regulating all genes. Thus, various aspects of the present disclosure, for example, provide engineered gene effectors that are not identical to the KRAB, yet comparably effective as the KRAB in reducing expression or activity level of one or more target genes, compositions thereof, and methods of use thereof. Examples of proteins ( or fragments thereof) that can be used as a fusion partner to decrease transcription include but are not limited to: transcriptional repressors such as the Krueppel associated box (KRAB or SKD); K0X1 repression domain; the Mad mSIN3 interaction domain (SID); the ERF repressor domain (ERD), or the SRDX repression domain (e g, for repression in plants), and the like. “Suppressor” and “repressor” are used interchangeably herein.

[0213] In some embodiments, “KRAB” denotes a transcriptional repressor domain having any one of the following amino acid sequences:SEQ ID NO: 34 (KRAB - 1)1 DAKSLTAWSR TLVTFKDVFV DFTREEWKLL DTAQQIVYRN VMLENYKNLV 51 SLGYQLTKPD VILRLEKGEE PSEQ ID NO: 35 (KRAB - 2)1 DAKSLTAWSR TLVTFKDVFV DFTREEWKLL DTAQQILYRN VMLENYKNLV 51 SLGYQLTKPD VILRLEKGEE PWLVEREIHQ ETHPDSETAF EIKSSV

[0214] In some embodiments, “EZH2” denotes a transcriptional repressor domain having the following amino acid sequence:SEQ ID NO: 36 (EZH2)1 MGQTGKKSEK GPVCWRKRVK SEYMRLRQLK RFRRADEVKT MFSSNRQKIL 51 ERTETLNQEW KQRRIQPVHI MTSVSSLRGT RECSVTSDLD FPAQVIPLKT 101 LNAVASVPIM YSWSPLQQNF MVEDETVLHN IPYMGDEVLD QDGTFIEELI 151 KNYDGKVHGD RECGFINDEI FVELVNALGQ YNDDDDDDDG DDPDEREEKQ 201 KDLEDNRDDK ETCPPRKFPA DKIFEAISSM FPDKGTAEEL KEKYKELTEQ 251 QLPGALPPEC TPNIDGPNAK SVQREQSLHS FHTLFCRRCF KYDCFLHPFH 301 ATPNTYKRKN TETALDNKPC GPQCYQHLEG AKEFAAALTA ERIKTPPKRP351 GGRRRGRLPN NSSRPSTPTI SVLESKDTDS DREAGTETGG ENNDKEEEEK 401 KDETSSSSEA NSRCQTPIKM KPNIEPPENV EWSGAEASMF RVLIGTYYDN 451 FCAIARLIGT KTCRQVYEFR VKESSIIAPV PTEDVDTPPR KKKRKHRLWA 501 AHCRKIQLKK DGSSNHVYNY QPCDHPRQPC DSSCPCVIAQ NFCEKFCQCS 551 SECQNRFPGC RCKAQCNTKQ CPCYLAVREC DPDLCLTCGA ADHWDSKNVS 601 CKNCSIQRGS KKHLLLAPSD VAGWGIFIKD PVQKNEFISE YCGEIISQDE 651 ADRRGKVYDK YMCSFLFNLN NDFWDATRK GNKIRFANHS VNPNCYAKVM 701 MVNGDHRIGI FAKRAIQTGE ELFFDYRYSQ ADALKYVGIE REMEIP

[0215] In some embodiments, “ZNF689” denotes a transcriptional repressor domain having the following amino acid sequence:SEQ ID NO: 37 (ZNF689)1 APPSAPLPAQ GPGKARPSRK RGRRPRALKF VDVAVYFSPE EWGCLRPAQR 51 ALYRDVMRET YGHLGALGCA GPKPALISWL ERNTD

[0216] In some embodiments, “AK9” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 38 (AK9)KPLVENRAS I FEKCHPI PAPLAQKMLT FT YKYI S S FGYWDPVKLSEGET IKPVENAE NPIYPVIHRQYIYFLSSKETKEKFMKNP

[0217] In some embodiments, “ZNF419” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 587 (ZNF419)AAAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLE NFTLLASLGLASSKTHEITQLESWEEPF

[0218] In some embodiments, “hvTR_Q9WT06” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 39 (hvTR_Q9WT06)VASQMCI FCALYKQNKLSLE YVSGDLKTSVFSPI I IKDCLCVQTTI STTQMLPGTKS SAI FPVYDLRKLLSALVI SEGSVRFDI *

[0219] In some embodiments,“cds NZ WFIY01000004.1 cds WP 155864260.1 2251” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 40 (cds_NZ_WFIY01000004. l _cds_WP_l 55864260.1_2251)AEEKFPGGLVKVQAKVQGKVI EAI L I T GDFFVE PKRAI YDLEARLKWS LVDD I EKEV RDWYSQVKIIGIKPEDLIKVIKEAVSK*

[0220] In some embodiments, “VGLL4.1” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 41 (VGLL4.1)LPSLGLEQPLALTKNSLDASRPAGLSPTLTPGERQQNRPSVITCASAGARNCNLSHC PIAHSGCAAPGPASYRRPPSAATTCDPV

[0221] In some embodiments, “VGLL4.2” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 42 (VGLL4.2)APTMSLHGSHLYTSLPSLGLEQPLALTKNSLDASRPAGLSPTLTPGERQQNRPSVIT CASAGARNCNLSHC P I AHS GCAAPGPAS

[0222] In some embodiments, “VGLL4.3” denotes a transcriptional suppressor domain having the following amino acid sequence:SEQ ID NO: 43 (VGLL4.3)GLEQPLALTKNSLDASRPAGLSPTLTPGERQQNRPSVITCASAGARNCNLSHCPIAH SGCAAPGPASYRRPPSAATTCDPWEEH

[0223] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N-terminus or C -terminus of the variant DNA methyltransferase, wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, ZNF419, hvTR_Q9WT06, cds_NZ_WFIY01000004. l_cds_WP_l 55864260.1 2251, VGLL4.1, VGLL4.2, VGLL4.3, and a variant thereof.

[0224] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N-terminus or C-terminus of the variant DNA methyltransferase, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34- 43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

[0225] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, isabout, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N- terminus or C-terminus of the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 38, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 39, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 40, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 41, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 42, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 43, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, whereinthe heterologous polypeptide comprises the amino acid sequence of SEQ ID NO: 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

[0226] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419, hvTR_Q9WT06, cds_NZ_WFIY01000004.1_cds_WP_155864260.1_2251, VGLL4.1, VGLL4.2, VGLL4.3, and a variant thereof, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98% or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N- terminus or C-terminus of the variant DNA methyltransferase, wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419, hvTR_Q9WT06, cds_NZ_WHY01000004.1_cds_WP_155864260.1_2251, VGLL4.1, VGLL4.2, VGLL4.3, and a variant thereof, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, and wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419, hvTR_Q9WT06, cds_NZ_WFIY01000004. l_cds_WP_l 55864260.1_2251, VGLL4.1, VGLL4.2, VGLL4.3, and a variant thereof, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is fused to an N-terminus or C- terminus of the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, and wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419, hvTR Q9WT06, cds_NZ_WFIY01000004.1_cds_WP_155864260.1_2251, VGLL4.1, VGLL4.2, VGLL4.3,and a variant thereof, and wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

[0227] In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the variant DNA methyltransferase is coupled to the C-terminus of the heterologous polypeptide. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, and wherein the variant DNA methyltransferase is coupled to the C-terminus of the heterologous polypeptide. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34- 43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto, and wherein the variant DNA methyltransferase is coupled to the C-terminus of the heterologous polypeptide. In some embodiments, the engineered gene effector includes a heterologous polypeptide coupled to the variant DNA methyltransferase, wherein the heterologous polypeptide is a transcriptional repressor, wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34-43 or 587, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto, and wherein the variant DNA methyltransferase is coupled to the C-terminus of the heterologous polypeptide.

[0228] In some embodiments, the engineered gene effector includes a KRAB domain fused at the C-terminus to a variant DNA methyltransferase of the present disclosure. In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in any one of SEQ ID NOs: 23-29. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in any one of SEQ ID NOs: 23-29.

[0229] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO:34, and the sequence set forth in SEQ ID NO: 23. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 23.

[0230] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 24. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 24.

[0231] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 25. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 25.

[0232] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 26. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 26.

[0233] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 27. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 27.

[0234] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 28. In some embodiments, the engineered geneeffector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 28.

[0235] In some embodiments, the engineered gene effector includes an amino acid sequence that includes, from N-terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, and the sequence set forth in SEQ ID NO: 29. In some embodiments, the engineered gene effector includes an amino acid sequence that consists of or consists essentially of, from N- terminus to C-terminus, the sequence set forth in SEQ ID NO: 34, the sequence set forth in SEQ ID NO: 32, and the sequence set forth in SEQ ID NO: 29.

[0236] In some embodiments, an engineered gene effector of the present disclosure having a variant DNA methyltransferase is capable of methylating a target gene to thereby suppress the target gene in a cell upon expressing the engineered gene effector to be targeted to a locus of the target gene. The engineered gene effector can be targeted to the target gene or locus using any suitable option. In some embodiments, the engineered gene effector is targeted to a locus of the target gene that includes a regulatory sequence (e.g., a promoter, enhancer, etc.) of the target gene. In some embodiments, the engineered gene effector is targeted to a locus of the target gene that includes a promoter region of the target gene. In some embodiments, the target gene is endogenous to the cell.HETEROLOGOUS ENDONUCLEASES

[0237] In some embodiments, an engineered gene effector of the present disclosure(e.g., having a variant methylcytosine dioxygenase or a variant DNA methyltransferase) includes a polypeptide that is coupled to a heterologous endonuclease (e.g., enzymatically active Cas protein, enzymatically deactivated Cas protein, etc.). In some embodiments, the engineered gene effector as disclosed herein, or a protein comprising the engineered gene effector (e.g., a protein comprising the engineered gene effector coupled to the heterologous endonuclease) can be referred to as an actuator moiety. The engineered gene effector and the heterologous endonuclease can be coupled to each other, e.g., directly or indirectly (e.g., via a linker). For example, the engineered gene effector and the heterologous endonuclease can be fused to each other, e.g., directly or indirectly (e.g., via the linker). In another example, the engineered gene effector and the heterologous endonuclease can be non-covalently coupled toeach other, e.g., via ionic bonds, hydrogen bonds, interactions mediated by oligomerization or dimerization domains, etc. In some cases, the engineered gene effector and the heterologous endonuclease can be part of a single polypeptide molecule (e.g., a chimeric or fusion polypeptide).

[0238] In a wide variety of organisms including diverse mammals, animals, plants, microbes, and yeast, a CRISPR / Cas system (e.g., modified and / or unmodified) can be utilized as a genome engineering tool, or can be modified to direct specific binding of engineered proteins to target loci as disclosed herein. A CRISPR / Cas system can comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid binding. An RNA-guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease) can specifically bind a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave the DNA.

[0239] Non-limiting examples of the heterologous endonuclease as disclosed and used herein can include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases including type I CRISPR-associated (Cas) polypeptides, type II CRISPR-associated (Cas) polypeptides, type in CRISPR-associated (Cas) polypeptides, type IV CRISPR- associated (Cas) polypeptides, type V CRISPR-associated (Cas) polypeptides, and type VI CRISPR-associated (Cas) polypeptides; zinc finger nucleases (ZFN); transcription activatorlike effector nucleases (TALEN); meganucleases; RNA-binding proteins (RBP); CRISPR- associated RNA binding proteins; recombinases; flippases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaeal Argonaute (aAgo), or eukaryotic Argonaute (eAgo)); or any derivative thereof; any variant thereof and any fragment thereof.

[0240] In some embodiments, the heterologous endonuclease as disclosed herein can be nuclease-deficient. In some embodiments, a heterologous endonuclease can be a nuclease-null DNA binding protein that does not induce transcriptional activation or repression of a target DNA sequence unless it is present in a complex with one or more engineered gene effectors of the disclosure. In some embodiments, the heterologous endonuclease can be a nuclease-null DNA binding protein that can induce transcriptional activation or repression of a target DNA sequence (e.g., which can be altered or augmented by the presence of an engineered gene effector as provided herein). In some cases, the Cas protein is mutated and / ormodified to yield a nuclease deficient protein or a protein with decreased nuclease activity relative to a wild-type Cas protein. A nuclease deficient protein can retain the ability to bind DNA, but may lack or have reduced nucleic acid cleavage activity. In some embodiments, the heterologous endonuclease is a nuclease-deactivated variant of a Cas protein (e.g., a Cas protein having a nuclease activity).

[0241] In some embodiments, the heterologous endonuclease as disclosed herein can be an RNA nuclease such as an engineered (e.g., programmable or targetable) RNA nuclease. In some embodiments, the heterologous endonuclease as disclosed herein can be a nuclease-null RNA binding protein that does not induce transcriptional activation or repression of a target RNA sequence unless it is present in a complex with one or more engineered gene effectors of the disclosure. In some embodiments, the heterologous endonuclease as disclosed herein can be a nuclease-null RNA binding protein that can induce transcriptional activation or repression of a target RNA sequence (e.g., which can be altered or augmented by the presence of an engineered gene effector as provided herein).

[0242] In some embodiments, the heterologous endonuclease can be a nucleic acid- guided targeting system. In some embodiments, the heterologous endonuclease can be a DNA- guided targeting system. In some embodiments, the heterologous endonuclease can be an RNA-guided targeting system. The nucleic acid-guided targeting system can comprise and utilize, for example, a guide nucleic acid sequence that facilitates specific binding of a CRISPR-Cas system (e.g., a nuclease deficient form thereof, such as dCas9 or dCasl4) to a target gene (e.g., target endogenous gene) or target gene regulatory sequence. For example, the target gene may be any one of the genes listed in Table 0.12, and the target gene regulatory sequence may be operatively coupled to any one of the genes listed in Table 0.12. Binding specificity can be determined by use of a guide nucleic acid, such as a single guide RNA (sgRNA) or a part thereof. In some embodiments, the use of different sgRNAs allows the compositions, combinations, systems, and methods of the disclosure to be used with (e.g., targeted to) different target genes (e.g., target endogenous genes) or target gene regulatory sequences.

[0243] In some embodiments, prokaryotic CRISPR-Cas (Clustered regularly interspaced short palindromic repeats-CRISPR associated) systems, for example, Class II CRISPR-Cas systems such as Cas9 and Cpfl, can be repurposed as a tool for regulation of geneexpression, epigenome editing, and chromatin looping in compositions, combinations, systems, and methods of the disclosure. In some embodiments, nuclease-deactivated Cas (dCas) proteins complexed with heterologous gene effectors can allow for regulation of expression of target genes (e.g., target endogenous genes) adjacent to a site bound by the dCas.

[0244] Any suitable CRISPR / Cas system can be used. A CRISPR / Cas system can be referred to using a variety of naming systems. A CRISPR / Cas system can be a type I, a type II, a type III, a type IV, a type V, a type VI system, or any other suitable CRISPR / Cas system. A CRISPR / Cas system as used herein can be a Class 1, Class 2, or any other suitably classified CRISPR / Cas system. Class 1 or Class 2 determination can be based upon the genes encoding the effector module. Class 1 systems generally have a multi-subunit crRNA-effector complex, whereas Class 2 systems generally have a single protein, such as Cast), Cpfl, C2cl, C2c2, C2c3 or a crRNA-effector complex. A Class 1 CRISPR / Cas system can use a complex of multiple Cas proteins to effect regulation. A Class 1 CRISPR / Cas system can comprise, for example, type I (e.g., I, LA, IB, IC, ID, IE, IF, or IU), type III (e.g., in, IDA, U1B, IHC, or IUD), and type IV (e.g., IV, IVA, or IVB) CRISPR / Cas type. A Class 2 CRISPR / Cas system can use a single large Cas protein to effect regulation. A Class 2 CRISPR / Cas systems can comprise, for example, type n (e.g., II, IIA, or IIB) and type V CRISPR / Cas type. CRISPR systems can be complementary to each other, and / or can lend functional units in trans to facilitate CRISPR locus targeting.

[0245] When a heterologous endonuclease includes a Cas protein or derivative thereof, the Cas protein or derivative thereof can be a Class 1 or a Class 2 Cas protein. A Cas protein can be a type I, type n, type III, type IV, type V Cas protein, or type VI Cas protein. A Cas protein can comprise one or more domains. Non-limiting examples of domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, or HNH), DNA binding domain, RNA binding domain, helicase domains, protein-protein interaction domains, or dimerization domains. A guide nucleic acid recognition and / or binding domain can interact with a guide nucleic acid. A nuclease domain can comprise catalytic activity for nucleic acid cleavage. A nuclease domain can lack catalytic activity to prevent nucleic acid cleavage. A Cas protein can be a chimeric Cas protein or fragment thereof that is fused to other proteins or polypeptides. A Cas protein can be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins.

[0246] Non-limiting examples of Cas proteins include c2cl, C2c2, c2c3, Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, CasSal, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csxl2), CaslO, CaslOd, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cul966, Casl3a, Casl3b, Casl3c, Casl3d, Casl3X, or Casl3 Y, and homologs or modified versions thereof.

[0247] In some embodiments, the Cas protein as disclosed herein may not and need not be Cas9 or Cas 12a. The Cas protein as disclosed herein can have a smaller size as compared to Cas9 or Casl 2a. The Cas protein as disclosed herein can be derived from UnlCasl2fl. In some embodiments, a heterologous endonuclease can comprise an amino acid sequence having at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 91%, at least or up to about 92%, at least or up to about 93%, at least or up to about 94%, at least or up to about 95%, at least or up to about 96%, at least or up to about 97%, at least or up to about 98%, at least or up to about 99%, or about 100% sequence identity to the polypeptide sequence of SEQ ID NO: 45 (e.g., CasMini). In some embodiments, a heterologous endonuclease can comprise an amino acid sequence having at least or up to about 50%, at least or up to about 55%, at least or up to about 60%, at least or up to about 65%, at least or up to about 70%, at least or up to about 75%, at least or up to about 80%, at least or up to about 85%, at least or up to about 90%, at least or up to about 91%, at least or up to about 92%, at least or up to about 93%, at least or up to about 94%, at least or up to about 95%, at least or up to about 96%, at least or up to about 97%, at least or up to about 98%, at least or up to about 99%, or about 100% sequence identity to the polypeptide sequence of SEQ ID NO: 54 (e.g., dCasMini). As disclosed herein, SEQ ID NO: 45 encodes the polypeptide sequence of UnlCasl2fl. As disclosed herein, SEQ ID NO: 54 encodes an engineered variant of UnlCasl2fl with reduced nuclease activity. As disclosed herein, SEQ ID NO: 55 encodes a non-limiting examples of a Casl2f variant suitable for use in the systems, compositions, combinations, and methods of the present disclosure. In some embodiments, the Casl2f variant as disclosed herein can comprise an amino acid sequence that is at least or at least about50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO: 55.SEQ ID NO: 45 (UnlCasl2f1, “CasMini”)1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSDVCYTRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKIGEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIDVGVK SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAD YNAALNISNP KLKSTKEEPSEQ ID NO: 54 (deactivated nuclease variant of UnlCasl2fl, “dCasMini”)1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSRVCYRRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKICEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIAVGVR SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAA YNAALNISNP KLKSTKERPSEQ ID NO: 55 (Casl2f variant)MAKNT I TKTLKLRI VRPYNSAEVEKI VADEKERRKQAGGTGELDDKFYQKLR GQFPDAVFWQE I SE I FRQLQKQAAE I YNQS LIE L YYE I F I KGKG I ANAS S VE HYLSRVCYRRAAELFKNAAIAGLRSKIKSNFRLKELKNMKSGLPTTKSDNFP I PLVKQKGGQYTGFEISNHNSDFI IKI PFGRWQVKKEIDKYRPWEKFDFEQV QKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSK I CEKSAWMLNLS I DVPKI DKGVDPS I I GGIAVGVRS PLVCAINNAFSRYS I S DNDLFHFNKKMFARRRI LLKKNRHKRAGHGAKNKLKPI T I LTEKSERFRKKL IERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNK IEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKC NFKENAAYNAALNI SNPKLKSTKERP

[0248] In some embodiments, the amino acid sequence of the heterologous endonuclease as disclosed herein can be mutated and / or modified to yield a nuclease deficient protein or a protein with decreased nuclease activity relative to a wild-type Cas protein. A nuclease deficient protein can retain the ability to bind a target gene (e.g., DNA), but may lack or have reduced nucleic acid cleavage activity. In some embodiments, a heterologous endonuclease can exhibit reduced nuclease activity (e.g., nuclease deficient or nuclease null) as compared to wild type UnlCasl2fl. The reduced nuclease activity can be at most about 95%, at most about 90%, at most about 80%, at most about 70%, at most about 60%, at most about 50%, at most about 40%, at most about 30%, at most about 20%, at most about 10%, at most about 5%, at most about 1%, at most about 0.5%, at most about 0.1%, or less than that of the wild type UnlCasl2fl.

[0249] In some cases, a Cas protein as provided herein may not be a Cas 14 protein.

[0250] A Cas protein or fragment or derivative thereof can be from any suitable organism. Non-limiting examples include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus psetidomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas nap hthalenivorans, Polaromonas sp., Crocosphaera "watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desuljbrudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidilhiobacillus caldus, Acidilhiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus "watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho ajricanus, Acaryochloris marina, Leptotrichia shahii, or Francisella novicida. In some aspects, the organism is Streptococcus pyogenes (S. pyogenesy In some aspects, the organism isStaphylococcus aureus (S. aureus). In some aspects, the organism is Streptococcus thermophilus (S'. thermophilus).

[0251] A Cas protein can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans. Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter muslelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasleurella multocida subsp. Multocida, Sutterella "wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, or Francisella novicida.

[0252] A Cas protein as used herein can be a wildtype or a modified form of a Cas protein. A Cas protein can be an active variant, inactive variant, or fragment of a wild type or modified Cas protein. A Cas protein can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein. A Cas protein can be a polypeptide with at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type Cas protein. A Cas protein can be a polypeptide with at most or at most about 5%, at most or at most about 10%, at most or at most about 20%, at most or at most about 30%, at most or at most about 40%, at most or at most about 50%, at most or at most about 60%, at most or at most about 70%, at most or at most about 80%, at most or at most about 90%, or at most or at most about 100% sequence identity and / or sequence similarity to a wild type exemplary Cas protein. Variants or fragments can comprise at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.

[0253] A Cas protein can comprise one or more nuclease domains, such as DNase domains. For example, a Cas9 protein can comprise a RuvC-like nuclease domain and / or an HNH-like 20 nuclease domain. The in a nuclease active form of Cas9, RuvC and HNH domains can each cut a different strand of double-stranded DNA to make a double-stranded break in theDNA. A Cas protein can comprise only one nuclease domain (e.g., Cpfl comprises RuvC domain but lacks HNH domain). In some embodiments, nuclease domains are absent. In some embodiments, nuclease domains are present but inactive or have reduced or minimal activity. In some embodiments, nuclease domains are present and active.

[0254] One or a plurality of the nuclease domains (e.g., RuvC, or HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. For example, in a Cas protein comprising at least two nuclease domains (e.g., Cas9), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, knownas a nickase, can generate a single-strand break at a CRISPR RNA (crRNA) recognition sequence within a double- stranded DNA but not a double-strand break. Such a nickase can cleave the complementary strand or the non-complementary strand, but may not cleave both. If all of the nuclease domains of a Cas protein (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are deleted or mutated, the resulting Cas protein can have a reduced or no ability to cleave both strands of a double-stranded DNA. An example of a mutation that can convert a Cas9 protein into a nickase is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from & pyogenes. H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of Cas9 from S', pyogenes can convert the Cas9 into a nickase. An example of a mutation that can convert a Cas9 protein into a dead Cas9 is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain and H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of Cas9 from S. pyogenes.

[0255] A nuclease dead Cas protein (e.g., one derived from any Cas protein, such as UnlCasl2fl) can comprise one or more mutations relative to a wild-type version of the protein. The mutation can result in no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity in one or more of the plurality of nucleic acid-cleaving domains of the wild-type Cas protein. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains retaining the ability to cleave the complementary strand of the target nucleic acid but reducing its ability to cleave the non-complementary strand of the target nucleic acid. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains retaining the ability to cleave the non-complementary strand of the target nucleic acid but reducing its ability to cleave the complementary strand of the target nucleic acid. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains lacking the ability to cleave the complementary strand and the non-complementary strand of the target nucleic acid. The residues to be mutated in a nuclease domain can correspond to one or more catalytic residues of the nuclease. For example, residues in the wild type exemplary & pyogenes CasS polypeptide such as AsplO, His840, Asn854 and Asn856 can be mutated to inactivate one ormore of the plurality of nucleic acid-cleaving domains (e.g., nuclease domains). The residues to be mutated in a nuclease domain of a Cas protein can correspond to residues Asp 10, His840, Asn854 and Asn856 in the wild type 5. pyogenes Cas9 polypeptide, for example, as determined by sequence and / or structural alignment.

[0256] As non-limiting examples, residues DIO, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 (or the corresponding mutations of any of the Cas proteins) can be mutated. For example, e.g., D 10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. Mutations other than alanine substitutions can be suitable.

[0257] A D10A mutation can be combined with one or more of H840A, N854A, or N856A mutations to produce a Cas9 protein substantially lacking DNA cleavage activity (e.g., a dead Cas9 protein). A H840A mutation can be combined with one or more of D10A, N854A, or N856A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity. An N854A mutation can be combined with one or more of H840A, D1OA, or N856A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity. A N856A mutation can be combined with one or more of H840A, N854A, or DI 0A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity.

[0258] In some embodiments, a Cas protein is a Class 2 Cas protein. In some embodiments, a Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, or derived from a Cas9 protein. For example, a Cas9 protein lacking cleavage activity. In some embodiments, the Cas9 protein is a Cas9 protein from 8. pyogenes (e.g., SwissProt accession number Q99ZW2). In some embodiments, the Cas9 protein is a Cas9 from S’. aureus (e.g., SwissProt accession number J7RUA5). In some embodiments, the Cas9 protein is a modified version of a Cas9 protein from S’. pyogenes or S’. Aureus. In some embodiments, the Cas9 protein is derived from a Cas9 protein from 8 pyogenes or 8. Aureus. For example, a S. pyogenes or 5. Aureus Cas9 protein lacking cleavage activity.

[0259] In some embodiments, Cas9 can generally refer to a polypeptide with at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or atleast about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, or about 100% sequence identity and / or sequence similarity to a wild type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyogenes). In some embodiments, Cas9 can refer to a polypeptide with at most about 5%, at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 70%, at most about 80%, at most about 90%, or about 100% sequence identity and / or sequence similarity to a wild type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 can refer to the wildtype or a modified form of the Cas9 protein that can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof.

[0260] A Cas protein can comprise an amino acid sequence having at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a nuclease domain (e.g., RuvC domain, or HNH domain) of a wild-type Cas protein.

[0261] A Cas protein, variant or derivative thereof can be modified to enhance regulation of gene expression by compositions, combinations, systems, and methods of the disclosure, e.g., as part of a complex disclosed herein. A Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, enzymatic activity, and / or binding to other factors, such as heterodimerization or oligomerization domains and induce ligands. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the desired function of the protein or complex. A Cas protein can be modified to modulate (e.g., enhance or reduce) the activity of the Cas protein for regulating gene expression by a complex of the disclosure that comprises a heterologous gene effector.

[0262] For example, a Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a heterologous gene effector (e.g., an epigenetic modification domain, a transcriptional activation domain, and / or a transcriptional repressor domain). A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to an oligomerization or dimerization domain as disclosed herein (e.g., a heterodimerization domain). A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a heterologous polypeptide that provides increased or decreased stability. A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a sequence that can facilitate degradation of the Cas protein or a complex containing the Cas protein, for example, a degron, such as an inducible degron (e.g., auxin inducible).

[0263] A Cas protein can be coupled (e.g., fused, covalently coupled, or non- covalently coupled) to any suitable number of partners, for example, at least one, at least two, at least three, at least four, or at least five, at least six, at least seven, or at least 8 partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to at most two, at most three, at most four, at most five, at most six, at most seven, at most eight, or at most ten partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to 1 - 5, 1 - 4, 1 - 3, 1 - 2, 2 - 5, 2 - 4, 2 - 3, 3 - 5, 3 - 4, or 4 - 5 partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to one partner. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to two partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to three partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to four partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to five partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to six partners.

[0264] A Cas protein can be a fusion protein. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.

[0265] A Cas protein can be provided in any form. For example, a Cas protein can be provided in the form of a protein, such as a Cas protein alone or complexed with a guide nucleic acid as a ribonucleoprotein. A Cas protein can be provided in a complex, for example, complexed with a guide nucleic acid and / or one or more heterologous gene effectors of the disclosure. A Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as an RNA (e.g., messenger RNA (mRNA)), or DNA. The nucleic acid encoding the Cas protein can be codon optimized for efficient translation into protein in a particular cell or organism.

[0266] Nucleic acids encoding Cas proteins, fragments, or derivatives thereof can be stably integrated in the genome of a cell. Nucleic acids encoding Cas proteins can be operably linked to a promoter, for example, a promoter that is constitutively or inducibly active in the cell. Nucleic acids encoding Cas proteins can be operably linked to a promoter in an expression construct. Expression constructs can include any nucleic acid constructs capable of directing expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and which can transfer such a nucleic acid sequence of interest to a target cell.

[0267] In some embodiments, a Cas protein, variant or derivative thereof is a nuclease dead Cas (dCas) protein. A dead Cas protein can be a protein that lacks nucleic acid cleavage activity.

[0268] A Cas protein can comprise a modified form of a wild type Cas protein. The modified form of the wild type Cas protein can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid-cleaving activity of the Cas protein. For example, the modified form of the Cas protein can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity of the wild-type Cas protein (e.g., Cas9 from S. pyogenes). The modified form of Cas protein can have no substantial nucleic acid-cleaving activity. When a Cas protein is a modified form that has no substantial nucleic acid-cleaving activity, it can be referred to as enzymatically inactive, “deactivated” and / or “dead” (abbreviated by “d”). A dead Cas protein (e.g., dCas, or dCas9) can bind to a target polynucleotide but may not cleave or minimally cleaves the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein.

[0269] A Cas protein can comprise a modified form of a wild type Cas protein. The modified form of the wild type Cas protein can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid-cleaving activity of the Cas protein. For example, the modified form of the Cas protein can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity of the wild-type Cas protein (e.g., Cas9 from S. pyogenes). The modified form of Cas protein can have no substantial nucleic acid-cleaving activity. When aCas protein is a modified form that has no substantial nucleic acid-cleaving activity, it can be referred to as enzymatically inactive, “deactivated” and / or “dead” (abbreviated by “d”). A deadCas protein (e.g., dCas, or dCas9) can bind to a target polynucleotide but may not cleave or minimally cleaves the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein.SEQ ID NO: 44 (deactivated nuclease variant of Cas9 or “dCas9”)1 MDKKYSIGLA IGTNSVGWAV ITDEYKVPSK KFKVLGNTDR HSIKKNLIGA 51 LLFDSGETAE ATRLKRTARR RYTRRKNRIC YLQEIFSNEM AKVDDSFFHR 101 LEESFLVEED KKHERHPIFG NIVDEVAYHE KYPTIYHLRK KLVDSTDKAD 151 LRLIYLALAH MIKFRGHFLI EGDLNPDNSD VDKLFIQLVQ TYNQLFEENP 201 INASGVDAKA ILSARLSKSR RLENLIAQLP GEKKNGLFGN LIALSLGLTP 251 NFKSNFDLAE DAKLQLSKDT YDDDLDNLLA QIGDQYADLF LAAKNLSDAI 301 LLSDILRVNT EITKAPLSAS MIKRYDEHHQ DLTLLKALVR QQLPEKYKEI 351 FFDQSKNGYA GYIDGGASQE EFYKFIKPIL EKMDGTEELL VKLNREDLLR 401 KQRTFDNGSI PHQIHLGELH AILRRQEDFY PFLKDNREKI EKILTFRIPY 451 YVGPLARGNS RFAWMTRKSE ETITPWNFEE WDKGASAQS FIERMTNFDK 501 NLPNEKVLPK HSLLYEYFTV YNELTKVKYV TEGMRKPAFL SGEQKKAIVD 551 LLFKTNRKVT VKQLKEDYFK KIECFDSVEI SGVEDRFNAS LGTYHDLLKI 601 IKDKDFLDNE ENEDILEDIV LTLTLFEDRE MIEERLKTYA HLFDDKVMKQ 651 LKRRRYTGWG RLSRKLINGI RDKQSGKTIL DFLKSDGFAN RNFMQLIHDD 701 SLTFKEDIQK AQVSGQGDSL HEHIANLAGS PAIKKGILQT VKWDELVKV 751 MGRHKPENIV IEMARENQTT QKGQKNSRER MKRIEEGIKE LGSQILKEHP 801 VENTQLQNEK LYLYYLQNGR DMYVDQELDI NRLSDYDVDA IVPQSFLKDD 851 SIDNKVLTRS DKNRGKSDNV PSEEWKKMK NYWRQLLNAK LITQRKFDNL 901 TKAERGGLSE LDKAGFIKRQ LVETRQITKH VAQILDSRMN TKYDENDKLI 951 REVKVITLKS KLVSDFRKDF QFYKVREINN YHHAHDAYLN AWGTALIKK 1001 YPKLESEFVY GDYKVYDVRK MIAKSEQEIG KATAKYFFYS NIMNFFKTEI 1051 TLANGEIRKR PLIETNGETG EIVWDKGRDF ATVRKVLSMP QVNIVKKTEV 1101 QTGGFSKESI LPKRNSDKLI ARKKDWDPKK YGGFDSPTVA YSVLWAKVE 1151 KGKSKKLKSV KELLGITIME RSSFEKNPID FLEAKGYKEV KKDLIIKLPK 1201 YSLFELENGR KRMLASAGEL QKGNELALPS KYVNFLYLAS HYEKLKGSPE 1251 DNEQKQLFVE QHKHYLDEII EQISEFSKRV ILADANLDKV LSAYNKHRDK1301 PIREQAENII HLFTLTNLGA PAAFKYFDTT IDRKRYTSTK EVLDATLIHQ 1351 SITGLYETRI DLSQLGGD

[0270] In some embodiments, “dCas-cA2” denotes a Cas protein having the following amino acid sequence:SEQ n> NO: 56 ( “dCas-cA2” or “cA2”)1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KERRKQAGGT GELDDKFYQK 51 LRGQFPDAVF WQEISEIFRQ LQKQAAEIYN QSLIELYYEI FIKGKGIANA 101 SSVEHYLSRV CYRRAAELFK NAAIASGLRS KIKSNFRLKE LKNMKSGLPT 151 TKSDNFPIPL VKQKGGQYTG FEISNHNSDF IIKIPFGRWQ VKKEIDKYRP 201 WEKFDFEQVQ KSPKPISLLL STQRRKRNKG WSKDEGTEAE IKKVMNGDYQ 251 TSYIEVKRGS KICEKSAWML NLSIDVPKID KGVDPSIIGG IAVGVRSPLV 301 CAINNAFSRY SISDNDLFHF NKKMFARRRI LLKKNRHKRA GHGAKNKLKP 351 ITILTEKSER FRKKLIERWA CEIADFFIKN KVGTVQMENL ESMKRKEDSY 401 FNIRLRGFWP YAEMQNKIEF KLKQYGIEIR KVAPNNTSKT CSKCGHLNNY 451 FNFEYRKKNK FPHFKCEKCN FKENAAYNAA LNISNPKLKS TKERP

[0271] In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease. In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the heterologous endonuclease is a Cas protein. In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the heterologous endonuclease has a length of, of at most, or of 450, 460, 470, 480, 490, 500, 520, 540, 560, 580, 600, 620, 640, 660, 680, 700 amino acids. In some embodiments, the heterologous endonuclease has a length in a range defined by any two of the preceding values (e.g., 450-700 amino acids, 480-600 amino acids, 500-530 amino acids, 500-600 amino acids, etc.). In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the heterologous endonuclease is a Cas protein, and wherein the heterologous endonuclease has a length of, of at most, or of 450, 460, 470, 480, 490, 500, 520, 540, 560, 580, 600, 620, 640, 660, 680, 700 amino acids. In some embodiments, the heterologous endonuclease has a length in a range defined by any two of the preceding values (e.g., 450-700 amino acids, 480-600 amino acids, 500-530 amino acids, 500-600 amino acids, etc.).

[0272] The engineered gene effector can be coupled with any suitable heterologous endonuclease, e.g., without limitation, a heterologous endonuclease having any one of the amino acid sequence set forth in Table 0.6, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the engineered geneeffector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NOs: 44-394, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous endonuclease includes the amino acid sequence of any one of SEQ ID NOs: 44-394. In some embodiments, the heterologous endonuclease consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 44-394. In some embodiments, the heterologous endonuclease includes the amino acid sequence of SEQ ID NO:44, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous endonuclease consists of or consists essentially of the amino acid sequence of SEQ ID NO:44. In some embodiments, the heterologous endonuclease includes the amino acid sequence of SEQ ID NO: 56, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous endonuclease consists of or consists essentially of the amino acid sequence of SEQ ID NO: 56. In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NOs: 56 or 64-245, or a sequence that is, is about, or is at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto. In some embodiments, the heterologous endonuclease includes the amino acid sequence of any one of SEQ ID NOs: 56 or 64-245. In some embodiments, the heterologous endonuclease consists of or consists essentially of the amino acid sequence of any one of SEQ ID NOs: 56 or 64-245.Table 0.6: Heterologous endonuclease amino acid sequences00

[0273] The polypeptide of the engineered gene effector can be coupled to the heterologous endonuclease in any suitable manner. In some embodiments, the heterologous endonuclease is fused to the heterologous endonuclease, e.g., a Cas protein. In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the polypeptide is fused to the N-terminus of heterologous endonuclease. In some embodiments, the engineered gene effector includes a polypeptide that is coupled to a heterologous endonuclease, wherein the polypeptide is fused to the C-terminus of heterologous endonuclease.

[0274] The polypeptide of the engineered gene effector can be coupled to the heterologous endonuclease can be linked to each other directly or indirectly (e.g., via a linker).

[0275] In some embodiments, the engineered gene effector is coupled or fused to the heterologous endonuclease via a linker. Any suitable linker, such as those described herein, can be used to couple the engineered gene effector to the heterologous endonuclease. In some embodiments, the engineered gene effector is coupled or fused to the heterologous endonuclease via a linker (e.g., peptide linker). Any suitable linker can be used. In some embodiments, the linker (e.g., peptide linker) is a flexible linker that has an amino acid sequence containing stretches of glycine and serine residues. The small size of the glycine and serine residues provides flexibility and allows for mobility of the connected functional domains. The incorporation of serine or threonine can in some embodiments maintain the stability of the linker (e.g., peptide linker) or linker in aqueous solutions by forming hydrogen bonds with the water molecules, thereby reducing unfavorable interactions between the linker and the linked moieties. Flexible linkers can also contain in some embodiments additional amino acids such as threonine and alanine to maintain flexibility, as well as polar amino acids such as lysine and glutamine to improve solubility. A rigid linker can have, for example, an alpha helix-structure. An alpha-helical rigid linker can act as a linker between protein domains. Non-limiting examples of linkers include the sequences in Table 0.7, and repeats thereof, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 repeats. SEQ ID NOs: 396-402 provide flexible linkers or subunits thereof. SEQ ID NOs: 403-406 provide rigid linkers or subunits thereof. In some embodiments, the linker is or includes SEQ ID NO: 396.

[0276] In some embodiments, a linker as disclosed herein can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29,30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length.

[0277] In some embodiments, a linker as disclosed herein can comprise at least 1, at least 2, at least 3, at least 5, at least 7, at least 9, at least 11, at least 13, at least 15, or at least 20 amino acids. In some embodiments, a linker can comprise at most 5, at most 7, at most 9, at most 11, at most 13, at most 15, at most 20, at most 25, at most 30, at most 40, or at most 50 amino acids.

[0278] In some embodiments, the linker includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine-serine (GS) linker(s). In some embodiments, the engineered gene effector is coupled to a heterologous endonuclease by a linker that includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine-serine (GS) linker(s). In some embodiments, the engineered gene effector is coupled to a heterologous endonuclease by a linker that includes at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about , at least about 9, at least about 10, at least about 11 , at least about 12, at least about 13, at least about 14, at least about 15, at least about 20, or more GS linkers).

[0279] In some embodiments, the linker includes at most about 20, at most about 15, at most about 14, at most about 13, at most about 12, at most about 11, at most about 10, at most about 9, at most about 8, at most about 7, at most about 6, at most about 5, at most about 4, at most about 3, at most about 2, or about 1 glycine (G) linker(s). In some embodiments, the engineered gene effector is coupled to a heterologous endonuclease by a linker that includes at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about , at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15...

Claims

1. WHAT IS CLAIMED IS:

1. An engineered gene effector comprising a variant methylcytosine dioxygenase that is (1) 400 to 700 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain; or (2) 750 to 1,000 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain.

2. The engineered gene effector of claim 1, wherein the variant methylcytosine dioxygenase is no more than 500 amino acids in length.

3. The engineered gene effector of claim 1 or 2, wherein the variant methylcytosine dioxygenase comprises one or more mutations relative to a native methylcytosine dioxygenase, optionally wherein the one or more mutations comprises a deletion outside a catalytic domain of the native methylcytosine dioxygenase.

4. The engineered gene effector of claim 3, wherein (a) the one or more mutations comprises a deletion of a low-complexity domain (LCD) of the native methylcytosine dioxygenase, or (b) the native methylcytosine dioxygenase comprises a C-terminal catalytic domain comprising a low-complexity domain (LCD), and wherein the one or more mutations comprises a deletion of at least a portion of the LCD.

5. The engineered gene effector of claim 3 or 4, wherein the variant methylcytosine dioxygenase is a variant ten-eleven translocation (TET) methylcytosine dioxygenase comprising one or more mutations relative to a native TET methylcytosine dioxygenase.

6. The engineered gene effector of any one of claims 3-5, wherein the native methylcytosine dioxygenase comprises: human TET1; human TET2; or human TET3 optionally wherein the native TET methylcytosine dioxygenase comprises: human TET1, isoform 1; human TET1, isoform 2; human TET2, isoform 1; human TET2, isoform 2; human TET2, isoform 3; human TET3, isoform 1; human TET3, isoform 2; or human TET3, isoform 3.

7. The engineered gene effector of any one of claims 3-5, wherein the native methylcytosine dioxygenase comprises: murine TET1; murine TET2; or murine TET3.

8. The engineered gene effector of any one of claims 3-5, wherein the native methylcytosine dioxygenase comprises: human TET1; human TET2; human TET3; murine TET1; murine TET2; or murine TET3.

9. The engineered gene effector of any one of claims 3-8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of any one of SEQ ID NOs: 1-6.

10. The engineered gene effector of any one of claims 3-6 or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 1, optionally wherein the variant methylcytosine dioxygenase comprises a sequence at least 85% identical to SEQ ID NO: 13 or 16.

11. The engineered gene effector of claim 10, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1419-1753 of SEQ ID NO: 1.

12. The engineered gene effector of claim 10 or 11, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1991-2136 of SEQ ID NO: 1, optionally wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1987-2136 of SEQ ID NO: 1.

13. The engineered gene effector of any one of claims 3-6 or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 2, optionally wherein the variant methylcytosine dioxygenase comprises a sequence at least 85% identical to any one of SEQ ID NOs: 17, 585, or 940-948.

14. The engineered gene effector of claim 13, wherein the variant methylcytosine dioxygenase comprises at least:(i) amino acid residues 1130- 1155 of SEQ ID NO: 2, at the N-terminus of the variant methylcytosine dioxygenase;00 amino acid residues 1182-1194 of SEQ ID NO: 2; and (iii) amino acid residues 1912-1935 of SEQ ID NO: 2.

15. The engineered gene effector of claim 13 or 14, wherein the variant methylcytosine comprises at least:(i) amino acid residues 1130-1460 and 1848-2002 of SEQ ID NO: 2;(ii) amino acid residues 1130-1463 and 1844-1968 of SEQ ID NO: 2;(iii) amino acid residues 1130-1463 and 1844-1953 of SEQ ID NO: 2;(iv) amino acid residues 1130-1463 and 1844-1935 of SEQ ID NO: 2;(v) amino acid residues 1130-1456 and 1852-1911 of SEQ ID NO: 2; or(vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 2.

16. The engineered gene effector of claim 13, wherein the variant methylcytosine comprises at least amino acid residues 1130-1463 of SEQ ID NO: 2.

17. The engineered gene effector of claim 13 or 14, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1844-2002 of SEQ ID NO: 2.

18. The engineered gene effector of any one of claims 3-6 or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 3, optionally wherein the variant methylcytosine dioxygenase comprises a sequence at least 85% identical to SEQ ID NO: 14 or 15.

19. The engineered gene effector of claim 18, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 825-1158 of SEQ ID NO: 3.

20. The engineered gene effector of claim 18 or 19, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1636-1795 of SEQ ID NO: 3.

21. The engineered gene effector of any one of claims 3-5, 7, or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 4, optionally wherein the variant methylcytosine dioxygenase comprises a sequence at least 85% identical to any one of SEQ ID NOs:4, 11, 12, or 931-939.

22. The engineered gene effector of claim 21, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1368-1383 of SEQ ID NO: 4, at the N- terminus of the variant methylcytosine dioxygenase, optionally wherein the variant methylcytosine dioxygenase comprises 1564-1582 of SEQ ID NO: 4.

23. The engineered gene effector of claim 21 or 22, wherein the variant methylcytosine comprises at least:(i) amino acid residues 1368-1723 and 1848-2002 of SEQ ID NO: 4;(i) amino acid residues 1368-1723 and 1844-1968 of SEQ ID NO: 4;(iii) amino acid residues 1368-1723 and 1844-1953 of SEQ ID NO: 4;(iv) amino acid residues 1368-1723 and 1844-1935 of SEQ ID NO: 4;(v) amino acid residues 1368-1723 and 1852-1911 of SEQ ID NO: 4;(vi) amino acid residues 1130-1460 and 1848-1935 of SEQ ID NO: 4;(vii) amino acid residues 1368-1718 and 1909-1972 of SEQ ID NO: 4;(viii) amino acid residues 1368-1718 and 1904-1972 of SEQ ID NO: 4; or (ix) amino acid residues 1368-1723 and 1902-1972 of SEQ ID NO: 4.

24. The engineered gene effector of claim 21, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1368-1732 of SEQ ID NO: 4.

25. The engineered gene effector of claim 21 or 22, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1902-2039 of SEQ ID NO: 4, optionally wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1368-2039 of SEQ ID NO: 4.

26. The engineered gene effector of any one of claims 3-5, 7, or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 5.

27. The engineered gene effector of claim 26, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1046-1377 of SEQ ID NO:

528. The engineered gene effector of claim 26 or 27, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1758-1912 of SEQ ID NO:

529. The engineered gene effector of any one of claims 3-5, 7, or 8, wherein the native methylcytosine dioxygenase comprises the amino acid sequence of SEQ ID NO: 6, optionally wherein the variant methylcytosine dioxygenase comprises a sequence at least 85% identical to SEQ ID NO: 9 or 10.

30. The engineered gene effector of claim 29, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 698-1031 of SEQ ID NO: 6.

31. The engineered gene effector of claim 29 or 30, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1509-1668 of SEQ ID NO: 6.

32. The engineered gene effector of any one of the preceding claims, wherein the variant methylcytosine dioxygenase comprises an amino acid sequence of any one of SEQ ID NOs: 10-13, 15-18, 585, 586, 932-939, 941, or 943-948, or a sequence that is at least 85% identical thereto.

33. The engineered gene effector of claim 1, wherein the variant methylcytosine dioxygenase is no more than 900 amino acids in length.

34. The engineered gene effector of claim 13, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1130-2002 of SEQ ID NO: 2.

35. The engineered gene effector of claim 26, wherein the variant methylcytosine dioxygenase comprises at least amino acid residues 1046-1912 of SEQ ID NO: 5.

36. The engineered gene effector of any one of claims 1-9, wherein the variant methylcytosine dioxygenase comprises an amino acid sequence of SEQ ID NO: 585 or 586, or a sequence that is at least 85% identical thereto.

37. An engineered gene effector comprising a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase comprises the amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023 or 1024, or a sequence at least 85% identical thereto, optionally wherein the variant methylcytosine dioxygenase is 400 to 700 amino acids in length.

38. The engineered gene effector of claim 37, wherein the variant methylcytosine dioxygenase comprises the amino acid sequence of any one of SEQ ID NOs: 7-18, 585, 586, 931-948, 1023 or 1024, with 0-3 amino acid residue mutations, optionally wherein any mutations thereof are conservative substitutions.

39. An engineered gene effector comprising a variant methylcytosine dioxygenase, wherein the variant methylcytosine dioxygenase is a variant murine methylcytosine dioxygenase that is 400 to 1700 amino acids in length and comprises a methylcytosine dioxygenase catalytic domain.

40. The engineered gene effector of claim 39, wherein the variant methylcytosine dioxygenase comprises an amino acid sequence that is at least 85% identical to the amino acid sequence of any one of SEQ ID NOs: 7-12, 18, 586, 931-939 or 1023.

41. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector demethylates a target gene and at least partially reactivates the target gene in a cell upon expressing the engineered gene effector to be targeted to a locus of the target gene, optionally wherein the locus comprises a promoter region of the target gene, optionally wherein the target gene is endogenous to the cell.

42. The engineered gene effector of claim 41 , wherein the target gene is a silenced gene.

43. The engineered gene effector of claim 41 or 42, wherein the target gene is a methylated gene, such as a hypermethylated gene.

44. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector reactivates expression of a target gene in a cell at a rate of, of about, or of at least 5% .

45. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector reactivates expression of a target gene in a cell, wherein the expression level of the target gene that is reactivated via the engineered gene effector persists for a duration of, of about, or of at least 5 days.

46. The engineered gene effector of any one of the preceding claims, comprising a heterologous polypeptide coupled to the variant methylcytosine dioxygenase, optionally wherein the heterologous polypeptide is fused to an N-terminus or C-terminus of the variant methylcytosine dioxygenase, optionally wherein the heterologous polypeptide is fused indirectly to the variant methylcytosine dioxygenase via a first peptide spacer.

47. The engineered gene effector of claim 46, wherein the heterologous polypeptide is a transcriptional activator, optionally wherein the heterologous polypeptide comprises one or more amino acid sequences selected from Tables 0.3 and / or 0.4.

48. The engineered gene effector of claim 46 or 47, wherein the heterologous polypeptide comprises the amino acid sequence of any one or more of SEQ ID NOs: 20-21, 684, 688, 735, 768, 780, and / or 781, or a sequence at least 85% identical thereto.

49. An engineered gene effector comprising a variant DNA methyltransferase, wherein the variant DNA methyltransferase is 180 to 360 amino acids in length and comprises a catalytic domain (CD)-like domain (e.g., a C-terminal CD-like domain).

50. The engineered gene effector of claim 49, wherein the variant DNA methyltransferase is at most 300 amino acids in length.

51. The engineered gene effector of claim 49 or 50, wherein the variant DNA methyltransferase is at most 230 amino acids in length.

52. The engineered gene effector of any one of claims 49-51, wherein the variant DNA methyltransferase comprises one or more mutations in the amino acid sequence of a native human DNA methyltransferase 3 like (DNMT3L).

53. The engineered gene effector of claim 52, wherein the one or more mutations in the amino acid sequence of the DNMT3L is or comprises a deletion of an N-terminal domain of the native human DNMT3L.

54. The engineered gene effector of claim 53, wherein the variant DNA methyltransferase comprises at least amino acid residues 178-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 164-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 153-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 86-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 68-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 57-380 of SEQ ID NO: 22, optionally wherein the variant DNA methyltransferase comprises at least amino acid residues 34-380 of SEQ ID NO: 22.

55. An engineered gene effector comprising a variant DNA methyltransferase, wherein the variant DNA methyltransferase comprises any one of SEQ ID NOs: 23-29, or a sequence at least 85% identical thereto, optionally wherein the variant DNA methyltransferase is 180 to 360 amino acids in length.

56. The engineered gene effector of claim 55, wherein the variant DNA methyltransferase comprises any one of SEQ ID NOs: 23-29 with 0-3 amino acid residue mutation, optionally wherein any mutations thereof are conservative substitutions.

57. The engineered gene effector of any one of claims 49-56, wherein the engineered gene effector is capable of methylating a target gene to thereby suppress the target gene in a cell upon expressing the engineered gene effector to be targeted to a locus of the target gene, optionally wherein the locus comprises a promoter region of the target gene, optionally wherein the target gene is endogenous to the cell.

58. The engineered gene effector of any one of claims 49-57, wherein the engineered gene effector suppresses expression of a target gene in a cell, and wherein the expression level of the target gene is suppressed by, by at least or by about 5, %.

59. The engineered gene effector of any one of claims 49-58, wherein the engineered gene effector suppresses expression of a target gene in a cell, wherein theexpression level of the target gene is suppressed via the engineered gene effector persists for a duration of, of about, or of at least 5 days.

60. The engineered gene effector of any one of claims 49-59, comprising a heterologous polypeptide coupled to the variant DNA methyltransferase, optionally wherein the heterologous polypeptide is fused to an N-terminus or C-terminus of the variant DNA methyltransferase, optionally wherein the heterologous polypeptide is fused indirectly to the variant DNA methyltransferase via a second peptide spacer.

61. The engineered gene effector of claim 60, wherein the heterologous polypeptide is a transcriptional repressor.

62. The engineered gene effector of claim 60 or 61, wherein the heterologous polypeptide is selected from the group consisting of KRAB, EZH2, ZNF689, AK9, ZNF419, hvTR_Q9WT06, cds_NZ_WHY01000004.1_cds_WP_155864260.1_2251, VGLL4.1, VGLL4.2, VGLL4.3, and a variant thereof.

63. The engineered gene effector of any one of claims 60-62, wherein the heterologous polypeptide comprises the amino acid sequence of any one of SEQ ID NO: 34- 43 or 587, or a sequence at least 85% identical thereto.

64. The engineered gene effector of any one of claims 60-63, wherein the variantDNA methyltransferase is coupled to the C-terminus of the heterologous polypeptide.

65. The engineered gene effector of any one of the preceding claims, wherein the engineered gene effector is coupled to a heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof.

66. The engineered gene effector of claim 65, wherein the engineered gene effector is fused to the heterologous endonuclease.

67. The engineered gene effector of claim 66, wherein the engineered gene effector is fused to the N-terminus of the heterologous endonuclease.

68. The engineered gene effector of claim 66, wherein the engineered gene effector is fused to the C-terminus of the heterologous endonuclease.

69. The engineered gene effector of any one of claims 1-45, wherein the variant methylcytosine dioxygenase is fused to a N-terminus of a heterologous endonuclease, and wherein a heterologous polypeptide is fused to a C-terminus of the heterologous endonuclease,optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof, optionally wherein the heterologous polypeptide is a transcriptional activator.

70. The engineered gene effector of claim 69, wherein the heterologous polypeptide comprises one or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto.

71. The engineered gene effector of claim 70, wherein the heterologous polypeptide comprises two or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto, optionally wherein the transcriptional activator comprises a peptide spacer between the two or more amino acid sequences, optionally wherein the transcriptional activator comprises SEQ ID NOs:735 and 684, or SEQ ID NOs:688 and 768.

72. The engineered gene effector of any one of claims 65-71, wherein the heterologous endonuclease has a length of, of at most, or of 700 amino acids, or wherein the heterologous endonuclease has a length in a range of 450-700 amino acids.

73. The engineered gene effector of any one of claims 65-72, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NOs: 44-394, or a sequence at least 85% identical thereto, optionally wherein the heterologous endonuclease comprises the amino acid sequence of SEQ ID NO:44 or 56.

74. The engineered gene effector of any one of claims 65-73, wherein the engineered gene effector is coupled or fused to the heterologous endonuclease via a linker.

75. The engineered gene effector of claim 74, wherein the linker is a human protein- derived linker.

76. The engineered gene effector of claim 74 or 75, wherein the linker comprises the amino acid sequence of any one of SEQ ID NO: 396-431 , or a sequence at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto, optionally wherein the linker comprises the amino acid sequence of SEQ ID NO 429, or a sequence at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

77. A fusion protein comprising: the engineered gene effector of any one of claims 1-64; and a heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof.

78. The fusion protein of claim 77, wherein the engineered gene effector is fused to the heterologous endonuclease.

79. The fusion protein of claim 77, wherein the engineered gene effector is fused to the N-terminus of the heterologous endonuclease.

80. The fusion protein of claim 77, wherein the engineered gene effector is fused to the C-terminus of the heterologous endonuclease.

81. The fusion protein of claim 77, comprising, from N- to C-terminus: the heterologous endonuclease; a human protein-derived linker; and the engineered gene effector, optionally, wherein the engineered gene effector comprises, from N- to C-terminus: (a) the variant methylcytosine dioxygenase; a first peptide spacer; and a transcriptional activator; or (b) a heterologous polypeptide; a second peptide spacer; and the variant DNA methyltransferase, optionally wherein the heterologous polypeptide is a transcriptional repressor.

82. A fusion protein comprising: the engineered gene effector of any one of claims 1-45; a heterologous polypeptide; and a heterologous endonuclease, wherein the variant methylcytosine dioxygenase is fused to a N-terminus of the heterologous endonuclease, and wherein the heterologous polypeptide is fused to a C-terminus of the heterologous endonuclease, optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof, optionally wherein the heterologous polypeptide is a transcriptional activator.

83. The fusion protein of claim 82, wherein the heterologous polypeptide comprises one or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto.

84. The fusion protein of claim 83, wherein the heterologous polypeptide comprises two or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto, optionally wherein the transcriptional activator comprises a peptide spacer between the two or more amino acid sequences, optionally wherein the transcriptional activator comprises, from N- to C-terminus, SEQ ID NOs:735 and 684, or SEQ ID NOs:688 and 768.

85. The fusion protein of any one of claims 77-84, wherein the engineered gene effector is coupled or fused to the heterologous endonuclease via a linker.

86. The fusion protein of claim 85, wherein the linker is a human protein-derived linker.

87. The fusion protein of claim 85 or 86, wherein the linker comprises the amino acid sequence of any one of SEQ ID NO: 396-431 , or a sequence at least 85%, 90%, 95% 97%, 98%, or 99% identical thereto, optionally wherein the linker comprises the amino acid sequence of SEQ ID NO 429, or a sequence at least 85% identical thereto.

88. The fusion protein of any one of claims 77-87, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NOs: 44-394, or a sequence at least 85% identical thereto, optionally wherein the heterologous endonuclease comprises the amino acid sequence of SEQ ID NO:44 or 56.

89. A fusion protein comprising an amino acid sequence that is at least 85%, 90%, 95%, 97%, 98%, 99% identical to, or is about 100% identical to the amino acid sequence of any one of SEQ ID NOs: 432-438, 783, 795-803, 813-865, 919-924, 967-993 or 994.

90. A polynucleotide comprising a nucleotide sequence encoding the engineered gene effector or the fusion protein of any one of the preceding claims.

91. A polynucleotide comprising a nucleotide sequence that is at least 85%, 90%, 95%, 97%, 98%, or 99% identical to the nucleotide sequence of any one of SEQ ID NOs: 442- 451, 453-459, 804-812, 86-918, 925-930, or 995-1021, or 1022.

92. A vector comprising the polynucleotide of claim 90 or 91.

93. A cell comprising the polynucleotide of claim 90 or 91, or the vector of claim92.

94. A system comprising: the engineered gene effector of any one of claims 1-64; a heterologous endonuclease coupled to the engineered gene effector; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein, optionally wherein the heterologous endonuclease is a Cas protein or a nuclease-deactivated variant thereof.

95. The system of claim 94, wherein the engineered gene effector is fused to the heterologous endonuclease.

96. The system of claim 94, wherein the engineered gene effector is fused to the N- terminus of the heterologous endonuclease.

97. The system of claim 94, wherein the engineered gene effector is fused to the C- terminus of the heterologous endonuclease.

98. A system comprising: the engineered gene effector of any one of claims 1-45; a heterologous polypeptide; a heterologous endonuclease, wherein the variant methylcytosine dioxygenase is fused to a N-terminus of the heterologous endonuclease, and wherein the heterologous polypeptide is fused to a C -terminus of the heterologous endonuclease; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein, optionally wherein the heterologous endonuclease is a Cas protein or a nucleasedeactivated variant thereof, optionally wherein the heterologous polypeptide is a transcriptional activator.

99. The system of claim 98, wherein the heterologous polypeptide comprises one or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto.

100. The system of claim 99, wherein the heterologous polypeptide comprises two or more amino acid sequences selected from Tables 0.3 and / or 0.4, or a sequence having no more than 5 substitutions thereto, optionally wherein the transcriptional activator comprises a peptide spacer between the two or more amino acid sequences, optionally wherein the transcriptional activator comprises, from N- to C-terminus, SEQ ID NOs:735 and 684, or SEQ ID NOs:688 and 768.

101. The system of any one of claims 94-100, wherein the heterologous endonuclease comprises the amino acid sequence of any one of SEQ ID NOs: 44-394, or a sequence at least 85% identical thereto, optionally wherein the heterologous endonuclease comprises the amino acid sequence of SEQ ID NO:44 or 56.

102. The system of any one of claims 94-101, wherein the engineered gene effector is coupled or fused to the heterologous endonuclease via a linker.

103. The system of claim 102, wherein the linker is a human protein-derived linker.

104. The system of claim 102 or 103, wherein the linker comprises the amino acid sequence of any one of SEQ ID NO: 407-431, or a sequence at least 85% identical thereto, optionally wherein the linker comprises the amino acid sequence of SEQ ID NO 429, or a sequence at least 85% identical thereto.

105. A system comprising: the fusion protein of any one of claims 77-89; and a guide nucleic acid capable of forming a complex with the heterologous endonuclease, wherein the complex exhibits specific binding to a target gene in a cell when the system is expressed therein.

106. The system of any one of claims 94-105, wherein the guide nucleic acid comprises: a guide nucleic acid spacer sequence exhibiting specific binding to the target gene; and a scaffold sequence for forming the complex with the heterologous endonuclease, wherein the scaffold sequence comprises a polynucleotide sequence of any one of SEQ ID NO: 488-581, or a sequence at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

107. The system of any one of claims 94-106, wherein the guide nucleic acid is 8 to 40 nucleotides in length.

108. The system of any one of claims 94-107, wherein the guide nucleic acid is a guide RNA.

109. The system of any one of claims 94-108, wherein the guide nucleic acid is a single guide RNA (sgRNA).

110. A polynucleotide encoding the engineered gene effector of any one of claims 1-64, or the fusion protein of any one of claims 77-89.

111. A polynucleotide or a combination of polynucleotides encoding the system of any one of claims 94-109, wherein the polynucleotide or the combination of polynucleotidesis configured to express the heterologous endonuclease coupled to the engineered gene effector and the guide nucleic acid in a cell.

112. A vector comprising the polynucleotide or the combination of polynucleotides of claim 110 or 111.

113. The vector of claim 112, that is a viral vector; optionally wherein the viral vector is an adeno-associated viral (AAV) vector.

114. A cell comprising the polynucleotide or the combination of polynucleotides of claim 110 or 111, or the vector of claim 112 or 113.

115. A kit comprising the engineered gene effector, fusion protein, combination, system, polynucleotide(s), vector, and / or cell of any one of the preceding claims.

116. A method of epigenetic regulation of a target gene in a cell, comprising contacting a cell with the system of any one of claims 94-109, or the polynucleotide or the combination of polynucleotides of claim 110 or 111, or the vector of claim 112 or 113.

117. The method of claim 116, wherein the target gene is endogenous to the cell.

118. The method of claim 116 or 117, wherein the contacting is performed in vitro or ex vivo.

119. The method of any one of claims 116-118, wherein the engineered gene effector comprises the variant methylcytosine dioxygenase, thereby causing reactivation of the target gene.

120. The method of claim 119, wherein the target gene is a silenced gene.

121. The method of claim 119 or 120, wherein the target gene is a methylated gene.

122. The method of any one of claims 119-121, wherein the target gene is reactivated in, in about, or in at least 5%of cells that are contacted.

123. The method of any one of claims 119-122, wherein the expression level of the target gene in the cell that is reactivated via the system, polynucleotide, combination of polynucleotides, or vector persists for a duration of, of about, or of at least 5days.

124. The method ofany one of claims 116-118, wherein the engineered gene effector comprises the variant DNA methyltransferase, thereby suppressing the target gene.

125. The method of claim 124, wherein the expression level of the target gene in the cell is suppressed by, by at least or by about 5%.

126. The method of claim 124 or 125, wherein the expression level of the target gene that is suppressed via the system, polynucleotide, combination of polynucleotides, or vector persists for a duration of, of about, or of at least 5 days.

127. The method of any one of claims 124- 126, wherein the target gene is suppressed in, in about, or in at least 5% of cells that are contacted.

128. The method of claim 127, wherein the suppression increases over a duration of, of about, or of at least 5 days.

129. A computer-implemented method of generating a linker sequence, comprising:(a) selecting a plurality of human protein-derived peptide sequences having a length of 25 to 150 amino acids based on a structural confidence metric, wherein at least two of the peptide sequences of the plurality of human protein-derived peptide sequences have different lengths, wherein each peptide sequence of the plurality of human protein-derived peptide sequences is associated with one of a plurality of length categories based on its amino acid sequence length, optionally wherein at least one peptide sequence of the plurality of human protein-derived peptide sequences comprises a nuclear localization signal (NLS);(b) generating a flexibility score for each peptide sequence of the plurality of human protein-derived peptide sequences by sequence modeling, optionally wherein the flexibility score is a normalized B-factor ranking; and(c) for each of the plurality of length categories, selecting one or more candidate linker sequences from the peptide sequences associated with the length category based on the flexibility score of each of the peptide sequences, optionally selecting the one or more candidate linker sequences from the peptide sequences associated with the length category based on one or more functional features associated with each of the peptide sequences, thereby generating one or more candidate human protein-derived linker sequences.

130. The method of claim 129, comprising in (c), selecting the one or more candidate linker sequences that do not comprise a transmembrane domain, a functional domain, and / or an antigenic sequence of the human protein from which the linker sequence is derived.

131. A system comprising a computing device comprising at least one processor and instructions executable by the at least one processor to perform the method of claim 129 or 130.

132. A non-transitory computer-readable medium having stored thereon computer- readable instructions that, when executed by a processor, cause the processor to execute the method of claim 129 or 130.

133. An engineered human protein-derived linker generated by the method of claim129 or 130.

134. An engineered human protein-derived linker comprising the amino acid sequence of any one of SEQ ID NO: 463-466, 468-471, 473-476, 478-481, 483-486, or a sequence at least 85%, 90%, 95%, 97%, 98%, or 99% identical thereto.

135. A polynucleotide comprising: a heterologous polynucleotide sequence comprising (i) a CRISPR target sequence and (ii) a CRISPR protospacer adjacent motif (PAM) sequence and an additional CRISPR PAM sequence that are different, wherein said CRISPR target sequence is flanked by said CRISPR PAM sequence and said additional CRISPR PAM sequence; a promoter region polynucleotide sequence from a target gene; and a reporter gene polynucleotide sequence, optionally wherein the reporter gene is a green fluorescent protein (GFP).

136. The polynucleotide of claim 135, wherein the heterologous polynucleotide sequences comprises SEQ ID NO: 584.

137. The polynucleotide of claim 135 or 136, wherein the target gene is selected from the group consisting of EFlα, mSnrpn, hSNRPN, EFS, and CD81.

138. The polynucleotide of any one of claims 135-137, wherein at least one nucleotide of the promotor region polynucleotide sequence of the polynucleotide is methylated.

139. A vector comprising the polynucleotide of any one of claims 135-137.

140. A cell comprising the polynucleotide of any one of claims 135-138, or the vector of claim 139.

141. A method of reactivating a target gene in a cell, comprising contacting a cell with the engineered gene effector of any one of claims 1-48, the fusion protein of any one of claims 77-89, the system of any one of claims 94-109, or the polynucleotide or the combination of polynucleotides of claim 110 or 111, wherein the engineered gene effector comprises the variant methylcytosine dioxygenase, and wherein the cell comprises a reporter polynucleotide comprising the polynucleotide of any one of claims 135-138.

142. The method of claim 141, wherein the reporter polynucleotide is integrated into the cell’s genome.

143. The method of claim 141 or 142, wherein the reporter polynucleotide is silenced.

144. The method of any one of claims 141-143, wherein the reporter polynucleotide is methylated.

145. The method of any one of claims 141-144, comprising, before the contacting, silencing or methylating the reporter polynucleotide.

Citation Information

Patent Citations

  • High-activity mTET2 enzyme mutant as well as encoding DNA and application thereof

    CN113846070A

  • Assay for the removal of methyl-cytosine residues from DNA

    WO2017208247A1

  • Fusion effector proteins and uses thereof

    WO2023064923A2

  • Concurrent sequencing of forward and reverse complement strands on concatenated polynucleotides for methylation detection

    WO2023175040A2

  • Epigenetic modulation of genomic targets to control expression of PWS-associated genes

    WO2024040253A1