Systems and methods for regulating aberrant gene expressions

EP4680266A2Pending Publication Date: 2026-01-21EPICRISPR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024771747
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-24
Filing Date
2024-03-14
Publication Date
2026-01-21

AI Technical Summary

Technical Problem

Aberrant gene expression, particularly in muscle cells, leads to diseases like facioscapulohumeral muscular dystrophy, where transient modifications are insufficient to treat or cure the condition, necessitating systems and methods to sustainably modify and regulate aberrant gene expression.

Method used

A system comprising a heterologous polypeptide with a nuclease and a guide nucleic acid molecule specifically binding to a D4Z4 repeat array in muscle cells, effectively modifying the expression and methylation levels of target genes, such as DUX4, while minimizing off-target effects and maintaining the modified expression for extended periods.

Benefits of technology

The system persistently modulates the expression and methylation levels of target genes, normalizes apoptosis levels, enhances muscle cell survival, and maintains minimal effects on non-target genes, offering a sustainable treatment for aberrant gene expression-related diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024019974_19092024_PF_FP_ABST
    Figure US2024019974_19092024_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are systems, compositions, methods for regulating aberrant expression of a target gene in cell (e.g., a muscle cell), to treat or ameliorate a disease or a condition in a subject (e.g., muscular dystrophy, such as Facioscapulohumeral Muscular Dystrophy (FSHD)).
Need to check novelty before this filing date? Find Prior Art

Description

EPICR.022WO PATENT SYSTEMS AND METHODS FOR REGULATING ABERRANT GENE EXPRESSIONS CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and benefit of U.S. Provisional Patent Application No. 63 / 490678, filed March 16, 2023; U.S. Provisional Patent Application No. 63 / 520253, filed August 17, 2023; and U.S. Provisional Patent Application No. 63 / 592882, filed October 24, 2023, each of which are expressly incorporated herein by reference. REFERENCE TO SEQUENCE LISTING

[0002] The present application is being filed along with a Sequence Listing XML in electronic format. The Sequence Listing XML is provided as a file entitled SequenceListing_EPICR022WO.xml, created March 14, 2024, which is 929,129 bytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety. BACKGROUND

[0003] Aberrant expression of one or more genes can lead to a disease or a condition. In some cases, aberrant expression of a germinal transcription factor in a muscle cell can in a subject can lead to muscular dystrophy. For example, aberrant expression of a transcription factor in a muscle cell (e.g., aberrant expression of DUX4 in a skeletal muscle cell) can lead to Facioscapulohumeral Muscular Dystrophy (FSHD). SUMMARY

[0004] Transiently modifying aberrant expression of a target gene in a cell may not be sufficient to treat or cure a disease that is manifested by the aberrant expression of the target gene. Thus, there remains a substantial need for systems and methods to modify the aberrant expression of the target gene and sustain the modified expression level of the target gene for an extended period of time.

[0005] In an aspect, the present disclosure provides a system for regulating aberrant expression of a target gene in a muscle cell, comprising: a heterologous polypeptide comprising a nuclease; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell, wherein, upon formation of the complex, the complex is capable of binding the target polynucleotide sequence, to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, and wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) the system is configured to effect reduced expression level of the target gene in the muscle cell by at least or at least about 80%, as compared to a control cell; and / or (3) the system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off-target effect below a predetermined threshold level in the muscle cell; and / or (4) the system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or(7) the system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell- specific gene and the target gene are not the same; and / or (8) the system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the system.

[0006] In some embodiments of any of the systems disclosed herein, upon formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least about 17 days. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least about 18 days.

[0007] In some embodiments of any of the systems disclosed herein, the muscle cell is in a subject having or is suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the systems disclosed herein, the target gene is Dux4.

[0008] In some embodiments of any of the systems disclosed herein, the nuclease has a length that is less than or equal to about 800 amino acids. In some embodiments of any of the systems disclosed herein, the nuclease has a length that is less than or equal to about 750 amino acids.

[0009] In some embodiments of any of the systems disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.

[0010] In some embodiments of any of the systems disclosed herein, the heterologous polypeptide further comprises a transcriptional regulator. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises at least one methyltransferases. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises at least one DNA Methyltransferases (DNMT). In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises (i) DNMT-L and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises DNMT-L (or DNMT3L) or KRAB or variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises a plurality of different transcriptional regulators.

[0011] In some embodiments of any of the systems disclosed herein, the modification of the expression level and / or the methylation level of the target gene effects downregulation of a downstream gene of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0012] In some embodiments of any of the systems disclosed herein, the modification of the expression level and / or the methylation level of the target gene effects downregulation of an apoptosis marker in the muscle cell. In some embodiments of any of the systems disclosed herein, the apoptosis marker comprises Caspase 3.

[0013] In some embodiments of any of the systems disclosed herein, the complex effects the modification of the expression level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, the modification of the expression level results in downregulation of the target gene.

[0014] In some embodiments of any of the systems disclosed herein, the complex effects the modification of the methylation level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, the modification of the methylation level results in downregulation of the target gene.

[0015] In some embodiments of any of the systems disclosed herein, the nuclease is a deactivated nuclease.

[0016] In another aspect, the present disclosure provides a composition comprising any of the systems disclosed herein.

[0017] In another aspect, the present disclosure provides a viral vector comprising any of the systems or any of the compositions disclosed herein.

[0018] In some embodiments of any of the viral vectors disclosed herein, the viral vector comprises an adeno-associated virus (AAVs), a retrovirus, a lentivirus, a poxvirus, or an adenovirus. In some embodiments of any of the viral vectors disclosed herein, the AAV comprises a AAV serotype RH74 AAV.

[0019] In another aspect, the present disclosure provides a method for regulating aberrant expression of a target gene in a muscle cell, the method comprising (a) contacting the muscle cell with a complex or system comprising (i) a heterologous polypeptide comprising a nuclease and (ii) a guide nucleic acid molecule exhibiting specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell; and (b) upon the contacting, binding the target gene with the complex or system to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, and wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or(2) the complex or system is configured to effect reduced expression level of the target gene in the muscle cell by at least or at least about 80%, as compared to a control cell; and / or (3) the complex or system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off- target effect below a predetermined threshold level in the muscle cell; and / or (4) the complex or system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the complex or system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the complex or system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the complex or system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the complex or system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the system.

[0020] In some embodiments of any of the methods disclosed herein, upon formation of the complex or system, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least or at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least or at least about 17 days. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is sustained for at least or at least about 18 days.

[0021] In some embodiments of any of the methods disclosed herein, the contacting comprises injecting a composition comprising the complex or system to a subject in need thereof, wherein the subject has or is suspected of having facioscapulohumeral musculardystrophy (FSHD). In some embodiments of any of the methods disclosed herein, the target gene is Dux4.

[0022] In some embodiments of any of the methods disclosed herein, the nuclease has a length that is less than or equal to 800 amino acids. In some embodiments of any of the methods disclosed herein, the nuclease has a length that is less than or equal to 750 amino acids.

[0023] In some embodiments of any of the methods disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof.

[0024] In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43. In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.

[0025] In some embodiments of any of the methods disclosed herein, the heterologous polypeptide further comprises a transcriptional regulator. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises at least one methyltransferases. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises at least one DNA Methyltransferases (DNMT). In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises DNMT-L (or DNMT3L) or KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises KRAB or a variant of KRAB. In some embodiments ofany of the methods disclosed herein, the transcriptional regulator comprises a plurality of different transcriptional regulators.

[0026] In some embodiments of any of the methods disclosed herein, the modification of the expression level and / or the methylation level of the target gene effects downregulation of a downstream gene of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0027] In some embodiments of any of the methods disclosed herein, the modification of the expression level and / or the methylation level of the target gene effects downregulation of an apoptosis marker in the muscle cell. In some embodiments of any of the methods disclosed herein, the apoptosis marker comprises Caspase 3.

[0028] In some embodiments of any of the methods disclosed herein, the complex or system effects the modification of the expression level of the target gene in the muscle gene.

[0029] In some embodiments of any of the methods disclosed herein, the modification of the expression level results in downregulation of the target gene.

[0030] In some embodiments of any of the methods disclosed herein, the complex or system effects the modification of the methylation level of the target gene in the muscle gene. In some embodiments of any of the methods disclosed herein, the modification of the methylation level results in downregulation of the target gene.

[0031] In some embodiments of any of the methods disclosed herein, the nuclease is a deactivated nuclease.

[0032] In some embodiments, any of the compositions or vectors disclosed herein are for use as a medicament, such as for the inhibition, amelioration, or treatment of a cancer, autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

[0033] In some embodiments, any of the compositions or vectors disclosed herein are for use in editing a gene in a cell in vitro or in vivo.

[0034] Also provided is a method of providing a cell with a system, which regulates an aberrant expression of a target gene, comprising introducing any of the system, the composition, or the vector disclosed herein into a cell, preferably a cell in a subject, such as a human having a cancer, autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

[0035] In some embodiments of any system, composition, viral vector disclosed herein, the system comprises: a heterologous polypeptide comprising a nuclease comprising an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs:43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865.

[0036] In some embodiments of any system, composition, viral vector disclosed herein, the system comprises: a heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to a SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865.

[0037] In some embodiments of any method disclosed herein, the complex or system comprises: a heterologous polypeptide comprising a nuclease comprising an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs:43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865.

[0038] In some embodiments of any method disclosed herein, the complex or system comprises: a heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to a SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865.

[0039] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. INCORPORATION BY REFERENCE

[0040] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forthillustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0042] FIG. 1 provides different target polynucleotide sequences (e.g., Rank #1 through Rank #91) between two CpG islands within a D4Z4 repeat array that encodes DUX4.

[0043] FIG. 2 provides regulation of DUX4 expression in a target cell population (e.g., lymphoblasts) by a heterologous actuator moiety coupled to a gene regulator (e.g., dCas- KRAB-DNMT3A-DNMT3L) that is complexed with various guide RNA molecules target polynucleotide sequences (e.g., Rank #1 through Rank #91) within the D4Z4 repeat array that encodes DUX4.

[0044] FIG. 3A depicts the gene expression of DUX4 and DUX4-target genes in immortalized patient-derived human FSHD skeletal myoblasts (SkM) cells (12ABIC / 12A and 15ABIC / 15A). The gene expression of DUX4 and DUX4-target genes is measured in 12ABIC and 15ABIC undifferentiated cells, 12ABIC and 15ABIC cells after 2 days of differentiation, and 12ABIC and 15ABIC cells after 7 days of differentiation. Each shade of gray on the graph depicts the gene expression for a different gene corresponding to the legend on the right. FIG. 3B depicts the proportion of apoptotic cells in FSHD myoblasts 12ABIC and 15ABIC (right column) compared to their healthy sibling control myoblasts, 12UBIC and 15VBIC, respectively (left column) after two days of differentiation. The white dots in the images on the left represent apoptotic cells. The graph on the right depicts the percentage of apoptotic cells in the 12ABIC, 15ABIC, 12UBIC, and 15VBIC cell cultures after two days of differentiation shown in the images on the left. DAPI stain is used to stain the for nuclei. FIG. 3C depicts the percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after seven days of differentiation. Percentage of apoptotic cells are measured on day 0, day 1, day 2, and day 7 of differentiation. FIG. 3D depicts the expression of MYHC in 12ABIC, 15ABIC, 12UBIC, and 15 VBIC cells after 7 days of differentiation. Myosin Heavy Chain (MYHC) is a marker for muscle cell differentiation. The white dots indicate expression of MYHC. FIG. 3E depicts the expression level of MYOG, MYH2, and MYMK in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. MYOG is a myogenic regulatory factor that regulates skeletal muscle differentiation and MyoMaker (MYMK) is a marker for muscle cell differentiation. DAPI stain is used to stain the for nuclei. Expression level for the 12ABIC and 15ABIC cells is measured on day 2 and day 7 of differentiation.12AUD and 15A UD: undifferentiated, proliferating control myoblasts. Dark gray bars depicts MYOG expression levels, light gray bars depicts MYH2 expression levels, and gray bars depicts MYMK expression levels.

[0045] FIG.4 depicts the design of multiple gRNAs in relation to the D4Z4 repeat region. The multiple DUX4-targeting gRNAs are designed to span across the D4Z4 repeat region. The D4Z4 repeat region and the DUX4 gene locations in relation to each other is shown at the bottom of FIG.4. The newly designed gRNAs are shown at the top of FIG.4.

[0046] FIG.5 depicts the Cas12f effector-modulator vector design. The expression of the Cas12f variant, KRAB domain, and DNMT3L domain are under the control of a muscle- specific promoter, CK8e. The expression of sgRNA spacer sequence with scaffold driven by RNA polymerase III is under the control of a human U6g promoter. The vector additionally includes a modified WPRE and polyadenylation regulatory sequences.

[0047] FIG. 6A depicts the relative expression level of DUX4 in 12ABIC FSHD myoblasts that stably express the Cas12f-KRAB effector-modulator after 78 gRNAs were nucleofected into the 12ABIC myoblasts. Following nucleofection, the cells are cultured in differentiation conditions for 7 days before the gene expression of DUX4 is measured. The 78 gRNAs tested are listed on the x-axis and the y-axis represents the relative fold expression of DUX4. The expression level of DUX4 was normalized with the expression of control gene HPRT1. FIG. 6B depicts the relative expression level of DUX4 in 12ABIC FSHD myoblasts that stably express the Cas12f-KRAB effector-modulator after 78 gRNAs were nucleofected into the 12ABIC myoblasts. Following nucleofection, the cells are cultured in differentiation conditions for 7 days before the gene expression of DUX4 and the DUX4-target gene, MBD3L2, is measured.

[0048] FIG. 7A depicts the repression of DUX4 and DUX4-target genes, DBET / DUX4, MBD3L2, and TRIM48, in immortalized patient-derived FSHD myoblasts transfected with six gRNAs and a Cas12f effector-modulator. The Cas12f effector-modulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLa domain. One of the six sgRNA is a control sgRNA (Empty / trcr) which did not target the D4Z4 repeat region. Expression level of MYOG is measured in the cells to assay if the differentiation ability of DUX4 sgRNA transfected cells is similar to control sgRNA transfected myoblasts. Expression level of DUX4, DUX4-target genes, and MYOG is measured 17 days post transfection. FIG. 7B depicts therepression of DUX4 and DUX4-target genes, DBET / DUX4, MBD3L2, and TRIM48, in immortalized patient-derived FSHD myoblasts transfected with six gRNAs and a Cas12f effector-modulator. The Cas12f effector-modulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLb domain. One of the six sgRNA is a control sgRNA (Empty) which did not target the D4Z4 repeat region. Expression level of MYOG is measured in the cells to assay if the differentiation ability of DUX4 sgRNA transfected cells is similar to control sgRNA transfected myoblasts. Expression level of DUX4, DUX4-target genes, and MYOG is measured 18 days post transfection.

[0049] FIGs. 8A and 8B depict the apoptosis level of FSHD-patient derived myoblasts transfected with Cas12f effector-modulator and DUX4-targeting gRNA. The percentage of apoptotic-positive cells is measured after two days of differentiation following transfection. The images in FIG. 8A depicts the proportion of apoptotic cells in control 12UBIC cells and 12ABIC cells transfected with the Cas12f effector-modulator and DUX4- targeting gRNA. The white dots depict apoptotic cells. The graphs in FIG. 8B depicts the percentage of apoptotic cells measured in the images on the left, as well as the percentage of apoptotic cells in 12ABIC cells transfected with either a DUX4-targeting gRNA or a control gRNA, which does not target DUX4. DAPI stain is used to stain the for nuclei.

[0050] FIGs.9A-9E show the effects of the exemplary system in 3D ex vivo FSHD organoid model. FIG.9A depicts the workflow for an ex vivo FSHD model. The ex vivo model cultures immortalized healthy sibling control cells and FSHD skeletal myoblasts and then engineered the cells into 3D tissues. The 3D tissues were contacted with either a control AAV or a AAV with the exemplary system described herein. The 3D tissues were then tested for phenotypic differences in mechanical force, tetanic force, and fatigue, in addition to measuring 3D tissue morphology and gene expression profile. FIG. 9B shows the mean active twitch force plotted over time in both GFP control(-) treated (top) & the exemplary system treated (bottom) 3D organoid tissue. FIG.9C shows the active forces at the endpoint (Day 46) in both GFP control(-) treated (top) and the exemplary system treated (bottom) 3D organoid tissue. FIG. 9D shows mean tetanus force plotted over time in both GFP control(-) treated (top) & the exemplary system treated (bottom) 3D organoid tissue. FIG. 9E shows the normalized tetanic force at the endpoint (Day 46) in both GFP control(-) treated (top) and the exemplary system (bottom) treated 3D organoid tissue.

[0051] FIGs. 10A-10G show the effects of the exemplary system in suppressing DUX4 and DUX-4 pathway genes in humanized mice. FIG.10A depicts the workflow for an in vivo xenograft model. The in vivo model was created with treating mice legs with irradiation and TA muscle cardiotoxin to prepare for the transplantation of human myoblast cells into the mice’s leg. Following transplantation, the mice were contacted with either a control AAV or a AAV with the exemplary system described herein. At designated time points, the mice were euthanized, and the xenograft and tissue samples were collected for analysis. The collected xenograft was fixed, sectioned, and stained with Hematoxylin and eosin. The remaining tissues were used for gene expression assays, as well as determining AAV tropism within the mice. FIG.10B shows mRNA expression of DUX4 in humanized TA muscle. FIG.10C shows gene expression of DUX4 pathway genes plotted as a composite score. FIG. 10D shows biodistribution of the exemplary system. FIGs. 10E and 10F show quantification of DUX4 protein (FIG.10E) and SLC34A2 protein (FIG.10F) staining. FIG.10G shows quantification of TUNEL staining.

[0052] FIG. 11 depicts the methylation of a target gene (e.g., D4Z4) via the exemplary system / method.

[0053] FIG.12 illustrates a proposed therapy for FSHD patients.

[0054] FIGs.13A and 13B show the unmethylated or methylated state at the target gene / location. FIG. 13A shows the methylation / unmethylation state of target sites via the exemplary system comprising various gene repressors at multiple time points. FIG.13B shows the percent of CpG methylation at target locus (D4Z4 locus) in healthy sibling control, FSHD patient derived myoblasts treated with control AAV or the exemplary system.

[0055] FIG.14 shows a screening method and results of high throughput screening for anti-DUX4 guide RNA spacer sequences.

[0056] FIG.15 shows a validation and in silico off-target analysis.

[0057] FIG.16 depicts a viral delivery cargo of the exemplary system.

[0058] FIG. 17 shows myogenic gene expression in a patient-derived FSHD myoblast.

[0059] FIG. 18 shows DUX4 and DUX4-downstream gene expression in patient- derived FSHD myoblasts by the exemplary system.

[0060] FIGs.19A – 19D show apoptotic cells in patient-derived FSHD myoblasts by the exemplary system. FIG. 19A shows the % of apoptotic cells in differentiated patient- derived FSHD myoblasts and patient-derived FSHD myoblasts contacted with the exemplary system. FIG. 19B shows mRNA expression of DUX4 and an exemplary system cargo. FIG. 19C shows caspase 3 / 7 stained live cell imaging analysis. FIG. 19D shows normalized intensity of total Caspase3 / 7 signal at the endpoint.

[0061] FIG.20 illustrates the effects of the exemplary system in skeletal muscle of humanized FSHD mice model.

[0062] FIG.21 shows a histology of TA muscle treated with the exemplary system.

[0063] FIG. 22 shows the expression of DUX4 and DUX4 pathway genes in a humanized FSHD mice model contacted with the exemplary system.

[0064] FIG. 23 illustrates a 6-month non-GLP toxicology study on immunocompetent mice.

[0065] FIG. 24 shows evaluation of clinical chemistry in system-treated animals. At each time point, the left bar represents control animals, and the right bar represents the system-treated animals.

[0066] FIG.25 shows a histopathology evaluation of animals 3-month post system administration.

[0067] FIG. 26 shows blood chemistry and hematology data of the non-human primate model contacted with the system.

[0068] FIG.27 shows pharmacokinetics study of the non-human primate model.

[0069] FIG. 28 shows tissue specific mRNA expression level and tropism of a candidate guide nucleic acid molecule.

[0070] FIGs. 29A-29E show the effects of an efficacious dose range of the exemplary system on the molecular and cellular phenotype in humanized mice (together with FIGs.10A-10G). FIG.29A depicts a schematic of an in vivo xenograft model of FSHD. FIGs. 29B-29C show qPCR (FIG.29B) and qRT-PCR (FIG.29C) analysis of the exemplary system in TA muscle biopsies harvested from FSHD humanized mice 24 days post-intravenous administration of the exemplary system or vehicle. FIG. 29B relates to FIG. 10D. FIGs. 29D-29E show quantitative analysis of the DUX4 cascade in TA muscle biopsies harvested from FSHD humanized mice 24 days post-intravenous administration of the exemplary systemor vehicle. FIG.29D shows qRT-PCR analysis of 4 DUX4-target genes plotted as composite score. FIG. 29E shows quantitative image analysis of SLC34A2 immunohistochemistry staining plotted as H Score.

[0071] FIGs. 30A-30H show the effects of a dose escalation of the exemplary system on FSHD patient myoblast derived 3-D engineered muscle tissues (EMTs). FIG. 30A shows a schematic of 3D EMT tissue casting and stimulation model. FIGs.30B-C show max active twitch force (FIG.30B) and tetanic force (FIG.30C) in untransduced control 3D EMT and exemplary system-transduced tissues from day 6 through day 46. FIGs. 30D-E show muscle contractile force normalized to untransduced 3D EMTs at day 46. The max active twitch force (FIG. 30D) and tetanic force (FIG. 30E) showed an increase in the normalized force up to 7.5E8 vg / tissue exemplary system dose level. Data is presented as the average ±SEM. FIGs. 30F-G show quantitative PCR evaluation of the exemplary system vector genome in 3D EMT (FIG. 30F) and qRT-PCR to evaluate mRNA expression (FIG. 30G). Data is presented as the average ±SEM. FIG.30H shows DUX4 and DUX4 and pathway gene expression in 3D EMT’s transduced with exemplary system and untransduced control tissues. Data is presented as the average ±SEM.

[0072] FIGs.31A-31H show the mechanism of action of the exemplary system as a gene-targeted therapy for FSHD directly modulating the epigenetic state of the DUX4 locus and thereby mitigating its pathological overexpression. FIG. 31A shows a schematic of the exemplary system mechanism of action assay using an on-target methylation assay in FSHD myoblasts. FIG. 31B shows DUX4 and six DUX4 genes, namely MBD3L2, ZSCAN4, LEUTX, TRIM43 & TRIM48, mRNA expression in patient-derived myoblasts. FIG 31C shows FSHD patient derived primary myoblasts (07ABIC) contacted with the exemplary system show increased methylation compared to (-) Control test article treated myoblasts. FIG. 31D shows DUX4 and six DUX4 genes, namely MBD3L2, ZSCAN4, LEUTX, TRIM43 & TRIM48, mRNA expression in patient-derived myoblasts (07ABIC) contacted with (-) Control and exemplary system assayed using RT-qPCR. Mean±SEM (n=3) across the experimental replicates is plotted. FIG. 31E show myoblasts (07ABIC) contacted with the exemplary system showed a significant decrease in apoptosis. FIG. 31F show unmethylated and methylated DNA standards were assayed for targeted methylation using custom primers (Table 17) after enzymatic conversion of the 100ng of DNA. The % methylation is calculatedusing QUMA analysis from at least 12 clones each, and median (Q2) is plotted (top panel of FIG. 31F). The number of CpG analyzed for clone is 22. The lollipop plot indicating methylation status of each CpG in each clone are also plotted (bottom panel of FIG. 31F). Each dark circle on the plot represents methylated CpG dinucleotides. FIG. 31G depicts unmethylated and methylated DNA standards (Sigma-Aldrich) assayed for targeted methylation in the next generation sequencing assay using the primers described in Table 17 after enzymatic conversion of the 100ng of DNA. FIG.31H shows next-generation sequencing on-target methylation assay analysis down sampling results. DETAILED DESCRIPTION

[0073] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.

[0074] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.

[0075] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.

[0076] The singular terms “a,” “an,” and “the” include plural referents unless context clearly indicates otherwise. Similarly, the word “or” is intended to include “and” unless the context clearly indicates otherwise. The abbreviation, “e.g.” is used herein to indicate a non-limiting example. Thus, the abbreviation “e.g.” is synonymous with the term “for example.” Numbers provided in ranges include overlapping ranges and integers in between;for example a range of 1-4 and 5-7 includes for example, 1-7, 1-6, 1-5, 2-5, 2-7, 4-7, 1, 2, 3, 4, 5, 6 and 7.

[0077] The term “about” or “approximately” generally mean within an acceptable error range for the particular value as determined by one of ordinary skill in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” can mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5-fold, and more preferably within 2- fold, of a value. Where particular values are described in the application and claims, unless otherwise stated, the term “about” meaning within an acceptable error range for the particular value should be assumed.

[0078] The use of the alternative (e.g., “or”) should be understood to mean either one, both, or any combination thereof of the alternatives. The term “and / or” should be understood to mean either one, or both of the alternatives.

[0079] The term “cell” generally refers to a biological cell. A cell can be the basic structural, functional and / or biological unit of a living organism. A cell can originate from any organism having one or more cells. Some non-limiting examples include: a prokaryotic cell, eukaryotic cell, a bacterial cell, an archaeal cell, a cell of a single-cell eukaryotic organism, a protozoa cell, a cell from a plant (e.g. cells from plant crops, fruits, vegetables, grains, soy bean, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, hornworts, liverworts, mosses), an algal cell, (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C. Agardh, and the like), seaweeds (e.g. kelp), a fungal cell (e.g., a yeast cell, a cell from a mushroom), an animal cell, a cell from an invertebrate animal (e.g. fruit fly, cnidarian, echinoderm, nematode, etc.), a cell from a vertebrate animal (e.g., fish, amphibian, reptile, bird, mammal), a cell from a mammal (e.g., a pig, a cow, a goat, a sheep, a rodent, a rat, a mouse, a non-human primate, a human, etc.), and etcetera. Sometimes a cell is not originating from a natural organism (e.g. a cell can be a synthetically made, sometimes termed an artificial cell).

[0080] The term “nucleotide,” as used herein, generally refers to a base-sugar- phosphate combination. A nucleotide can comprise a synthetic nucleotide. A nucleotide can comprise a synthetic nucleotide analog. Nucleotides can be monomeric units of a nucleic acid sequence (e.g. deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide can include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such GHULYDWLYHV^ FDQ^ LQFOXGH^^ IRU^ H[DPSOH^^ >Į6@G$73^^ ^-deaza-dGTP and 7-deaza-dATP, and nucleotide derivatives that confer nuclease resistance on the nucleic acid molecule containing them. The term nucleotide as used herein can refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of dideoxyribonucleoside triphosphates can include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. A nucleotide may be unlabeled or detectably labeled by well-known techniques. Labeling can also be carried out with quantum dots. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels and enzyme labels. Fluorescent labels of nucleotides may include but are not limited fluorescein, 5- FDUER[\IOXRUHVFHLQ^ ^)$0^^^ ^ƍ^ƍ-dimethoxy-^ƍ^-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R^*^^^ 1^1^1ƍ^1ƍ-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-^^ƍGLPHWK\ODPLQRSKHQ\OD]R^^ EHQ]RLF^ DFLG^ (DABCYL), Cascade Blue, Oregon Green, Texas Red, Cyanine and 5-^^ƍ- aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS). Specific examples of fluorescently ODEHOHG^ QXFOHRWLGHV^ FDQ^ LQFOXGH^ >5^*@G873^^ >7$05$@G873^^ >5^^^@G&73^^ >5^*@^ G&73^^ >7$05$@^ G&73^^ >-2(@^ GG$73^^ >5^*@^ GG$73^^ >)$0@^ GG&73^^ >5^^^@GG&73^^ >7$05$@GG*73^^ >52;@GG773^^ >G5^*@GG$73^^ >G5^^^@GG&73^^ >G7$05$@GG*73^^ DQG^ >G52;@GG773^ Dvailable from Perkin Elmer, Foster City, Calif. FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X- dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein-15-dATP, Fluorescein-12-dUTP, Tetramethyl-rodamine- 6-dUTP, IR770-9-dATP, Fluorescein-12-ddUTP, Fluorescein-12-UTP, and Fluorescein-15-^ƍ- dATP available from Boehringer Mannheim, Indianapolis, Ind.; and Chromosome Labeled Nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, fluorescein-12-UTP, fluorescein-12-dUTP, Oregon Green 488-5- dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, tetramethylrhodamine-6-UTP, tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12- dUTP available from Molecular Probes, Eugene, Oreg. Nucleotides can also be labeled or marked by chemical modification. A chemically-modified single nucleotide can be biotin- dNTP. Some non-limiting examples of biotinylated dNTPs can include, biotin-dATP (e.g., bio- N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g. biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0081] The term “polynucleotide,” “oligonucleotide,” or “nucleic acid,” as used interchangeably herein, generally refers to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, either in single-, double-, or multi- stranded form. A polynucleotide can be exogenous or endogenous to a cell. A polynucleotide can exist in a cell-free environment. A polynucleotide can be a gene or fragment thereof. A polynucleotide can be DNA. A polynucleotide can be RNA. A polynucleotide can have any three dimensional structure, and can perform any function, known or unknown. A polynucleotide can comprise one or more analogs (e.g. altered backbone, sugar, or nucleobase). If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, florophores (e.g. rhodamine or fluorescein linked to the sugar), thiol containing nucleotides, biotin linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudourdine, dihydrouridine, queuosine, and wyosine. Non- limiting examples of polynucleotides include coding or non-coding regions of a gene or gene fragment, loci (locus) defined from linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short-hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA(cfRNA), nucleic acid probes, and primers. The sequence of nucleotides can be interrupted by non-nucleotide components.

[0082] The term “sequence identity” generally refers to an exact nucleotide-to- nucleotide or amino acid-to-amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Typically, techniques for determining sequence identity include determining the nucleotide sequence of a polynucleotide and / or determining the amino acid sequence encoded thereby, and comparing these sequences to a second nucleotide or amino acid sequence. Two or more sequences (polynucleotide or amino acid) can be compared by determining their “percent identity.” The percent identity of two sequences, whether nucleic acid or amino acid sequences, is the number of exact matches between two aligned sequences divided by the length of the longer sequence and multiplied by 100. Percent identity may also be determined, for example, by comparing sequence information using the advanced BLAST computer program, including version 2.2.9, available from the National Institutes of Health. The BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87:2264-2268 (1990) and as discussed in Altschul, et al., J. Mol. Biol., 215:403-410 (1990); Karlin And Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993); and Altschul et al., Nucleic Acids Res., 25:3389-3402 (1997). The program may be used to determine percent identity over the entire length of the proteins being compared. Default parameters are provided to optimize searches with short query sequences in, for example, with the blastp program. The program also allows use of an SEG filter to mask-off segments of the query sequences as determined by the SEG program of Wootton and Federhen, Computers and Chemistry 17:149-163 (1993). Ranges of desired degrees of sequence identity are approximately 50% to 100% and integer values therebetween. In general, this disclosure encompasses sequences with at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% sequence identity with any sequence provided herein.

[0083] The term “gene” generally refers to a nucleic acid (e.g., DNA such as genomic DNA and cDNA) and its corresponding nucleotide sequence that is involved in encoding an RNA transcript. The term as used herein with reference to genomic DNA includes intervening, non-FRGLQJ^UHJLRQV^DV^ZHOO^DV^UHJXODWRU\^UHJLRQV^DQG^FDQ^LQFOXGH^^ƍ^DQG^^ƍ^HQGV^^ ,Q^VRPH^XVHV^^WKH^WHUP^HQFRPSDVVHV^WKH^WUDQVFULEHG^VHTXHQFHV^^LQFOXGLQJ^^ƍ^DQG^^ƍ^XQWUDQVODWHG^UHJLRQV^^^ƍ-875^DQG^^ƍ-UTR), exons and introns. In some genes, the transcribed region will contain “open reading frames” that encode polypeptides. In some uses of the term, a “gene” comprises only the coding sequences (e.g., an “open reading frame” or “coding region”) necessary for encoding a polypeptide. In some cases, genes do not encode a polypeptide, for example, ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term “gene” includes not only the transcribed sequences, but in addition, also includes non- transcribed regions including upstream and downstream regulatory regions, enhancers and promoters. For example, a gene can refer to a portion of the gene that is near or adjacent to a transcription start site (TSS) of the gene. The gene (e.g., that is targeted as disclosed herein) can be at least, up to, at least about, or up to about 2,000 nucleobases, at least, up to, at least about, or up to about 1,800 nucleobases, at least, up to, at least about, or up to about 1,600 nucleobases, at least, up to, at least about, or up to about 1,500 nucleobases, at least, up to, at least about, or up to about 1,400 nucleobases, at least, up to, at least about, or up to about 1,200 nucleobases, at least, up to, at least about, or up to about 1,000 nucleobases, at least, up to, at least about, or up to about 900 nucleobases, at least, up to, at least about, or up to about 800 nucleobases, at least, up to, at least about, or up to about 700 nucleobases, at least, up to, at least about, or up to about 600 nucleobases, at least, up to, at least about, or up to about 500 nucleobases, at least, up to, at least about, or up to about 400 nucleobases, at least, up to, at least about, or up to about 300 nucleobases, at least, up to, at least about, or up to about 200 nucleobases, at least, up to, at least about, or up to about 100 nucleobases, or at least, up to, at least about, or up to about 50 nucleobases away from the TSS of the gene.

[0084] A gene can refer to an “endogenous gene” or a native gene in its natural location in the genome of an organism. A gene can refer to an “exogenous gene” or a non- native gene. A non-native gene can refer to a gene not normally found in the host organism but which is introduced into the host organism by gene transfer. A non-native gene can also refer to a gene not in its natural location in the genome of an organism. A non-native gene can also refer to a naturally occurring nucleic acid or polypeptide sequence that comprises mutations, insertions and / or deletions (e.g., non-native sequence).

[0085] The term “expression” generally refers to one or more processes by which a polynucleotide is transcribed from a DNA template (such as into an mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated intopeptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as “gene product.” If the polynucleotide is derived from genomic DNA, expression can include splicing of the mRNA in a eukaryotic cell. “Up-regulated,” with reference to expression, generally refers to an increased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression level in a wild-type state while “down-regulated” generally refers to a decreased expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression in a wild- type state. Expression of a transfected gene can occur transiently or stably in a cell. During “transient expression” the transfected gene is not transferred to the daughter cell during cell division. Since its expression is restricted to the transfected cell, expression of the gene is lost over time. In contrast, stable expression of a transfected gene can occur when the gene is co- transfected with another gene that confers a selection advantage to the transfected cell. Such a selection advantage may be a resistance towards a certain toxin that is presented to the cell.

[0086] The term “expression profile” generally refers to quantitative (e.g., abundance) and qualitative expression of one or more genes in a sample (e.g., a cell). The one or more genes can be expressed and ascertained in the form of a nucleic acid molecule (e.g., an mRNA or other RNA transcript). Alternatively or in addition to, the one or more genes can be expressed and ascertained in the form of a polypeptide (e.g., a protein measured via Western blot). An expression profile of a gene may be defined as a shape of an expression level of the gene over a time period (e.g., at least, up to, at least about, or up to about 1 hour, at least, up to, at least about, or up to about 2 hours, at least, up to, at least about, or up to about 3 hours, at least, up to, at least about, or up to about 4 hours, at least, up to, at least about, or up to about 5 hours, at least, up to, at least about, or up to about 6 hours, at least, up to, at least about, or up to about 7 hours, at least, up to, at least about, or up to about 8 hours, at least, up to, at least about, or up to about 9 hours, at least, up to, at least about, or up to about 10 hours, at least, up to, at least about, or up to about 11 hours, at least, up to, at least about, or up to about 12 hours, at least, up to, at least about, or up to about 16 hours, at least, up to, at least about, or up to about 18 hours, at least, up to, at least about, or up to about 24 hours, at least, up to, at least about, or up to about 36 hours, at least, up to, at least about, or up to about 48 hours, at least, up to, at least about, or up to about 3 days, at least, up to, at least about, or up to about 4 days, at least, up to, at least about, or up to about 5 days, at least, up to, at least about, or up to about6 days, at least, up to, at least about, or up to about 7 days, at least, up to, at least about, or up to about 8 days, at least, up to, at least about, or up to about 9 days, at least, up to, at least about, or up to about 10 days, at least, up to, at least about, or up to about 11 days, at least, up to, at least about, or up to about 12 days, at least, up to, at least about, or up to about 13 days, at least, up to, at least about, or up to about 14 days, etc.). Alternatively, an expression profile of a gene may be defined as an expression level of the gene at a time point of interest (e.g., the expression level of the gene measured at least, up to, at least about, or up to about 1 hour, at least, up to, at least about, or up to about 2 hours, at least, up to, at least about, or up to about 3 hours, at least, up to, at least about, or up to about 4 hours, at least, up to, at least about, or up to about 5 hours, at least, up to, at least about, or up to about 6 hours, at least, up to, at least about, or up to about 7 hours, at least, up to, at least about, or up to about 8 hours, at least, up to, at least about, or up to about 9 hours, at least, up to, at least about, or up to about 10 hours, at least, up to, at least about, or up to about 11 hours, at least, up to, at least about, or up to about 12 hours, at least, up to, at least about, or up to about 16 hours, at least, up to, at least about, or up to about 18 hours, at least, up to, at least about, or up to about 24 hours, at least, up to, at least about, or up to about 36 hours, at least, up to, at least about, or up to about 48 hours, at least, up to, at least about, or up to about 3 days, at least, up to, at least about, or up to about 4 days, at least, up to, at least about, or up to about 5 days, at least, up to, at least about, or up to about 6 days, at least, up to, at least about, or up to about 7 days, at least, up to, at least about, or up to about 8 days, at least, up to, at least about, or up to about 9 days, at least, up to, at least about, or up to about 10 days, at least, up to, at least about, or up to about 11 days, at least, up to, at least about, or up to about 12 days, at least, up to, at least about, or up to about 13 days, or at least, up to, at least about, or up to about 14 days after treating a cell to induce such expression level.)

[0087] The term “peptide,” “polypeptide,” or “protein,” as used interchangeably herein, generally refers to a polymer of at least two amino acid residues joined by peptide bond(s). This term does not connote a specific length of polymer, nor is it intended to imply or distinguish whether the peptide is produced using recombinant techniques, chemical or enzymatic synthesis, or is naturally occurring. The terms apply to naturally occurring amino acid polymers as well as amino acid polymers comprising at least one modified amino acid. In some cases, the polymer can be interrupted by non-amino acids. The terms include amino acidchains of any length, including full length proteins, and proteins with or without secondary and / or tertiary structure (e.g., domains). The terms also encompass an amino acid polymer that has been modified, for example, by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation such as conjugation with a labeling component. The terms “amino acid” and “amino acids,” as used herein, generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogues. Modified amino acids can include natural amino acids and non-natural amino acids, which have been chemically modified to include a group or a chemical moiety not naturally present on the amino acid. Amino acid analogues can refer to amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0088] The term “derivative,” “variant,” or “fragment,” as used herein with reference to a polypeptide, generally refers to a polypeptide related to a wild type polypeptide, for example either by amino acid sequence, structure (e.g., secondary and / or tertiary), activity (e.g., enzymatic activity) and / or function. Derivatives, variants and fragments of a polypeptide can comprise one or more amino acid variations (e.g., mutations, insertions, and deletions), truncations, modifications, or combinations thereof compared to a wild type polypeptide.

[0089] The term “engineered,” “chimeric,” or “recombinant,” as used herein with respect to a polypeptide molecule (e.g., a protein), generally refers to a polypeptide molecule having a heterologous amino acid sequence or an altered amino acid sequence as a result of the application of genetic engineering techniques to nucleic acids which encode the polypeptide molecule, as well as cells or organisms which express the polypeptide molecule. The term “engineered” or “recombinant,” as used herein with respect to a polynucleotide molecule (e.g., a DNA or RNA molecule), generally refers to a polynucleotide molecule having a heterologous nucleic acid sequence or an altered nucleic acid sequence as a result of the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning technologies; transfection, transformation and other gene transfer technologies; homologous recombination; site-directed mutagenesis; and gene fusion. In some cases, an engineered or recombinant polynucleotide (e.g., a genomic DNA sequence) can be modified or altered by a gene editing moiety.

[0090] The terms “engineered” and “modified” are used interchangeably herein. The terms “engineering” and “modifying” are used interchangeably herein. The terms “engineered cell” or “modified cell” are used interchangeably herein. The terms “engineered characteristic” and “modified characteristic” are used interchangeably herein.

[0091] The term “enhanced expression,” “increased expression,” or “upregulated expression” generally refers to production of a moiety of interest (e.g., a polynucleotide or a polypeptide) to a level that is above a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression can be substantially zero (or null) or higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. The moiety of interest can comprise a heterologous gene or polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced expression of the polypeptide of interest in the host strain.

[0092] The term “enhanced activity,” “increased activity,” or “upregulated activity” generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is above a normal level of activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity can be substantially zero (or null) or higher than zero. The moiety of interest can comprise a polypeptide construct of the host strain. The moiety of interest can comprise a heterologous polypeptide construct that is introduced to or into the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked-in (KI) to a genome of the host strain for enhanced activity of the polypeptide of interest in the host strain.

[0093] The term “reduced expression,” “decreased expression,” or “downregulated expression” generally refers to a production of a moiety of interest (e.g., a polynucleotide or a polypeptide) to a level that is below a normal level of expression of the moiety of interest in a host strain (e.g., a host cell). The normal level of expression is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked-out or knocked-down in the host strain. In some examples, reduced expression of the moiety of interest can include a complete inhibition of such expression in the host strain.

[0094] The term “reduced activity,” “decreased activity,” or “downregulated activity” generally refers to activity of a moiety of interest (e.g., a polynucleotide or a polypeptide) that is modified to a level that is below a normal level of activity of the moiety of interest in a host strain (e.g., a host cell). The normal level of activity is higher than zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest can be knocked-out or knocked-down in the host strain. In some examples, reduced activity of the moiety of interest can include a complete inhibition of such activity in the host strain.

[0095] The term “subject,” “individual,” or “patient,” as used interchangeably herein, generally refers to a vertebrate, preferably a mammal such as a human. Mammals include, but are not limited to, murines, simians, humans, farm animals, sport animals, and pets. Tissues, cells and their progeny of a biological entity obtained in vivo or cultured in vitro are also encompassed.

[0096] The term “treatment” or “treating” generally refers to an approach for obtaining beneficial or desired results including but not limited to a therapeutic benefit and / or a prophylactic benefit. For example, a treatment can comprise administering a system or cell population disclosed herein. By therapeutic benefit is meant any therapeutically relevant improvement in or effect on one or more diseases, conditions, or symptoms under treatment. For prophylactic benefit, a composition can be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject reporting one or more of the physiological symptoms of a disease, even though the disease, condition, or symptom may not have yet been manifested.

[0097] The term “effective amount” or “therapeutically effective amount” generally refers to the quantity of a composition, for example a composition comprising heterologous polypeptides, heterologous polynucleotides, and / or modified cells (e.g., modified stem cells), that is sufficient to result in a desired activity upon administration to a subject in need thereof. Within the context of the present disclosure, the term “therapeutically effective” generally refers to that quantity of a composition that is sufficient to delay the manifestation, arrest the progression, relieve or alleviate at least one symptom of a disorder treated by the methods of the present disclosure.

[0098] The term “muscle cell” as used herein generally refers to any cell which contributes to muscle tissue. Myoblasts, satellite cells, myotubes, and myofibril tissues are all included in the term "muscle cells". Muscle cell effects may be induced within skeletal, cardiac and smooth muscles.

[0099] Aberrant expression of one or more genes can lead to a disease or a condition. The aberrant expression can be characterized by aberrantly low expression level of the gene(s). Alternatively, the aberrant expression can be characterized by aberrantly high expression level of the gene(s). In some cases, the gene(s) can be genetically modified (e.g., via action of endonucleases, such as CRISPR-Cas enzymes) to reverse the aberrant expression (e.g., for treatment of Duchenne muscular dystrophy (DMD)). Alternatively, the aberrant expression can be transiently modified without genetically modifying such gene(s) of interest, e.g., by targeting the gene(s) with gene effectors (e.g., deactivated CRISPR-Cas enzyme that is coupled to a gene effector). Transiently modifying aberrant expression of a target gene in a cell may not be sufficient to treat or cure a disease that is manifested by the aberrant expression of the target gene. Thus, in some embodiments, the present disclosure provides systems and methods for modifying the aberrant expression of the target gene, such that the modified expression level of the target gene may be sustained for an extended period of time. MODIFICATION OF ABERRANT EXPRESSION OF A TARGET GENE

[0100] The present disclosure provides compositions, systems, and methods thereof for regulating aberrant expression of a target gene in a cell (e.g., a muscle cell). For example, the target gene can be within a D4Z4 repeat array. The target gene can encode at least a portion of DUX4. The compositions, systems, and methods disclosed herein can utilize at least a heterologous polypeptide (e.g., a heterologous actuator moiety, optionally with a heterologous polynucleotide such as a guide nucleic acid molecule) to modify an expression level and / or a epigenetic modification level (e.g., methylation level) of the target gene. For example, the compositions, systems, and methods disclosed herein can utilize a heterologous actuator moiety that is operatively coupled (e.g., covalently or non-covalently coupled) to a heterologous gene effector or regulator (e.g., gene actuator, gene repressor, etc.) to modify an expression level and / or a epigenetic modification level of the target gene. In some embodiments, the heterologous polypeptide includes a heterologous actuator moiety that isoperatively coupled (e.g., covalently or non-covalently coupled) to a heterologous gene effector or regulator (e.g., gene actuator, gene repressor, etc.). In some embodiments, the heterologous polypeptide includes a heterologous actuator moiety containing a nuclease and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide. In some embodiments, the guide nucleic acid molecule configured to form a complex with the heterologous polypeptide includes (i) a scaffold sequence configured to complex with the heterologous polypeptide, and (ii) a spacer sequence exhibiting specific binding to the target polynucleotide sequence, e.g., targeting at least a portion of DUX4, as described herein. In some embodiments, the nuclease is fused (directly or indirectly) to the heterologous gene effector or regulator (e.g., gene actuator, gene repressor, etc.). In some embodiments, the nuclease is fused at the C-terminus to the heterologous gene effector or regulator (e.g., gene actuator, gene repressor, etc.).

[0101] In some embodiments a system of the present disclosure includes a heterologous polypeptide comprising a nuclease; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell, wherein, upon formation of the complex, the complex is capable of binding the target polynucleotide sequence, to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array. In some embodiments, (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) the system is configured to effect reduced expression level of the target gene in the muscle cell by at least or at least about 80% as compared to a control cell; and / or (3) the system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off-target effect below a predetermined threshold level in themuscle cell; and / or (4) the system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the system is configured to exhibit minimal effect on at least one health indication of a subject upon administration of the system to the subject. In some embodiments, the system includes a heterologous polypeptide comprising a nuclease (e.g., Cas12f or variant thereof) containing an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs:43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising a nuclease (e.g., Cas12f or variant thereof) containing the amino acid sequence of any one of SEQ ID NOs: 43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having the sequence of any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease (e.g., Cas12f or variant thereof) operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to any one of SEQ ID NOs:43, 44, and 728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs: 43, 44, and 727; and a guide nucleic acid molecule configured to form a complex withthe heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease (e.g., Cas12f or variant thereof) operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of any one of SEQ ID NOs:43, 44, and 728, and the heterologous gene effector contains the amino acid sequence of any one of SEQ ID NOs: 15-42 and 727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease (e.g., Cas12f or variant thereof) operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of any one of SEQ ID NOs:43, 44, and 728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease (e.g., Cas12f or variant thereof) operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to a SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to any one of: SEQ ID NO: 800, SEQID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease (e.g., Cas12f or variant thereof) operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865. In any system or complex of the present disclosure, the guide nucleic acid molecule can include (i) a scaffold sequence configured to complex with the nuclease and (ii) a spacer sequence as described herein. In any system or complex of the present disclosure, in some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In any system or complex of the present disclosure, in some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease. In any system or complex of the present disclosure, in some embodiments, the heterologous gene effector is fused internally within the nuclease. In some embodiments, the heterologous polypeptide includes a Cas12f-KRAB- DNMT3L modulator, where a nuclease, Cas12f (or a variant thereof), as described herein, is operatively coupled (e.g., fused at the C-terminus, with or without a linker) to a heterologous gene effector that includes KRAB and DNMT3L, as described herein.

[0102] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 800. In some embodiments, the system includes a heterologouspolypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO: 727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 800. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0103] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 867. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 867. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0104] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 836. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 836. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0105] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO: 727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 851. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by apolynucleotide sequence of SEQ ID NO: 851. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0106] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 874. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 874. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0107] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequencehaving a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 841. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 841. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0108] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 830. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 830. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0109] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 833. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and the heterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 833. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0110] In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO:728, and the heterologous gene effector contains an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or of about 100% to SEQ ID NO: 865. In some embodiments, the system includes a heterologous polypeptide comprising: a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease contains the amino acid sequence of SEQ ID NO:728, and theheterologous gene effector contains the amino acid sequence of SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO: 865. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0111] In some cases, the cell can be a muscle cell. A muscle cell as disclosed herein can be any classification of muscle cells at any state of development. The muscle cell can comprise undifferentiated muscle cells (e.g., mononucleated cells, such as muscle stem cells, muscle satellite cells, myoblasts, etc.). Alternatively or in addition to, the muscle cell can comprise differentiated muscle cells (e.g. multinucleated muscle cells, such as myotubes). The muscle cell can be a skeletal muscle cell, a cardiac muscle cell, or a smooth muscle cell. For example, the skeletal muscle cell can be a primary myoblast (e.g., an immortalized primary myoblast cell line). In some cases, the cell can be a non-muscle cell, such as a lymphoblast.

[0112] In some cases, the target gene can be in chromosome number 4 of the cell as disclosed herein. In some cases, the target gene can be in chromosome number 10 of the cell, such as a distal portion of the q (long) arm of the chromosome number 10 of the cell.

[0113] Prior to the modification of the target gene as disclosed herein, the aberrant expression of the target gene can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is higher than that in a control cell (e.g., a healthy cell in a healthy subject) by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more. The aberrant expression can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is higher than that in a control cell (e.g., a healthy cell in a healthy subject) by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less.

[0114] Prior to the modification of the target gene as disclosed herein, the aberrant expression of the target gene can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is lower than that in a control cell (e.g., a healthy cell in a healthy subject) by at least or at least about 1%, 5%, 10%, 20%,30%, 40%, 50%, 70%, 99%, or more. The aberrant expression can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is lower than that in a control cell (e.g., a healthy cell in a healthy subject) by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.

[0115] Prior to the modification of the target gene as disclosed herein, the aberrant expression of the target gene can be characterized by a duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is longer than that in a control cell (e.g., a healthy cell in a healthy subject) by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more. The aberrant expression can be characterized by a duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is longer than that in a control cell (e.g., a healthy cell in a healthy subject) by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less.

[0116] Prior to the modification of the target gene as disclosed herein, the aberrant expression of the target gene can be characterized by a duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is shorter than that in a control cell (e.g., a healthy cell in a healthy subject) by at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more. The aberrant expression can be characterized by a duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is shorter than that in a control cell (e.g., a healthy cell in a healthy subject) by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.

[0117] Subsequent to the modification of the target gene as disclosed herein, modification of the aberrant expression of the target gene can be characterized by an increased expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more, as compared to a control (e.g., without the modification). The modification of the aberrant expression of the target gene can be characterized by an increased expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less, as compared to a control.

[0118] Subsequent to the modification of the target gene as disclosed herein, modification of the aberrant expression of the target gene can be characterized by a decreased expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more, as compared to a control (e.g., without the modification). The modification of the aberrant expression of the target gene can be characterized by a decreased expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.

[0119] Subsequent to the modification of the target gene as disclosed herein, modification of the aberrant expression of the target gene can be characterized by an increased duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more, as compared to a control (e.g., without the modification). The modification of the aberrant expression of the target gene can be characterized by an increased duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less, as compared to a control.

[0120] Subsequent to the modification of the target gene as disclosed herein, modification of the aberrant expression of the target gene can be characterized by a decreased duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more, as compared to a control (e.g., without the modification). The modification of the aberrant expression of the target gene can be characterized by a decreased duration of an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.

[0121] The system as provided herein or uses thereof can persistently modulate (e.g., reduce) expression level of at least one downstream gene of the target gene (e.g., ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, and / or RFLP2, which are downstream genes of Dux4), subsequent to the modification of the target gene, either in vitro, ex vivo, or in vivo. In some cases, the modulation of the expression level of the at least one downstream gene can be a downregulation by at least, up to, at least about, or up to about 10%, at least, upto, at least about, or up to about 20%, at least, up to, at least about, or up to about 30%, at least, up to, at least about, or up to about 40%, at least, up to, at least about, or up to about 50%, at least, up to, at least about, or up to about 60%, at least, up to, at least about, or up to about 70%, at least, up to, at least about, or up to about 75%, at least, up to, at least about, or up to about 80%, at least, up to, at least about, or up to about 85%, at least, up to, at least about, or up to about 90%, at least, up to, at least about, or up to about 95%, at least, up to, at least about, or up to about 99%, or substantially about 100%, as compared to that of a control cell lacking the system.

[0122] Subsequent to the modification of the target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the aberrantly expressed target gene) or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) can be sustained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 2 weeks, at least or at least about 3 weeks, at least or at least about 4 weeks, at least or at least about 2 months, at least or at least about 3 months, at least or at least about 4 months, at least or at least about 5 months, at least or at least about 6 months, at least or at least about 7 months, at least or at least about 8 months, at least or at least about 9 months, at least or at least about 10 months, at least or at least about 11 months, at least or at least about 12 months, at least or at least about 2 years, at least or at least about 3 years, at least or at least about 4 years, at least or at least about 5 years, or more. Subsequent to the modification of the target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the aberrantly expressed target gene) or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) can be sustained for at most or at most about 5 years, 4 years, 3 years, 2 years, 12 months, 11 months, 10 months, 9 months, 8 months, 7 months, 6 months, 5 months, 4 months, 3 months, 2 months, 4 weeks, 3 weeks, 2 weeks, at most or at most about 13 days, at most or at most about 12 days, at most or at most about 11 days, at most or at most about 10 days, at most or at most about 9 days, at most or atmost about 8 days, at most or at most about 7 days, at most or at most about 6 days, at most or at most about 5 days, at most or at most about 4 days, at most or at most about 3 days, at most or at most about 2 days, at most or at most about 1 day, or less.

[0123] Subsequent to the modification of the target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the aberrantly expressed target gene) or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) can be sustained for at least or at least about 1 cell division, at least or at least about 2 cell divisions, at least or at least about 3 cell divisions, at least or at least about 4 cell divisions, at least or at least about 5 cell divisions, at least or at least about 6 cell divisions, at least or at least about 7 cell divisions, at least or at least about 8 cell divisions, at least or at least about 9 cell divisions, at least or at least about 10 cell divisions, at least or at least about 15 cell divisions, at least or at least about 20 cell divisions, at least or at least about 25 cell divisions, at least or at least about 30 cell divisions, at least or at least about 40 cell divisions, at least or at least about 50 cell divisions, or at least or at least about 100 cell divisions. Subsequent to the modification of the target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the aberrantly expressed target gene) or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) can be sustained for at most or at most about 100 cell divisions, at most or at most about 50 cell divisions, at most or at most about 40 cell divisions, at most or at most about 30 cell divisions, at most or at most about 25 cell divisions, at most or at most about 20 cell divisions, at most or at most about 15 cell divisions, at most or at most about 10 cell divisions, at most or at most about 9 cell divisions, at most or at most about 8 cell divisions, at most or at most about 7 cell divisions, at most or at most about 6 cell divisions, at most or at most about 5 cell divisions, at most or at most about 4 cell divisions, at most or at most about 3 cell divisions, at most or at most about 2 cell division, or at most or at most about 1 cell division.

[0124] Subsequent to the modification of the target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the aberrantly expressed target gene) or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) can be measured after at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days,at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 2 weeks, at least or at least about 3 weeks, or at least or at least about 4 weeks.

[0125] As disclosed herein, non-limiting examples of the epigenetic modification can include methylation, acetylation, phosphorylation, ADP-ribosylation, glycosylation, SUMOylation, ubiquitination, modification of histone structure (e.g., via an ATP hydrolysis- dependent process). For example, the epigenetic modification can result in a modified methylation level of one or more target genes or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.).

[0126] As disclosed herein, the sustained modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (or any other gene of interest) can be characterized by maintaining at least or at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the modified expression level and / or methylation level of the target gene (or the any other gene of interest). The sustained modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (or the any other gene of interest) can be characterized by maintaining at most or at most about 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% of the modified expression level and / or methylation level of the target gene (or the any other gene of interest).

[0127] The system provided herein or uses thereof can have minimal effect (e.g., substantially no effect) on the expression profile of at least one cell type-specific gene in the target cell, which cell type-specific gene is not the target cell. For example, the system provided herein or uses thereof can have minimal change (e.g., comparable) in expression profile of at least one myogenic gene in a muscle cell, e.g., upon differentiation from a myoblast to a myotube or a myofiber. Non-limiting examples of the at least one myogenic gene can include p21, MyoD, Ezh2, Notch1, MyoG, MyHC, MEF2, Smad4, IGF2, Sirt1, myostatin, Myh2, Myh4, Myh1, myomixer, myomaker, and Mrf4. The expression profile of the at least one myogenic gene of a muscle cell comprising the system can be at least, up to, at least about, or up to about 80%, at least, up to, at least about, or up to about 85%, at least, up to, at least about, or up to about 90%, at least, up to, at least about, or up to about 91%, at least,up to, at least about, or up to about 92%, at least, up to, at least about, or up to about 93%, at least, up to, at least about, or up to about 94%, at least, up to, at least about, or up to about 95%, at least, up to, at least about, or up to about 96%, at least, up to, at least about, or up to about 97%, at least, up to, at least about, or up to about 98%, at least, up to, at least about, or up to about 99%, at least, up to, at least about, or up to about 100%, at least, up to, at least about, or up to about 105%, at least, up to, at least about, or up to about 110%, or at least, up to, at least about, or up to about 120% as compared to that of a control cell lacking the system. Subsequent to the modification of the target gene as provided herein, the minimized effect on the expression profile of the at least one cell type-specific gene can be sustained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, or at least or at least about 4 weeks. The expression profile of the at least one cell type- specific gene can be measured by various methods, such as, but not limited to, reverse transcription-polymerase chain reaction (RT-PCR or real-time quantitative RT-PCR) or Western blot.

[0128] The system provided herein or uses thereof in a cell (e.g., a diseased cell, such as FSHD muscle cell) can normalize apoptosis level of the cell, e.g., adjust expression level of at least one stress-related marker as provided herein (e.g., caspase proteins) in the cell to a level that is similar with, closer to, or comparable to a “normal” level of a control cell (e.g., a healthy muscle cell at a comparable differentiation state). Subsequent to the modification of the target gene in the cell as provided herein, the adjusted expression level of the at least one stress-related marker in the cell can be at least, up to, at least about, or up to about 80%, at least, up to, at least about, or up to about 85%, at least, up to, at least about, or up to about 90%, at least, up to, at least about, or up to about 91%, at least, up to, at least about, or up to about 92%, at least, up to, at least about, or up to about 93%, at least, up to, at least about, or up to about 94%, at least, up to, at least about, or up to about 95%, at least, up to, at least about, or up to about 96%, at least, up to, at least about, or up to about 97%, at least, up to, at least about, or up to about 98%, at least, up to, at least about, or up to about 99%, at least, up to, atleast about, or up to about 100%, at least, up to, at least about, or up to about 105%, at least, up to, at least about, or up to about 110%, or at least, up to, at least about, or up to about 120% as compared to the normal level of the control cell. The normalized expression level of the at least one stress-related marker can be sustained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, at least or at least about 4 weeks, at least or at least about 2 months, at least or at least about 6 months, or at least or at least about 1 year. Subsequent to the modification of the target gene in the cell as provided herein in a diseased cell, the expression level of the at least one stress-related marker can be less than that in a control diseased cell (e.g., lacking the system) by at least or at least about 5%, at least or at least about 10%, at least or at least about 15%, at least or at least about 20%, at least or at least about 25%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, or at least or at least about 90%.

[0129] The system provided herein or uses thereof in a cell (e.g., a diseased cell, such as FSHD muscle cell) can enhance survival of the cell in vivo. The enhanced survival can be at least or at least about 5%, at least or at least about 10%, at least or at least about 15%, at least or at least about 20%, at least or at least about 25%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, or at least or at least about 90%, as compared to that of a control cell lacking the system (e.g., a control diseased cell). The enhanced survival can be measured, for example, by duration of viability in vivo. The enhanced survival can be measured after at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, at least or at least about 4 weeks, at leastor at least about 2 months, at least or at least about 6 months, or at least or at least about 1 year subsequent to the modulation of the target gene in the cell by the system. The cell can be contacted by the system (e.g., to effect the modulation of the target gene) prior to, during, or subsequent to administration (e.g., transplantation to a muscle tissue) to a subject in need thereof. Alternatively, the system can be administered to the subject in need thereof, to contact the cell in vivo to effect the modulation of the target gene or any other gene of interest (e.g., downstream gene(s) of the target gene, cell type-specific gene(s), etc.) in the cell in vivo.

[0130] The system provided herein or uses thereof in a cell (e.g., a diseased cell, such as FSHD muscle cell) can have minimal effect (e.g., substantially no effect) on one or more health indications of a subject. For example, one or more health indications can be measured subsequent to administration of the system to the subject, to effect modification of the target gene in a cell in the subject. Non-limiting examples of the health conditions can include food intake (e.g., grams per subject per day), body weight (e.g., grams), body dimension (e.g., length, height, or circumference in meters), body mass index (e.g., grams per centimeter squared), body fat (e.g., milligrams of fat per grams of body weight), organ weight (e.g., grams of liver, kidney, spleen, small intestine, etc. per grams of body weight), clinical chemistry, tissue function or integrity (e.g., structure as ascertained by biopsy and / or histology), body injury, distress, grooming, etc. Clinical chemistry can be ascertained by measuring one or more markers from a bodily fluid (e.g., blood), and non-limiting examples of the marker(s) for clinical chemistry can include electrolytes (e.g., sodium, potassium, chloride, bicarbonate), kidney markers (e.g., creatinine, blood urea nitrogen), liver function markers (e.g., albumin, globulin, albumin / globulin ratio, bilirubin, aspartate transaminase (AST), alanine transaminase (ALT), gamma-glutamyl transpeptidase (GGT), alkaline phosphatase (ALP)), cardiac markers (e.g., H-FABP, troponin, myoglobin, CK-MB, B-type natriuretic peptide (BNP)), minerals (e.g., calcium, magnesium, phosphate, potassium), blood disorder markers (e.g., iron, transferrin, TIBC, vitamin B12, vitamin D, folic acid), and others (e.g., glucose, C-reactive protein, glycated hemoglobin, uric acid, arterial blood gases, adrenocorticotropic hormone (ACTH), neuron-specific enolase (NSE), fecal occult blood test (FOBT), etc.). For example, a panel of clinical chemistry markers can be used, including one or more members consisting of sodium, potassium, chloride, bicarbonate, blood urea nitrogen (BUN), creatinine, glucose, calcium, total protein, albumin, alkaline phosphatase (ALP),alanine amino transferase (ALT), aspartate amino transferase (AST), or bilirubin. The health indication(s) of the subject receiving the system or method as provided herein can be at least, up to, at least about, or at least about 80%, at least, up to, at least about, or at least about 85%, at least, up to, at least about, or at least about 90%, at least, up to, at least about, or at least about 91%, at least, up to, at least about, or at least about 92%, at least, up to, at least about, or at least about 93%, at least, up to, at least about, or at least about 94%, at least, up to, at least about, or at least about 95%, at least, up to, at least about, or at least about 96%, at least, up to, at least about, or at least about 97%, at least, up to, at least about, or at least about 98%, at least, up to, at least about, or at least about 99%, at least, up to, at least about, or at least about 100%, at least, up to, at least about, or at least about 105%, at least, up to, at least about, or at least about 110%, or at least, up to, at least about, or at least about 120% as compared to that of a control subject without such treatment. Subsequent to the treatment, the minimized effect on the health condition(s) can be sustained for (and / or measured on) at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, or at least or at least about 4 weeks.

[0131] The systems, compositions, and methods as disclosed herein can be used to treat, inhibit, or ameliorate a disease (e.g., muscular dystrophy, such as Facioscapulohumeral Muscular Dystrophy (FSHD)) of a subject. HETEROLOGOUS POLYPEPTIDES

[0132] The heterologous polypeptide as disclosed herein, either alone or in conjunction with one or more co-agents such as a heterologous polynucleotide (e.g., a guide nucleic acid) can be configured to specifically bind a target polynucleotide sequence, to modulate an expression level and / or an epigenetic level of the target gene (e.g., the D4Z4 repeat array) in the target cell, as disclosed herein. The target polynucleotide sequence can be at (e.g., within) the target gene. Alternatively, the target polynucleotide sequence can be adjacent to the target gene. For example, the target polynucleotide sequence can be adjacent to an end(e.g., a 5’ end or a 3’ end) of the target gene. The target polynucleotide sequence can be at least or at least about 5 nucleobases, at least or at least about 10 nucleobases, at least or at least about 20 nucleobases, at least or at least about 30 nucleobases, at least or at least about 40 nucleobases, at least or at least about 50 nucleobases, at least or at least about 100 nucleobases, at least or at least about 150 nucleobases, at least or at least about 200 nucleobases, at least or at least about 250 nucleobases, at least or at least about 300 nucleobases, at least or at least about 400 nucleobases, at least or at least about 500 nucleobases, at least or at least about 1,000 nucleobases, at least or at least about 1,500 nucleobases, at least or at least about 2,000 nucleobases, at least or at least about 3,000 nucleobases, at least or at least about 4,000 nucleobases, or at least or at least about 5,000 nucleobases away from the end of the target gene. The target polynucleotide sequence can be at most or at most about 5,000 nucleobases, at most or at most about 4,000 nucleobases, at most or at most about 3,000 nucleobases, at most or at most about 2,000 nucleobases, at most or at most about 1,500 nucleobases, at most or at most about 1,000 nucleobases, at most or at most about 500 nucleobases, at most or at most about 400 nucleobases, at most or at most about 300 nucleobases, at most or at most about 200 nucleobases, at most or at most about 150 nucleobases, at most or at most about 100 nucleobases, at most or at most about 50 nucleobases, at most or at most about 40 nucleobases, at most or at most about 30 nucleobases, at most or at most about 20 nucleobases, at most or at most about 10 nucleobases, or at most or at most about 5 nucleobases away from the end of the target gene.

[0133] Without wishing to be bound by theory, when the target polynucleotide sequence is not within the target gene, the target polynucleotide sequence can interact (e.g., via direct or indirect binding) with at least a portion of the target gene (e.g., a promoter sequence of the target gene), such that binding or targeting of the target polynucleotide sequence by at least the heterologous polypeptide (e.g., by a complex comprising the heterologous polypeptide and the heterologous polynucleotide as disclosed herein) can target the at least the portion of the target gene (e.g., the promoter sequence), to effect the modulation of the expression level and / or the epigenetic level of the target gene in the cell.

[0134] In some cases, the target polynucleotide sequence can comprise a plurality of target polynucleotide sequences. The plurality of target polynucleotide sequences can be within the target gene. Alternatively, the plurality of target polynucleotide sequences can beoutside but adjacent to the target gene, as disclosed herein. Yet in another alternative, the plurality of target polynucleotide sequences can comprise at least one target polynucleotide sequence within the target gene (e.g., within the D4Z4 repeat domain) and at least one additional target polynucleotide sequence that is outside of but adjacent to the target gene. In such case, targeting both the at least one target polynucleotide sequence and the at least one additional target polynucleotide sequence may yield a greater effect (e.g., greater degree of modulation of the expression and / or epigenetic level of the target gene) (e.g., by at least 0.1- fold, at least 0.5-fold, at least 1-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5- fold, at least 10-fold, at least 15-fold, at least 20-fold, or more) as compared to that by targeting only one of the at least one target polynucleotide sequence and the at least one additional target polynucleotide sequence.

[0135] The heterologous polypeptide as disclosed herein can comprise one or more heterologous gene effectors (e.g., gene effectors that are heterologous to a cell comprising the gene effectors and / or another component in a complex of the disclosure). Heterologous gene effectors can comprise domains that are capable of, or are candidates for, modulating expression of a target gene (e.g., a target endogenous gene), for example, activating, repressing, upregulating, downregulating, or stabilizing an expression level or activity level of the gene. Heterologous gene effectors can be heterologous with respect to another component that is present in a complex, for example, a guide moiety (e.g., nuclease and / or guide nucleic acid, as disclosed herein). In some cases, heterologous gene effectors can be heterologous with respect to a host cell they are introduced to.

[0136] A heterologous gene effector can be or can comprise a sequence from any suitable source, for example, an amino acid sequence from a human protein, viral protein, or other protein as disclosed herein. A heterologous gene effector can be or can comprise a sequence from a protein that primarily localized to the nucleus, for example, a member of the human nuclear proteome. A heterologous gene effector can be or can comprise one or more natural amino acid residues. A heterologous gene effector can be or can comprise one or more synthetic amino acid residues.

[0137] A heterologous gene effector can be or can comprise a sequence from a mammalian protein. A heterologous gene effector can be or can comprise a sequence from a human protein. A heterologous gene effector can be or can comprise a sequence from a viralprotein. A heterologous gene effector can be or can comprise a sequence from a non-human primate protein. A heterologous gene effector can be or can comprise a sequence from a non- human mammal protein. A heterologous gene effector can be or can comprise a sequence from a non-rodent mammal protein. A heterologous gene effector can be or can comprise a sequence from a plant protein. A heterologous gene effector can be or can comprise a sequence from a pig protein. A heterologous gene effector can be or can comprise a sequence from a lagomorph protein. A heterologous gene effector can be or can comprise a sequence from a canine protein. A heterologous gene effector can be or can comprise a sequence from an avian protein. A heterologous gene effector can be or can comprise a sequence from a reptilian protein. A heterologous gene effector can be or can comprise a sequence from a bacterial protein. A heterologous gene effector can be or can comprise a sequence from an archaeal protein.

[0138] For example, the amino acid sequence of the heterologous gene effector as disclosed herein may not and need not be derived from a bacterial protein (e.g., may be derived from an archaeal protein). Without wishing to be bound by theory, a subject in need thereof may be treated with a composition comprising the non-bacterial protein-derived heterologous gene effector, such that the composition may not (i) induce a bacterial stimulus in the subject and / or (ii) elicit a bacterial immune response in the subject.

[0139] The heterologous actuator moiety can comprise a nuclease (e.g., an endonuclease). For example, the nuclease can be a CRISPR / Cas protein. The nuclease can have a length that is less than a threshold length. The threshold length can be at most or at most about 1,000 amino acids, at most or at most about 950 amino acids, at most or at most about 900 amino acids, at most or at most about 850 amino acids, at most or at most about 800 amino acids, at most or at most about 750 amino acids, at most or at most about 700 amino acids, at most or at most about 650 amino acids, at most or at most about 600 amino acids, at most or at most about 550 amino acids, at most or at most about 500 amino acids, at most or at most about 450 amino acids, at most or at most about 400 amino acids, at most or at most about 350 amino acids, or at most or at most about 300 amino acids. The threshold length can be at least or at least about 300 amino acids, at least or at least about 350 amino acids, at least or at least about 400 amino acids, at least or at least about 450 amino acids, at least or at least about 500 amino acids, at least or at least about 550 amino acids, at least or at least about 600 amino acids, at least or at least about 650 amino acids, at least or at least about 700 aminoacids, at least or at least about 750 amino acids, at least or at least about 800 amino acids, at least or at least about 850 amino acids, at least or at least about 900 amino acids, at least or at least about 950 amino acids, or at least or at least about 1,000 amino acids.

[0140] Without wishing to be bound by theory, using a size of the nuclease to be less than the threshold length can have one or advantages over using a control nuclease having a size greater than the threshold length. When using a delivery vehicle having a limited size (e.g., a limited physical size to entrap the nuclease or a limited expression cassette size, such as a viral genome) can leave sufficient room (or sufficient space within the expression cassette) for one or more co-agents, such as one or more gene regulators (e.g., transcriptional regulator) and / or one or more heterologous polynucleotides (e.g., one or more guide nucleic acid molecules). Alternatively or in addition to, using the nuclease having a size less than or equal to the threshold size can elicit a greater effect on the modulation of the expression level and / or the epigenetic level of the target gene, as compared to the effect on the modulation of the expression level and / or the epigenetic level of the target gene by a control nuclease having a size greater than the threshold size.

[0141] In some examples, the degree of modulation (e.g., increase or decrease) of the expression level and / or the epigenetic level of the target gene by the nuclease as disclosed herein (e.g., having a size less than or equal to the threshold size) can be greater than that by the control nuclease (e.g., having a size greater than the threshold size) by at least, up to, at least about, or up to about 0.1-fold, at least, up to, at least about, or up to about 0.5-fold, at least, up to, at least about, or up to about 1-fold, at least, up to, at least about, or up to about 2- fold, at least, up to, at least about, or up to about 3-fold, at least, up to, at least about, or up to about 4-fold, at least, up to, at least about, or up to about 5-fold, at least, up to, at least about, or up to about 6-fold, at least, up to, at least about, or up to about 7-fold, at least, up to, at least about, or up to about 8-fold, at least, up to, at least about, or up to about 9-fold, at least, up to, at least about, or up to about 10-fold, at least, up to, at least about, or up to about 15-fold, at least, up to, at least about, or up to about 20-fold, at least, up to, at least about, or up to about 25-fold, at least, up to, at least about, or up to about 30-fold, at least, up to, at least about, or up to about 40-fold, at least, up to, at least about, or up to about 50-fold, at least, up to, at least about, or up to about 60-fold, at least, up to, at least about, or up to about 70-fold, at least, upto, at least about, or up to about 80-fold, at least, up to, at least about, or up to about 90-fold, or at least, up to, at least about, or up to about 100-fold.

[0142] In some examples, the degree of modulation (e.g., increase or decrease) of the expression level and / or the epigenetic level of the target gene by the nuclease as disclosed herein (e.g., having a size less than or equal to the threshold size) can persist longer than that by the control nuclease (e.g., having a size greater than the threshold size) by at least, up to, at least about, or up to about 0.1-fold, at least, up to, at least about, or up to about 0.5-fold, at least, up to, at least about, or up to about 1-fold, at least, up to, at least about, or up to about 2- fold, at least, up to, at least about, or up to about 3-fold, at least, up to, at least about, or up to about 4-fold, at least, up to, at least about, or up to about 5-fold, at least, up to, at least about, or up to about 6-fold, at least, up to, at least about, or up to about 7-fold, at least, up to, at least about, or up to about 8-fold, at least, up to, at least about, or up to about 9-fold, at least, up to, at least about, or up to about 10-fold, at least, up to, at least about, or up to about 15-fold, at least, up to, at least about, or up to about 20-fold, at least, up to, at least about, or up to about 25-fold, at least, up to, at least about, or up to about 30-fold, at least, up to, at least about, or up to about 40-fold, at least, up to, at least about, or up to about 50-fold, at least, up to, at least about, or up to about 60-fold, at least, up to, at least about, or up to about 70-fold, at least, up to, at least about, or up to about 80-fold, at least, up to, at least about, or up to about 90-fold, or at least, up to, at least about, or up to about 100-fold.

[0143] In some examples, the degree of modulation (e.g., increase or decrease) of the expression level and / or the epigenetic level of the target gene by the nuclease as disclosed herein (e.g., having a size less than or equal to the threshold size) can persist longer than (or sustained longer than) that by the control nuclease (e.g., having a size greater than the threshold size) by at least, up to, at least about, or up to about 1 cell division, at least, up to, at least about, or up to about 2 cell divisions, at least, up to, at least about, or up to about 3 cell divisions, at least, up to, at least about, or up to about 4 cell divisions, at least, up to, at least about, or up to about 5 cell divisions, at least, up to, at least about, or up to about 6 cell divisions, at least, up to, at least about, or up to about 7 cell divisions, at least, up to, at least about, or up to about 8 cell divisions, at least, up to, at least about, or up to about 9 cell divisions, at least, up to, at least about, or up to about 10 cell divisions, at least, up to, at least about, or up to about 11 cell divisions, at least, up to, at least about, or up to about 12 cell divisions, at least, up to, at leastabout, or up to about 13 cell divisions, at least, up to, at least about, or up to about 14 cell divisions, at least, up to, at least about, or up to about 15 cell divisions, at least, up to, at least about, or up to about 16 cell divisions, at least, up to, at least about, or up to about 17 cell divisions, at least, up to, at least about, or up to about 18 cell divisions, at least, up to, at least about, or up to about 19 cell divisions, at least, up to, at least about, or up to about 20 cell divisions, at least, up to, at least about, or up to about 25 cell divisions, at least, up to, at least about, or up to about 30 cell divisions, at least, up to, at least about, or up to about 40 cell divisions, at least, up to, at least about, or up to about 50 cell divisions, or at least about 100 cell divisions. As disclosed herein, a cell division can be characterized by a division of a parent cell into two daughter cells with substantially the same genetic material as the parent cell.

[0144] A heterologous gene effector can be or can comprise a sequence from a chromatic regulator (CR). Chromatin regulators include functional domains from various classes of histone and DNA modifying enzymes (e.g., DNMTs, HATs, or HMTs, etc.). In some embodiments, the heterologous gene effector is a DNMT that includes DNMT-A or DNMT-L. In some embodiments, “DNMT-L” is or includes DNMT3L. In some embodiments, “DNMT-A” is or includes DNMT3A. In some embodiments, the heterologous gene effector is a DNMT that includes DNMT3A or DNMT3L.

[0145] A heterologous gene effector can comprise two or more domains from chromatin regulators, e.g., located at a C-terminus, an N-terminus, or within a polypeptide sequence, in tandem or separate.

[0146] In some embodiments, a heterologous gene effector facilitates heterochromatin formation. Non-limiting examples of proteins that can facilitate KHWHURFKURPDWLQ^IRUPDWLRQ^LQFOXGH^+3^Į^^+3^ȕ^^.$3^^^.5$%^^689^^+^^^or G9a.

[0147] In some embodiments, a heterologous gene effector modulates histones through methylation. In some embodiments, a heterologous gene effector modulates histones through acetylation. In some embodiments, a heterologous gene effector modulates histones through phosphorylation. In some embodiments, a heterologous gene effector modulates histones through ADP-ribosylation. In some embodiments, a heterologous gene effector modulates histones through glycosylation. In some embodiments, a heterologous gene effector modulates histones through SUMOylation. In some embodiments, a heterologous gene effector modulates histones through ubiquitination. In some embodiments, a heterologous gene effectormodulates histones by remodeling histone structure, e.g., via an ATP hydrolysis-dependent process.

[0148] In some embodiments, a heterologous gene effector facilitates spatial positioning of proteins on or near the target polynucleotide, e.g., transcriptional repressors, transcription factors, or histones, etc. In some embodiments, a heterologous gene effector is useful for manipulating the spatiotemporal organization of genomic DNA and RNA components in the nucleus and / or cytoplasm, e.g., for regulating diverse cellular functions.

[0149] In some embodiments, a heterologous gene effector is from a histone acetyltransferase. Non-limiting examples of histone acetyltransferases include GNAT subfamily, MYST subfamily, p300 / CBP subfamily, HAT1 subfamily, GCN5, PCAF, Tip60, MOZ, MORF, MOF, HBO1, p300, CBP, HAT1, ATF-2, SRC1, or TAFII250.

[0150] In some embodiments, a heterologous gene effector is from a histone lysine methyltransferase. Non-limiting examples of histone lysine methyltransferases include EZH subfamily, Non-SET subfamily, Other SET subfamily, PRDM subfamily, SET1 subfamily, SET2 subfamily, SUV39 subfamily, SYMD subfamily, ASH1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, NSD2, NSD3, PRDM1, PRDM10, PRDM11, PRDM12, PRDM13, PRDM14, PRDM15, PRDM16, PRDM2, PRDM4, PRDM5, PRDM6, PRDM7, PRDM8, PRDM9, SET1, SET1L, SET2L, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETDB1, SETDB2, SETMAR, SUV39H1, SUV39H2, SUV420H1, SUV420H2, SYMD1, SYMD2, SYMD3, SYMD4, or SYMD5.

[0151] In some embodiments, a heterologous gene effector is from a component of a chromatin remodeling complex. In some embodiments, a heterologous gene effector is a component of BAF, for example, Actin, ARIDA / B, BAF155, BAF170, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRG1 / BRM, INI1, or SS18.

[0152] In some embodiments, a heterologous gene effector is from a component of PBAF, for example, Actin, ARID2, BAF155, BAF170, BAF180, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRD7, BRG1, or INI1.

[0153] In some embodiments, a heterologous gene effector is from a component of an ISWI family chromatin remodeling complex, for example, ACF subfamily, RSF subfamily, CERF subfamily, CHRAC subfamily, NURF subfamily, NoRC subfamily, WICH subfamily, b-WICH subfamily, ACF1, ATPase, BPTF, CECR2, CHRAC15, CHRAC17, CSB, DEK,MYBBP1A, NM1, RBAP46 / 48, RHII / Gua, RSF1, SAP155, SNF2H, SNF2H / L, SNF2L, TIP5, or WSTF.

[0154] In some embodiments, a heterologous gene effector is from a component of a CHD family complex, for example, a NuRD complex, NuRD-like complex, or CHD complex. In some embodiments, a heterologous gene effector is from CHD1 / 2 / 6 / 7 / 8 / 9, CHD3 / 4, CHD5, GATAD2 A / B, GATAD2 B, HDAC1, HDAC2, HDAC2, MBD2 / 3, MTA1 / 2 / 3, MTA3, or RBAP46, RBAP46 / 48.

[0155] In some embodiments, a heterologous gene effector is from a component of an INO80 family complex, for example, from an INO80 complex, Tip60 / p400 complex, SRCAP complex, AMIDA, ARP6, BAF53, BAF53, BAF53A, BRD8, DMAP1, DMAP1, EPC1 / 2, FLJ11730, GAS41, GAS41, IES2, IES6, ING3, INO80, INO80E, MCRS1, MRG15, MRGBP, MRGX, NFRKB, p400, RUVBL1 / 2, RUVBL1 / 2, RUVBL1 / 2, SRCAP, Tip60, TRRAP, UCH37, YL-1, YL-1, YY1, or ZnF-HIT1.

[0156] A heterologous gene effector can be or can comprise a sequence from a transcriptional regulator (TR). TR gene effectors include transcriptional regulatory domains from various families of transcription factors (e.g. KRAB, p65, MED, or GTFs, etc.).

[0157] A heterologous gene effector can comprise a transcriptional activator domain. A heterologous gene effector can comprise two or more tandem transcriptional activation domains, e.g., located at a C-terminus, an N-terminus, or within a polypeptide sequence.

[0158] Non-limiting examples of transcriptional activation domains include GAL4, herpes simplex activation domain VP16, VP64 (a Tetramer of the herpes simplex activation domain VP16), NF-KB p65 subunit, or Epstein-Barr virus R transactivator (Rta). In some embodiments, such transcriptional activation domains are used as controls in methods of the disclosure. In some embodiments, such transcriptional activation domains are used as one heterologous gene effector in a complex that comprises at least one additional heterologous gene effector (e.g., a different effector).

[0159] A heterologous gene effector can comprise a transcriptional repressor domain. A heterologous gene effector can comprise two or more transcriptional repressor domains, e.g., located at a C-terminus, an N-terminus, or within a polypeptide sequence, in tandem or separate.

[0160] Non-limiting examples of transcriptional repressor domains include the KRAB (Kruppel-associated box) domain of Koxl, the Mad mSIN3 interaction domain (SID), or ERF repressor domain (ERD). In some embodiments, such transcriptional repressor domains are used as controls in methods of the disclosure. In some embodiments, such transcriptional repressor domains are used as one heterologous gene effector in a complex that comprises at least one additional heterologous gene effector (e.g., a different effector).

[0161] In some embodiments, a heterologous gene effector is from a gene product that is a transcription factor.

[0162] In some embodiments, a heterologous gene effector is from a gene product that is a hematopoietic stem cell transcription factor. Non-limiting examples of hematopoietic stem cell transcription factors include AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1 alpha / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT Activators, STAT Inhibitors, STAT3, STAT4, STAT5a, STAT6, or TSC22.

[0163] In some embodiments, a heterologous gene effector is from a gene product that is a mesenchymal stem cell transcription factor. Non-limiting examples of mesenchymal stem cell transcription factors include DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF- 3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, Myocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT Activators, STAT Inhibitors, STAT1, STAT3, TBX18, Twist-1, or Twist-2.

[0164] In some embodiments, a heterologous gene effector is from a gene product that is an embryonic stem cell transcription factor. Non-limiting examples of embryonic stem cell transcription factors include Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3 alpha / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB Activators, NFkB / IkB Inhibitors, NFkB1, NFkB2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8,Snail, SOX2, SOX7, SOX15, SOX17, STAT Activators, STAT Inhibitors, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, or ZNF281.

[0165] In some embodiments, a heterologous gene effector is from a gene product that is an induced pluripotent stem cell (iPSC) transcription factor. Non-limiting examples of iPSC transcription factors include KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, or TBX18.

[0166] In some embodiments, a heterologous gene effector is from a gene product that is an epithelial stem cell transcription factor. Non-limiting examples of epithelial stem cell transcription factors include ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4 alpha / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT Activators, STAT Inhibitors, STAT3, SUZ12, TCF-3 / E2A, or TCF7 / TCF1.

[0167] In some embodiments, a heterologous gene effector is from a gene product that is a cancer stem cell transcription factor. Non-limiting examples of cancer stem cell transcription factors include Androgen R / NR3C4, AP-2 gamma, beta-Catenin, beta-Catenin Inhibitors, Brachyury, CREB, ER alpha / NR3A1, ER beta / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI-2, GLI-3, HIF-1 alpha / HIF1A, HIF-2 alpha / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB Activators, NFkB / IkB Inhibitors, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT Activators, STAT Inhibitors, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, or ZEB1.

[0168] In some embodiments, a heterologous gene effector is from a gene product that is a cancer-related transcription factor. Non-limiting examples of cancer-related transcription factors include ASCL1 / Mash1, ASCL2 / Mash2, ATF1, ATF2, ATF4, BLIMP1 / PRDM1, CDX2, CDX4, DLX5, DNMT1, E2F-1, EGR1, ELF3, Ets-1, FosB / G0S3, FoxC1, FoxC2, FoxF1, GADD153, GATA-2, HMGA2, HMGB1 / HMG-1, HNF-3 alpha / FoxA1, HNF-6 / ONECUT1, HSF1, ID1, ID2, JunD, KLF10, KLF12, KLF17, LMO2, MEF2C, MYCL1 / L-Myc, NFkB2, Oct-1, p63 / TP73L, Pax3, PITX2, Prox1, RAP80, Rex- 1 / ZFP42, RUNX1 / CBFA2, RUNX3 / CBFA3, SALL4, SCL / Tal1, Sirtuin 2 / SIRT2, Smad3,Smad4, Smad5, SOX11, STAT5a / b, STAT5a, STAT5b, TCF7 / TCF1, TORC1, TORC2, TRIM32, TRPS1, or TSC22.

[0169] In some embodiments, a heterologous gene effector is from a gene product that is an immune cell transcription factor. Non-limiting examples of immune cell transcription factors include AP-1, Bcl6, E2A, EBF, Eomes, FoxP3, GATA3, Id2, Ikaros, IRF, IRF1, IRF2, IRF3, IRF3, IRF7, NFAT, NFkB, Pax5, PLZF, PU.1, ROR-gamma-T, STAT, STAT1, STAT2, STAT3, STAT4, STAT5, STAT5A, STAT5B, STAT6, T-bet, TCF7, or ThPOK.

[0170] In some embodiments, a heterologous gene effector is from a gene product that is a RNA polymerase related protein. In some embodiments, a heterologous gene effector is from a transcription factor with a basic domain. In some embodiments, a heterologous gene effector is from a transcription factor with a zinc-coordinated DNA binding domain. In some embodiments, a heterologous gene effector is from a transcription factor with a helix-turn- helix domain. In some embodiments, a heterologous gene effector is from a transcription factor with an alpha helical DNA binding domain. In some embodiments, a heterologous gene effector is from a transcription factor with an alpha helix exposed by beta structures. In some embodiments, a heterologous gene effector is from a transcription factor with an immunoglobulin fold. In some embodiments, a heterologous gene effector is from a transcription factor with a with a beta-Hairpin exposed by an alpha / beta-scaffold. In some embodiments, a heterologous gene effector is from a transcription factor with a beta sheet binding to DNA. In some embodiments, a heterologous gene effector is from a transcription factor with a beta barrel DNA binding domain.

[0171] In some embodiments, a heterologous gene effector is from a gene product that is a nuclear receptor, for example, a nuclear hormone receptor. Non-limiting examples of nuclear hormone receptors include those encoded by NR0B1, NR0B2, NR1A1, NR1A2, NR1B1, NR1B2, NR1B3, NR1C1, NR1C2, NR1C3, NR1D1, NR1D2, NR1F1, NR1F2, NR1F3, NR1H4, NR1H5, NR1H3, NR1H2, NR1I1, NR1I2, NR1I3, NR2A1, NR2A2, NR2B1, NR2B2, NR2B3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3A1, NR3A2, NR3B1, NR3B2, NR3B3, NR3C4, NR3C1, NR3C2, NR3C3, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, or NR6A1.

[0172] In some embodiments, a heterologous gene effector is from a gene product that is involved in nucleosome assembly. In some embodiments, a heterologous gene effectoris from a gene product that is involved in DNA metabolism. In some embodiments, a heterologous gene effector is from a gene product that is involved in nucleotide metabolism. In some embodiments, a heterologous gene effector is from a gene product that is involved in ribosome biogenesis. In some embodiments, a heterologous gene effector is from a gene product that is involved in protein folding. In some embodiments, a heterologous gene effector is from a gene product that is involved in translation. In some embodiments, a heterologous gene effector is from a gene product that is involved in signaling. In some embodiments, a heterologous gene effector is from a gene product that is involved in proteolysis. In some embodiments, a heterologous gene effector is from a gene product that is involved in negative regulation of endopeptidase activity.

[0173] In some embodiments, a heterologous gene effector or gene regulator, as used interchangeably herein, can comprise a polypeptide sequence that exhibits at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity to any of the heterologous gene effector amino acid sequences provided in Table 3.

[0174] Table 3. Heterologous gene effector amino acid sequences. SEQ ID Heterologous gene effector amino acid sequences NO: 15 NNSQGRVTFEDVTVNFTQGEWQRLNPEQRNLYRDVMLENYSNLVSVGQGETT KPDVILRLEQGKEPWLEEEEVLGSGRAEKNGDI 16 SGHPGSWEMNSVAFEDVAVNFTQEEWALLDPSQKNLYRDVMQETFRNLASIG NKGEDQSIEDQYKNSSRNLRHIISHSGNNPYGC 17 AAATLRTPTQGTVTFEDVAVHFSWEEWGLLDEAQRCLYRDVMLENLALLTSL DVHHQKQHLGEKHFRSNVGRALFVKTCTFHVSG 18 TTFKEAMTFKDVAVVFTEEELGLLDLAQRKLYRDVMLENFRNLLSVGHQAFH RDTFHFLREEKIWMMKTAIQREGNSGDKIQTEM 19 VPAETSSSGLLEEQKMMKSQGLVSFKDVAVDFTQEEWQQLDPSQRTLYRDVM LENYSHLVSMGYPVSKPDVISKLEQGEEPWIIK 20 MKSQGLVSFKDVAVDFTQEEWQQLDPSQRTLYRDVMLENYSHLVSMGYPVSK PDVISKLEQGEEPWIIKGDISNWIYPDEYQADGSEQ ID Heterologous gene effector amino acid sequences NO: 21 AEGSVMFSDVSIDFSQEEWDCLDPVQRDLYRDVMLENYGNLVSMGLYTPKPQ VISLLEQGKEPWMVGRELTRGLCSDLESMCETK 22 AAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYRN VMLENFTLLASLGKVLTPHPSILSWARLFLLFL 23 AAAALRDPAQVPVAADLLTDHEEQGYVTFEDVAVYFSQEEWRLLDDAQRLLY RNVMLENFTLLASLGLASSKTHEITQLESWEEP 24 AAAALRDPAQVPVAADLLTDHEEQGYVTFEDVAVYFSQEEWRLLDDAQRLLY RNVMLENFTLLASLGCWHGAEAEEAPEQIASVG 25 AAAALRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLAS LGCWHGAEAEEAPEQIASVGLLSSNIQQHQKQH 26 AAAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYR NVMLENFTLLASLGKVLTPHPSILSWARLFLLF 27 YVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEIT QLESWEEPFMPAWEVVTSAIPRGSWWVELREV 28 AAAALRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASL GLASSKTHEITQLESWEEPFMPAWEVVTSAIPR 29 AAAALRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASL GCWHGAEAEEAPEQIASVGLLSSNIQQHQKQHC 30 AAAALRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLAS LGLASSKTHEITQLESWEEPFMPAWEVVTSAIP 31 AAAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYR NVMLENFTLLASLGLASSKTHEITQLESWEEPF 32 VTCAHLGRRARLPAAQPSACPGTCFSQEERMAAGYLPRWSQELVTFEDVSMD FSQEEWELLEPAQKNLYREVMLENYRNVVSLEA 33 LVTFEDVSMDFSQEEWELLEPAQKNLYREVMLENYRNVVSLEALKNQCTDVG IKEGPLSPAQTSQVTSLSSWTGYLLFQPVASSH 34 KNATIVMSVRREQGSSSGEGSLSFEDVAVGFTREEWQFLDQSQKVLYKEVML ENYINLVSIGYRGTKPDSLFKLEQGEPPGIAEG 35 SSGEGSLSFEDVAVGFTREEWQFLDQSQKVLYKEVMLENYINLVSIGYRGTK PDSLFKLEQGEPPGIAEGAAHSQICPDADFLE 36 GPLQFRDVAIEFSLEEWHCLDTAQRNLYRNVMLENYSNLVFLGITVSKPDLI TCLEQGRKPLTMKRNEMIAKPSVSFLQVHSESQ 37 GPLQFRDVAIEFSLEEWHCLDTAQRNLYRNVMLENYSNLVFLGITVSKPDLI TCLEQGRKPLTMKRNEMIAKPSVMCSHFAQDLW 38 APPSAPLPAQGPGKARPSRKRGRRPRALKFVDVAVYFSPEEWGCLRPAQRAL YRDVMRETYGHLGALGCAGPKPALISWLERNTD 39 QTNTKDWTVTPEHVLPESQSLLTFEEVAMYFSQEEWELLDPTQKALYNDVMQ ENYETVISLALFVLPKPKVISCLEQGEEPWVQV 40 AAATLRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLAS LGLASSKTHEITQLESWEEPFMPAWEVVTSAIL 41 AAATLRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASL GLASSKTHEITQLESWEEPFMPAWEVVTSAILRSEQ ID Heterologous gene effector amino acid sequences NO: 42 DSVAFEDVAVNFTQEEWALLDPSQKNLYREVMQETLRNLTSIGKKWNNQYIE DEHQNPRRNLRRLIGERLSESKESHQHGEVLTQ

[0175] In some embodiments, a heterologous gene effector or gene regulator, as used interchangeably herein, can comprise a polypeptide sequence that exhibits at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity to SEQ ID NO:727, shown below: RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQ LTKPDVILRLEKGEEPGKESGSVGGSGGSSEQLAQFRSLDGMAAIPAL DPEAEPSMDVILVGSSELSSSVSPGTGRDLIAYEVKANQRNIEDICICCG SLQVHTQHPLFEGGICAPCKDKFLDALFLYDDDGYQSYCSICCSGETL LICGNPDCTRCYCFECVDSLVGPGTSGKVHAMSNWVCYLCLPSSRSG LLQRRRKWRSQLKAFYDRESENPLEMFETVPVWRRQPVRVLSLFEDI KKELTSLGFLESGSDPGQLKHVVDVTDTVRKDVEEWGPFDLVYGATP PLGHTCDRPPSWYLFQFHRLLQYARPKPGSPRPFFWMFVDNLVLNKE DLDVASRFLEMEPVTIPDVHGGSLQNAVRVWSNIPAIRSRHWALVSEE ELSLLAQNKQSSKLAAKWPTKLVKNCFLPLREYFKYFST (SEQ ID NO: 727) In some embodiments, the heterologous polypeptide comprises a Cas12f-KRAB-DNMT3L modulator. In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a heterologous gene effector comprising KRAB and DNMT3L. In some embodiments, the the Cas12f-KRAB-DNMT3L modulator comprises a heterologous gene effector comprising SEQ ID NO:727, or a sequence with, with about, or with at least 80, 85, 90, 95, 97, 98, 99% identity or about 100% identity, or a percent identity in a range defined by any two of the preceding values (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.) to SEQ ID NO:727.

[0176] The heterologous polynucleotide as disclosed herein can comprise one or more guide moieties (e.g., one or more guide nucleic acid molecules) to direct a heterologous gene effector to a target gene (e.g., target endogenous gene) or a target gene regulatory sequence. A guide moiety can confer an ability to recognize and specifically bind to the target gene or the target gene regulatory sequence. The guide moiety can be configured to form a complex with the heterologous polypeptide (e.g., a guide nucleic acid forming a complex with a nuclease, such as a CRISPR / Cas protein), and the complex can be configured to exhibit specific binding to the target polypeptide sequence as disclosed herein, to modify the expression level and / or the epigenetic modification level of the target gene.

[0177] A guide moiety can comprise a guide nucleic acid. A guide moiety can comprise a nuclease and a guide nucleic acid as disclosed herein. A guide moiety can comprise a nuclease or a part thereof, for example, an endonuclease, such as a heterologous endonuclease. The nuclease can be, e.g., a DNA nuclease and / or RNA nuclease, a modified nuclease that is nuclease-deficient or has reduced nuclease activity compared to a wild-type nuclease, a derivative thereof, a variant thereof, or a fragment thereof. In some embodiments, the guide moiety has minimal nuclease activity.

[0178] Any suitable nuclease, fragment or derivative thereof can be used in a guide moiety. Suitable nucleases include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases including type I CRISPR-associated (Cas) polypeptides, type II CRISPR- associated (Cas) polypeptides, type III CRISPR-associated (Cas) polypeptides, type IV CRISPR-associated (Cas) polypeptides, type V CRISPR-associated (Cas) polypeptides, and type VI CRISPR-associated (Cas) polypeptides; zinc finger nucleases (ZFN); transcription activator-like effector nucleases (TALEN); meganucleases; RNA-binding proteins (RBP); CRISPR-associated RNA binding proteins; recombinases; flippases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaeal Argonaute (aAgo), or eukaryotic Argonaute (eAgo)); or any derivative thereof; or any variant thereof or any functional fragment thereof.

[0179] In some embodiments, the guide moiety comprises a DNA nuclease such as an engineered (e.g., programmable or targetable) DNA nuclease that is nuclease-deficient. In some embodiments, the guide moiety comprises a nuclease-null DNA binding protein derived from a DNA nuclease that does not induce transcriptional activation or repression of a targetDNA sequence unless it is present in a complex with one or more heterologous gene effectors of the disclosure. In some embodiments, the guide moiety comprises a nuclease-null DNA binding protein derived from a DNA nuclease that can induce transcriptional activation or repression of a target DNA sequence (e.g., which can be altered or augmented by the presence of a heterologous gene effector of the disclosure).

[0180] In some embodiments, the guide moiety comprises an RNA nuclease such as an engineered (e.g., programmable or targetable) RNA nuclease. In some embodiments, the guide moiety comprises a nuclease-null RNA binding protein derived from an RNA nuclease that does not induce transcriptional activation or repression of a target RNA sequence unless it is present in a complex with one or more heterologous gene effectors of the disclosure. In some embodiments, the guide moiety comprises a nuclease-null RNA binding protein derived from a RNA nuclease that can induce transcriptional activation or repression of a target RNA sequence (e.g., which can be altered or augmented by the presence of a heterologous gene effector of the disclosure).

[0181] In some embodiments, the guide moiety comprises a nucleic acid-guided targeting system. In some embodiments, the guide moiety comprises a DNA-guided targeting system. In some embodiments, the guide moiety comprises an RNA-guided targeting system. A guide moiety can comprise and utilize, for example, a guide nucleic acid sequence that facilitates specific binding of a CRISPR-Cas system (e.g., a nuclease deficient form thereof, such as dCas9) to a target gene (e.g., target endogenous gene) or target gene regulatory sequence. Binding specificity can be determined by use of a guide nucleic acid, such as a single guide RNA (sgRNA) or a part thereof. In some embodiments, the use of different sgRNAs allows the compositions and methods of the disclosure to be used with (e.g., targeted to) different target genes (e.g., target endogenous genes) or target gene regulatory sequences.

[0182] Prokaryotic CRISPR-Cas (Clustered regularly interspaced short palindromic repeats-CRISPR associated) systems, for example, Class II CRISPR-Cas systems such as Cas9 and Cpfl, can be repurposed as a tool for regulation of gene expression, epigenome editing, and chromatin looping in compositions and methods of the disclosure. Nuclease-deactivated Cas (dCas) proteins complexed with heterologous gene effectors can allow for regulation of expression of target genes (e.g., target endogenous genes) adjacent to a site bound by the dCas.

[0183] In some embodiments, the guide moiety comprises a CRISPR-associated (Cas) protein or a Cas nuclease that functions in a non-naturally occurring CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated) system. In bacteria, this system can provide adaptive immunity against foreign DNA.

[0184] In a wide variety of organisms including diverse mammals, animals, plants, microbes, and yeast, a CRISPR / Cas system (e.g., modified and / or unmodified) can be utilized as a genome engineering tool, or can be modified to direct specific binding of engineered proteins to target loci as disclosed herein. A CRISPR / Cas system can comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid binding. An RNA-guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease) can specifically bind a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave the DNA.

[0185] In some cases, the Cas protein is mutated and / or modified to yield a nuclease deficient protein or a protein with decreased nuclease activity relative to a wild-type Cas protein. A nuclease deficient protein can retain the ability to bind DNA, but may lack or have reduced nucleic acid cleavage activity.

[0186] In some embodiments, the guide moiety comprises a Cas protein that forms a complex with a guide nucleic acid, such as a guide RNA or a part thereof. In some embodiments, the guide moiety comprises a Cas protein that forms a complex with a single guide nucleic acid, such as a single guide RNA (sgRNA). In some embodiments, the guide moiety comprises a RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), which is able to form a complex with a Cas protein. In some embodiments, the guide moiety comprises a nuclease-null DNA binding protein derived from a DNA nuclease that can induce transcriptional activation or repression of a target DNA sequence. In some embodiments, the guide moiety comprises a nuclease-null RNA binding protein derived from a RNA.

[0187] In some embodiments, a guide nucleic acid used in compositions and methods of the disclosure can be, for example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotide(s).

[0188] In some embodiments, a guide nucleic acid used in compositions and methods of the disclosure is at most at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1 nucleotide(s).

[0189] A guide nucleic acid can be a guide RNA or a part thereof.

[0190] Any suitable CRISPR / Cas system can be used. A CRISPR / Cas system can be referred to using a variety of naming systems. A CRISPR / Cas system can be a type I, a type II, a type III, a type IV, a type V, a type VI system, or any other suitable CRISPR / Cas system. A CRISPR / Cas system as used herein can be a Class 1, Class 2, or any other suitably classified CRISPR / Cas system. Class 1 or Class 2 determination can be based upon the genes encoding the effector module. Class 1 systems generally have a multi-subunit crRNA-effector complex, whereas Class 2 systems generally have a single protein, such as Cas9, Cpfl, C2c1, C2c2, C2c3 or a crRNA-effector complex. A Class 1 CRISPR / Cas system can use a complex of multiple Cas proteins to effect regulation. A Class 1 CRISPR / Cas system can comprise, for example, type I (e.g., I, IA, IB, IC, ID, IE, IF, or IU), type III (e.g., III, IIIA, IIIB, IIIC, or IIID), and type IV (e.g., IV, IVA, or IVB) CRISPR / Cas type. A Class 2 CRISPR / Cas system can use a single large Cas protein to effect regulation. A Class 2 CRISPR / Cas systems can comprise, for example, type II (e.g., II, IIA, or IIB) and type V CRISPR / Cas type. CRISPR systems can be complementary to each other, and / or can lend functional units in trans to facilitate CRISPR locus targeting.

[0191] When a guide moiety comprises a Cas protein or derivative thereof, the Cas protein or derivative thereof can be a Class 1 or a Class 2 Cas protein. A Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein. A Cas protein can comprise one or more domains. Non-limiting examples of domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, or HNH), DNA binding domain, RNA binding domain, helicase domains, protein-protein interaction domains, and dimerization domains. A guide nucleic acid recognition and / or binding domain can interact with a guide nucleic acid. A nuclease domain can comprise catalytic activity for nucleic acid cleavage. A nuclease domain can lack catalytic activity to prevent nucleic acid cleavage. A Cas protein can be a chimeric Cas protein or fragment thereof that is fused to other proteins or polypeptides. A Cas protein can be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins.

[0192] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4, Cul966, Cas13a, Cas13b, Cas13c, Cas13d, Cas13X, or Cas13Y, and homologs or modified versions thereof.

[0193] In some cases, the Cas protein as disclosed herein may not and need not be Cas9 or Cas12a. The Cas protein as disclosed herein can have a smaller size as compared to Cas9 or Cas12a. The Cas protein as disclosed herein can be derived from Un1Cas12f1. For example, the Cas protein as disclosed herein can comprise an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO. 43 In another example, the Cas protein as disclosed herein can comprise an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO: 44 As disclosed herein, SEQ ID NO: 43 encodes the polypeptide sequence of Un1Cas12f1. As disclosed herein, SEQ ID NO: 44 encodes an engineered variant of Un1Cas12f1 with reduced nuclease activity. As disclosed herein, SEQ ID NO: 728 encodes a non-limiting examples of a Cas12f variant suitable for use in the systems, compositions, and methods of the present disclosure. In some embodiments, the Cas12f variant as disclosed herein can comprise an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%,at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO: 728. SEQ ID NO: 43 (Un1Cas12f1) 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSDVCYTRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKIGEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIDVGVK SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAD YNAALNISNP KLKSTKEEP SEQ ID NO: 44 (deactivated nuclease variant of Un1Cas12f1) 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSRVCYRRAA 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKICEKS 301 AWMLNLSIDV PKIDKGVDPS IIGGIAVGVR SPLVCAINNA FSRYSISDND 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC 501 EKCNFKENAA YNAALNISNP KLKSTKERP SEQ ID NO: 728 (Cas12f variant) MAKNTITKTLKLRIVRPYNSAEVEKIVADEKERRKQAGGTGELDDKFYQKLR GQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASSVE HYLSRVCYRRAAELFKNAAIAGLRSKIKSNFRLKELKNMKSGLPTTKSDNFP IPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQV QKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGDYQTSYIEVKRGSKICEKSAWMLNLSIDVPKIDKGVDPSIIGGIAVGVRSPLVCAINNAFSRYSIS DNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSERFRKKL IERWACEIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNK IEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKC NFKENAAYNAALNISNPKLKSTKERP In some embodiments, the heterologous polypeptide includes a Cas12f-KRAB-DNMT3L modulator. In some embodiments, the Cas12f-KRAB-DNMT3L modulator includes a nuclease comprising Cas12f or a variant thereof. In some embodiments, the Cas12f-KRAB- DNMT3L modulator comprises a nuclease having the amino acid sequence of SEQ ID NO:44, or a sequence with, with about, or with at least 80, 85, 90, 95, 97, 98, 99% identity or about 100% identity, or a percent identity in a range defined by any two of the preceding values (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.) to SEQ ID NO:44. In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a nuclease having the amino acid sequence of SEQ ID NO:728, or a sequence with, with about, or with at least 80, 85, 90, 95, 97, 98, 99% identity or about 100% identity, or a percent identity in a range defined by any two of the preceding values (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.) to SEQ ID NO:728.

[0194] In some cases, a Cas protein as provided herein may not be a Cas14 protein. In some cases, a dead Cas protein (dCas) as provided herein may not be a dead Cas protein.

[0195] A Cas protein or fragment or derivative thereof can be from any suitable organism. Non-limiting examples include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas nap hthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp.,Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, or Francisella novicida. In some aspects, the organism is Streptococcus pyogenes (S. pyogenes). In some aspects, the organism is Staphylococcus aureus (S. aureus). In some aspects, the organism is Streptococcus thermophilus (S. thermophilus).

[0196] A Cas protein can be derived from a variety of bacterial species including, but not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, or Francisella novicida.

[0197] A Cas protein as used herein can be a wildtype or a modified form of a Cas protein. A Cas protein can be an active variant, inactive variant, or fragment of a wild type or modified Cas protein. A Cas protein can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein. A Cas protein can be a polypeptide with at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type Cas protein. A Cas protein can be a polypeptide with at most or at most about 5%, at most or at most about 10%, at most or at most about 20%, at most or at most about 30%, at most or at most about 40%, at most or at most about 50%, at most or at most about 60%, at most or at most about 70%, at most or at most about 80%, at most or at most about 90%, or at most or at most about 100% sequence identity and / or sequence similarity to a wild type exemplary Cas protein. Variants or fragments can comprise at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a wild type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.

[0198] A Cas protein can comprise one or more nuclease domains, such as DNase domains. For example, a Cas9 protein can comprise a RuvC-like nuclease domain and / or an HNH-like 20 nuclease domain. The in a nuclease active form of Cas9, RuvC and HNH domains can each cut a different strand of double-stranded DNA to make a double-stranded break in the DNA. A Cas protein can comprise only one nuclease domain (e.g., Cpfl comprises RuvCdomain but lacks HNH domain). In some embodiments, nuclease domains are absent. In some embodiments, nuclease domains are present but inactive or have reduced or minimal activity. In some embodiments, nuclease domains are present and active.

[0199] One or a plurality of the nuclease domains (e.g., RuvC, or HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. For example, in a Cas protein comprising at least two nuclease domains (e.g., Cas9), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, known as a nickase, can generate a single-strand break at a CRISPR RNA (crRNA) recognition sequence within a double- stranded DNA but not a double-strand break. Such a nickase can cleave the complementary strand or the non-complementary strand, but may not cleave both. If all of the nuclease domains of a Cas protein (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are deleted or mutated, the resulting Cas protein can have a reduced or no ability to cleave both strands of a double-stranded DNA. An example of a mutation that can convert a Cas9 protein into a nickase is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of Cas9 from S. pyogenes. H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of Cas9 from S. pyogenes can convert the Cas9 into a nickase. An example of a mutation that can convert a Cas9 protein into a dead Cas9 is a D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain and H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of Cas9 from S. pyogenes.

[0200] A nuclease dead Cas protein (e.g., one derived from any Cas protein, such as Un1Cas12f1) can comprise one or more mutations relative to a wild-type version of the protein. The mutation can result in no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity in one or more of the plurality of nucleic acid-cleaving domains of the wild-type Cas protein. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains retaining the ability to cleave the complementary strand of the target nucleic acid but reducing its ability to cleave the non-complementary strand of the target nucleic acid. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains retainingthe ability to cleave the non-complementary strand of the target nucleic acid but reducing its ability to cleave the complementary strand of the target nucleic acid. The mutation can result in one or more of the plurality of nucleic acid-cleaving domains lacking the ability to cleave the complementary strand and the non-complementary strand of the target nucleic acid. The residues to be mutated in a nuclease domain can correspond to one or more catalytic residues of the nuclease. For example, residues in the wild type exemplary S. pyogenes Cas9 polypeptide such as Asp10, His840, Asn854 and Asn856 can be mutated to inactivate one or more of the plurality of nucleic acid-cleaving domains (e.g., nuclease domains). The residues to be mutated in a nuclease domain of a Cas protein can correspond to residues Asp10, His840, Asn854 and Asn856 in the wild type S. pyogenes Cas9 polypeptide, for example, as determined by sequence and / or structural alignment.

[0201] As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 (or the corresponding mutations of any of the Cas proteins) can be mutated. For example, e.g., D 10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. Mutations other than alanine substitutions can be suitable.

[0202] A D10A mutation can be combined with one or more of H840A, N854A, or N856A mutations to produce a Cas9 protein substantially lacking DNA cleavage activity (e.g., a dead Cas9 protein). A H840A mutation can be combined with one or more of D10A, N854A, or N856A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity. An N854A mutation can be combined with one or more of H840A, D1OA, or N856A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity. A N856A mutation can be combined with one or more of H840A, N854A, or D10A mutations to produce a site-directed polypeptide substantially lacking DNA cleavage activity.

[0203] In some embodiments, a Cas protein is a Class 2 Cas protein. In some embodiments, a Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, or derived from a Cas9 protein. For example, a Cas9 protein lacking cleavage activity. In some embodiments, the Cas9 protein is a Cas9 protein from S. pyogenes (e.g., SwissProt accession number Q99ZW2). In some embodiments, the Cas9 protein is a Cas9 from S. aureus (e.g., SwissProt accession numberJ7RUA5). In some embodiments, the Cas9 protein is a modified version of a Cas9 protein from S. pyogenes or S. Aureus. In some embodiments, the Cas9 protein is derived from a Cas9 protein from S. pyogenes or S. Aureus. For example, a S. pyogenes or S. Aureus Cas9 protein lacking cleavage activity.

[0204] In some embodiments, Cas9 can generally refer to a polypeptide with at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, or about 100% sequence identity and / or sequence similarity to a wild type exemplary Cas9 polypeptide (e.g., Cas9 from S. pyogenes). In some embodiments, Cas9 can refer to a polypeptide with at most about 5%, at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 70%, at most about 80%, at most about 90%, or about 100% sequence identity and / or sequence similarity to a wild type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 can refer to the wildtype or a modified form of the Cas9 protein that can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof.

[0205] A Cas protein can comprise an amino acid sequence having at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to a nuclease domain (e.g., RuvC domain, or HNH domain) of a wild-type Cas protein.

[0206] A Cas protein, variant or derivative thereof can be modified to enhance regulation of gene expression by compositions and methods of the disclosure, e.g., as part of a complex disclosed herein. A Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, enzymatic activity, and / or binding to other factors, such as heterodimerization or oligomerization domains and induce ligands. Casproteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the desired function of the protein or complex. A Cas protein can be modified to modulate (e.g., enhance or reduce) the activity of the Cas protein for regulating gene expression by a complex of the disclosure that comprises a heterologous gene effector.

[0207] For example, a Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a heterologous gene effector (e.g., an epigenetic modification domain, a transcriptional activation domain, and / or a transcriptional repressor domain). A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to an oligomerization or dimerization domain as disclosed herein (e.g., a heterodimerization domain). A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a heterologous polypeptide that provides increased or decreased stability. A Cas protein can be coupled (e.g., fused, covalently coupled, or non-covalently coupled) to a sequence that can facilitate degradation of the Cas protein or a complex containing the Cas protein, for example, a degron, such as an inducible degron (e.g., auxin inducible).

[0208] A Cas protein can be coupled (e.g., fused, covalently coupled, or non- covalently coupled) to any suitable number of partners, for example, at least one, at least two, at least three, at least four, or at least five, at least six, at least seven, or at least 8 partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to at most two, at most three, at most four, at most five, at most six, at most seven, at most eight, or at most ten partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to 1 – 5, 1 – 4, 1 – 3, 1 – 2, 2 – 5, 2 – 4, 2 – 3, 3 – 5, 3 – 4, or 4 – 5 partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to one partner. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to two partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to three partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to four partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, ornon-covalently coupled) to five partners. In some embodiments, a Cas protein of the disclosure is coupled (e.g., fused, covalently coupled, or non-covalently coupled) to six partners.

[0209] A Cas protein can be a fusion protein. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.

[0210] A Cas protein can be provided in any form. For example, a Cas protein can be provided in the form of a protein, such as a Cas protein alone or complexed with a guide nucleic acid as a ribonucleoprotein. A Cas protein can be provided in a complex, for example, complexed with a guide nucleic acid and / or one or more heterologous gene effectors of the disclosure. A Cas protein can be provided in the form of a nucleic acid encoding the Cas protein, such as an RNA (e.g., messenger RNA (mRNA)), or DNA. The nucleic acid encoding the Cas protein can be codon optimized for efficient translation into protein in a particular cell or organism.

[0211] Nucleic acids encoding Cas proteins, fragments, or derivatives thereof can be stably integrated in the genome of a cell. Nucleic acids encoding Cas proteins can be operably linked to a promoter, for example, a promoter that is constitutively or inducibly active in the cell. Nucleic acids encoding Cas proteins can be operably linked to a promoter in an expression construct. Expression constructs can include any nucleic acid constructs capable of directing expression of a gene or other nucleic acid sequence of interest (e.g., a Cas gene) and which can transfer such a nucleic acid sequence of interest to a target cell.

[0212] In some embodiments, a Cas protein, variant or derivative thereof is a nuclease dead Cas (dCas) protein. A dead Cas protein can be a protein that lacks nucleic acid cleavage activity.

[0213] A Cas protein can comprise a modified form of a wild type Cas protein. The modified form of the wild type Cas protein can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid-cleaving activity of the Cas protein. For example, the modified form of the Cas protein can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity of the wild-type Cas protein (e.g., Cas9 from S. pyogenes). The modified form of Cas protein can have no substantial nucleic acid-cleaving activity. When aCas protein is a modified form that has no substantial nucleic acid-cleaving activity, it can be referred to as enzymatically inactive, “deactivated” and / or “dead” (abbreviated by “d”). A dead Cas protein (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave or minimally cleaves the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein.

[0214] A dCas9 polypeptide can associate with a single guide RNA (sgRNA) to activate or repress transcription of a target gene (e.g., target endogenous gene), for example, in combination with heterologous gene effector(s) disclosed herein. sgRNAs can be introduced into cells expressing the Cas or guide moiety component of the disclosure. In some cases, such cells can contain one or more different sgRNAs that target the same target gene (e.g., target endogenous gene) or target gene regulatory sequence. In other cases, the sgRNAs target different nucleic acids in the cell (e.g., different target genes, different target gene regulatory sequences, or different sequences within the same target gene or target gene regulatory sequence).

[0215] Enzymatically inactive can refer to a nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner, but will not cleave a target polynucleotide or will cleave it at a substantially reduced frequency. An enzymatically inactive guide moiety can comprise an enzymatically inactive domain (e.g. nuclease domain). Enzymatically inactive can refer to no activity. Enzymatically inactive can refer to substantially no activity. Enzymatically inactive can refer to essentially no activity. Enzymatically inactive can refer to an activity no more than 1%, no more than 2%, no more than 3%, no more than 4%, no more than 5%, no more than 6%, no more than 7%, no more than 8%, no more than 9%, or no more than 10% activity compared to a comparable wild-type activity (e.g., nucleic acid cleaving activity, wild-type Cas9 activity).

[0216] In some embodiments, the guide moiety does not contain a nucleic acid- guided targeting system. For example, guide moieties can include proteins that bind to a target gene (e.g., target endogenous gene) or target gene regulatory sequence based on protein structural features, such as certain nucleases disclosed herein.

[0217] In some embodiments, a guide moiety comprises a zinc finger nuclease (ZFN) or a variant, fragment, or derivative thereof. ZFN can refer to a fusion between a cleavage domain, such as a cleavage domain of Fokl, and at least one zinc finger motif (e.g.,at least 2, at least 3, at least 4, or at least 5 zinc finger motifs) which can bind polynucleotides such as DNA and RNA. In some embodiments, a ZFN is used in a targeting moiety of the disclosure to bind a polynucleotide (e.g., target gene or target gene regulatory sequence), but the ZFN does not cleave or substantially does not cleave the polynucleotide, e.g., a nuclease dead ZFN. A ZFN or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure.

[0218] The heterodimerization at certain positions in a polynucleotide of two individual ZFNs in certain orientation and spacing can lead to cleavage of the polynucleotide in nuclease-active ZFN. For example, a ZFN binding to DNA can induce a double-strand break in the DNA. In order to allow two cleavage domains to dimerize and cleave DNA, two individual ZFNs can bind opposite strands of DNA with their C-termini at a certain distance apart. In some cases, linker sequences between the zinc finger domain and the cleavage domain can require the 5' edge of each binding site to be separated by 5, 6, or 7 base pairs. In some cases, a cleavage domain is fused to the C-terminus of each zinc finger domain.

[0219] In some embodiments, the cleavage domain of a guide moiety comprising a ZFN comprises a modified form of a wild type cleavage domain. The modified form of the cleavage domain can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid-cleaving activity of the cleavage domain. For example, the modified form of the cleavage domain can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity of the corresponding wild-type cleavage domain. The modified form of the cleavage domain can have no substantial nucleic acid-cleaving activity. In some embodiments, the cleavage domain is enzymatically inactive.

[0220] In some embodiments, a guide moiety comprises a “TALEN” or “TAL- effector nuclease” or a variant, fragment, or derivative thereof. TALENs refer to engineered transcription activator-like effector nucleases that generally contain a central domain of DNA- binding tandem repeats and a cleavage domain. TALENs can be produced by fusing a TAL effector DNA binding domain to a DNA cleavage domain. In some cases, a DNA-binding tandem repeat comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. Atranscription activator-like effector (TALE) protein can be fused to a nuclease such as a wild- type or mutated Fok1 endonuclease or the catalytic domain of Fok1. In some embodiments, a TALEN is used in a targeting moiety of the disclosure to bind a polynucleotide (e.g., target gene or target gene regulatory sequence), but the TALEN does not cleave or substantially does not cleave the polynucleotide, e.g., a nuclease dead TALEN. A TALEN or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure.

[0221] In some embodiments, a TALEN is engineered for reduced nuclease activity. In some embodiments, the nuclease domain of a TALEN comprises a modified form of a wild type nuclease domain. The modified form of the nuclease domain can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid- cleaving activity of the nuclease domain. For example, the modified form of the nuclease domain can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid-cleaving activity of the wild- type nuclease domain. The modified form of the nuclease domain can have no substantial nucleic acid-cleaving activity. In some embodiments, the nuclease domain is enzymatically inactive. A TALEN or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure.

[0222] Several mutations to Fok1 have been made for its use in TALENs, which, for example, improve cleavage specificity or activity. Such TALENs can be engineered to bind any desired DNA sequence. TALENs can be used to generate gene modifications (e.g., nucleic acid sequence editing) by creating a double-strand break in a target DNA sequence, which in turn, undergoes NHEJ or HDR.

[0223] A TALE or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure. In some embodiments, the transcription activator-like effector (TALE) protein is fused to a heterologous gene effector and does not comprise a nuclease. In some embodiments, a TALEN does not cleave or substantially does not cleave the polynucleotide, e.g., a nuclease dead TALE. A TALE or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure.

[0224] In some embodiments, the complex of the transcription activator-like effector (TALE) protein and the heterologous gene effector is designed to function as a transcriptional activator. In some embodiments, the complex of the transcription activator-like effector (TALE) protein and the heterologous gene effector is designed to function as a transcriptional repressor. For example, the DNA-binding domain of the transcription activator- like effector (TALE) protein can be fused (e.g., linked) to one or more heterologous gene effectors that comprise transcriptional activation domains, or to one or more heterologous gene effectors that comprise transcriptional repression domains.

[0225] In some embodiments, a guide moiety comprises a meganuclease. Meganucleases generally refer to rare-cutting endonucleases or homing endonucleases that can be highly sequence specific. Meganucleases can recognize DNA target sites ranging from at least 12 base pairs in length, e.g., from 12 to 40 base pairs, 12 to 50 base pairs, or 12 to 60 base pairs in length. Meganucleases can be modular DNA-binding nucleases such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA binding domain or protein specifying a nucleic acid target sequence. The DNA-binding domain can contain at least one motif that recognizes single- or double-stranded DNA. A nuclease- active meganuclease can generate a double-stranded break. In some embodiments, a meganuclease is used in a targeting moiety of the disclosure to bind a polynucleotide (e.g., target gene or target gene regulatory sequence), but the meganuclease does not cleave or substantially does not cleave the polynucleotide, e.g., a nuclease dead meganuclease. A meganuclease or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form a complex of the disclosure.

[0226] The meganuclease can be monomeric or dimeric. In some embodiments, the meganuclease is naturally-occurring (found in nature) or wild-type, and in other instances, the meganuclease is non-natural, artificial, engineered, synthetic, rationally designed, or man- made. In some embodiments, the meganuclease of the present disclosure includes an I-CreI meganuclease, I-CeuI meganuclease, I-Msol meganuclease, or I-SceI meganuclease, variants thereof, derivatives thereof, or functional fragments thereof.

[0227] In some embodiments, the nuclease domain of a meganuclease comprises a modified form of a wild type nuclease domain. The modified form of the nuclease domain can comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces oreliminates the nucleic acid-cleaving activity of the nuclease domain. For example, the modified form of the nuclease domain can have no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the nucleic acid- cleaving activity of the wild-type nuclease domain. The modified form of the nuclease domain can have no substantial nucleic acid-cleaving activity. In some embodiments, the nuclease domain is enzymatically inactive. In some embodiments, a meganuclease can bind DNA but cannot cleave the DNA. In some embodiments, a nuclease-inactive meganuclease is fused to or associated with one or more heterologous gene effectors to generate a complex of the disclosure.

[0228] In some embodiments, the guide moiety can regulate expression and / or activity of a target gene (e.g., target endogenous gene). In some embodiments, the guide moiety can edit the sequence of a nucleic acid (e.g., a gene and / or gene product). A nuclease-active Cas protein can edit a nucleic acid sequence by generating a double-stranded break or single- stranded break in a target polynucleotide.

[0229] In some embodiments, a guide moiety comprising a nuclease can generate a double-strand break in a target polynucleotide, such as DNA. A double-strand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). In some embodiments, a nuclease induces site-specific single-strand DNA breaks or nicks, thus resulting in HDR.

[0230] A double-strand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor DNA repair template or template polynucleotide that contains homology arms flanking sites of the target DNA can be provided.

[0231] In some embodiments, a guide moiety or complex comprising a nuclease does not generate a double-strand break in a target polynucleotide, such as DNA.

[0232] Disclosed herein, in some aspects, are one or more complexes or systems that comprise a heterologous polypeptide and a heterologous polynucleotide. In some cases, a complex can comprise a heterologous gene effector and a guide moiety, for example, a guidenucleic acid and / or a nuclease, such as an endonuclease that lacks or substantially lacks cleavage activity.

[0233] Complexes of the disclosure can be useful, for example, for bringing one or more heterologous gene effectors into close proximity with a target gene (e.g., target endogenous gene) or target gene regulatory sequence, thereby facilitating modulation of an expression, epigenetic modification, or activity level of the target gene.

[0234] In some embodiments, a complex of the disclosure binds to DNA, e.g., genomic DNA. In some embodiments, a complex of the disclosure binds to RNA, e.g., mRNA, microRNA, siRNA, or non-coding RNA. In some embodiments, a complex of the disclosure binds to DNA and RNA.

[0235] In some embodiments, a complex can modulate (e.g., increase or decrease) expression and / or activity of a target gene (e.g., target endogenous gene) by physical obstruction of a polynucleotide sequence (e.g., a promoter, enhancer, repressor, operator, or silencer, insulator, cis-regulatory element, trans-regulatory element, epigenetic modification (e.g., DNA methylation) site, coding sequence).

[0236] In some embodiments, a complex can modulate (e.g., increase or decrease) expression and / or activity of a target gene (e.g., target endogenous gene) by recruitment of additional factors effective to suppress or enhance expression of the target gene.

[0237] In some embodiments, complexes of the disclosure are used for introducing epigenetic modifications to a target gene (e.g., target endogenous gene) or target gene regulatory sequence (e.g., promoter, enhancer, silencer, insulator, cis-regulatory element, trans-regulatory element, or epigenetic modification (e.g., DNA methylation) site). In some embodiments, complexes of the disclosure are used for producing three-dimensional structures, topologically associating domains, or genomic boundaries comprising a target gene or target gene regulatory sequence (e.g., distal or proximal gene from the target gene).

[0238] In some embodiments, a complex or system comprises a heterologous gene effector and a guide moiety. In some embodiments, a complex or system comprises one heterologous gene effector and one guide moiety. In some embodiments, a complex or system comprises two heterologous gene effectors and one guide moiety. In some embodiments, a complex or system comprises three or more heterologous gene effectors and one guide moiety.

[0239] In some embodiments, a complex or system comprises a heterologous gene effector and a guide nucleic acid. In some embodiments, a complex or system comprises one heterologous gene effector and one guide nucleic acid. In some embodiments, a complex or system comprises two heterologous gene effectors and one guide nucleic acid. In some embodiments, a complex or system comprises three or more heterologous gene effectors and one guide nucleic acid.

[0240] Two components present in a complex or system can be covalently linked, for example, present in a fusion protein, or cross-linked, e.g., treated with a crosslinking agent, or joined by a peptide or non-peptide linker as disclosed herein.

[0241] In some embodiments, two components present in a complex or system are part of the same fusion protein. Components can optionally be joined by a linker, such as a peptide linker or a non-peptide linker.

[0242] In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is joined to a heterologous gene effector by a linker. In some embodiments the guide moiety or part thereof is further joined to a second heterologous gene effector by a second linker that is the same or different. In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is fused to a heterologous gene effector without a linker.

[0243] In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is joined to an oligomerization domain or dimerization (e.g., heterodimerization) domain by a linker. In some embodiments the guide moiety or part thereof is further joined to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by a second linker that is the same or different. In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is fused to a second oligomerization domain or dimerization (e.g., heterodimerization) domain without a linker.

[0244] In some embodiments, heterologous gene effector is joined to a second heterologous gene effector by a linker. In some embodiments the heterologous gene effector is further joined to a third heterologous gene effector by a second linker that is the same or different. In some embodiments, a heterologous gene effector is fused to a second heterologous gene effector without a linker.

[0245] In some embodiments, heterologous gene effector is joined to an oligomerization domain or dimerization (e.g., heterodimerization) domain by a linker. In someembodiments the heterologous gene effector is further joined to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by a second linker that is the same or different. In some embodiments, a heterologous gene effector is fused to a second oligomerization domain or dimerization (e.g., heterodimerization) domain without a linker.

[0246] Any suitable linker can be used. A flexible linker can have a sequence containing stretches of glycine and serine residues. The small size of the glycine and serine residues provides flexibility, and allows for mobility of the connected functional domains. The incorporation of serine or threonine can maintain the stability of the linker in aqueous solutions by forming hydrogen bonds with the water molecules, thereby reducing unfavorable interactions between the linker and protein moieties. Flexible linkers can also contain additional amino acids such as threonine and alanine to maintain flexibility, as well as polar amino acids such as lysine and glutamine to improve solubility. A rigid linker can have, for example, an alpha helix-structure. An alpha-helical rigid linker can act as a spacer between protein domains.

[0247] A linker sequence can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length.

[0248] In some embodiments, a linker is at least 1, at least 2, at least 3, at least 5, at least 7, at least 9, at least 11, at least 13, at least 15, or at least 20 amino acids. In some embodiments, a linker is at most 5, at most 7, at most 9, at most 11, at most 13, at most 15, at most 20, at most 25, at most 30, at most 40, or at most 50 amino acids.

[0249] In some embodiments, non-peptide linkers are used. A non-peptide linker can be, for example a chemical linker. Two parts of a complex or system of the disclosure can be connected by a chemical linker. Each chemical linker of the disclosure can be alkylene, alkenylene, alkynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, any of which is optionally substituted. In some embodiments, a chemical linker of the disclosure can be an ester, ether, amide, thioether, or polyethyleneglycol (PEG). In some embodiments, a linker can reverse the order of the amino acids sequence in a compound, for example, so that the amino acid sequences linked by the linked are head-to-head, rather than head-to-tail. Non-limiting examples of such linkers include diesters of dicarboxylic acids, such as oxalyl diester, malonyl diester, succinyl diester, glutaryl diester, adipyl diester, pimetyldiester, fumaryl diester, maleyl diester, phthalyl diester, isophthalyl diester, and terephthalyl diester. Non-limiting examples of such linkers include diamides of dicarboxylic acids, such as oxalyl diamide, malonyl diamide, succinyl diamide, glutaryl diamide, adipyl diamide, pimetyl diamide, fumaryl diamide, maleyl diamide, phthalyl diamide, isophthalyl diamide, and terephthalyl diamide. Non-limiting examples of such linkers include diamides of diamino linkers, such as ethylene diamine, 1,2-di(methylamino)ethane, 1,3-diaminopropane, 1,3- di(methylamino)propane, 1,4-di(methylamino)butane, 1,5-di(methylamino)pentane, 1,6- di(methylamino)hexane, or pipyrizine. Non-limiting examples of optional substituents include hydroxyl groups, sulfhydryl groups, halogens, amino groups, nitro groups, nitroso groups, cyano groups, azido groups, sulfoxide groups, sulfone groups, sulfonamide groups, carboxyl groups, carboxaldehyde groups, imine groups, alkyl groups, halo-alkyl groups, alkenyl groups, halo-alkenyl groups, alkynyl groups, halo-alkynyl groups, alkoxy groups, aryl groups, aryloxy groups, aralkyl groups, arylalkoxy groups, heterocyclyl groups, acyl groups, acyloxy groups, carbamate groups, amide groups, ureido groups, epoxy groups, or ester groups.

[0250] Two components present in a complex or system can be non-covalently coupled, for example, by ionic bonds, hydrogen bonds, interactions mediated by oligomerization or dimerization domains disclosed herein, etc.

[0251] In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is joined to a heterologous gene effector by non-covalent coupling. In some embodiments the guide moiety or part thereof is further joined to a second heterologous gene effector by non-covalent coupling. In some embodiments the guide moiety or part thereof is joined to a first heterologous gene effector covalently (e.g., as a fusion protein, optionally with a linker), and the guide moiety or part thereof is further joined to a second heterologous gene effector by non-covalent coupling.

[0252] In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is joined to an oligomerization domain or dimerization (e.g., heterodimerization) domain by non-covalent coupling. In some embodiments the guide moiety or part thereof is further joined to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by non-covalent coupling. In some embodiments, a guide moiety or a part thereof (e.g., nuclease, such as dCas9) is fused to a first oligomerization domain or dimerization (e.g., heterodimerization) domain by covalent coupling (e.g., fused, optionally by a linker) and isjoined to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by non-covalent coupling.

[0253] In some embodiments, a first component of a guide moiety (e.g., a guide nucleic acid) is joined to a second component of the guide moiety (e.g., nuclease) non- covalently. In some embodiments, a first component of a guide moiety (e.g., a guide nucleic acid) is joined to a second component of the guide moiety (e.g., nuclease) covalently.

[0254] Any combination of covalent and non-covalent coupling can be used in a complex or system of the disclosure, for example, one or more heterologous gene effectors can be fused to a guide moiety non-covalently, and one or more oligomerization domains can be bound to a component of the complex or system (e.g., nuclease) covalently.

[0255] In some embodiments, a polypeptide providing increased or decreased stability is fused to or otherwise associated with a component of a complex or system of the disclosure, e.g., a guide moiety or a heterologous gene effector. The fused polypeptide can be located at the N-terminus, the C-terminus, or internally within the fusion protein.

[0256] In some embodiments, one or more components of a complex or system of the disclosure is fused to a domain the directs desirable sub-cellular localization, for example, a nuclear localization signal or a protein for targeting to the inner nuclear membrane, outer nuclear membrane, Cajal body, nuclear speckle, nuclear pore complex, PML body, nucleolus, P granule, GW body, stress granule, sponge body, endoplasmic reticulum, mitochondria, etc.

[0257] In some embodiments, a complex or system of the disclosure comprises a first protein linked to a first oligomerization (e.g., dimerization) domain, and a second protein linked to a second oligomerization (e.g., dimerization) domain. In some embodiments, an oligomerization domain or a dimerization domain can comprise a peptide interaction domain, for example, systems utilizing sgRNA2.0, SAM, SunTag, RAB, FLAG-biotin, or inducible oligomerization (e.g., dimerization) systems disclosed herein. DELIVERY

[0258] One or more genes encoding any of the heterologous polypeptide (e.g., the heterologous gene effectors) and any additional molecule operatively coupled thereto (e.g., the heterologous polynucleotide, such as one or more guide nucleic acid molecules), as disclosed herein, can be integrated into a genome of the cell, in which the aberrant expression of a targetgene is to be modified. Alternatively, the one or more genes may not and need not be integrated into the genome of the cell. The one or more genes can be a single gene (e.g., a single expression cassette) or a plurality of genes (e.g., a plurality of expression cassettes). The one or more genes can be heterologous to the cell(s).

[0259] Any of the heterologous polypeptide (e.g., the heterologous gene effectors) and any additional molecule operatively coupled thereto (e.g., the heterologous polynucleotide, such as one or more guide nucleic acid molecules) can be introduced (e.g., delivered, expressed, etc.) to a cell by various methods, e.g., viral and non-viral delivery methods. Viral vector delivery systems can include DNA and RNA viruses, which can have either episomal or integrated genomes after delivery to the cell. Non-viral vector delivery systems can include DNA plasmids, RNA (e.g. a transcript of a vector described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle, such as a liposome.

[0260] The one or more genes can further comprise one or more promoters to control expression of the systems. A promoter as disclosed herein can be active in a eukaryotic, mammalian, non-human mammalian or human cell. The promoter can be an inducible or constitutively active promoter. Alternatively or additionally, the promoter can be tissue or cell specific. Non-limiting examples of suitable eukaryotic promoters (i.e. promoters functional in a eukaryotic cell) can include those from cytomegalovirus (CMV) immediate early, herpes simplex virus (HSV) thymidine kinase, early and late SV40, long terminal repeats (LTRs) from retrovirus, human elongation factor-1 promoter (EF1), a hybrid construct comprising the cytomegalovirus (CMV) enhancer fused to the chicken beta-active promoter (CAG), murine stem cell virus promoter (MSCV), phosphoglycerate kinase-1 locus promoter (PGK) or mouse metallothionein-I. The promoter can be a fungi promoter. The promoter can be a plant promoter. A database of plant promoters can be found (e.g., PlantProm). The expression vector may also contain a ribosome binding site for translation initiation and a transcription terminator. The expression vector may also include appropriate sequences for amplifying expression. In some cases, a promoter as disclosed herein can be a promoter specific for any of the tissues provided herein, or a promoter specific for any of the cell types provided herein.

[0261] The single gene as provided herein (e.g., encoding the systems of the present disclosure) can have a size of at least or up to about 2.5 kilobases, at least or up to about 2.6 kilobases, at least or up to about 2.7 kilobases, at least or up to about 2.8 kilobases, at least orup to about 2.9 kilobases, at least or up to about 3.0 kilobases, at least or up to about 3.1 kilobases, at least or up to about 3.2 kilobases, at least or up to about 3.3 kilobases, at least or up to about 3.4 kilobases, at least or up to about 3.5 kilobases, at least or up to about 3.6 kilobases, at least or up to about 3.7 kilobases, at least or up to about 3.8 kilobases, at least or up to about 3.9 kilobases, at least or up to about 4.0 kilobases, at least or up to about 4.1 kilobases, at least or up to about 4.2 kilobases, at least or up to about 4.3 kilobases, at least or up to about 4.4 kilobases, at least or up to about 4.5 kilobases, at least or up to about 4.6 kilobases, at least or up to about 4.7 kilobases, at least or up to about 4.8 kilobases, at least or up to about 4.9 kilobases, at least or up to about 5.0 kilobases, at least or up to about 5.5 kilobases, at least or up to about 6.0 kilobases, at least or up to about 6.5 kilobases, at least or up to about 7.0 kilobases, at least or up to about 7.5 kilobases, at least or up to about 8.0 kilobases, at least or up to about 9.0 kilobases, or at least or up to about 10 kilobases. In some cases, the single gene can have a size of between about 3 kilobases and about 5 kilobases, between about 3 kilobases and about 4.8 kilobases, between about 3 kilobases and about 4.6 kilobases, between about 3 kilobases and about 4.4 kilobases, between about 3 kilobases and about 4.2 kilobases, between about 3 kilobases and about 4.0 kilobases, between about 3 kilobases and about 3.5 kilobases, between about 3.5 kilobases and about 5 kilobases, between about 3.5 kilobases and about 4.8 kilobases, between about 3.5 kilobases and about 4.6 kilobases, between about 3.5 kilobases and about 4.4 kilobases, between about 3.5 kilobases and about 4.2 kilobases, between about 3.5 kilobases and about 4 kilobases, between about 4 kilobases and about 5 kilobases, between about 4 kilobases and about 4.9 kilobases, between about 4 kilobases and about 4.8 kilobases, between about 4 kilobases and about 4.7 kilobases, between about 4 kilobases and about 4.6 kilobases, between about 4 kilobases and about 4.5 kilobases, between about 4 kilobases and about 4.4 kilobases, between about 4 kilobases and about 4.3 kilobases, between about 4 kilobases and about 4.2 kilobases, or between about 4 kilobases and about 4.1 kilobases.

[0262] RNA or DNA viral based systems can be used to target specific cells and traffic the viral payload to the nucleus of the cell. Viral vectors can be used to contact cells in vitro, and the modified cells can optionally be administered (ex vivo) to a subject, such as a human. Alternatively, viral vectors can be administered directly (in vivo) to the subject. Viral based systems can include retroviral, lentivirus, adenoviral, adeno-associated or herpessimplex virus vectors for gene transfer. Integration in the host genome can occur with the retrovirus, lentivirus, or adeno-associated virus gene transfer methods, which can result in long term expression of the inserted transgene.

[0263] In some embodiments, the compositions and systems provided herein are delivered to a subject using a viral vector. In some cases, the viral vector is an adeno-associated viral (AAV) vector. The term “AAV” is an abbreviation for adeno-associated virus, and may be used to refer to the virus itself or a derivative thereof. The term covers all serotypes, subtypes, and both naturally occurring and recombinant forms, except where required otherwise. The abbreviation “rAAV” refers to recombinant adeno-associated virus, also referred to as a recombinant AAV vector (or “rAAV vector”). The term “AAV” includes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, rh10, and hybrids thereof, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, or ovine AAV. The genomic sequences of various serotypes of AAV, as well as the sequences of the native terminal repeats (TRs), Rep proteins, and capsid subunits are known in the art. Such sequences may be found in the literature or in public databases such as GenBank. An “rAAV vector” as used herein refers to an AAV vector comprising a polynucleotide sequence not of AAV origin (i.e., a polynucleotide heterologous to AAV), typically a sequence of interest for the genetic transformation of a cell. In general, the heterologous polynucleotide is flanked by at least one, and generally by two, AAV inverted terminal repeat sequences (ITRs). The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids. An rAAV vector may either be single-stranded (ssAAV) or self-complementary (scAAV). An “AAV virus” or “AAV viral particle” or “rAAV vector particle” refers to a viral particle composed of at least one AAV capsid protein and an encapsidated polynucleotide rAAV vector. If the particle comprises a heterologous polynucleotide (i.e., a polynucleotide other than a wild-type AAV genome such as a transgene to be delivered to a mammalian cell), it is typically referred to as an “rAAV vector particle” or simply an “rAAV vector”. Thus, production of rAAV particle necessarily includes production of rAAV vector, as such a vector is contained within an rAAV particle. In some cases, the AAV vector is selected based on the tropism of viral vector. In some embodiments, an AAV vector with tropism for the tissue of interest may be used (e.g., AAV2 for muscle tissue) to deliver polynucleotides encoding the compositions and systems provided herein to the tissue.

[0264] RNA or DNA viral based systems can be used to target specific cells in the body and trafficking the viral payload to the nucleus of the cell. Viral vectors can be administered directly (in vivo) or they can be used to contact cells in vitro, and the modified cells can optionally be administered (ex vivo) to a subject, such as a human. Viral based systems can include retroviral, lentivirus, adenoviral, adeno-associated or herpes simplex virus vectors for gene transfer. Integration in the host genome can occur with the retrovirus, lentivirus, and adeno-associated virus gene transfer methods, which can result in long term expression of the inserted transgene. High transduction efficiencies can be observed in many different cell types and target tissues.

[0265] The tropism of a retrovirus can be altered by incorporating foreign envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and produce high viral titers. Selection of a retroviral gene transfer system can depend on the target tissue. Retroviral vectors can comprise cis-acting long terminal repeats with packaging capacity for up to 6-10 kb of foreign sequence. The minimum cis-acting LTRs can be sufficient for replication and packaging of the vectors, which can be used to integrate the therapeutic gene into the target cell to provide permanent transgene expression. Retroviral vectors can include those based upon murine leukemia virus (MuLV), gibbon ape leukemia virus (GaLV), Simian Immuno deficiency virus (SIV), or human immuno deficiency virus (HIV), and combinations thereof.

[0266] An adenoviral-based systems can be used. Adenoviral-based systems can lead to transient expression of the transgene. Adenoviral based vectors can have high transduction efficiency in cells and may not require cell division. High titer and levels of expression can be obtained with adenoviral based vectors. Adeno-associated virus (“AAV”) vectors can be used to transduce cells with target nucleic acids, e.g., in the in vitro production of nucleic acids and peptides, and for in vivo and ex vivo gene therapy procedures.

[0267] Packaging cells can be used to form virus particles capable of infecting a host cell. Such cells can include 293 cells, (e.g., for packaging adenovirus), or Psi2 cells or PA317 cells (e.g., for packaging retrovirus). Viral vectors can be generated by producing a cell line that packages a nucleic acid vector into a viral particle. The vectors can contain the minimal viral sequences required for packaging and subsequent integration into a host. The vectors can contain other viral sequences being replaced by an expression cassette for thepolynucleotide(s) to be expressed. The missing viral functions can be supplied in trans by the packaging cell line. For example, AAV vectors can comprise ITR sequences from the AAV genome which are required for packaging and integration into the host genome. Viral DNA can be packaged in a cell line, which can contain a helper plasmid encoding the other AAV genes, namely rep and cap, while lacking ITR sequences. The cell line can also be infected with adenovirus as a helper. The helper virus can promote replication of the AAV vector and expression of AAV genes from the helper plasmid. Contamination with adenovirus can be reduced by, e.g., heat treatment to which adenovirus is more sensitive than AAV.

[0268] A host cell can be transiently or non-transiently transfected with one or more vectors described herein. A cell can be transfected as it naturally occurs in a subject. A cell can be taken or derived from a subject and transfected. A cell can be derived from cells taken from a subject, such as a cell line. In some embodiments, a cell transfected with one or more vectors described herein is used to establish a new cell line comprising one or more vector-derived sequences. In some embodiments, a cell transiently transfected with the compositions of the disclosure (such as by transient transfection of one or more vectors, or transfection with RNA), and modified through the activity of an actuator moiety such as a CRISPR complex, is used to establish a new cell line comprising cells containing the modification but lacking any other exogenous sequence.

[0269] Any suitable vector compatible with the host cell can be used with the methods of the disclosure. Non-limiting examples of vectors for eukaryotic host cells include pXT1, pSG5 (Stratagene™), pSVK3, pBPV, pMSG, or pSVLSV40 (Pharmacia™).

[0270] In some embodiments, the additional ingredient of the composition as disclosed herein can comprise an excipient. Non-limiting examples of the excipient can include solvents, dispersion media, diluents, or other liquid vehicles, dispersion or suspension aids, surface active agents, isotonic agents, thickening or emulsifying agents, preservatives, lipidoids, liposomes, lipid nanoparticles, polymers, lipoplexes, core-shell nanoparticles, peptides, proteins, hyaluronidase, nanoparticle mimics, inert diluents, buffering agents, lubricating agents, or oils, and combinations thereof. In some examples, the composition as disclosed herein can include one or more excipients, each in an amount that together increases the stability of (i) the heterologous polypeptide or the heterologous gene encoding thereof and / or (ii) cells or modified cells.

[0271] In some aspects, the present disclosure provides a kit comprising such composition and instructions directing (i) contacting the cell with the composition (e.g., in vitro, ex vivo, or in vivo), or (ii) administration of cells comprising any one of the compositions disclosed herein to a subject. The subject may have or may be suspected of having a condition, such as a hereditary disease.

[0272] In some embodiments, any of the compositions as disclosed herein, can be administered to the subject via orally, intraperitoneally, intravenously, intraarterially, transdermally, intramuscularly, liposomally, via local delivery by catheter or stent, subcutaneously, intraadiposally, or intrathecally. In particular aspects, the compositions and systems provided herein (including polynucleotides encoding said compositions and systems, e.g., contained in an AAV vector) can be administered to a subject via intravenous administration.

[0273] Non-limiting examples of viral vectors that can be utilized to deliver the heterologous polypeptide and / or heterologous polynucleotide (or one or more genes encoding thereof) can include, but are not limited to, retroviral vectors, lentiviral vectors, adenovirus vectors, poxvirus vectors, herpesvirus vectors, or adeno-associated virus (AAV) vectors. Non- limiting examples of AAV vectors can include AAV1, AAV10, AAV106.1 / hu.37, AAV11, AAV114.3 / hu.40, AAV12, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV128.1 / hu.43, AAV128.3 / hu.44, AAV130.4 / hu.48, AAV145.1 / hu.53, AAV145.5 / hu.54, AAV145.6 / hu.55, AAV16.12 / hu.11, AAV16.3, AAV16.8 / hu.10, AAV161.10 / hu.60, AAV161.6 / hu.61, AAV1- 7 / rh.48, AAV1-8 / rh.49, AAV2, AAV2.5T, AAV2-15 / rh.62, AAV223.1, AAV223.2, AAV223.4, AAV223.5, AAV223.6, AAV223.7, AAV2-3 / rh.61, AAV24.1, AAV2-4 / rh.50, AAV2-5 / rh.51, AAV27.3, AAV29.3 / bb.1, AAV29.5 / bb.2, AAV2G9, AAV-2-pre-miRNA- 101, AAV3, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-11 / rh.53, AAV3-3, AAV33.12 / hu. 17, AAV33.4 / hu.15, AAV33.8 / hu.16, AAV3-9 / rh.52, AAV3a, AAV3b, AAV4, AAV4-19 / rh.55, AAV42.12, AAV42-10, AAV42-11, AAV42-12, AAV42-13, AAV42-15, AAV42-1b, AAV42-2, AAV42-3a, AAV42-3b, AAV42-4, AAV42-5a, AAV42-5b, AAV42-6b, AAV42- 8, AAV42-aa, AAV43-1, AAV43-12, AAV43-20, AAV43-21, AAV43-23, AAV43-25, AAV43-5, AAV4-4, AAV44.1, AAV44.2, AAV44.5, AAV46.2 / hu.28, AAV46.6 / hu.29, AAV4-8 / r11.64, AAV4-8 / rh.64, AAV4-9 / rh.54, AAV5, AAV52.1 / hu.20, AAV52 / hu. 19, AAV5-22 / rh.58, AAV5-3 / rh.57, AAV54.1 / hu.21, AAV54.2 / hu.22, AAV54.4R / hu.27,AAV54.5 / hu.23, AAV54.7 / hu.24, AAV58.2 / hu.25, AAV6, AAV6.1, AAV6.1.2, AAV6.2, AAV7, AAV7.2, AAV7.3 / hu.7, AAV8, AAV-8b, AAV-8h, AAV9, AAV9.11, AAV9.13, AAV9.16, AAV9.24, AAV9.45, AAV9.47, AAV9.61, AAV9.68, AAV9.84, AAV9.9, AAVA3.3, AAVA3.4, AAVA3.5, AAVA3.7, AAV-b, AAVC1, AAVC2, AAVC5, AAVCh.5, AAVCh.5R1, AAVcy.2, AAVcy.3, AAVcy.4, AAVcy.5, AAVCy.5R1, AAVCy.5R2, AAVCy.5R3, AAVCy.5R4, AAVcy.6, AAV-DJ, AAV-DJ8, AAVF3, AAVF5, AAV-h, AAVH-1 / hu.1, AAVH2, AAVH-5 / hu.3, AAVH6, AAVhE1.1, AAVhER1.14, AAVhEr1.16, AAVhEr1.18, AAVhER1.23, AAVhEr1.35, AAVhEr1.36, AAVhEr1.5, AAVhEr1.7, AAVhEr1.8, AAVhEr2.16, AAVhEr2.29, AAVhEr2.30, AAVhEr2.31, AAVhEr2.36, AAVhEr2.4, AAVhEr3.1, AAVhu.1, AAVhu.10, AAVhu.11, AAVhu.11, AAVhu.12, AAVhu.13, AAVhu.14 / 9, AAVhu.15, AAVhu.16, AAVhu.17, AAVhu.18, AAVhu.19, AAVhu.2, AAVhu.20, AAVhu.21, AAVhu.22, AAVhu.23.2, AAVhu.24, AAVhu.25, AAVhu.27, AAVhu.28, AAVhu.29, AAVhu.29R, AAVhu.3, AAVhu.31, AAVhu.32, AAVhu.34, AAVhu.35, AAVhu.37, AAVhu.39, AAVhu.4, AAVhu.40, AAVhu.41, AAVhu.42, AAVhu.43, AAVhu.44, AAVhu.44R1, AAVhu.44R2, AAVhu.44R3, AAVhu.45, AAVhu.46, AAVhu.47, AAVhu.48, AAVhu.48R1, AAVhu.48R2, AAVhu.48R3, AAVhu.49, AAVhu.5, AAVhu.51, AAVhu.52, AAVhu.53, AAVhu.54, AAVhu.55, AAVhu.56, AAVhu.57, AAVhu.58, AAVhu.6, AAVhu.60, AAVhu.61, AAVhu.63, AAVhu.64, AAVhu.66, AAVhu.67, AAVhu.7, AAVhu.8, AAVhu.9, AAVhu.t 19, AAVLG-10 / rh.40, AAVLG-4 / rh.38, AAVLG-9 / hu.39, AAVLG-9 / hu.39, AAV-LKO1, AAV-LK02, AAVLK03, AAV-LK03, AAV-LK04, AAV-LK05, AAV-LKO6, AAV-LK07, AAV-LK08, AAV-LK09, AAV-LK10, AAV-LK11, AAV-LK12, AAV-LK13, AAV-LK14, AAV-LK15, AAV-LK17, AAV-LK18, AAV-LK19, AAVN721-8 / rh.43, AAV-PAEC, AAV-PAEC11, AAV-PAEC12, AAV-PAEC2, AAV-PAEC4, AAV-PAEC6, AAV-PAEC7, AAV-PAEC8, AAVpi.1, AAVpi.2, AAVpi.3, AAVrh.10, AAVrh.12, AAVrh.13, AAVrh.13R, AAVrh.14, AAVrh.17, AAVrh.18, AAVrh.19, AAVrh.2, AAVrh.20, AAVrh.21, AAVrh.22, AAVrh.23, AAVrh.24, AAVrh.25, AAVrh.2R, AAVrh.31, AAVrh.32, AAVrh.33, AAVrh.34, AAVrh.35, AAVrh.36, AAVrh.37, AAVrh.37R2, AAVrh.38, AAVrh.39, AAVrh.40, AAVrh.43, AAVrh.44, AAVrh.45, AAVrh.46, AAVrh.47, AAVrh.48, AAVrh.48, AAVrh.48.1, AAVrh.48.1.2, AAVrh.48.2, AAVrh.49, AAVrh.50, AAVrh.51, AAVrh.52, AAVrh.53, AAVrh.54, AAVrh.55, AAVrh.56, AAVrh.57, AAVrh.58, AAVrh.59, AAVrh.60, AAVrh.61, AAVrh.62,AAVrh.64, AAVrh.64R1, AAVrh.64R2, AAVrh.65, AAVrh.67, AAVrh.68, AAVrh.69, AAVrh.70, AAVrh.72, AAVrh.73, AAVrh.74, AAVrh.8, AAVrh.8R, AAVrh8R, AAVrh8R A586R mutant, AAVrh8R R533A mutant, BAAV, BNP61 AAV, BNP62 AAV, BNP63 AAV, bovine AAV, caprine AAV, Japanese AAV 10, true type AAV (ttAAV), UPENN AAV 10, AAV-LK16, AAAV, AAV Shuffle 100-1, AAV Shuffle 100-2, AAV Shuffle 100-3, AAV Shuffle 100-7, AAV Shuffle 10-2, AAV Shuffle 10-6, AAV Shuffle 10-8, AAV SM 100-10, AAV SM 100-3, AAV SM 10-1, AAV SM 10-2, or AAV SM 10-8. For example, AAVrh.74 can be used as a viral vector to deliver a polynucleotide sequence encoding the heterologous polypeptide and the heterologous polynucleotide (e.g., Cas protein-gene effector fusion and one or more guide nucleic acid molecules).

[0274] Methods of non-viral delivery of nucleic acids can include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, lipid nanoparticles (LNPs), naked DNA, artificial virions, or agent-enhanced uptake of DNA. Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides can be used.

[0275] Any of the compositions disclosed herein (or one or more genes encoding any portion of the compositions), such as the heterologous gene effector(s) and / or the guide nucleic acid molecule(s), can be administered by any suitable administration route, including but not limited to, parenteral (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intracerebroventricular, intra-articular, intraperitoneal, or intracranial), intranasal, buccal, sublingual, oral, or rectal administration routes. In some instances, the pharmaceutical composition is formulated for parenteral (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intracerebroventricular, intra-articular, intraperitoneal, or intracranial) administration.

[0276] The compositions (e.g., pharmaceutical compositions) as disclosed herein can be suitable for administration to humans. In addition, such compositions can be suitable for administration to any other animal, e.g., to non-human animals, e.g. non-human mammals. Modification of pharmaceutical compositions suitable for administration to humans in order to render the compositions suitable for administration to various animals is well understood, and the ordinarily skilled veterinary pharmacologist can design and / or perform such modification with merely ordinary, if any, experimentation. Subjects to which administration of thepharmaceutical compositions is contemplated include, but are not limited to, humans and / or other primates; mammals, including commercially relevant mammals such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats; and / or birds, including commercially relevant birds such as poultry, chickens, ducks, geese, and / or turkeys. TARGET GENE(S)

[0277] The disclosure provides compositions, methods, and systems for modulating expression of target genes (e.g., target endogenous genes). For example, disclosed herein are complexes or systems that comprise a guide moiety and one or more heterologous gene effectors that can increase or decrease an activity or expression level of a target gene.

[0278] In some embodiments, a target gene or regulatory sequence thereof is endogenous to a subject, for example, present in the subject’s genome. In some embodiments, a target gene or regulatory sequence thereof is not part of an engineered reporter system.

[0279] In some embodiments, a target gene is exogenous to a host subject, for example, a pathogen target gene or an exogenous gene expressed as a result of a therapeutic intervention, such as a gene therapy and / or cell therapy. In some embodiments, a target gene is an exogenous reporter gene. In some embodiments, a target gene is an exogenous synthetic gene.

[0280] In some embodiments, a target gene (e.g., target endogenous gene) is a gene that is over-expressed or under-expressed in a disease or condition. In some embodiments, a target gene is a gene that is over-expressed or under-expressed in a heritable genetic disease.

[0281] In some embodiments, a target gene (e.g., target endogenous gene) is a gene that is over-expressed or under-expressed in an autoimmune disease. In some embodiments, a target gene is a gene that is over-expressed or under-expressed in Acute disseminated encephalomyelitis, Acute motor axonal neuropathy, Addison's disease, Adiposis dolorosa, Adult-onset Still's disease, Alopecia areata, Ankylosing Spondylitis, Anti-Glomerular Basement Membrane nephritis, Anti-neutrophil cytoplasmic antibody-associated vasculitis, Anti-N-Methyl-D-Aspartate Receptor Encephalitis, Antiphospholipid syndrome, Antisynthetase syndrome, Aplastic anemia, Autoimmune Angioedema, Autoimmune Encephalitis, Autoimmune enteropathy, Autoimmune hemolytic anemia, Autoimmune hepatitis, Autoimmune inner ear disease, Autoimmune lymphoproliferative syndrome,Autoimmune neutropenia, Autoimmune oophoritis, Autoimmune orchitis, Autoimmune pancreatitis, Autoimmune polyendocrine syndrome, Autoimmune polyendocrine syndrome type 2, Autoimmune polyendocrine syndrome type 3, Autoimmune progesterone dermatitis, Autoimmune retinopathy, Autoimmune thrombocytopenic purpura, Autoimmune thyroiditis, Autoimmune urticaria, Autoimmune uveitis, Balo concentric sclerosis, Behçet's disease, Bickerstaff's encephalitis, Bullous pemphigoid, Celiac disease, Chronic fatigue syndrome, Chronic inflammatory demyelinating polyneuropathy, Churg-Strauss syndrome, Cicatricial pemphigoid, Cogan syndrome, Cold agglutinin disease, Complex regional pain syndrome, CREST syndrome, Crohn's disease, Dermatitis herpetiformis, Dermatomyositis, Diabetes mellitus type 1, Discoid lupus erythematosus, Endometriosis, Enthesitis, Enthesitis-related arthritis, Eosinophilic esophagitis, Eosinophilic fasciitis, Epidermolysis bullosa acquisita, Erythema nodosum, Essential mixed cryoglobulinemia, Evans syndrome, Felty syndrome, Fibromyalgia, Gastritis, Gestational pemphigoid, Giant cell arteritis, Goodpasture syndrome, Graves' disease, Graves ophthalmopathy, Guillain–Barré syndrome, Hashimoto's Encephalopathy, Hashimoto Thyroiditis, Henoch-Schonlein purpura, Hidradenitis suppurativa, Idiopathic dilated cardiomyopathy, Idiopathic inflammatory demyelinating diseases, IgA nephropathy, IgG4-related systemic disease, Inclusion body myositis, Inflamatory Bowel Disease (IBD), Intermediate uveitis, Interstitial cystitis, Juvenile Arthritis, Kawasaki's disease, Lambert-Eaton myasthenic syndrome, Leukocytoclastic vasculitis, Lichen planus, Lichen sclerosus, Ligneous conjunctivitis, Linear IgA disease, Lupus nephritis, Lupus vasculitis, Lyme disease, Ménière's disease, Microscopic colitis, Microscopic polyangiitis, Mixed connective tissue disease, Mooren's ulcer, Morphea, Mucha-Habermann disease, Multiple sclerosis, Myasthenia gravis, Myocarditis, Myositis, Neuromyelitis optica, Neuromyotonia, Opsoclonus myoclonus syndrome, Optic neuritis, Ord's thyroiditis, Palindromic rheumatism, Paraneoplastic cerebellar degeneration, Parry Romberg syndrome, Parsonage-Turner syndrome, Pediatric Autoimmune Neuropsychiatric Disorder Associated with Streptococcus, Pemphigus vulgaris, Pernicious anemia, Pityriasis lichenoides et varioliformis acuta, POEMS syndrome, Polyarteritis nodosa, Polymyalgia rheumatica, Polymyositis, Postmyocardial infarction syndrome, Postpericardiotomy syndrome, Primary biliary cirrhosis, Primary immunodeficiency, Primary sclerosing cholangitis, Progressive inflammatory neuropathy, Psoriasis, Psoriatic arthritis, Pure red cell aplasia, Pyodermagangrenosum, Raynaud’s phenomenon, Reactive arthritis, Relapsing polychondritis, Restless leg syndrome, Retroperitoneal fibrosis, Rheumatic fever, Rheumatoid arthritis, Rheumatoid vasculitis, Sarcoidosis, Schnitzler syndrome, Scleroderma, Sjogren's syndrome, Stiff person syndrome, Subacute bacterial endocarditis, Susac's syndrome, Sydenham chorea, Sympathetic ophthalmia, Systemic Lupus Erythematosus, Systemic scleroderma, Thrombocytopenia, Tolosa-Hunt syndrome, Transverse myelitis, Ulcerative colitis, Undifferentiated connective tissue disease, Urticaria, Urticarial vasculitis, Vasculitis, or Vitiligo.

[0282] In some embodiments, a target gene (e.g., target endogenous gene) is a gene that is over-expressed or under-expressed in a cancer, for example, acute leukemia, astrocytomas, biliary cancer (cholangiocarcinoma), bone cancer, breast cancer, brain stem glioma, bronchioloalveolar cell lung cancer, cancer of the adrenal gland, cancer of the anal region, cancer of the bladder, cancer of the endocrine system, cancer of the esophagus, cancer of the head or neck, cancer of the kidney, cancer of the parathyroid gland, cancer of the penis, cancer of the pleural / peritoneal membranes, cancer of the salivary gland, cancer of the small intestine, cancer of the thyroid gland, cancer of the ureter, cancer of the urethra, carcinoma of the cervix, carcinoma of the endometrium, carcinoma of the fallopian tubes, carcinoma of the renal pelvis, carcinoma of the vagina, carcinoma of the vulva, cervical cancer, chronic leukemia, colon cancer, colorectal cancer, cutaneous melanoma, ependymoma, epidermoid tumors, Ewings sarcoma, gastric cancer, glioblastoma, glioblastoma multiforme, glioma, hematologic malignancies, hepatocellular (liver) carcinoma, hepatoma, Hodgkin's Disease, intraocular melanoma, Kaposi sarcoma, lung cancer, lymphomas, medulloblastoma, melanoma, meningioma, mesothelioma, multiple myeloma, muscle cancer, neoplasms of the central nervous system (CNS), neuronal cancer, small cell lung cancer, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pediatric malignancies, pituitary adenoma, prostate cancer, rectal cancer, renal cell carcinoma, sarcoma of soft tissue, schwanoma, skin cancer, spinal axis tumors, squamous cell carcinomas, stomach cancer, synovial sarcoma, testicular cancer, uterine cancer, or tumors and their metastases, including refractory versions of any of the above cancers, or a combination thereof.

[0283] In some embodiments, a target gene (e.g., target endogenous gene) is a differentiation-associated gene, for example, SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1- 85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56,CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30, CD50, AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1 alpha / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT Activators, STAT Inhibitors, STAT3, STAT4, STAT5a, STAT6, TSC22, DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, Myocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT Activators, STAT Inhibitors, STAT1, STAT3, TBX18, Twist-1, Twist-2, Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA- 2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3 alpha / FoxA1, c-Jun, KLF2, KLF4, KLF5, c- Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB Activators, NFkB / IkB Inhibitors, NFkB1, NFkB2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT Activators, STAT Inhibitors, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, ZNF281, KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct- 3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, TBX18, ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4 alpha / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT Activators, STAT Inhibitors, STAT3, SUZ12, TCF-3 / E2A, TCF7 / TCF1, Androgen R / NR3C4, AP-2 gamma, beta-Catenin, beta-Catenin Inhibitors, Brachyury, CREB, ER alpha / NR3A1, ER beta / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI-2, GLI-3, HIF-1 alpha / HIF1A, HIF-2 alpha / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB Activators, NFkB / IkB Inhibitors, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT Activators, STAT Inhibitors, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, or ZEB1.

[0284] In some embodiments, modulation of the expression level and / or epigenetic level (e.g., methylation level) of the target gene in the target cell (e.g., muscle cell) can effectmodification (e.g., upregulation or downregulation) of a downstream gene (e.g., one or more downstream genes) of the target gene. In some cases, the target gene can be encoded by the D4Z4 repeat array (e.g., target gene being DUX4), and the downstream genes that are in turn modified in their expressions (e.g., downregulated) can include, but are not limited to, ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, DEFB103, ZFN217, RNASEL, EIF2AK2, BMP2, SP1 P21, MYC, MURF1, ATROGIN1, CRYM, PRAMEF1, RFPL2, KHDC1, SPRYD5, TPRX1, HSPA2, FGFR3, SLC2A14, ID2, PVRL3, SFRS2B, THOC4, ZNHIT6, DBR1, TFIP11, FBXO33, USP29, TRIM23, SLC34A2, CSAG3, and / or PNMA6B.

[0285] In some embodiments, modulation of the expression level and / or epigenetic level (e.g., methylation level) of the target gene in the target cell can effect apoptosis of the target cell (e.g., muscle cell). In some cases, such modulation of the target gene can reduce stress in the target cell. For example, the modulation of the target gene (e.g., DUX4) can effect downregulation of one or more stress-related markers in the target cell. Non-limiting examples of the one or more stress-related markers can include ACTH, glucocorticoid receptor, CRHR- 1 / 2, POMC, prolactin, arginine vasopressin receptor V1a, superoxide dismutase 1, superoxide dismutase 2, peroxiredoxin-3, CCR5, iNOS, eNOS, heme oxygenase-2, cyclooxygenase-2, HSP27, HSP40, HSP60, HSP70, HSP70i, HSP90, HSP110, GRP78 / BIP, AIF, annexin II, annexin IV, caspase 1, caspase 2, caspase 3, caspase 6, cytokeratin, E-cadherin, and / or Annexin V, caspase 5, caspase 7, caspase 8, caspase 9, caspase 10, BAD, BAX, BAK, BCL2, BID, PARP-1, NOXA, PUMA, RIPK3, RIPK1, FADD, APAF1, DFF40, DFF45, or ROCK. The one or more stress-related markers as disclosed herein can be an apoptotic marker. Without wishing to be bound by theory, a higher expression level of the one or more stress- related markers provided herein in a cell can indicate a higher apoptosis level in the cell.

[0286] In some embodiments, a heterologous gene effector is from a gene product that is a hematopoietic stem cell transcription factor. In some embodiments, a target gene is a mesenchymal stem cell transcription factor. In some embodiments, a target gene is an embryonic stem cell transcription factor. In some embodiments, a target gene is an induced pluripotent stem cell (iPSC) transcription factor. In some embodiments, a target gene is an epithelial stem cell transcription factor. In some embodiments, a target gene is a cancer stem cell transcription factor.

[0287] In some embodiments, a target gene is an age-related gene. In some embodiments, a target gene is a senescence-associated protein. In some embodiments, a target gene is a drug target.

[0288] In some embodiments, a target gene (e.g., target endogenous gene) is a cancer-related gene. Non-limiting examples of cancer-related genes include A1CF, ABI1, ABL1, ABL2, ACKR3, ACSL3, ACSL6, ACVR1, ACVR2A, AFDN, AFF1, AFF3, AFF4, AKAP9, AKT1, AKT2, AKT3, ALDH2, ALK, AMER1, ANK1, APC, APOBEC3B, AR, ARAF, ARHGAP26, ARHGAP5, ARHGEF10, ARHGEF10L, ARHGEF12, ARID1A, ARID1B, ARID2, ARNT, ASPSCR1, ASXL1, ASXL2, ATF1, ATIC, ATM, ATP1A1, ATP2B3, ATR, ATRX, AXIN1, AXIN2, B2M, BAP1, BARD1, BAX, BAZ1A, BCL10, BCL11A, BCL11B, BCL2, BCL2L12, BCL3, BCL6, BCL7A, BCL9, BCL9L, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BIRC6, BLM, BMP5, BMPR1A, BRAF, BRCA1, BRCA2, BRD3, BRD4, BRIP1, BTG1, BTK, BUB1B, C15orf65, CACNA1D, CALR, CAMTA1, CANT1, CARD11, CARS, CASP3, CASP8, CASP9, CBFA2T3, CBFB, CBL, CBLB, CBLC, CCDC6, CCNB1IP1, CCNC, CCND1, CCND2, CCND3, CCNE1, CCR4, CCR7, CD209, CD274, CD28, CD74, CD79A, CD79B, CDC73, CDH1, CDH10, CDH11, CDH17, CDK12, CDK4, CDK6, CDKN1A, CDKN1B, CDKN2A, CDKN2C, CDX2, CEBPA, CEP89, CHCHD7, CHD2, CHD4, CHEK2, CHIC2, CHST11, CIC, CIITA, CLIP1, CLP1, CLTC, CLTCL1, CNBD1, CNBP, CNOT3, CNTNAP2, CNTRL, COL1A1, COL2A1, COL3A1, COX6C, CPEB3, CREB1, CREB3L1, CREB3L2, CREBBP, CRLF2, CRNKL1, CRTC1, CRTC3, CSF1R, CSF3R, CSMD3, CTCF, CTNNA2, CTNNB1, CTNND1, CTNND2, CUL3, CUX1, CXCR4, CYLD, CYP2C8, CYSLTR2, DAXX, DCAF12L2, DCC, DCTN1, DDB2, DDIT3, DDR2, DDX10, DDX3X, DDX5, DDX6, DEK, DGCR8, DICER1, DNAJB1, DNM2, DNMT3A, DROSHA, DUX4L1 (or Dux4), EBF1, ECT2L, EED, EGFR, EIF1AX, EIF3E, EIF4A2, ELF3, ELF4, ELK4, ELL, ELN, EML4, EP300, EPAS1, EPHA3, EPHA7, EPS15, ERBB2, ERBB3, ERBB4, ERC1, ERCC2, ERCC3, ERCC4, ERCC5, ERG, ESR1, ETNK1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT1, EXT2, EZH2, EZR, FAM131B, FAM135B, FAM47C, FANCA, FANCC, FANCD2, FANCE, FANCF, FANCG, FAS, FAT1, FAT3, FAT4, FBLN2, FBXO11, FBXW7, FCGR2B, FCRL4, FEN1, FES, FEV, FGFR1, FGFR1OP, FGFR2, FGFR3, FGFR4, FH, FHIT, FIP1L1, FKBP9, FLCN, FLI1, FLNA, FLT3, FLT4, FNBP1, FOXA1, FOXL2, FOXO1, FOXO3, FOXO4, FOXP1, FOXR1, FSTL3, FUBP1,FUS, GAS7, GATA1, GATA2, GATA3, GLI1, GMPS, GNA11, GNAQ, GNAS, GOLGA5, GOPC, GPC3, GPC5, GPHN, GRIN2A, GRM3, H3F3A, H3F3B, HERPUD1, HEY1, HIF1A, HIP1, HIST1H3B, HIST1H4I, HLA-A, HLF, HMGA1, HMGA2, HMGN2P46, HNF1A, HNRNPA2B1, HOOK3, HOXA11, HOXA13, HOXA9, HOXC11, HOXC13, HOXD11, HOXD13, HRAS, HSP90AA1, HSP90AB1, ID3, IDH1, IDH2, IGF2BP2, IGH, IGK, IGL, IKBKB, IKZF1, IL2, IL21R, IL6ST, IL7R, IRF4, IRS4, ISX, ITGAV, ITK, JAK1, JAK2, JAK3, JAZF1, JUN, KAT6A, KAT6B, KAT7, KCNJ5, KDM5A, KDM5C, KDM6A, KDR, KDSR, KEAP1, KIAA1549, KIF5B, KIT, KLF4, KLF6, KLK2, KMT2A, KMT2C, KMT2D, KNL1, KNSTRN, KRAS, KTN1, LARP4B, LASP1, LATS1, LATS2, LCK, LCP1, LEF1, LEPROTL1, LHFPL6, LIFR, LMNA, LMO1, LMO2, LPP, LRIG3, LRP1B, LSM14A, LYL1, LZTR1, MACC1, MAF, MAFB, MALAT1, MALT1, MAML2, MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MAX, MB21D2, MDM2, MDM4, MDS2, MECOM, MED12, MEN1, MET, MGMT, MITF, MLF1, MLH1, MLLT1, MLLT10, MLLT11, MLLT3, MLLT6, MN1, MNX1, MPL, MRTFA, MSH2, MSH6, MSI2, MSN, MTCP1, MTOR, MUC1, MUC16, MUC4, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, MYH11, MYH9, MYO5A, MYOD1, N4BP2, NAB2, NACA, NBEA, NBN, NCKIPSD, NCOA1, NCOA2, NCOA4, NCOR1, NCOR2, NDRG1, NF1, NF2, NFATC2, NFE2L2, NFIB, NFKB2, NFKBIE, NIN, NKX2-1, NONO, NOTCH1, NOTCH2, NPM1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1, NTRK1, NTRK3, NUMA1, NUP214, NUP98, NUTM1, NUTM2B, NUTM2D, OLIG2, OMD, P2RY8, PABPC1, PAFAH1B2, PALB2, PATZ1, PAX3, PAX5, PAX7, PAX8, PBRM1, PBX1, PCBP1, PCM1, PDCD1LG2, PDE4DIP, PDGFB, PDGFRA, PDGFRB, PER1, PHF6, PHOX2B, PICALM, PIK3CA, PIK3CB, PIK3R1, PIM1, PLAG1, PLCG1, PML, PMS1, PMS2, POLD1, POLE, POLG, POLQ, POT1, POU2AF1, POU5F1, PPARG, PPFIBP1, PPM1D, PPP2R1A, PPP6C, PRCC, PRDM1, PRDM16, PRDM2, PREX2, PRF1, PRKACA, PRKAR1A, PRKCB, PRPF40B, PRRX1, PSIP1, PTCH1, PTEN, PTK6, PTPN11, PTPN13, PTPN6, PTPRB, PTPRC, PTPRD, PTPRK, PTPRT, PWWP2A, QKI, RABEP1, RAC1, RAD17, RAD21, RAD51B, RAF1, RALGDS, RANBP2, RAP1GDS1, RARA, RB1, RBM10, RBM15, RECQL4, REL, RET, RFWD3, RGPD3, RGS7, RHOA, RHOH, RMI2, RNF213, RNF43, ROBO2, ROS1, RPL10, RPL22, RPL5, RPN1, RSPO2, RSPO3, RUNX1, RUNX1T1, S100A7, SALL4, SBDS, SDC4, SDHA, SDHAF2, SDHB, SDHC, SDHD, 44444, 44445, 44448, SET, SETBP1, SETD1B,SETD2, SETDB1, SF3B1, SFPQ, SFRP4, SGK1, SH2B3, SH3GL1, SHTN1, SIRPA, SIX1, SIX2, SKI, SLC34A2, SLC45A3, SMAD2, SMAD3, SMAD4, SMARCA4, SMARCB1, SMARCD1, SMARCE1, SMC1A, SMO, SND1, SNX29, SOCS1, SOX2, SOX21, SPECC1, SPEN, SPOP, SRC, SRGAP3, SRSF2, SRSF3, SS18, SS18L1, SSX1, SSX2, SSX4, STAG1, STAG2, STAT3, STAT5B, STAT6, STIL, STK11, STRN, SUFU, SUZ12, SYK, TAF15, TAL1, TAL2, TBL1XR1, TBX3, TCEA1, TCF12, TCF3, TCF7L2, TCL1A, TEC, TENT5C, TERT, Tet1, Tet2, TFE3, TFEB, TFG, TFPT, TFRC, TGFBR2, THRAP3, TLX1, TLX3, TMEM127, TMPRSS2, TNC, TNFAIP3, TNFRSF14, TNFRSF17, TOP1, TP53, TP63, TPM3, TPM4, TPR, TRA, TRAF7, TRB, TRD, TRIM24, TRIM27, TRIM33, TRIP11, TRRAP, TSC1, TSC2, TSHR, U2AF1, UBR5, USP44, USP6, USP8, VAV1, VHL, VTI1A, WAS, WDCP, WIF1, WNK2, WRN, WT1, WWTR1, XPA, XPC, XPO1, YWHAE, ZBTB16, ZCCHC8, ZEB1, ZFHX3, ZMYM2, ZMYM3, ZNF331, ZNF384, ZNF429, ZNF479, ZNF521, ZNRF3, or ZRSR2.

[0289] In some embodiments, the target gene as provided herein can be Dux4. The target polynucleotide sequence that can be targeted (e.g., bound) by the systems of the present disclosure (e.g., Cas or dCas protein and / or guide nucleic acid molecule) can be at or adjacent to a D4Z4 repeat array. For example, a guide nucleic acid molecule (e.g., a guide RNA) can comprise (i) a scaffold sequence configured to complex with a Cas protein and (ii) a spacer sequence exhibiting specific binding to the target polynucleotide sequence. Any suitable scaffold sequence can be used to complex with a Cas protein. Non-liming examples of suitable scaffold sequences are disclosed in, e.g., International Publication No. WO2023 / 168242, the entirety of which is hereby incorporated by reference. In another example, a nucleic acid that is heterologous to the cell and that functions in absence of a Cas / dCas protein (e.g., antisense oligonucleotides, small interfering RNAs, ribozymes, etc.) can exhibit specific binding to the target polynucleotide sequence. In a different example, non-Cas / dCas protein (e.g., ZFN, Talen, etc.) can exhibit specific binding to the target polynucleotide sequence.

[0290] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, atleast or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to one or more polynucleotide sequences of Table 1 (e.g., one or more polynucleotide sequences of SEQ ID NOs: 1-14) or a complementary sequence thereof. In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to one or more polynucleotide sequences of Table 2 (e.g., one or more polynucleotide sequences of SEQ ID NOs: 45-726) or a complementary sequence thereof. In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to one or more polynucleotide sequences of Table 4 (e.g., one or more polynucleotide sequences of SEQ ID NOs: 800-881) or a complementary sequence thereof.

[0291] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to effect reduced expression of Dux4 in the muscle cell by a threshold level of at least or at least about 50% as compared to the control cell. In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or atleast about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity to the polynucleotide sequence of one or more members selected from the group consisting of SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 854, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 824, SEQ ID NO: 829, SEQ ID NO: 840, SEQ ID NO: 848, SEQ ID NO: 860, SEQ ID NO: 825, SEQ ID NO: 828, SEQ ID NO: 826, SEQ ID NO: 864, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865, and complementary sequence(s) thereof.

[0292] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to effect reduced expression of the target gene (e.g., Dux4) by a threshold level (e.g., a predetermined threshold level) in a muscle cell (e.g., FSHD muscle cell) as compared to a control cell (e.g., lacking the system for suppressing the target gene). The threshold level can be at least or at least about 40%, at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 99%, or substantially about 100%. For example, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to effect reduced expression of Dux4 in the muscle cell by a threshold level of at least or at least about 50% as compared to the control cell. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to the polynucleotide sequence of one or more members selected from the group consisting of SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851,SEQ ID NO: 854, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 841, and complementary sequence(s) thereof.

[0293] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to effect reduced expression of the target gene (e.g., Dux4) while exhibiting off-target effect below a threshold level (e.g., a predetermined threshold level) in a muscle cell (e.g., FSHD muscle cell) as compared to a control cell (e.g., lacking the system for suppressing the target gene). The off-target effect can be determined in vitro, ex vivo, in vivo, or in silico. For example, in an in silico off-target prediction analysis, an off-target site may be identified when a similar genomic sequence having less than a predetermined edit distance or “edit distance” (e.g., a sum of mismatch, deletion, and insertion identified) (e.g., less than 3 edit distances) is (i) found in a target cell of interest (e.g., a muscle cell, such as an adult skeletal muscle cell), (ii) found in a non-active or “quiescent” chromatin region of the target cell, and / or (iii) is not found in a promoter or exonic region (e.g., not intergenic, not intronic, etc.) of the chromosome of the target cell. Based on identification of such off-target site(s), an off-target activity level (e.g., a predicted off-target activity level) can be determined (e.g., based one or more methods described in Muhammad Naeem et al., Cells, 9(7), 16008 (2020), which is incorporated herein by reference in its entirety. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected for having no more than about 5, no more than about 4, no more than about 3, no more than about 2, no more than about 1, or substantially no off-target site(s) identified. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected for exhibiting an off-target activity level of no more than about 30%, no more than about 25%, no more than about 20%, no more than about 15%, no more than about 14%, no more than about 13%, no more than about 12%, no more than about 11%, no more than about 10%, no more than about 9%, no more than about 8%, no more than about 7%, no more than about 6%, no more than about 5%, no more than about 4%, no more than about 3%, no more than about 2%, no more than about 1%, or substantially about 0%. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or atleast about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to the polynucleotide sequence of one or more members selected from the group consisting of SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 824, SEQ ID NO: 829, SEQ ID NO: 840, SEQ ID NO: 825, SEQ ID NO: 828, and complementary sequence(s) thereof.

[0294] In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected for having at most about one or substantially no off-target site, which off-target site having an edit distance of two. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected for having at most about one or substantially no off-target site, which off-target site having an edit distance of one. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected for having at most about one or substantially no off-target site, which off-target site having an edit distance of zero.

[0295] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to (i) effect reduced expression of the target gene (e.g., Dux4) by a threshold level as abovementioned (e.g., at least or at least about 95% or at least or at least about 98% reduction in expression level of the target gene) and (ii) exhibit off-target effect below a threshold level as abovementioned (e.g., no more than about 10% or substantially 0% off-target activity). In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to the polynucleotide sequence of one or more members selected from thegroup consisting of SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865, and complementary sequence(s) thereof. In some examples, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can comprise (or the spacer sequence is encoded by) a polynucleotide sequence that exhibits at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity to the polynucleotide sequence of one or more members selected from the group consisting of SEQ ID NO: 800, SEQ ID NO: 836, SEQ ID NO: 851, and complementary sequence(s) thereof. TABLE 4. Examples Of A Target Polynucleotide Sequence Or Guide Nucleic Acid Spacer Sequence. SEQ Guide nucleic acid spacerSEQGuide nucleic acid spacer ID sequence (without scaffold ID sequence (without scaffold NO: sequence) NO: sequence) 800 TTTATTTTTTCACCCAGAACAGTAACT 841 TTTGCCTTCAACTTCACCTTAACAACA801 TTTAAAGAGATCTGGGGATCTATACAG 842 TTTGTATGTTGTTAAGGTGAAGTTGAA802 TTTAACTTGGAAACACAGCGAAGTCCA 843 TTTAATTTTTCAACATAATTAACTCTC803 TTTGCACTGGAGCAGAGATGACCACAG 844 TTTAAAAATAACGAATGAGCAAAATAT804 TTTGCCTGTGAGTTCGAATGCACTTTA 845 TTTAGATTCTATTGTCTATTTTCTTCC805 TTTGGGAATGTGTTTGTGAAGCACCTA 846 TTTGTGAAGCACCTAGAATCTATAGCC806 TTTGGCTTTTTGATAAATTGTCTAATG 847 TTTGATAAATTGTCTAATGACTAGATT807 TTTATCAAAAAGCCAAACATTTCAACA 848 TTTGCTCATTCGTTATTTTTAAATTTC808 TTTGATGAAGTCTGGCTTACAGCCTGT 849 TTTAAGTTCTCCATCAGATATGCAAAA809 TTTACTTCCGTCACTTTCTTAACATTA 850 TTTGCCTAGACAGCGTCGGAAGGTGGG810 TTTGTAATGTTAAGAAAGTGACGGAAG 851 TTTATAAATTCACTACAGAGACACAAC811 TTTAAGATTCTGGGAGGGAGAGAAAAA 852 TTTGCCCTGGGGACCTTAGCAATGGGC812 TTTATATATGATTTGTATTTTCACAGA 853 TTTGTATTTTCACAGAGATTTAAGAAT813 TTTAATGCACCATTATAGTAGAAAATT 854 TTTAAAAACCCAACAGAAATCATAGAG814 TTTGAAATCTGGAAAGTTCTTAGCATC 855 TTTAAAAAAAAAAAATCACAAGGCACA815 TTTAAAAAGAATAGAGGGAGAAAATGG 856 TTTGCCCGCTTCCTGGCTAGACCTGCG816 TTTAAAGAATGGGAAAATTACGGGTGA 857 TTTATGATGCTGTCCAGCATCATTTAA817 TTTAAAATATTAGTTTCCAGGACTCAA 858 TTTACATTGAACAGAGAGCTTTATATTSEQ Guide nucleic acid spacerSEQGuide nucleic acid spacer ID sequence (without scaffold ID sequence (without scaffold NO: sequence) NO: sequence)818 TTTATCTCTTTGTTGATATTTTGCTCA 859 TTTGGCATTGCTTTTGGGGATCTGGGA819 TTTGTTGATATTTTGCTCATTCGTTAT 860 TTTATGTTCTCACAAGATTCTGGGAGG820 TTTAAATTTCACTCAGTTGTCTCTTTC 861 TTTGGGGATCTGGGAAAATCTGTGCAC821 TTTATGTTTTTCTTCCAATGGGGAATA 862 TTTGGTTTCCGCGTGGCTTTGCCCTCC822 TTTAAAGACTGGCTCAGTAAAGGGGGA 863 TTTGTCCCGGAGGAAACCGCCCACTCC823 TTTGCTACAGCACTAGTGAAACTGCAA 864 TTTACAAGGGCGGCTGGCTGGCTGGCT824 TTTAATTCTCTCCTGAAGGAGATACTG 865 TTTGCCCTCCGCAAGGCGGCCTGTTGC825 TTTGAATATACTGTGGTCATCTCTGCT 866 TTTGCTCCCGGAGCTCTGCGGGCACCC826 TTTATAAATAATGGCATGACAAGGGTC 867 TTTGGAACCTGGCAAGGAGAGCGAAGG827 TTTAGCATTTTTTTTCCTAGGGTTATT 868 TTTGAGAAGGATCGCTTTCCAGGCATC828 TTTGCTCACTGAGAATGCATAAGATGA 869 TTTGAGCGGAACCCGTACCCGGGCATC829 TTTGATGAGTGCTGTATAGATCCCCAG 870 TTTGGTTTCAGAATGAGAGGTCACGCC830 TTTGTGTCTGCTGAGAAGAAAGATGAG 871 TTTGGACCCCGAGCCAAAGCGAGGCCC831 TTTAAGAATTTAATGCACCATTATAGT 872 TTTGGCTCGGGGTCCAAACGAGTCTCC832 TTTGAAATACAGTATTTCCCAGATCAA 873 TTTAGGACGCGGGGTTGGGACGGGGTC833 TTTATGCCATTTTCTCCCTCTATTCTT 874 TTTAGGACGCGGGGTTGGGACGGGGTC834 TTTACTGAGCCAGTCTTTAAATGCTAG 875 TTTATATTTTCATGTGGTTTTATGATG835 TTTAAATGCTAGATTTGATGAGTGCTG 876 TTTACAAGAGAAAAACAAAAAACCCTA836 TTTAAAAGACTCTATCTCTGAATGTAT 877 TTTAATAGGGTTTTTTGTTTTTCTCTT837 TTTGCATATCTGATGGAGAACTTAAAA 878 TTTCACCCAGAACAGTAACT838 TTTATTTGTTAAAATTCAGTTTCTGAA 879 CCCAGAACAGTAACT839 TTTATAAATCTATTGTGCCTCAAGTCA 880 AACAGTAACT840 TTTGACCGCCAGGCGCTCCGTGCTGGC 881 TAACTCELLS

[0296] Compositions, methods, and systems of the disclosure can be applied to cells of various types, and populations thereof. For example, a complex or system of the disclosure can be used to elicit changes in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in cells of a particular type, or populations thereof. Methods of the disclosure can be used to identify complexes that are capable of eliciting changes in the expression or activity of target genes (e.g., target endogenous genes) in cells of a particular type, or populations thereof.

[0297] In some embodiments, a complex or system or a heterologous gene effector identified by methods of the disclosure effects a desirable change in expression of a target gene (e.g., target endogenous gene) that is specific to a particular cell type. In some embodiments,a complex or system or a heterologous gene effector identified by methods of the disclosure effects a desirable change in expression of a target gene (e.g., target endogenous gene) that is applicable to two or more cell types. In some embodiments, a complex or system or a heterologous gene effector identified by methods of the disclosure effects a desirable change in expression of a target gene (e.g., target endogenous gene) that is applicable to three or more cell types. In some embodiments, a complex or system or a heterologous gene effector identified by methods of the disclosure effects a desirable change in expression of a target gene (e.g., target endogenous gene) that is applicable to a class of cell types, for example, cell types with overlapping functional roles, that are present in similar tissues, or that are from the same or similar differentiation lineages, e.g., stem cells, immune cells, T cells, T effector cells, etc. In some embodiments, a complex or system or a heterologous gene effector identified by methods of the disclosure effects a desirable change in expression of a target gene (e.g., target endogenous gene) that is broadly applicable to a wide variety of cell types, for example, elicits an expression level of a target gene that is above or below a certain threshold for multiple target cell types when introduced to the cells using suitable methods.

[0298] In some embodiments, a composition, complex, system, or method of the disclosure is used to effect a change in the expression, epigenetic modification, or activity level of a target gene in a primary cell. In some embodiments, a composition, complex, system, or method of the disclosure is used to effect a change in the expression, epigenetic modification, or activity level of a target gene in a cell line. In some embodiments, a composition, complex, system, or method of the disclosure is used to effect a change in the expression, epigenetic modification, or activity level of a target gene in an immortalized cell.

[0299] In some embodiments, a composition, complex, system, or method of the disclosure is used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a mammalian cell, for example, a human cell, non-human primate cell, non-rodent mammal cell, non-human mammal cell, swine cell, lagomorph cell, canine cell, etc. In some embodiments, a composition, complex, system, or method of the disclosure is used to effect a change in the expression, epigenetic modification, or activity level of a target gene in a plant cell, an avian cell, a reptilian cell, a bacterial cell, or an archaeal cell.

[0300] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a human cell.

[0301] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a stem cell.

[0302] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a differentiated cell.

[0303] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a disease-associated cell.

[0304] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a cancer cell.

[0305] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a non-cancer cell.

[0306] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a lymphoid cell, such as a B cell, a T cell (Cytotoxic T cell, Natural Killer T cell, Regulatory T cell, T helper cell), Natural killer cell, cytokine induced killer (CIK) cells (see e.g. US20080241194); myeloid cells, such as granulocytes (Basophil granulocyte, Eosinophil granulocyte, Neutrophil granulocyte / Hypersegmented neutrophil), Monocyte / Macrophage, Red blood cell, Reticulocyte, Mast cell, Thrombocyte / Megakaryocyte, Dendritic cell; cells from the endocrine system, including thyroid (Thyroid epithelial cell, Parafollicular cell), parathyroid (Parathyroid chief cell, Oxyphil cell), adrenal (Chromaffin cell), pineal (Pinealocyte) cells; cells of the nervous system, including glial cells (Astrocyte, Microglia), Magnocellular neurosecretory cell, Stellate cell, Boettcher cell, and pituitary (Gonadotrope, Corticotrope, Thyrotrope, Somatotrope, Lactotroph); cells of the Respiratory system, including Pneumocyte (Type I pneumocyte, TypeII pneumocyte), Clara cell, Goblet cell, Dust cell; cells of the circulatory system, including Myocardiocyte, Pericyte; cells of the digestive system, including stomach (Gastric chief cell, Parietal cell), Goblet cell, Paneth cell, G cells, D cells, ECL cells, I cells, K cells, S cells; enteroendocrine cells, including enterochromaffm cell, APUD cell, liver cells (e.g., Hepatocyte, or Kupffer cell), Cartilage / bone / muscle; bone cells, including Osteoblast, Osteocyte, Osteoclast, teeth cells, (Cementoblast, Ameloblast); cartilage cells, including Chondroblast, Chondrocyte; skin cells, including Trichocyte, Keratinocyte, Melanocyte (Nevus cell); muscle cells, including Myocyte; urinary system cells, including Podocyte, Juxtaglomerular cell, Intraglomerular mesangial cell / Extraglomerular mesangial cell, Kidney proximal tubule brush border cell, Macula densa cell; reproductive system cells, including Spermatozoon, Sertoli cell, Leydig cell, Ovum; and other cells, including Adipocyte, Fibroblast, Tendon cell, Epidermal keratinocyte, Epidermal basal cell, Keratinocyte of fingernails and toenails, Nail bed basal cell, Medullary hair shaft cell, Cortical hair shaft cell, Cuticular hair shaft cell, Cuticular hair root sheath cell, Hair root sheath cell of Huxley's layer, Hair root sheath cell of Henle's layer, External hair root sheath cell, Hair matrix cell, Wet stratified barrier epithelial cells, Surface epithelial cell of stratified squamous epithelium of cornea, tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, basal cell of epithelia of cornea, tongue, oral cavity, esophagus, anal canal, distal urethra and vagina, Urinary epithelium cell, Exocrine secretory epithelial cells, Salivary gland mucous cell, Salivary gland serous cell, Von Ebner's gland cell in tongue, Mammary gland cell, Lacrimal gland cell, Ceruminous gland cell in ear, Eccrine sweat gland dark cell, Eccrine sweat gland clear cell. Apocrine sweat gland cell, Gland of Moll cell in eyelid, Sebaceous gland cell, Bowman's gland cell in nose, Brunner's gland cell in duodenum, Seminal vesicle cell, Prostate gland cell, Bulbourethral gland cell, Bartholin's gland cell, Gland of Littre cell, Uterus endometrium cell, Isolated goblet cell of respiratory and digestive tracts, Stomach lining mucous cell, Gastric gland zymogenic cell, Gastric gland oxyntic cell, Pancreatic acinar cell, Paneth cell of small intestine, Type II pneumocyte of lung, Clara cell of lung, Hormone secreting cells, Anterior pituitary cells, Somatotropes, Lactotropes, Thyrotropes, Gonadotropes, Corticotropes, Intermediate pituitary cell, Magnocellular neurosecretory cells, Gut and respiratory tract cells, Thyroid gland cells, thyroid epithelial cell, parafollicular cell, Parathyroid gland cells, Parathyroid chief cell, Oxyphil cell, Adrenal gland cells, chromaffincells, Ley dig cell of testes, Theca interna cell of ovarian follicle, Corpus luteum cell of ruptured ovarian follicle, Granulosa lutein cells, Theca lutein cells, Juxtaglomerular cell, Macula densa cell of kidney, Metabolism and storage cells, Barrier function cells (e.g., Lung, Gut, Exocrine Glands and Urogenital Tract), Kidney, Type I pneumocyte, Pancreatic duct cell (centroacinar cell), Nonstriated duct cell (of sweat gland, salivary gland, mammary gland, etc.), Duct cell (of seminal vesicle, prostate gland, etc.), Epithelial cells lining closed internal body cavities, Ciliated cells with propulsive function, Extracellular matrix secretion cells, Contractile cells; Skeletal muscle cells, stem cell, Heart muscle cells, Blood and immune system cells, Erythrocyte, Megakaryocyte, Monocyte, Connective tissue macrophage (various types), Epidermal Langerhans cell, Osteoclast, Dendritic cell, Microglial cell, Neutrophil granulocyte, Eosinophil granulocyte, Basophil granulocyte, Mast cell, Helper T cell, Suppressor T cell, Cytotoxic T cell, Natural Killer T cell, B cell, Natural killer cell, Reticulocyte, Stem cells and committed progenitors for the blood and immune system (various types), Pluripotent stem cells, Totipotent stem cells, Induced pluripotent stem cells, adult stem cells, Sensory transducer cells, neurons, Autonomic neuron cells, Sense organ and peripheral neuron supporting cells, Central nervous system neurons and glial cells, Lens cells, Pigment cells, Melanocyte, Retinal pigmented epithelial cell, Germ cells, Oogonium / Oocyte, Spermatid, Spermatocyte, Spermatogonium cell, Spermatozoon, Nurse cells, Ovarian follicle cell, Sertoli cell, Thymus epithelial cell, Interstitial cells, Interstitial kidney cells, common myeloid progenitors, common lymphoid progenitors, or stem cells that are differentiated into or are to be differentiated into any cell type disclosed herein.

[0307] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a stem cell, for example, an isolated stem cell (e.g., an ESC) or an induced stem cell (e.g., an iPSC).

[0308] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in a hematopoietic stem cell, for example, a hematopoietic stem cell from a subject, for example, from bone marrow, or peripheral blood (e.g., a mobilized peripheral blood apheresis product, for example, mobilized by administration of GCSF, GM- CSF, mozobil, or a combination thereof).

[0309] In some cases, pluripotency of stem cells (e.g., ESCs or iPSCs) can be determined, in part, by assessing pluripotency characteristics of the cells. Pluripotency characteristics can include, but are not limited to: pluripotent stem cell morphology; the potential for unlimited self-renewal; expression of pluripotent stem cell markers including, but not limited to SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30 and / or CD50; ability to differentiate to all three somatic lineages (ectoderm, mesoderm and endoderm); ability to form teratomas comprising the three somatic lineages; and / or (vi) formation of embryoid bodies comprising cells from the three somatic lineages.

[0310] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of a target gene (e.g., target endogenous gene) in an immune cell, for example, lymphocytes, T cells, CD4+ T cells, CD8+ T cells, alpha-beta T cells, gamma-delta T cells, T regulatory cells (Tregs), cytotoxic T lymphocytes, Th1 cells, Th2 cells, Th17 cells, Th9 cells, naïve T cells, memory T cells, effector T cells, effector-memory T cells (TEM), central memory T cells (TCM), resident memory T cells (TRM), follicular helper T cells (TFH), Natural killer T cells (NKTs), tumor- infiltrating lymphocytes (TILs), Natural killer cells (NKs), Innate Lymphoid Cells (ILCs), ILC1 cells, ILC2 cells, ILC3 cells, lymphoid tissue inducer (LTi) cells, B cells, B1 cells, B1a cells, B1b cells, B2 cells, plasma cells, B regulatory cells, memory B cells, marginal zone B cells, follicular B cells, germinal center B cells, antigen presenting cells (APCs), monocytes, macrophages, M1 macrophages, M2 macrophages, tissue-associated macrophages, dendritic cells, plasmacytoid dendritic cells, neutrophils, mast cells, basophils, eosinophils, common myeloid progenitors, common lymphoid progenitors, or any combination thereof.

[0311] A composition, complex, system, or method of the disclosure can be used to effect a change in the expression, epigenetic modification, or activity level of an engineered cell that is used to manufacture a biologic, for example, an antibody or other protein-based therapeutic.Additional Embodiments

[0312] Non-limiting embodiments of the present disclosure are also provided by the following numbered options. 1. A system for regulating aberrant expression of a target gene in a muscle cell, comprising: a heterologous polypeptide comprising a nuclease; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell, wherein, upon formation of the complex, the complex is capable of binding the target polynucleotide sequence, to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, and wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) the system is configured to effect reduced expression level of the target gene in the muscle cell by at least about 80% as compared to a control cell; and / or (3) the system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off-target effect below a predetermined threshold level in the muscle cell; and / or (4) the system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or(6) the system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the system. 2. The system of option 1, wherein upon formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days. 3. The system of option 2, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. 4. The system of option 2, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 17 days. 5. The system of option 2, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 18 days. 6. The system of option 1, wherein the muscle cell is in a subject having or is suspected of having facioscapulohumeral muscular dystrophy (FSHD). 7. The system of option 1, wherein the target gene is Dux4. 8. The system of option 1, wherein the nuclease has a length that is less than or equal to about 800 amino acids. 9. The system of option 1, wherein the nuclease has a length that is less than or equal to about 750 amino acids. 10. The system of option 1, wherein the nuclease is Un1Cas12f1 or a modified variant thereof. 11. The system of option 1, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43.12. The system of option 1, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44. 13. The system of option 1, wherein the heterologous polypeptide further comprises a transcriptional regulator. 14. The system of option 13, wherein the transcriptional regulator comprises at least one methyltransferases. 15. The system of option 14, wherein the transcriptional regulator comprises at least one DNA Methyltransferases (DNMT). 16. The system of option 15, wherein the transcriptional regulator comprises DNMT-A or DNMT-L. 17. The system of option 14, wherein the transcriptional regulator comprises (i) DNMT-A or DNMT-L and (ii) KRAB or a variant of KRAB. 18. The system of option 14, wherein the transcriptional regulator comprises (i) DNMT-L and (ii) KRAB or a variant of KRAB. 19. The system of option 14, wherein the transcriptional regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB. 20. The system of option 13, wherein the transcriptional regulator comprises DNMT-L or KRAB or variant of KRAB. 21. The system of option 13, wherein the transcriptional regulator comprises KRAB or a variant of KRAB. 22. The system of option13, wherein the transcriptional regulator comprises a plurality of different transcriptional regulators. 23. The system of option 1, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of a downstream gene of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43. 24. The system of option 1, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of an apoptosis marker in the muscle cell. 25. The system of option 24, wherein the apoptosis marker comprises Caspase 3.26. The system of option 1, wherein the complex effects the modification of the expression level of the target gene in the muscle gene. 27. The system of option 26, wherein the modification of the expression level results in downregulation of the target gene. 28. The system of option 1, wherein the complex effects the modification of the methylation level of the target gene in the muscle gene. 29. The system of option 28, wherein the modification of the methylation level results in downregulation of the target gene. 30. The system of option 1, wherein the nuclease is a deactivated nuclease. 31. A composition comprising the system of any one of the preceding options. 32. A viral vector comprising the system of any one of the preceding options. 33. The viral vector of option 32, wherein the viral vector comprises an adeno-associated virus (AAVs), a retrovirus, a lentivirus, a poxvirus, or an adenovirus. 34. The viral vector of option 33, wherein the AAV comprises a AAV serotype RH74 AAV. 35. A method for regulating aberrant expression of a target gene in a muscle cell, comprising: (a) contacting the muscle cell with a complex comprising (i) a heterologous polypeptide comprising a nuclease and (ii) a guide nucleic acid molecule exhibiting specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell; and (b) upon the contacting, binding the target gene with the complex to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or(2) the system is configured to effect reduced expression level of the target gene in the muscle cell by at least about 80% as compared to a control cell; and / or (3) the system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off-target effect below a predetermined threshold level in the muscle cell; and / or (4) the system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the system. 36. The method of option 35, wherein upon formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days. 37. The method of option 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. 38. The method of option 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 17 days. 39. The method of option 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 18 days. 40. The method of option 35, wherein the contacting comprises injecting a composition comprising the complex to a subject in need thereof, wherein the subject has or is suspected of having facioscapulohumeral muscular dystrophy (FSHD). 41. The method of option 35, wherein the target gene is Dux4.42. The method of option 35, wherein the nuclease has a length that is less than or equal to about 800 amino acids. 43. The method of option 35, wherein the nuclease has a length tha...

Claims

WHAT IS CLAIMED IS:

1. A system for regulating aberrant expression of a target gene in a muscle cell, comprising: a heterologous polypeptide comprising a nuclease; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell, wherein, upon formation of the complex, the complex is capable of binding the target polynucleotide sequence, to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, and wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) the system is configured to effect reduced expression level of the target gene in the muscle cell by at least about 80% as compared to a control cell; and / or (3) the system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off- target effect below a predetermined threshold level in the muscle cell; and / or(4) the system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the system.

2. The system of claim 1, wherein upon formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days.

3. The system of claim 2, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months.

4. The system of claim 2 or 3, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 17 days.

5. The system of claim 2 or 3, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 18 days.

6. The system of any one of claims 1-5, wherein the muscle cell is in a subject having or is suspected of having facioscapulohumeral muscular dystrophy (FSHD). The system of any one of claims 1-6, wherein the target gene is Dux4.

8. The system of any one of claims 1-7, wherein the nuclease has a length that is less than or equal to about 800 amino acids.

9. The system of any one of claims 1-8, wherein the nuclease has a length that is less than or equal to about 750 amino acids.

10. The system of any one of claims 1-9, wherein the nuclease is Un1Cas12f1 or a modified variant thereof.

11. The system of any one of claims 1-10, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:

43.

12. The system of any one of claims 1-10, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:

44.

13. The system of any one of claims 1-12, wherein the heterologous polypeptide further comprises a transcriptional regulator.

14. The system of claim 13, wherein the transcriptional regulator comprises at least one methyltransferases.

15. The system of claim 14, wherein the transcriptional regulator comprises at least one DNA Methyltransferases (DNMT).

16. The system of claim 15, wherein the transcriptional regulator comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).

17. The system of claim 14, wherein the transcriptional regulator comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.

18. The system of claim 14, wherein the transcriptional regulator comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.

19. The system of claim 14, wherein the transcriptional regulator comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant of KRAB.

20. The system of claim 13, wherein the transcriptional regulator comprises DNMT-L (or DNMT3L) or KRAB or variant of KRAB.

21. The system of claim 13, wherein the transcriptional regulator comprises KRAB or a variant of KRAB.

22. The system of claim13, wherein the transcriptional regulator comprises a plurality of different transcriptional regulators.

23. The system of any one or more of claims 1-22, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of a downstream gene of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

24. The system of any one of claims 1-23, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of an apoptosis marker in the muscle cell.

25. The system of claim 24, wherein the apoptosis marker comprises Caspase 3.

26. The system of claim 1, wherein the complex effects the modification of the expression level of the target gene in the muscle gene.

27. The system of claim 26, wherein the modification of the expression level results in downregulation of the target gene.

28. The system of any one of claims 1-27, wherein the complex effects the modification of the methylation level of the target gene in the muscle gene.

29. The system of claim 28, wherein the modification of the methylation level results in downregulation of the target gene.

30. The system of any one of claims 1-29, wherein the nuclease is a deactivated nuclease.

31. A composition comprising the system of any one of the preceding claims.

32. A viral vector comprising the system of any one of the preceding claims.

33. The viral vector of claim 32, wherein the viral vector comprises an adeno- associated virus (AAVs), a retrovirus, a lentivirus, a poxvirus, or an adenovirus.

34. The viral vector of claim 33, wherein the AAV comprises a AAV serotype RH74 AAV.

35. A method for regulating aberrant expression of a target gene in a muscle cell, comprising: (a) contacting the muscle cell with a complex or system comprising (i) a heterologous polypeptide comprising a nuclease and (ii) a guide nucleic acid molecule exhibiting specific binding to a target polynucleotide sequence at or adjacent to, a D4Z4 repeat array in the muscle cell; and (b) upon the contacting, binding the target gene with the complex or system to effect modification of an expression level and / or a methylation level of the target gene in the muscle cell, wherein the target gene is within the D4Z4 repeat array, wherein: (1) the guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, orat least about 95% sequence identity to the polynucleotide sequence of one or more members from Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that exhibits at least about 80%, at least about 90%, or at least about 95% sequence identity to the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) the complex or system is configured to effect reduced expression level of the target gene in the muscle cell by at least about 80% as compared to a control cell; and / or (3) the complex or system is configured to modify the expression level and / or the methylation level of the target gene in the muscle cell while exhibiting off-target effect below a predetermined threshold level in the muscle cell; and / or (4) the complex or system is configured to persistently modulate expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; and / or (5) the complex or system is configured to normalize apoptosis level of the muscle cell to a level that is comparable to a healthy muscle cell; and / or (6) the complex or system is configured to enhance survival of the muscle cell in vitro or in vivo; and / or (7) the complex or system is configured to exhibit minimal effect on expression profile of at least one muscle cell-specific gene in the muscle cell, wherein the muscle cell-specific gene and the target gene are not the same; and / or (8) the complex or system is configured to exhibit minimal effect on at least one health indication of a subject upon treatment of the subject with the complex or system.

36. The method of claim 35, wherein upon formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days.

37. The method of claim 35 or 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months.

38. The method of claim 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 17 days.

39. The method of claim 36, wherein the modified expression level and / or methylation level of the target gene is sustained for at least about 18 days.

40. The method of any one of claims 35-39, wherein the contacting comprises injecting a composition comprising the complex to a subject in need thereof, wherein the subject has or is suspected of having facioscapulohumeral muscular dystrophy (FSHD).

41. The method of any one of claims 35-40, wherein the target gene is Dux4.

42. The method of any one of claims 35-41, wherein the nuclease has a length that is less than or equal to about 800 amino acids.

43. The method of claim 42, wherein the nuclease has a length that is less than or equal to about 750 amino acids.

44. The method of any one of claims 35-43, wherein the nuclease is Un1Cas12f1 or a modified variant thereof.

45. The method of any one of claims 35-44, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:

43.

46. The method of any one of claims 35-45, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.

47. The method of any one of claims 35-46, wherein the heterologous polypeptide further comprises a transcriptional regulator.

48. The method of claim 47, wherein the transcriptional regulator comprises at least one methyltransferases.

49. The method of claim 48, wherein the transcriptional regulator comprises at least one DNA Methyltransferases (DNMT).

50. The method of claim 49, wherein the transcriptional regulator comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).

51. The method of claim 49, wherein the transcriptional regulator comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.

52. The method of claim 49, wherein the transcriptional regulator comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.

53. The method of claim 49, wherein the transcriptional regulator comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or variant of KRAB.

54. The method of claim 47, wherein the transcriptional regulator comprises DNMT-L (or DNMT3L) or KRAB or a variant of KRAB.

55. The method of claim 47, wherein the transcriptional regulator comprises KRAB or a variant of KRAB.

56. The method of claim 47, wherein the transcriptional regulator comprises a plurality of different transcriptional regulators.

57. The method of any one of claims 35-56, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of adownstream gene of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

58. The method of any one of claims 35-57, wherein the modification of the expression level and / or the methylation level of the target gene effects downregulation of an apoptosis marker in the muscle cell.

59. The method of claim 58, wherein the apoptosis marker comprises Caspase 3.

60. The method of any one of claims 35-59, wherein the complex effects the modification of the expression level of the target gene in the muscle gene.

61. The method of any one of claims 35-60, wherein the modification of the expression level results in downregulation of the target gene.

62. The method of any one of claims 35-61, wherein the complex effects the modification of the methylation level of the target gene in the muscle gene.

63. The method of claim 62, wherein the modification of the methylation level results in downregulation of the target gene.

64. The method of any one of claims 35-63, wherein the nuclease is a deactivated nuclease.

65. The system of any one of claims 1-30, the composition of claim 31, or the vector of any one of claims 32-34 for use as a medicament, such as for the inhibition, amelioration, or treatment of a cancer, autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

66. The system of any one of claims 1-30, the composition of claim 31, or the vector of any one of claims 32-34 for use in editing a gene in a cell in vitro or in vivo.

67. A method of providing a cell with a system, which regulates an aberrant expression of a target gene, comprising introducing the system of any one of claims 1-30, the composition of claim 31, or the vector of any one of claims 32-34 into a cell, preferably a cell in a subject, such as a human having a cancer, autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

68. The system, composition, viral vector of any one of the preceding claims, wherein the system comprises: a heterologous polypeptide comprising a nuclease comprising an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs:43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO:

865.

69. The system, composition, viral vector of any one of the preceding claims, wherein the system comprises: a heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to a SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO:836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO:

865.

70. The method of any one of the preceding claims, wherein the complex or system comprises: a heterologous polypeptide comprising a nuclease comprising an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of SEQ ID NOs:43, 44, and 728; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO:

865.

71. The method of any one of the preceding claims, wherein the complex or system comprises: a heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to a SEQ ID NO:727; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence identity of, of about, or of at least 80, 85, 90, 95, 97, 98, 99%, or about 100% to any one of: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833, and SEQ ID NO: 865.