Systems and methods for regulating abnormal gene expression

A system using a nuclease and guide nucleic acid complex targets the D4Z4 repeat array to sustainably modify gene expression and methylation in muscle cells, effectively treating FSHD by reducing Dux4 expression and improving cell survival.

JP2026510894APending Publication Date: 2026-04-10EPICRISPR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing treatments for diseases caused by abnormal gene expression, such as facioscapulohumeral muscular dystrophy (FSHD), are inadequate as they fail to sustainably alter the expression of target genes, leading to insufficient therapeutic effects.

Method used

A system comprising a heterologous polypeptide with a nuclease and a guide nucleic acid molecule is used to form a complex that targets the D4Z4 repeat array in muscle cells, modifying the expression and methylation levels of target genes like Dux4, achieving at least 80% reduction and maintaining the altered state for extended periods while minimizing off-target effects.

Benefits of technology

The system effectively reduces target gene expression by at least 80%, normalizes apoptosis levels, and improves muscle cell survival, providing a sustainable therapeutic effect comparable to healthy cells, with minimal impact on non-target genes or health conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026510894000001_ABST
    Figure 2026510894000001_ABST
Patent Text Reader

Abstract

This specification provides systems, compositions, and methods for treating or improving diseases or conditions in a subject (e.g., muscular dystrophy, e.g., facioscapulohumeral muscular dystrophy (FSHD)) by regulating the abnormal expression of target genes within cells (e.g., muscle cells).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Cross-reference of related applications This application claims priority and benefits under U.S. Provisional Patent Application No. 63 / 490678, filed March 16, 2023; U.S. Provisional Patent Application No. 63 / 520253, filed August 17, 2023; and U.S. Provisional Patent Application No. 63 / 592882, filed October 24, 2023, all of which are expressly incorporated herein by reference.

[0002] Sequence listing reference This application is filed together with a sequence listing in an electronic XML file. This XML sequence listing is provided as a 929,129-byte file created on March 14, 2024, with the filename SequenceListing_EPICR022WO.xml. The information contained in this electronic sequence listing is incorporated herein by reference in its entirety. [Background technology]

[0003] Diseases or conditions can result from the abnormal expression of one or more genes. In some cases, abnormal expression of embryonic transcription factors in muscle cells can lead to the development of muscular dystrophy. For example, abnormal expression of transcription factors in muscle cells (e.g., abnormal expression of DUX4 in skeletal muscle cells) can cause facioscapulohumeral muscular dystrophy (FSHD). [Overview of the Initiative] [Means for solving the problem]

[0004] Transiently altering the abnormal expression of target genes in cells may be insufficient for treating or curing diseases caused by this abnormal expression. Therefore, there is still a strong need for systems and methods to alter the abnormal expression of target genes and maintain this altered expression for extended periods.

[0005] In one aspect, the present disclosure is a system for regulating abnormal expression of a target gene in muscle cells, comprising a heterologous polypeptide comprising a nuclease, and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence present in or adjacent to the D4Z4 repeat array in the muscle cells, when the complex is formed, the complex binds to the target polynucleotide sequence and can modify the expression level and / or methylation level of the target gene in the muscle cells, the target gene is present within the D4Z4 repeat array, (1) the guide nucleic acid molecule comprises a spacer sequence that exhibits the specific binding, and the spacer sequence comprises a polynucleotide sequence having at least about 80%, at least about 90% or at least about 95% sequence identity with the polynucleotide sequence of one or more members described in Table 4, or is encoded by this polynucleotide sequence, and the spacer sequence may comprise a polynucleotide sequence having at least about 80%, at least about 90% or at least about 95% sequence identity with the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836 or SEQ ID NO: 851, or may be encoded by this polynucleotide sequence; (2) the system is configured to reduce the expression level of the target gene in the muscle cells by at least 80% or at least about 80% compared to control cells; (3) the system is configured to modify the expression level and / or methylation level of the target gene in the muscle cells while showing an off-target effect below a predetermined threshold in the muscle cells; (4) the system is configured to continuously regulate the expression level and / or methylation level of a downstream gene of the target gene for at least about 5 days; (5) The system is configured to normalize the apoptosis level of the muscle cells to a level comparable to that of healthy muscle cells, (6) The system is configured to improve the survival of the muscle cells in vitro or in vivo, (7) The system is configured to show a minimal effect on the expression profile of at least one muscle cell-specific gene different from the target gene in the muscle cells, and / or (8) The system is configured to show a minimal effect on at least one health condition symptom of a subject treated with the system. Provide a system.

[0006] In some embodiments of any of the systems disclosed herein, after the complex is formed, the modified expression level and / or methylation level of the target gene in the muscle cells persists for at least about 2 days. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least about 3 days, at least about 4 days, at least about 5 days, at least about 6 days, at least about 1 week, at least about 2 weeks, at least about 2 weeks, at least about 4 weeks, or at least about 2 months. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least about 17 days. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least about 18 days.

[0007] In some embodiments of any of the systems disclosed herein, the muscle cells are present in a subject having facioscapulohumeral muscular dystrophy (FSHD) or a subject suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the systems disclosed herein, the target gene is Dux4.

[0008] In some embodiments of the systems disclosed herein, the length of the nuclease is about 800 amino acids or less. In some embodiments of the systems disclosed herein, the length of the nuclease is about 750 amino acids or less.

[0009] In some embodiments of the systems disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof. In some embodiments of the systems disclosed herein, the nuclease comprises an amino acid sequence having at least 80% or at least about 80%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 99% or at least about 99% identity with the polypeptide sequence of SEQ ID NO: 43. In some embodiments of the systems disclosed herein, the nuclease comprises an amino acid sequence having at least 80% or at least about 80%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 99% or at least about 99% identity with the polypeptide sequence of SEQ ID NO: 44.

[0010] In some embodiments of the systems disclosed herein, the heterologous polypeptide further comprises a transcription factor. In some embodiments of the systems disclosed herein, the transcription factor comprises at least one methyltransferase. In some embodiments of the systems disclosed herein, the transcription factor comprises at least one DNA methyltransferase (DNMT). In some embodiments of the systems disclosed herein, the transcription factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of the systems disclosed herein, the transcription factor comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof. In some embodiments of the systems disclosed herein, the transcription factor comprises (i) DNMT-L and (ii) KRAB or a variant thereof. In some embodiments of the systems disclosed herein, the transcription factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant thereof. In some embodiments of the systems disclosed herein, the transcription factor comprises DNMT-L (or DNMT3L) or KRAB or a variant thereof. In some embodiments of the systems disclosed herein, the transcription factor comprises KRAB or a variant thereof. In some embodiments of the systems disclosed herein, the transcription factor comprises a plurality of different transcription factors.

[0011] In some embodiments of the systems disclosed herein, modification of the expression level and / or methylation level of the target gene results in downregulation of a downstream gene of the target gene, the downstream gene comprising one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0012] In some embodiments of the systems disclosed herein, modification of the expression level and / or methylation level of the target gene results in downregulation of an apoptosis marker in the muscle cell. In some embodiments of the systems disclosed herein, the apoptosis marker includes caspase 3.

[0013] In some embodiments of the systems disclosed herein, the complex causes modification of the expression level of the target gene within the muscle gene. In some embodiments of the systems disclosed herein, the modification of the expression level causes downregulation of the target gene.

[0014] In some embodiments of the systems disclosed herein, the complex causes modification of the methylation level of the target gene within the muscle gene. In some embodiments of the systems disclosed herein, the modification of the methylation level causes downregulation of the target gene.

[0015] In some embodiments of the systems disclosed herein, the nuclease is an inactive nuclease.

[0016] In one embodiment, the present disclosure provides a composition comprising any of the systems disclosed herein.

[0017] In one embodiment, the present disclosure provides a viral vector comprising any of the systems or compositions disclosed herein.

[0018] In some embodiments of the viral vectors disclosed herein, the viral vector comprises an adeno-associated virus (AAV), a retrovirus, a lentivirus, a poxvirus, or an adenovirus. In some embodiments of the viral vectors disclosed herein, the AAV comprises AAV serotype RH74 AAV.

[0019] In one embodiment, the present disclosure relates to a method for regulating the abnormal expression of a target gene in muscle cells, (a) A step of contacting muscle cells with a complex or system comprising (i) a heterologous polypeptide containing a nuclease and (ii) a guide nucleic acid molecule that exhibits specific binding to a target polynucleotide sequence present in or adjacent to a D4Z4 repeat array within muscle cells, and (b) A step following the contact step, to attach a target gene to the complex or system to modify the expression level and / or methylation level of the target gene in the muscle cells. Includes, The target gene is located within the D4Z4 repeat array. (1) The guide nucleic acid molecule includes a spacer sequence exhibiting the specific binding, and the spacer sequence includes or is encoded by a polynucleotide sequence that exhibits at least 80% or at least about 80%, at least 90% or at least about 90%, or at least 95% or at least about 95% sequence identity with one or more member polynucleotide sequences listed in Table 4, and the spacer sequence may also include or be encoded by a polynucleotide sequence that exhibits at least 80% or at least about 80%, at least 90% or at least about 90%, or at least 95% or at least about 95% sequence identity with the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851, (2) The complex or system is configured to reduce the expression level of the target gene in the muscle cells by at least 80% or at least about 80% compared to control cells. (3) The complex or system is configured to modify the expression level and / or methylation level of the target gene in the muscle cell while exhibiting off-target effects below a predetermined threshold in the muscle cell, (4) The complex or system is configured to continuously regulate the expression level and / or methylation level of downstream genes of the target gene for at least about 5 days, (5) The complex or system is configured to normalize the apoptosis level of the muscle cells to a level equivalent to that of healthy muscle cells, (6) The complex or system is configured to improve the survival of the muscle cells in vitro or in vivo, (7) The complex or system is configured to have minimal effect on the expression profile of at least one muscle cell-specific gene different from the target gene in the muscle cell, and / or (8) The complex or system is configured to have minimal effect on at least one sign of a health condition in the subject treated by the system. Provide a method.

[0020] In some embodiments of the methods disclosed herein, after the complex or system is formed, the modified expression level and / or methylation level of the target gene in the muscle cells persists for at least about 2 days. In some embodiments of the methods disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 1 week, at least 2 weeks, at least 2 weeks, at least 4 weeks, or at least 2 months, or at least about 3 days, at least about 4 days, at least about 5 days, at least about 6 days, at least about 1 week, at least about 2 weeks, at least about 2 weeks, at least about 4 weeks, or at least about 2 months. In some embodiments of the methods disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least 17 days or at least about 17 days. In some embodiments of the methods disclosed herein, the modified expression level and / or methylation level of the target gene persists for at least 18 days or at least about 18 days.

[0021] In some embodiments of the methods disclosed herein, the contact step comprises injecting a composition comprising the complex or system into a subject who requires the modulation of abnormal expression of a target gene in the muscle cells, the subject having or suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of the methods disclosed herein, the target gene is Dux4.

[0022] In some embodiments of the methods disclosed herein, the length of the nuclease is about 800 amino acids or less. In some embodiments of the methods disclosed herein, the length of the nuclease is about 750 amino acids or less.

[0023] In some embodiments of the methods disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof.

[0024] In some embodiments of the methods disclosed herein, the nuclease comprises an amino acid sequence having at least 80% or at least about 80%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 99% or at least about 99% identity with the polypeptide sequence of SEQ ID NO: 43. In some embodiments of the methods disclosed herein, the nuclease comprises an amino acid sequence having at least 80% or at least about 80%, at least 90% or at least about 90%, at least 95% or at least about 95%, or at least 99% or at least about 99% identity with the polypeptide sequence of SEQ ID NO: 44.

[0025] In some embodiments of the methods disclosed herein, the heterologous polypeptide further comprises a transcription factor. In some embodiments of the methods disclosed herein, the transcription factor comprises at least one methyltransferase. In some embodiments of the methods disclosed herein, the transcription factor comprises at least one DNA methyltransferase (DNMT). In some embodiments of the methods disclosed herein, the transcription factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of the methods disclosed herein, the transcription factor comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof. In some embodiments of the methods disclosed herein, the transcription factor comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof. In some embodiments of the methods disclosed herein, the transcription factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant thereof. In some embodiments of the methods disclosed herein, the transcription factor comprises DNMT-L (or DNMT3L) or KRAB or a variant thereof. In some embodiments of the methods disclosed herein, the transcription factor comprises KRAB or a variant thereof. In some embodiments of the methods disclosed herein, the transcription factor comprises a plurality of different transcription factors.

[0026] In some embodiments of the methods disclosed herein, modification of the expression level and / or methylation level of the target gene results in downregulation of a downstream gene of the target gene, the downstream gene comprising one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0027] In some embodiments of the methods disclosed herein, modification of the expression level and / or methylation level of the target gene results in downregulation of an apoptosis marker in the muscle cell. In some embodiments of the methods disclosed herein, the apoptosis marker includes caspase 3.

[0028] In some embodiments of the methods disclosed herein, the complex or system causes modification of the expression level of the target gene within the muscle gene.

[0029] In some embodiments of the methods disclosed herein, the modification of the expression level results in the downregulation of the target gene.

[0030] In some embodiments of the methods disclosed herein, the complex or system causes modification of the methylation level of the target gene within the muscle gene. In some embodiments of the methods disclosed herein, the modification of the methylation level causes downregulation of the target gene.

[0031] In some embodiments of the methods disclosed herein, the nuclease is an inactive nuclease.

[0032] In some embodiments, the compositions or vectors disclosed herein are intended for use as pharmaceuticals, such as pharmaceuticals for the suppression, improvement, or treatment of cancer, autoimmune diseases, or facioscapulohumeral muscular dystrophy (FSHD).

[0033] In some embodiments, the compositions or vectors disclosed herein are intended for use in intracellular gene editing in vitro or in vivo.

[0034] Furthermore, the present invention provides a method for providing cells with a system for regulating the abnormal expression of a target gene, comprising the step of introducing any of the systems, compositions, and vectors disclosed herein into cells, preferably the cells being cells in a subject such as a human having cancer, an autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

[0035] In some embodiments of the systems, compositions, or viral vectors disclosed herein, the system is A heterogeneous polypeptide containing a nuclease having an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity.

[0036] In some embodiments of the systems, compositions, or viral vectors disclosed herein, the system is A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity.

[0037] In some embodiments of the methods disclosed herein, the complex or system is A heterogeneous polypeptide containing a nuclease having an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity.

[0038] In some embodiments of the methods disclosed herein, the complex or system is A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity.

[0039] Those skilled in the art will readily understand from the following detailed description, which presents and describes only exemplary embodiments of the disclosure, further aspects and advantages of the disclosure. It will also readily understand that the disclosure implements various other embodiments, and that the details thereof can be modified in various distinct ways without departing from the disclosure. Therefore, the drawings and detailed description are considered exemplary in nature and are not limited thereto.

[0040] Citation All publications, patents, and patent applications described herein are incorporated herein by reference to the same extent as each individual publication, patent, or patent application is specifically described herein. [Brief explanation of the drawing]

[0041] The novel features of this disclosure are specifically described in the appended claims. The features and advantages of the present invention can be better understood by referring to the following detailed description of exemplary embodiments in which the principles of this disclosure are utilized, and by referring to the appended drawings.

[0042] [Figure 1] This shows various target polynucleotide sequences (e.g., Rank #1 to Rank #91) located between two CpG islands within the D4Z4 repeat array encoding DUX4.

[0043] [Figure 2] We demonstrate that DUX4 expression in a target cell population (e.g., lymphoblasts) can be regulated by complexing a heterogeneous actuator moiety linked to a gene regulator (e.g., dCas-KRAB-DNMT3A-DNMT3L) with various guide RNA molecules that target polynucleotide sequences (e.g., Rank #1 to Rank #91) within the D4Z4 repeat array encoding DUX4.

[0044] [Figure 3]Figure 3A shows the gene expression of DUX4 and DUX4 target genes in patient-derived immortalized human FSHD skeletal myoblasts (SkM) (12ABIC / 12A and 15ABIC / 15A). Gene expression of DUX4 and DUX4 target genes was measured in undifferentiated 12ABIC and undifferentiated 15ABIC cells, 12ABIC and 15ABIC cells after 2 days of differentiation, and 12ABIC and 15ABIC cells after 7 days of differentiation. Each shade of gray in the graph represents the gene expression of various genes corresponding to the legend on the right. Figure 3B shows a comparison of the percentage of apoptotic cells between FSHD myoblasts 12ABIC and 15ABIC (right column) and their corresponding healthy sibling cells, control myoblasts 12UBIC and 15VBIC (left column), after 2 days of differentiation. White dots in the left image represent apoptotic cells. The graph on the right shows the percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cell cultures after 2 days of differentiation, as shown in the image on the left. Nuclei were stained using DAPI staining. Figure 3C shows the percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. The percentage of apoptotic cells was measured on days 0, 1, 2, and 7 of differentiation. Figure 3D shows the expression of MYHC in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. Myosin heavy chain (MYHC) is a muscle cell differentiation marker. White dots indicate MYHC expression. Figure 3E shows the expression levels of MYOG, MYH2, and MYMK in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. MYOG is a myogenic regulator that controls skeletal muscle differentiation, and MyoMaker (MYMK) is a myocyte differentiation marker. Nuclei were stained using DAPI staining. Expression levels of 12ABIC and 15ABIC cells were measured on days 2 and 7 of differentiation. 12A UD and 15A UD: undifferentiated proliferative control myoblasts.The dark gray bars indicate the expression level of MYOG, the light gray bars indicate the expression level of MYH2, and the gray bars indicate the expression level of MYMK.

[0045] [Figure 4] This section shows the design of multiple gRNAs associated with the D4Z4 repeat region. Multiple DUX4 target gRNAs were designed to extend across the entire D4Z4 repeat region. The relationship between the D4Z4 repeat region and its position within the DUX4 gene is shown at the bottom of Figure 4. The newly designed gRNAs are shown at the top of Figure 4.

[0046] [Figure 5] The design of a Cas12f effector-modulator vector is presented. Expression of the Cas12f variant, KRAB domain, and DNMT3L domain is regulated by the muscle-specific promoter CK8e. Expression of an sgRNA spacer sequence with a scaffold utilized by RNA polymerase III is regulated by the human U6g promoter. This vector further contains a modified WPRE and a polyadenylation regulatory sequence.

[0047] [Figure 6] Figure 6A shows the relative expression levels of DUX4 in 12ABIC FSHD myoblasts that stably expressed the Cas12f-KRAB effector-modulator after nucleofection of 12ABIC myoblasts with one of 78 gRNAs. After nucleofection, the cells were cultured under differentiation conditions for 7 days, and DUX4 gene expression was measured. The 78 gRNAs tested are shown on the X axis, and the relative expression ratio of DUX4 is shown on the Y axis. The expression level of DUX4 was normalized to the expression level of the control gene HPRT1. Figure 6B shows the relative expression levels of DUX4 in 12ABIC FSHD myoblasts that stably expressed the Cas12f-KRAB effector-modulator after nucleofection of 12ABIC myoblasts with one of 78 gRNAs. After nucleofection, the cells were cultured under differentiation conditions for 7 days, and the gene expression of DUX4 and the DUX4 target gene (MBD3L2) was measured.

[0048] [Figure 7] Figure 7A shows the repression of DUX4 and its target genes DBET / DUX4, MBD3L2, and TRIM48 in patient-derived immortalized FSHD myoblasts transfected with one of six gRNAs and a Cas12f effector-modulator. The Cas12f effector-modulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLa domain. One of the six sgRNAs is a control sgRNA (empty / trcr) that does not target the D4Z4 repeat region. To analyze whether the differentiation potential of cells transfected with DUX4 sgRNA is comparable to that of control myoblasts transfected with sgRNA, the expression levels of MYOG in the cells were measured. The expression levels of DUX4, DUX4 target genes, and MYOG were measured 17 days after transfection. Figure 7B shows the repression of DUX4 and its target genes DBET / DUX4, MBD3L2, and TRIM48 in patient-derived immortalized FSHD myoblasts transfected with one of six gRNAs and a Cas12f effector-modulator. The Cas12f effector-modulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLb domain. One of the six sgRNAs is a control sgRNA (empty) that does not target the D4Z4 repeat region. To analyze whether the differentiation potential of cells transfected with DUX4 sgRNA is comparable to that of control myoblasts transfected with sgRNA, the expression levels of MYOG in the cells were measured. The expression levels of DUX4, DUX4 target genes, and MYOG were measured 18 days after transfection.

[0049] [Figure 8]Figures 8A and 8B show the apoptosis levels of myoblasts derived from FSHD patients transfected with Cas12f effector-modulator and DUX4-targeting gRNA. After transfection and differentiation for two days, the percentage of apoptosis-positive cells was measured. The image in Figure 8A shows the percentage of apoptotic cells in control 12UBIC cells and 12ABIC cells transfected with Cas12f effector-modulator and DUX4-targeting gRNA. White dots indicate apoptotic cells. The graph in Figure 8B shows the percentage of apoptotic cells measured in the image on the left and the percentage of apoptotic cells in 12ABIC cells transfected with DUX4-targeting gRNA or a control gRNA that does not target DUX4. The nuclei were stained using DAPI staining.

[0050] [Figures 9A-9E]The effects of the exemplary system of this disclosure in an ex vivo 3D FSHD organoid model are shown. Figure 9A shows the workflow of the ex vivo FSHD model. In this ex vivo model, immortalized healthy sibling control cells and FSHD skeletal myoblasts were cultured, and then these cells were incorporated into a 3D tissue. This 3D tissue was brought into contact with a control AAV or an AAV containing the exemplary system described herein. The 3D tissue was then tested for differences in phenotypic characteristics of mechanical force, typological force, and fatigue, and the morphology and gene expression profiles of the 3D tissue were measured. Figure 9B shows the mean typological force generated by action potentials over time for 3D organoid tissue treated with GFP control (-) (upper panel) and 3D organoid tissue treated with the exemplary system of this disclosure (lower panel). Figure 9C shows the typological force generated by action potentials at the endpoint (day 46) for 3D organoid tissue treated with GFP control (-) (upper panel) and 3D organoid tissue treated with the exemplary system of this disclosure (lower panel). Figure 9D shows the average tyrosin force over time for 3D organoid tissue treated with GFP control (-) (upper panel) and 3D organoid tissue treated with the exemplary system of this disclosure (lower panel). Figure 9E shows the normalized tyrosin force at the endpoint (day 46) for 3D organoid tissue treated with GFP control (-) (upper panel) and 3D organoid tissue treated with the exemplary system of this disclosure (lower panel).

[0051] [Figure 10A-10G]This document demonstrates the effect of an exemplary system of this disclosure on the repression of DUX4 and DUX-4 pathway genes in humanized mice. Figure 10A shows the workflow of an in vivo xenograft model. An in vivo model was prepared by irradiating the leg of a mouse and treating the tibialis anterior muscle with cardiotoxicity, and human myoblasts were transplanted into the leg of this mouse. After transplantation, the mice were exposed to an AAV containing either a control AAV or an exemplary system described herein. The mice were euthanized at predetermined time points, and xenograft and tissue samples were collected and analyzed. The collected xenografts were fixed, sectioned, and stained with hematoxylin and eosin. The remaining tissues were used for gene expression assays and analysis of AAV directionality in the mouse body. Figure 10B shows the mRNA expression of DUX4 in the humanized tibialis anterior muscle. Figure 10C shows the gene expression of DUX4 pathway genes plotted as an overall score. Figure 10D shows the in vivo distribution of the exemplary system of this disclosure. Figures 10E and 10F show the quantification results of DUX4 protein staining (Figure 10E) and SLC34A2 protein staining (Figure 10F). Figure 10G shows the quantification results of TUNEL staining.

[0052] [Figure 11] This disclosure illustrates methylation of a target gene (e.g., D4Z4) via an exemplary system / method.

[0053] [Figure 12] This document presents proposed treatments for FSHD patients.

[0054] [Figures 13A-13B] This shows the unmethylated or methylated state of the target gene / location. Figure 13A shows the methylated / unmethylated state of the target site at multiple time points via an exemplary system of the disclosure, including various gene repressors. Figure 13B shows the CpG methylation rate of the target locus (D4Z4 locus) in healthy sibling cell controls or FSHD patient-derived myoblasts treated with control AAV or an exemplary system of the disclosure.

[0055] [Figure 14]This paper describes the screening method and results of high-throughput screening of anti-DUX4 guide RNA spacer sequences.

[0056] [Figure 15] Verification and in silico off-target analysis are presented.

[0057] [Figure 16] This disclosure illustrates a virus delivery cargo for an exemplary system.

[0058] [Figure 17] This shows the expression of myogenic genes in patient-derived FSHD myoblasts.

[0059] [Figure 18] This disclosure shows the expression of DUX4 and DUX4 downstream genes in patient-derived FSHD myoblasts treated with the exemplary system described herein.

[0060] [Figures 19A-19D] Figure 19A shows apoptotic cells in patient-derived FSHD myoblasts treated with the exemplary system of this disclosure. Figure 19A shows the percentage of apoptotic cells in differentiated patient-derived FSHD myoblasts and patient-derived FSHD myoblasts exposed to the exemplary system of this disclosure. Figure 19B shows mRNA expression of DUX4 and cargo of the exemplary system. Figure 19C shows live-cell imaging analysis stained for caspase 3 / 7. Figure 19D shows the normalized intensity of total caspase 3 / 7 signal at the endpoint.

[0061] [Figure 20] This paper demonstrates the effects of an exemplary system of this disclosure in the skeletal muscle of a humanized FSHD mouse model.

[0062] [Figure 21] Histological analysis of the tibialis anterior muscle processed by the exemplary system of this disclosure is shown.

[0063] [Figure 22] This disclosure shows the expression of DUX4 and DUX4 pathway genes in a humanized FSHD mouse model exposed to an exemplary system.

[0064] [Figure 23] This shows a 6-month non-GLP toxicity study in immunocompetent mice.

[0065] [Figure 24] This shows the evaluation of the clinical chemical analysis of animals treated with the system of this disclosure. At each point in time, the bar on the left represents the control animal, and the bar on the right represents the animal treated with the system of this disclosure.

[0066] [Figure 25] This disclosure presents an evaluation of histopathological analysis three months after administration of the system to animals.

[0067] [Figure 26] This disclosure shows blood chemistry and hematological data from a non-human primate model that was exposed to the system described herein.

[0068] [Figure 27] This paper presents pharmacokinetic studies using non-human primate models.

[0069] [Figure 28] This shows the tissue-specific mRNA expression levels and directivity of candidate guide nucleic acid molecules.

[0070] [Figures 29A-29E]The effects of the effective dose range of the exemplary system of this disclosure on the molecular and cellular phenotypes of humanized mice are shown (see also Figures 10A–10G). Figure 29A shows a schematic diagram of an in vivo xenograft model of FSHD. Figures 29B–29C show qPCR analysis (Figure 29B) and qRT-PCR analysis (Figure 29C) of the exemplary system of this disclosure in tibialis anterior muscle biopsy samples taken from FSHD humanized mice 24 days after intravenous administration of the exemplary system or solvent of this disclosure. Figure 29B is related to Figure 10D. Figures 29D–29E show quantitative analysis of the DUX4 cascade in tibialis anterior muscle biopsy samples taken from FSHD humanized mice 24 days after intravenous administration of the exemplary system or solvent of this disclosure. Figure 29D shows qRT-PCR analysis of four DUX4 target genes plotted as a total score. Figure 29E shows quantitative image analysis of SLC34A2 immunohistochemical staining plotted as H-scores.

[0071] [Figures 30A-30H]This report demonstrates the effects of dose-escalation of the exemplary system of this disclosure on three-dimensional recombinant muscle tissue (EMT) derived from myoblasts of FSHD patients. Figure 30A shows a schematic diagram of the casting and stimulation model of the three-dimensional EMT tissue. Figures 30B–C show the action-potential-generated maximal unilateral contraction force (Figure 30B) and maximal stoichiometric contraction force (Figure 30C) from day 6 to day 46 in untransduced control three-dimensional EMT or tissue transduced with the exemplary system of this disclosure. Figures 30D–E show the muscle contraction force normalized to the untransduced three-dimensional EMT at day 46. The action-potential-generated maximal unilateral contraction force (Figure 30D) and maximal stoichiometric contraction force (Figure 30E) show a normalized increase in contraction force up to an exemplary system dose level of up to 7.5 × 10⁸ vg per tissue. Data are shown as mean ± SEM. Figures 30F-30G show quantitative PCR evaluation of the vector genome of the exemplary system of this disclosure in 3D EMT (Figure 30F) and evaluation of mRNA expression by qRT-PCR (Figure 30G). Data are shown as mean ± SEM. Figure 30H shows the expression of DUX4 and DUX4 pathway genes in 3D EMT transduced or untransduced control tissues of the exemplary system of this disclosure. Data are shown as mean ± SEM.

[0072] [Figure 31A-31H]This illustrates the mechanism of action of an exemplary system of the disclosure as a gene-targeted therapy for FSHD that mitigates pathological overexpression by directly modulating the epigenetic state of the DUX4 locus. Figure 31A shows a schematic diagram of the mechanism of action assay of the exemplary system of the disclosure using an on-target methylation assay in FSHD myoblasts. Figure 31B shows DUX4 and the mRNA expression of six DUX4 genes, namely MBD3L2, ZSCAN4, LEUTX, TRIM43, and TRIM48, in patient-derived myoblasts. Figure 31C shows that primary myoblasts from FSHD patients (07ABIC) exposed to the exemplary system of the disclosure showed increased methylation compared to myoblasts treated with the (-) control test substance. Figure 31D shows the results of RT-qPCR analysis of DUX4 and mRNA expression of six DUX4 genes, namely MBD3L2, ZSCAN4, LEUTX, TRIM43, and TRIM48, in patient-derived myoblasts (07ABIC) exposed to (-) control or the exemplary system of this disclosure. The mean ± SEM (n=3) of the replicated experiments is plotted. Figure 31E shows that myoblasts (07ABIC) exposed to the exemplary system of this disclosure showed a significant reduction in apoptosis. Figure 31F shows the results of analysis of unmethylated and methylated DNA standards for targeted methylation using custom primers (Table 17) after enzymatic conversion of 100 ng of DNA. Methylation rates were calculated using QUMA analysis from each of at least 12 clones, and the median (Q2) is plotted (upper panel of Figure 31F). The number of CpGs analyzed in each clone was 22. Lollipop plots showing the methylation status of each CpG in each clone are also plotted (lower panel of Figure 31F). The black circles on the plot each represent a methylated CpG dinucleotide. Figure 31G shows the results of analyzing unmethylated and methylated DNA standards (Sigma-Aldrich) for targeted methylation using a next-generation sequencing assay with primers listed in Table 17 after enzymatic conversion of 100 ng of DNA. Figure 31H shows the results of downsampling of on-target methylation assay analysis by next-generation sequencing. [Modes for carrying out the invention]

[0073] While various embodiments of the present invention are shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Those skilled in the art will also understand that various variations, modifications, and substitutions are possible without departing from the present invention. Various other aspects of the embodiments of the present invention described herein may be adopted.

[0074] When the terms “at least,” “greater than,” or “greater than” appear before or after the first or last number in a set of two or more numbers, these terms always apply to each of the numbers in that set. For example, “1, 2 or 3 or more” is the same as “1 or more, 2 or more, or 3 or more.”

[0075] When the terms “within,” “less than,” or “less than or equal to” follow the last number in a sequence of two or more numbers, these terms always apply to each of the numbers in that sequence. For example, “1, 2, or 3 or less” is the same as “1 or less, 2 or less, or 3 or less.”

[0076] The singular terms “a,” “an,” and “the” include plural nouns unless otherwise specified. Similarly, the term “or” is intended to include “and” unless otherwise specified. In this specification, the abbreviation “example” is used to indicate non-restrictive examples. Thus, “example” is synonymous with “for example.” Numerical ranges provided include overlapping ranges and integers between numerical values; for example, the ranges 1–4 and 5–7 include, for example, 1–7, 1–6, 1–5, 2–5, 2–7, 4–7, 1, 2, 3, 4, 5, 6, and 7.

[0077] The term "approximately" usually means within the permissible margin of error of a particular numerical value as determined by those skilled in the art, and this permissible margin of error depends in part on how the numerical value was measured or determined, i.e., in part on the limitations of the measurement system. For example, "approximately" in the art may mean within or greater than one standard deviation per measurement. Alternatively, "approximately" may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given numerical value. Or, particularly with respect to biological systems or biological processes, the term "approximately" may mean within one digit of a given numerical value, preferably within five times, more preferably within two times. Where a specific numerical value is described in this application and claims, unless otherwise stated, the term "approximately" is considered to mean within the permissible margin of error of that particular numerical value.

[0078] The use of options (e.g., indicated by "or") means one, both, or a combination of the options. The term "and / or" means one or both of the options.

[0079] The term "cell" usually refers to a living cell. A cell may be the basic structural unit, functional unit, and / or biological unit of a living organism. A cell may originate from any organism that has one or more cells. Some examples include cells from prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells obtained from plants (e.g., cereal plants, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkin, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, clubmosses, hornwort, liverworts or mosses), algal cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens C.) Examples of cells include, but are not limited to, those derived from natural organisms (e.g., Agardh), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells obtained from mushrooms), animal cells, cells obtained from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells obtained from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells obtained from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). In some cases, the cells do not have to be derived from natural organisms (e.g., cells may be synthetically produced and may be called "artificial cells" as appropriate).

[0080] In this specification, the term “nucleotide” usually means a combination of a base, a sugar, and a phosphate group. A nucleotide may be a synthetic nucleotide. A nucleotide may be a synthetic nucleotide analog. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) or ribonucleic acid (RNA)). The term “nucleotide” may include ribonucleoside triphosphates such as adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), and guanosine triphosphate (GTP), and may also include deoxyribonucleoside triphosphates such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Examples of such derivatives include [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules in which they themselves are contained. Furthermore, in this specification, "nucleotide" may mean dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Specific examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or labeled in a manner detectable by known techniques. Labeling may also be performed using quantum dots. Examples of detectable labels include radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzymatic labels. Examples of fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include, from PerkinElmer Corporation (Foster City, California): [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP; and from Amersham Corporation (Arlington Heights, Illinois): FluoroLink deoxyribonucleotides such as FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink Cy5-dUTP; Fluorescein-15-dATP, Fluorescein-12-dUTP, Tetramethyl-Rhodamine-6-dUTP, IR770-9-dATP, Fluorescein-12-ddUTP, Fluorescein-12-UTP, and Fluorescein-15-2'-dATP, available from Boehringer Mannheim (Indianapolis, Indiana); and Molecular Examples of chromosome-labeled nucleotides available from Probes, Inc. (Eugene, Oregon) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled by chemical modification. The chemically modified single nucleotide may also be biotin-dNTP.Some examples of biotinylated dNTPs include, but are not limited to, biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).

[0081] In this specification, the terms “polynucleotide,” “oligonucleotide,” and “nucleic acid” are used interchangeably and usually refer to a polymeric form of a nucleotide of a certain length, which may be a deoxyribonucleotide or a ribonucleotide, or an analog thereof, and may be single-stranded, double-stranded, or multi-stranded. Polynucleotides may be exogenous or endogenous to cells. Polynucleotides may exist in a cell-free environment. Polynucleotides may be genes or fragments thereof. Polynucleotides may be DNA. Polynucleotides may be RNA. Polynucleotides may have any three-dimensional structure and may perform any known or unknown function. Polynucleotides may contain one or more analogues (e.g., modified skeletons, sugars, or nucleic acid bases). If the nucleotide structure involves modifications, these modifications may be conferred before or after the assembly of the nucleotide. Some examples of analogs include, but are not limited to, 5-bromouracil, peptide nucleic acids, xeno nucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorescent dyes (e.g., sugar-linked rhodamine or fluorescein), nucleotide-containing thiols, biotin-labeled nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, quosin, and waiosin.Examples of polynucleotides include, but are not limited to, coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, DNA isolated from a sequence, RNA isolated from a sequence, cell-free polynucleotides such as cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. Non-nucleotide components can be inserted into nucleotide sequences.

[0082] The term "sequence identity" typically refers to the complete correspondence between nucleotides in two polynucleotide sequences, or between amino acids in two polypeptide sequences. Generally, techniques for determining sequence identity include determining the nucleotide sequence of a polynucleotide and / or the amino acid sequence encoded by that nucleotide sequence, and then comparing these sequences to a second nucleotide sequence or a second amino acid sequence. Two or more sequences (polynucleotide sequences or amino acid sequences) can be compared by determining their "percentage of identity." Whether nucleic acid sequences or amino acid sequences, the percentage of identity of two sequences is calculated by dividing the number of perfectly matching residues between the two aligned sequences by the length of the longer sequence and multiplying by 100. For example, the percentage of identity may be determined by comparing sequence information using an advanced BLAST computer program (such as version 2.2.9) available from the National Institutes of Health. The BLAST program is based on the alignment method described by Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87:2264-2268 (1990), which has been studied in Altschul, et al., J. Mol. Biol., 215:403-410 (1990); Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993); and Altschul et al., Nucleic Acids Res., 25:3389-3402 (1997). The BLAST program can also be used to determine the percentage of identity relative to the full length of the comparison protein. Default parameters are set to optimize searches using short query sequences, for example, with the blastp program.The BLAST program may also use a SEG filter to mask specific segments of the query sequence determined by the SEG program described in Wootton and Federhen, Computers and Chemistry 17:149-163 (1993). The desired range of sequence identity is approximately 50% to 100%, including integer values ​​within that range. Typically, this disclosure includes sequences having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% sequence identity with the sequences provided herein.

[0083] The term “gene” typically refers to nucleic acids (e.g., DNA such as genomic DNA and cDNA) and the corresponding nucleotide sequences involved in coding RNA transcripts. As used herein with respect to genomic DNA, this term also includes intervening non-coding and regulatory regions, as well as the 5' and 3' ends. In some uses, the term “gene” encompasses transcriptional sequences including the 5' and 3' uncoding regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcriptional region includes an “open reading frame” that codes for a polypeptide. In some uses of the term “gene,” this term includes only the coding sequence required for coding a polypeptide (e.g., the “open reading frame” or “coding region”). In some cases, a gene does not code for a polypeptide, such as a ribosomal RNA gene (rRNA) or a transfer RNA (tRNA) gene. In some cases, the term “gene” includes not only the transcriptional sequence but also non-coding regions, including upstream and downstream regulatory regions such as enhancers and promoters. For example, a gene may refer to a portion of a gene located near its transcription start site (TSS), or a portion of a gene adjacent to its transcription start site (TSS). Such a gene (e.g., a target-directed gene disclosed herein) is one whose distance from its TSS to the gene is at least 2,000 nucleic acid base lengths, up to 2,000 nucleic acid base lengths, at least about 2,000 nucleic acid base lengths, or up to about 2,000 nucleic acid base lengths, at least 1,800 nucleic acid base lengths, up to 1,800 nucleic acid base lengths, at least about 1,800 nucleic acid base lengths, or up to about 1,800 nucleic acid base lengths, at least 1,600 nucleic acid base lengths. , up to 1,600 nucleic acid base lengths, at least approximately 1,600 nucleic acid base lengths, or up to approximately 1,600 nucleic acid base lengths, at least 1,500 nucleic acid base lengths, up to 1,500 nucleic acid base lengths, at least approximately 1,500 nucleic acid base lengths, or up to approximately 1,500 nucleic acid base lengths, at least 1,400 nucleic acid base lengths, up to 1,400 nucleic acid base lengths, at least approximately 1,400 nucleic acid base lengths, or up to approximately 1,400 nucleic acid base lengths, at least 1,200 nucleic acid base lengths, up to 1,200 nucleic acid base lengths, at least approximately 1,200 nucleic acid base lengths, or up to approximately 1,200 nucleic acid base lengths, at least 1,000 nucleic acid base lengths, up to 1,000 nucleic acid base lengths, at least approximately 1,000 nucleic acid base lengths, or up to approximately 1,000 nucleic acid base lengths, at least 900 nucleic acid base lengths, up to 900 nucleic acid base lengths, at least approximately 900 nucleic acid base lengths, or up to approximately 900 nucleic acid base lengths, at least 800 nucleic acid base lengths, up to 800 nucleic acid base lengths, at least approximately 800 nucleic acid base lengths, or up to approximately 800 nucleic acid base lengths, at least 700 nucleic acid base lengths, up to 700 nucleic acid base lengths, at least approximately 700 nucleic acid base lengths, or up to approximately 700 nucleic acid base lengths, at least 600 nucleic acid base lengths, up to 600 nucleic acid base lengths, at least approximately 600 nucleic acid base lengths, or up to approximately 600 nucleic acid base lengths, at least It may be 500 nucleic acid base lengths, a maximum of 500 nucleic acid base lengths, at least approximately 500 nucleic acid base lengths, or a maximum of approximately 500 nucleic acid base lengths, at least 400 nucleic acid base lengths, a maximum of 400 nucleic acid base lengths, at least approximately 400 nucleic acid base lengths, or a maximum of approximately 400 nucleic acid base lengths, at least 300 nucleic acid base lengths, a maximum of 300 nucleic acid base lengths, at least approximately 300 nucleic acid base lengths, or a maximum of approximately 300 nucleic acid base lengths, at least 200 nucleic acid base lengths, a maximum of 200 nucleic acid base lengths, at least approximately 200 nucleic acid base lengths, or a maximum of approximately 200 nucleic acid base lengths, at least 100 nucleic acid base lengths, a maximum of 100 nucleic acid base lengths, at least approximately 100 nucleic acid base lengths, or a maximum of approximately 100 nucleic acid base lengths, or at least 50 nucleic acid base lengths, a maximum of 50 nucleic acid base lengths, at least approximately 50 nucleic acid base lengths, or a maximum of approximately 50 nucleic acid base lengths.

[0084] "Genes" may mean "endogenous genes" or native genes that are located in their natural positions in the genome of an organism. "Genes" may also mean "exogenous genes" or non-native genes. "Non-native genes" may mean genes that are not normally found in a host organism but have been introduced into the host organism through gene transfer. Furthermore, "non-native genes" may mean genes that do not exist in their natural positions in the genome of an organism. "Non-native genes" may also mean native nucleic acid sequences or polypeptide sequences (e.g., non-native sequences) that include mutations, insertions, and / or deletions.

[0085] The term “expression” typically refers to one or more processes in which polynucleotides (e.g., mRNA or other RNA transcripts) are transcribed from a DNA template, and / or the process in which the transcribed mRNA is translated into peptides, polypeptides, or proteins. Transcripts and the polypeptides they encode can be collectively referred to as “gene products.” If the polynucleotides originate from genomic DNA, “expression” may include the splicing of mRNA in eukaryotic cells. “Upregulation” of expression typically refers to an increase in the expression level of polynucleotide sequences (e.g., RNA such as mRNA) and / or polypeptide sequences compared to wild-type expression levels, while “downregulation” typically refers to a decrease in the expression level of polynucleotide sequences (e.g., RNA such as mRNA) and / or polypeptide sequences compared to wild-type expression levels. Expression of a transfected gene may occur transiently or stably in a cell. In “transient expression,” the transfected gene is not transferred to daughter cells during cell division. In transient expression, gene expression occurs only in transfected cells, and thus gene expression is lost over time. In contrast, stable expression of a transfected gene can occur when the gene is transfected simultaneously with another gene that confers a selective advantage to the transfected cells. Such a selective advantage may be resistance to a specific toxin to which the cells are exposed.

[0086] The term "expression profile" typically refers to the quantitative (e.g., abundance) and qualitative expression of one or more genes in a sample (e.g., cells). These one or more genes can be expressed and identified in the form of nucleic acid molecules (e.g., mRNA or other RNA transcripts). In addition, or in another aspect, these one or more genes can be expressed and identified in the form of polypeptides (e.g., proteins measured by Western blotting). A gene expression profile may be defined as the amount of gene expression within a given time frame (for example, at least 1 hour, up to 1 hour, at least about 1 hour, or up to about 1 hour, at least 2 hours, up to 2 hours, at least about 2 hours, or up to about 2 hours, at least 3 hours, up to 3 hours, at least about 3 hours, or up to about 3 hours, at least 4 hours, up to 4 hours, at least about 4 hours, or up to about 4 hours, at least 5 hours, up to 5 hours, at least about 5 hours, or up to about 5 hours, at least 6 hours, up to 6 hours, at least about 6 hours, or up to about 6 hours, at least 7 hours, up to 7 hours, at least about 7 hours, or up to about 7 hours, at least 8 hours, up to 8 hours, at least about 8 hours, or up to about 8 hours, at least 9 hours, up to 9 hours, or at least about 9 hours up to approximately 9 hours, at least 10 hours, up to 10 hours, at least approximately 10 hours, or up to approximately 10 hours, at least 11 hours, up to 11 hours, at least approximately 11 hours, or up to approximately 11 hours, at least 12 hours, up to 12 hours, at least approximately 12 hours, or up to approximately 12 hours, at least 16 hours, up to 16 hours, at least approximately 16 hours, or up to approximately 16 hours, at least 18 hours, up to 18 hours, at least approximately 18 hours, or up to approximately 18 hours, at least 24 hours, up to 24 hours, at least approximately 24 hours, or up to approximately 24 hours, at least 36 hours, up to 36 hours, at least approximately 36 hours, or up to approximately 36 hours, at least 48 hours, up to 48 hours, at least approximately 48 hours, or up to approximately 48 hours, at least 3 days, up to 3 days, at least approximately 3 days, or up to approximately 3 days,At least 4 days, up to 4 days, at least approximately 4 days, or up to approximately 4 days, at least 5 days, up to 5 days, at least approximately 5 days, or up to approximately 5 days, at least 6 days, up to 6 days, at least approximately 6 days, or up to approximately 6 days, at least 7 days, up to 7 days, at least approximately 7 days, or up to approximately 7 days, at least 8 days, up to 8 days, at least approximately 8 days, or up to approximately 8 days, at least 9 days, up to 9 days, at least approximately 9 days, or up to (For example, approximately 9 days, at least 10 days, up to 10 days, at least about 10 days, or up to about 10 days, at least 11 days, up to 11 days, at least about 11 days, or up to about 11 days, at least 12 days, up to 12 days, at least about 12 days, or up to about 12 days, at least 13 days, up to 13 days, at least about 13 days, or up to about 13 days, or at least 14 days, up to 14 days, at least about 14 days, or up to about 14 days, etc.) Alternatively, the gene expression profile may be defined as the amount of gene expression at a point in time of interest (for example, at least 1 hour, up to 1 hour, at least about 1 hour, or up to about 1 hour, at least 2 hours, up to 2 hours, at least about 2 hours, or up to about 2 hours, at least 3 hours, up to 3 hours, at least about 3 hours, or up to about 3 hours, at least 4 hours, up to 4 hours, at least about 4 hours, or up to about 4 hours, at least 5 hours, up to 5 hours, at least about 5 hours) , or at most about 5 hours later, at least 6 hours later, at most 6 hours later, at least about 6 hours later, or at most about 6 hours later, at least 7 hours later, at most 7 hours later, at least about 7 hours later, or at most about 7 hours later, at least 8 hours later, at most 8 hours later, at least about 8 hours later, or at most about 8 hours later, at least 9 hours later, at most about 9 hours later, or at most about 9 hours later, at least 10 hours later, at most about 10 hours later, or at most about 10 hours later, at least 11 hours later, at most about 11 hours later, or at most about 11 hours later,At least 12 hours later, up to 12 hours later, at least approximately 12 hours later, or up to approximately 12 hours later, at least 16 hours later, up to 16 hours later, at least approximately 16 hours later, or up to approximately 16 hours later, at least 18 hours later, up to 18 hours later, at least approximately 18 hours later, or up to approximately 18 hours later, at least 24 hours later, up to 24 hours later, at least approximately 24 hours later, or up to approximately 24 hours later, at least 36 hours later, up to 36 hours later, at least approximately 36 hours later, or up to approximately 36 hours later, at least 48 hours later, up to 48 hours later, at least approximately 48 hours later, or up to approximately 48 hours later, at least 3 days later, up to 3 days later, at least approximately 3 days later, or up to approximately 3 days later, at least 4 days later, up to 4 days later, at least approximately 4 days later, or up to approximately 4 days later, at least 5 days later, up to 5 days later, at least approximately 5 days later, or up to approximately 5 days Then, at least 6 days later, up to 6 days later, at least approximately 6 days later, or up to approximately 6 days later, at least 7 days later, up to 7 days later, at least approximately 7 days later, or up to approximately 7 days later, at least 8 days later, up to 8 days later, at least approximately 8 days later, or up to approximately 8 days later, at least 9 days later, up to 9 days later, at least approximately 9 days later, or up to approximately 9 days later, at least 10 days later, up to 10 days later, at least approximately 10 days later, or up to approximately 10 days (The gene expression levels may also be measured at least 11 days later, up to 11 days later, at least approximately 11 days later, or up to approximately 11 days later, at least 12 days later, up to approximately 12 days later, or up to approximately 12 days later, at least 13 days later, up to approximately 13 days later, or up to approximately 13 days later, or at least 14 days later, up to approximately 14 days later, or up to approximately 14 days later.)

[0087] In this specification, the terms “peptide,” “polypeptide,” and “protein” are used interchangeably and typically refer to polymers consisting of at least two amino acid residues linked by peptide bonds. These terms do not imply polymers of a specific length, nor do they imply or distinguish whether a peptide is produced using genetic engineering, chemical synthesis, or enzymatic synthesis, or whether it is naturally occurring. These terms are also used for amino acid polymers containing at least one modified amino acid, as well as naturally occurring amino acid polymers. Non-amino acids may be inserted into these polymers as needed. These terms may also include amino acid chains of any length, and may include full-length proteins, proteins with secondary and / or tertiary structures (e.g., domains), or proteins without secondary and / or tertiary structures. These terms also include modified amino acid polymers, including those modified by, for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, or other operations such as complex formation with labeling components. In this specification, the term “amino acid” usually means natural and non-natural amino acids, and includes, but is not limited to, modified amino acids and amino acid analogs. Modified amino acids include natural and non-natural amino acids that have been chemically modified to include a group or chemical part not present in natural amino acids. Amino acid analogs may mean amino acid derivatives. The term “amino acid” includes both D-amino acids and L-amino acids.

[0088] In this specification, the terms “derivative,” “variant,” or “fragment” as used with respect to a polypeptide generally mean a polypeptide relating to the wild-type polypeptide, for example, a polypeptide relating to the wild-type polypeptide in terms of amino acid sequence, structure (e.g., secondary and / or tertiary structure), activity (e.g., enzymatic activity), and / or function. Polypeptide derivatives, variants, and fragments may include variations (e.g., mutations, insertions, and deletions), cleavage, modification, or combinations thereof of one or more amino acids compared to the wild-type polypeptide.

[0089] In this specification, the terms “recombinant,” “chimeric,” or “recombinant” as used with respect to polypeptide molecules (e.g., proteins) usually mean polypeptide molecules having heterologous or modified amino acid sequences as a result of applying genetic engineering techniques to the nucleic acid encoding the polypeptide molecule, as well as cells or organisms that express such polypeptide molecules. In this specification, the terms “recombinant” or “recombinant” as used with respect to polynucleotide molecules (e.g., DNA molecules or RNA molecules) usually mean polynucleotide molecules having heterologous or modified nucleic acid sequences as a result of applying genetic engineering techniques. Genetic engineering techniques include, but are not limited to, PCR and DNA cloning techniques; transfection, transformation, and other gene transfer techniques; homologous recombination; site-directed mutagenesis; and gene fusion. In some cases, recombinant polynucleotides or recombinant polynucleotides (e.g., genomic DNA sequences) can be modified or altered by gene editing regions.

[0090] In this specification, the terms “recombinant” and “modified” are used interchangeably. In this specification, the terms “to recombinate” and “to modify” are used interchangeably. In this specification, the terms “recombinant cell” and “modified cell” are used interchangeably. In this specification, the terms “recombinant property” and “modified property” are used interchangeably.

[0091] The terms “enhanced expression,” “increased expression,” or “upregulated expression” typically refer to the production of a target molecule (e.g., a polynucleotide or polypeptide) at an expression level exceeding its normal expression level in a host cell. Normal expression levels may be substantially zero (or null) or greater than zero. The target molecule may include an endogenous gene or endogenous polypeptide construct from the host cell. The target molecule may also include a heterologous gene or heterologous polypeptide construct introduced into the host cell. For example, to enhance the expression of a target polypeptide in a host cell, a heterologous gene encoding the target polypeptide can be knocked into the host cell genome (KI).

[0092] The terms “enhanced activity,” “increased activity,” or “upregulated activity” typically refer to the activity of a target portion (e.g., a polynucleotide or polypeptide) that has been modified to exceed its normal activity level in a host cell. The normal activity level may be substantially zero (or null) or greater than zero. The target portion may contain a polypeptide construct from the host cell. The target portion may also contain a heterologous polypeptide construct introduced into the host cell. For example, to enhance the activity of a target polypeptide in a host cell, a heterologous gene encoding the target polypeptide can be knocked into the host cell genome (KI).

[0093] The terms “repression of expression,” “reduction of expression,” or “downregulated expression” typically refer to the production of a target molecule (e.g., a polynucleotide or polypeptide) at an expression level below its normal level in a host cell. Normal expression levels are above zero. The target molecule may include an endogenous gene or endogenous polypeptide construct in the host cell. In some cases, the target molecule can be knocked out or knocked down in the host cell. In some examples, repression of the expression of the target molecule may include complete inhibition of such expression in the host cell.

[0094] The terms “suppressed activity,” “reduced activity,” or “downregulated activity” typically refer to the activity of a target portion (e.g., a polynucleotide or polypeptide) that has been modified to be below its normal activity level in a host cell. Normal activity levels are above zero. The target portion may include an endogenous gene or endogenous polypeptide construct in the host cell. In some cases, the target portion can be knocked out or knocked down in the host cell. In some examples, the reduction in the activity of the target portion may include complete inhibition of such activity in the host cell.

[0095] In this specification, the terms “subject,” “individual,” or “patient” are used interchangeably and usually mean vertebrates, preferably mammals such as humans. Mammals include, but are not limited to, mice, monkeys, humans, livestock, sports animals, and pets. They also include tissues, cells, and their offspring of biological objects obtained in vivo or in vitro culture.

[0096] The terms “treatment” or “to treat” typically mean an approach to obtain beneficial or desired outcomes, such as therapeutic benefits and / or preventive benefits, but are not limited to the above. For example, “treatment” may include administering the systems or cell populations disclosed herein. “Therapeutic benefits” means any improvement related to treatment of one or more diseases, conditions or symptoms being treated, or an effect on one or more diseases, conditions or symptoms being treated. To obtain preventive benefits, the composition may be administered to subjects at risk of developing a particular disease, condition or symptom, or to subjects complaining of one or more physiological symptoms of a disease, even before the disease, condition or symptoms appear.

[0097] The terms “effective dose” or “therapeutic dose” typically mean an amount of a composition sufficient to obtain the desired activity when administered to a subject requiring the administration of the composition, and this composition is, for example, a composition comprising heterologous polypeptides, heterologous polynucleotides and / or recombinant cells (e.g., modified stem cells). In the context of this disclosure, the term “therapeutically effective” typically means an amount of a composition sufficient to delay the onset of a disease treated by the method of this disclosure, halt its progression, or alleviate or reduce at least one of its symptoms.

[0098] In this specification, the term “muscle cell” generally refers to any cell that contributes to muscle tissue. Myoblasts, satellite cells, myotubes, and myofibrils are all included in the term “muscle cell.” Effects on muscle cells may be induced in skeletal muscle, cardiac muscle, and smooth muscle.

[0099] Diseases or conditions can result from the abnormal expression of one or more genes. Such abnormal expression may be characterized by abnormally low gene expression levels, or by abnormally high gene expression levels. In some cases, such abnormal expression can be reversed by genetic modification of these genes (for example, through the action of an endonuclease such as the CRISPR-Cas enzyme) (for example, for the treatment of Duchenne muscular dystrophy (DMD)). Alternatively, abnormal expression can be transiently modified without genetic modification of the target gene, for example, by targeting the gene with a gene effector (for example, by using an inactivated CRISPR-Cas enzyme linked to a gene effector). Transient modification of the abnormal expression of a target gene in cells may be insufficient to treat or cure diseases caused by the abnormal expression of the target gene. Therefore, in some embodiments, this disclosure provides systems and methods for modifying the abnormal expression of a target gene and maintaining the modified expression level of the target gene for a long period of time.

[0100] Modification of abnormal expression of target genes This disclosure provides compositions, systems, and methods for regulating the abnormal expression of a target gene in cells (e.g., muscle cells). For example, the target gene may be located within a D4Z4 repeat array. The target gene may encode at least a portion of DUX4. The compositions, systems, and methods disclosed herein can modify the expression level and / or epigenetic modification level (e.g., methylation level) of a target gene by utilizing at least a heterologous polypeptide (e.g., a heterologous actuator moiety which may contain heterologous polynucleotides such as guide nucleic acid molecules). For example, the compositions, systems, and methods disclosed herein can modify the expression level and / or epigenetic modification level of a target gene by utilizing a heterologous actuator moiety operably linked (e.g., via covalent or non-covalent bonds) to a heterologous gene effector or heterologous gene regulator (e.g., a gene actuator or gene repressor). In some embodiments, the heterologous polypeptide includes a heterologous actuator moiety operably linked (e.g., via covalent or non-covalent bonds) to a heterologous gene effector or heterologous gene regulator (e.g., a gene actuator or gene repressor). In some embodiments, the heterologous polypeptide includes a heterologous actuator moiety comprising a nuclease and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide. In some embodiments, the guide nucleic acid molecule configured to form a complex with the heterologous polypeptide includes (i) a scaffold sequence configured to form a complex with the heterologous polypeptide and (ii) a spacer sequence exhibiting specific binding to the target polynucleotide sequence, for example, a spacer sequence targeting at least a portion of DUX4 as described herein. In some embodiments, the nuclease is fused (directly or indirectly) to a heterologous gene effector or heterologous gene regulator (e.g., a gene actuator or gene repressor).In some embodiments, the nuclease is fused to the C-terminus of a heterogeneous effector or heterogeneous regulator (e.g., a gene actuator or gene repressor).

[0101] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The aforementioned guide nucleic acid molecule exhibits specific binding to target polynucleotide sequences present in or adjacent to the D4Z4 repeat array within muscle cells. Once the complex is formed, it can bind to the target polynucleotide sequence and modify the expression level and / or methylation level of the target gene within the muscle cell. The target gene is located within the D4Z4 repeat array. In some embodiments, (1) The guide nucleic acid molecule includes a spacer sequence exhibiting the specific binding, the spacer sequence includes or is encoded by a polynucleotide sequence that exhibits at least 80% or at least about 80%, at least 90% or at least about 90%, or at least 95% or at least about 95% sequence identity with one or more member polynucleotide sequences listed in Table 4, the spacer sequence may also include or be encoded by a polynucleotide sequence that exhibits at least 80% or at least about 80%, at least 90% or at least about 90%, or at least 95% or at least about 95% sequence identity with the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851, (2) The system is configured to reduce the expression level of the target gene in the muscle cells by at least 80% or at least about 80% compared to control cells. (3) The system is configured to modify the expression level and / or methylation level of the target gene in the muscle cell while exhibiting off-target effects below a predetermined threshold in the muscle cell, (4) The system is configured to continuously regulate the expression level and / or methylation level of downstream genes of the target gene for at least about 5 days, (5) The system is configured to normalize the apoptosis level of the muscle cells to a level similar to that of healthy muscle cells. (6) The system is configured to improve the survival of the muscle cells in vitro or in vivo, (7) The system is configured to have minimal effect on the expression profile of at least one muscle cell-specific gene different from the target gene in the muscle cell, and / or (8) The system is configured to have minimal effect on at least one sign of a health condition in a subject who receives the system. In some embodiments, the system of this disclosure A heterogeneous polypeptide comprising a nuclease (e.g., Cas12f or its variant) containing an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterogeneous polypeptide comprising a nuclease (e.g., Cas12f or its variant) containing the amino acid sequence shown in any one of SEQ ID NOs: 43, 44, and 728, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having a sequence represented by any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865. In some embodiments, the system of this disclosure A heterologous polypeptide comprising a nuclease (e.g., Cas12f or its variant) operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease comprises an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous polypeptide comprises an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 43, 44, and 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide comprising a nuclease (e.g., Cas12f or its variant) operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease comprises the amino acid sequence shown in any one of SEQ ID NOs: 43, 44, and 728. The heterogeneous gene effector comprises an amino acid sequence represented by any one of sequence numbers 15-42 and 727. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence represented by any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865. In some embodiments, the system of this disclosure A heterologous polypeptide comprising a nuclease (e.g., Cas12f or its variant) operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease comprises the amino acid sequence shown in any one of SEQ ID NOs: 43, 44, and 728. The aforementioned heterogeneous effector includes the amino acid sequence shown in SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence represented by any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865. In some embodiments, the system of this disclosure A heterologous polypeptide comprising a nuclease (e.g., Cas12f or its variant) operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide comprising a nuclease (e.g., Cas12f or its variant) operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence represented by any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865. In the systems or complexes of the present disclosure, the guide nucleic acid molecule may include (i) a scaffold sequence configured to form a complex with a nuclease, and (ii) a spacer sequence as described herein. In some embodiments of the systems or complexes of the present disclosure, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments of the systems or complexes of the present disclosure, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease. In some embodiments of the systems or complexes of the present disclosure, the heterogeneous effector is fused inside the nuclease. In some embodiments, the heterogeneous polypeptide comprises a Cas12f-KRAB-DNMT3L modulator, and the nuclease, which is Cas12f (or a variant thereof) as described herein, is operably linked to the heterogeneous effector comprising KRAB and DNMT3L as described herein (e.g., fused to the C-terminus via or without a linker).

[0102] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 800, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 800. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0103] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 867, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 867. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0104] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 836, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 836. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0105] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 851, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 851. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0106] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 874, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 874. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0107] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 841, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 841. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0108] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 830, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 830. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0109] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 833, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The aforementioned guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 833. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0110] In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterologous gene effector includes an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. In some embodiments, the system of this disclosure A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains the amino acid sequence of SEQ ID NO: 728, The aforementioned heterogeneous effector includes the amino acid sequence of SEQ ID NO: 727, The guide nucleic acid molecule includes a spacer sequence encoded by the polynucleotide sequence of sequence number 865. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterogeneous effector is fused (directly or indirectly) to the N-terminus of the nuclease.

[0111] In some cases, the cells may be muscle cells. The muscle cells disclosed herein may be muscle cells of any developmental stage or classification. The muscle cells may include undifferentiated muscle cells (e.g., mononuclear cells such as muscle stem cells, muscle satellite cells, and myoblasts). In addition to this, or in another embodiment, the muscle cells may include differentiated muscle cells (e.g., multinucleated muscle cells such as myotubes). The muscle cells may be skeletal muscle cells, cardiomyocytes, or smooth muscle cells. For example, skeletal muscle cells may be primary myoblasts (e.g., immortalized primary myoblast cell lines). In some cases, the cells may be non-muscle cells, such as lymphoblasts.

[0112] In some cases, the target gene may be located on chromosome 4 of the cells disclosed herein. In some cases, the target gene may be located on chromosome 10 of the cells disclosed herein, for example, in the distal portion of the q arm (long arm) of chromosome 10 of the cells disclosed herein.

[0113] The unmodified abnormal expression of the target gene disclosed herein may be characterized by the fact that the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene is at least 1%, at least 5%, at least 10%, at least 20%, at least 50%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or more than that of a control cell (e.g., a healthy cell of a healthy subject). The abnormal expression of the target gene may be characterized by the fact that the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene is higher than that of control cells (e.g., healthy cells of a healthy subject) by up to 500%, up to 400%, up to 300%, up to 200%, up to 150%, up to 100%, up to 50%, up to 20%, up to 10%, up to 5%, up to 1%, or less, or by up to approximately 500%, up to approximately 400%, up to approximately 300%, up to approximately 200%, up to approximately 150%, up to approximately 100%, up to approximately 50%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0114] The abnormal expression of a target gene prior to modification disclosed herein may be characterized by the fact that the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene is lower by at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 70%, at least 99%, or more than that of a control cell (e.g., a healthy cell of a healthy subject). The abnormal expression of the target gene may be characterized by the fact that the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene is lower than that of control cells (e.g., healthy cells of a healthy subject) by up to 100%, up to 70%, up to 50%, up to 40%, up to 30%, up to 20%, up to 10%, up to 5%, up to 1%, or less, or by up to approximately 100%, up to approximately 70%, up to approximately 50%, up to approximately 40%, up to approximately 30%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0115] The abnormal expression of a target gene prior to modification disclosed herein may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being at least 1%, at least 5%, at least 10%, at least 20%, at least 50%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or more than that of a control cell (e.g., a healthy cell of a healthy subject). The abnormal expression of the target gene may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being longer by up to 500%, up to 400%, up to 300%, up to 200%, up to 150%, up to 100%, up to 50%, up to 20%, up to 10%, up to 5%, or less than that of control cells (e.g., healthy cells of a healthy subject), or by up to approximately 500%, up to approximately 400%, up to approximately 300%, up to approximately 200%, up to approximately 150%, up to approximately 100%, up to approximately 50%, up to approximately 20%, up to approximately 10%, up to approximately 5%, or less than that.

[0116] The abnormal expression of a target gene prior to modification disclosed herein may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 70%, at least 99%, or more than that of a control cell (e.g., a healthy cell of a healthy subject). The abnormal expression of the target gene may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being shorter by up to 100%, up to 70%, up to 50%, up to 40%, up to 30%, up to 20%, up to 10%, up to 5%, up to 1%, or less, compared to control cells (e.g., healthy cells of a healthy subject), or by up to approximately 100%, up to approximately 70%, up to approximately 50%, up to approximately 40%, up to approximately 30%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0117] Modified abnormal expression of a target gene after it has been modified as disclosed herein may be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least 1%, at least 5%, at least 10%, at least 20%, at least 50%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or more, compared to a control cell (e.g., an unmodified control cell), or at least about 1%, at least about 5%, at least about 10%, at least about 20%, at least about 50%, at least about 100%, at least about 150%, at least about 200%, at least about 300%, at least about 400%, at least about 500%, or more. The modification of the abnormal expression of the target gene may be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by up to 500%, up to 400%, up to 300%, up to 200%, up to 150%, up to 100%, up to 50%, up to 20%, up to 10%, up to 5%, up to 1%, or less, compared to control cells, by up to approximately 500%, up to approximately 400%, up to approximately 300%, up to approximately 200%, up to approximately 150%, up to approximately 100%, up to approximately 50%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0118] Modified abnormal expression of a target gene after it has been modified as disclosed herein may be characterized by a reduction in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 70%, at least 99%, or more, compared to a control cell (e.g., an unmodified control cell). The modification of the abnormal expression of the target gene may be characterized by a reduction in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by a percentage of up to 100%, up to 70%, up to 50%, up to 40%, up to 30%, up to 20%, up to 10%, up to 5%, up to 1%, or less, or by a percentage of up to approximately 100%, up to approximately 70%, up to approximately 50%, up to approximately 40%, up to approximately 30%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0119] The modified abnormal expression of the target gene after it has been modified as disclosed herein may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being at least 1%, at least 5%, at least 10%, at least 20%, at least 50%, at least 100%, at least 150%, at least 200%, at least 300%, at least 400%, at least 500%, or more, compared to a control cell (e.g., an unmodified control cell). The modification of the abnormal expression of the target gene may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being longer by up to 500%, up to 400%, up to 300%, up to 200%, up to 150%, up to 100%, up to 50%, up to 20%, up to 10%, up to 5%, up to 1%, or less, compared to control cells, by up to approximately 500%, up to approximately 400%, up to approximately 300%, up to approximately 200%, up to approximately 150%, up to approximately 100%, up to approximately 50%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0120] The modified abnormal expression of the target gene after it has been modified as disclosed herein may be characterized by the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene being shorter by at least 1%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 70%, at least 99%, or more compared to control cells (e.g., unmodified control cells). The modification of the abnormal expression of the target gene may be characterized by a shortening of the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by up to 100%, up to 70%, up to 50%, up to 40%, up to 30%, up to 20%, up to 10%, up to 5%, up to 1%, or less, or by up to approximately 100%, up to approximately 70%, up to approximately 50%, up to approximately 40%, up to approximately 30%, up to approximately 20%, up to approximately 10%, up to approximately 5%, up to approximately 1%, or less.

[0121] The systems or their use provided herein can sustainably regulate (e.g., reduce) the expression levels of at least one downstream gene of a target gene (e.g., ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, and / or RFLP2, which are downstream genes of Dux4) after the target gene has been modified in vitro, ex vivo, or in vivo. In some cases, the regulation of the expression level of at least one downstream gene can result in an at least 10%, up to 10%, at least about 10%, or up to about 10%, at least 20%, up to 20%, at least about 20%, or up to about 20%, at least 30%, up to 30%, at least about 30%, or up to about 30%, at least 40%, up to 40%, at least about 40%, or up to about 40%, at least 50%, up to 50%, at least about 50%, or up to about 50%, at least 60%, up to 60%, at least about 60%, or up to about 60%, at least 70% , or it may be a downregulation of up to 70%, at least about 70%, or up to about 70%, at least 75%, up to 75%, at least about 75%, or up to about 75%, at least 80%, up to 80%, at least about 80%, or up to about 80%, at least 85%, up to 85%, at least about 85%, or up to about 85%, at least 90%, up to 90%, at least about 90%, or up to about 90%, at least 95%, up to 95%, at least about 95%, or up to about 95%, at least 99%, up to 99%, at least about 99%, or up to about 99%, or substantially about 100%.

[0122] The modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the abnormally expressed target gene) or other gene of interest (e.g., a downstream gene of the target gene, such as a cell type-specific gene) after the target gene has been modified as disclosed herein shall last for at least 1 day or at least about 1 day, at least 2 days or at least about 2 days, at least 3 days or at least about 3 days, at least 4 days or at least about 4 days, at least 5 days or at least about 5 days, at least 6 days or at least about 6 days, at least 7 days or at least about 7 days, at least 8 days or at least about 8 days, at least 9 days or at least about 9 days, at least 10 days or at least about 10 days, at least 11 days or at least about 11 days, at least 12 days or at least about 12 days, at least 13 days or at least about 1 It can be made to last for 3 days, at least 2 weeks or at least about 2 weeks, at least 3 weeks or at least about 3 weeks, at least 4 weeks or at least about 4 weeks, at least 2 months or at least about 2 months, at least 3 months or at least about 3 months, at least 4 months or at least about 4 months, at least 5 months or at least about 5 months, at least 6 months or at least about 6 months, at least 7 months or at least about 7 months, at least 8 months or at least about 8 months, at least 9 months or at least about 9 months, at least 10 months or at least about 10 months, at least 11 months or at least about 11 months, at least 12 months or at least about 12 months, at least 2 years or at least about 2 years, at least 3 years or at least about 3 years, at least 4 years or at least about 4 years, at least 5 years or at least about 5 years, or longer.The modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the abnormally expressed target gene) or other target gene (e.g., a downstream gene of the target gene, such as a cell type-specific gene) after the target gene has been modified as disclosed herein may be maintained for up to 5 years or approximately 5 years, up to 4 years or 4 years, up to 3 years or 3 years, up to 2 years or 2 years, up to 12 months or 12 months, up to 11 months or 11 months, up to 10 months or 10 months, up to 9 months or 9 months, up to 8 months or 8 months, up to 7 months or 7 months, up to 6 months or 6 months, up to 5 months or 5 months, up to 4 months or It can be sustained for a maximum of 4 months, a maximum of 3 months or a maximum of 3 months, a maximum of 2 months or a maximum of 2 months, a maximum of 4 weeks or a maximum of 4 weeks, a maximum of 3 weeks or a maximum of 3 weeks, a maximum of 2 weeks or a maximum of 2 weeks, a maximum of 13 days or a maximum of approximately 13 days, a maximum of 12 days or a maximum of approximately 12 days, a maximum of 11 days or a maximum of approximately 11 days, a maximum of 10 days or a maximum of approximately 10 days, a maximum of 9 days or a maximum of approximately 9 days, a maximum of 8 days or a maximum of approximately 8 days, a maximum of 7 days or a maximum of approximately 7 days, a maximum of 6 days or a maximum of approximately 6 days, a maximum of 5 days or a maximum of approximately 5 days, a maximum of 4 days or a maximum of approximately 4 days, a maximum of 3 days or a maximum of approximately 3 days, a maximum of 2 days or a maximum of approximately 2 days, a maximum of 1 day or a maximum of approximately 1 day, or less.

[0123] After the target gene has been modified as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the abnormally expressed target gene) or other gene of interest (e.g., a downstream gene of the target gene, such as a cell type-specific gene) is such that it can withstand at least one or about one cell division, at least two or about two cell divisions, at least three or about three cell divisions, at least four or about four cell divisions, at least five or about five cell divisions, at least six or about six cell divisions, at least seven or fewer cell divisions. It can be sustained through at least about 7 cell divisions, at least 8 or at least about 8 cell divisions, at least 9 or at least about 9 cell divisions, at least 10 or at least about 10 cell divisions, at least 15 or at least about 15 cell divisions, at least 20 or at least about 20 cell divisions, at least 25 or at least about 25 cell divisions, at least 30 or at least about 30 cell divisions, at least 40 or at least about 40 cell divisions, at least 50 or at least about 50 cell divisions, or at least 100 or at least about 100 cell divisions.The modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the abnormally expressed target gene) or other target gene (e.g., a downstream gene of the target gene, such as a cell type-specific gene) after the target gene has been modified as disclosed herein is up to 100 or up to about 100 cell divisions, up to 50 or up to about 50 cell divisions, up to 40 or up to about 40 cell divisions, up to 30 or up to about 30 cell divisions, up to 25 or up to about 25 cell divisions, up to 20 or up to It can be sustained through approximately 20 cell divisions at its largest, up to 15 or up to approximately 15 cell divisions, up to 10 or up to approximately 10 cell divisions, up to 9 or up to approximately 9 cell divisions, up to 8 or up to approximately 8 cell divisions, up to 7 or up to approximately 7 cell divisions, up to 6 or up to approximately 6 cell divisions, up to 5 or up to approximately 5 cell divisions, up to 4 or up to approximately 4 cell divisions, up to 3 or up to approximately 3 cell divisions, up to 2 or up to approximately 2 cell divisions, or up to 1 or up to approximately 1 cell division.

[0124] The modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., the abnormally expressed target gene) or other target gene (e.g., a downstream gene of the target gene, such as a cell type-specific gene) after the target gene has been modified as disclosed herein shall remain unchanged for at least 1 day or at least about 1 day, at least 2 days or at least about 2 days, at least 3 days or at least about 3 days, at least 4 days or at least about 4 days, at least 5 days or at least about 5 days, at least 6 days or at least Measurements can also be taken after approximately 6 days, at least 7 days or at least about 7 days, at least 8 days or at least about 8 days, at least 9 days or at least about 9 days, at least 10 days or at least about 10 days, at least 11 days or at least about 11 days, at least 12 days or at least about 12 days, at least 13 days or at least about 13 days, at least 2 weeks or at least about 2 weeks, at least 3 weeks or at least about 3 weeks, or at least 4 weeks or at least about 4 weeks.

[0125] Examples of epigenetic modifications, as disclosed herein, include, but are not limited to, methylation, acetylation, phosphorylation, ADP-ribosylation, glycosylation, SUMOylation, ubiquitination, and histone structure modification (e.g., via ATP hydrolysis-dependent processes). For example, epigenetic modifications may alter the methylation level of one or more target genes or other genes of interest (e.g., downstream genes of target genes, such as cell-type-specific genes).

[0126] As disclosed herein, the persistence of a modified expression level and / or epigenetic modification level (e.g., methylation level) of a target gene (or other gene of purpose) can be assessed by the maintenance of the modified expression level and / or methylation level of the target gene (or other gene of purpose) at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100%, or at least about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100%. The persistence of a modified expression level and / or epigenetic modification level (e.g., methylation level) of a target gene (or other gene of interest) can be assessed by whether the modified expression level and / or methylation level of the target gene (or other gene of interest) is maintained at a maximum of 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70%, or at a maximum of approximately 100%, approximately 99%, approximately 98%, approximately 97%, approximately 96%, approximately 95%, approximately 90%, approximately 85%, approximately 80%, approximately 75%, or approximately 70%.

[0127] The systems or their use provided herein may have minimal effect (e.g., substantially no effect) on the expression profile of at least one cell type-specific gene in cells other than target cells. For example, the systems or their use provided herein may result in minimal changes (e.g., no changes) in the expression profile of at least one myogenic gene in muscle cells after differentiation from myoblasts to myotubes or muscle fibers. Examples of the at least one myogenic gene include, but are not limited to, p21, MyoD, Ezh2, Notch1, MyoG, MyHC, MEF2, Smad4, IGF2, Sirt1, myostatin, Myh2, Myh4, Myh1, myomixer, myomaker, and Mrf4.The expression profile of at least one myogenic gene in muscle cells including the system of this disclosure is at least 80%, up to 80%, at least about 80% or up to about 80%, at least 85%, up to 85%, at least about 85% or up to about 85%, at least 90%, up to 90%, at least about 90% or up to about 90%, at least 91%, up to 91%, at least about 91% or up to about 91%, at least 92%, up to 92%, at least about 92% or up to about 92%, at least 93%, up to 93%, at least about 93% or up to about 93%, at least 94%, up to 94%, at least about 94% or up to about 94%, at least 95%, up to 95%, less It may be at least approximately 95% or up to approximately 95%, at least 96%, up to 96%, at least approximately 96% or up to approximately 96%, at least 97%, up to 97%, at least approximately 97% or up to approximately 97%, at least 98%, up to 98%, at least approximately 98% or up to approximately 98%, at least 99%, up to 99%, at least approximately 99% or up to approximately 99%, at least 100%, up to 100%, at least approximately 100% or up to approximately 100%, at least 105%, up to 105%, at least approximately 105% or up to approximately 105%, at least 110%, up to 110%, at least approximately 110% or up to approximately 110%, or at least 120%, up to 120%, at least approximately 120% or up to approximately 120%.Following modification of the target gene provided herein, the minimum effect on the expression profile of at least one cell type-specific gene may last for at least 1 day or about 1 day, at least 2 days or about 2 days, at least 3 days or about 3 days, at least 4 days or about 4 days, at least 5 days or about 5 days, at least 6 days or about 6 days, at least 7 days or about 7 days, at least 8 days or about 8 days, at least 9 days or about 9 days, at least 10 days or about 10 days, at least 11 days or about 11 days, at least 12 days or about 12 days, at least 13 days or about 13 days, at least 14 days or about 14 days, at least 3 weeks or about 3 weeks, or at least 4 weeks or about 4 weeks. The expression profile of at least one cell type-specific gene can be measured by a variety of methods, including, but not limited to, reverse transcription polymerase chain reaction (RT-PCR or real-time quantitative RT-PCR) or Western blotting.

[0128] In cells (e.g., diseased cells such as FSHD muscle cells), the systems provided herein or their use can normalize the apoptotic level of such cells, for example, by adjusting the expression level of at least one stress-related marker provided herein (e.g., caspase protein) in the cells to a level similar to, close to, or equivalent to, the "normal" level in control cells (e.g., healthy muscle cells in a similar state of differentiation). As described herein, the regulated expression level of at least one stress-related marker in cells after modification of the target gene of the cell is at least 80%, up to 80%, at least about 80% or up to about 80%, at least 85%, up to 85%, at least about 85% or up to about 85%, at least 90%, up to 90%, at least about 90% or up to about 90%, at least 91%, up to 91%, at least about 91% or up to about 91%, at least 92%, up to 92%, at least about 92% or up to about 92%, at least 93%, up to 93%, at least about 93% or up to about 93%, at least 94%, up to 94%, at least about 94% or up to about 94%, at least 95%, up to It may be 95%, at least about 95% or up to about 95%, at least 96%, up to 96%, at least about 96% or up to about 96%, at least 97%, up to 97%, at least about 97% or up to about 97%, at least 98%, up to 98%, at least about 98% or up to about 98%, at least 99%, up to 99%, at least about 99% or up to about 99%, at least 100%, up to 100%, at least about 100% or up to about 100%, at least 105%, up to 105%, at least about 105% or up to about 105%, at least 110%, up to 110%, at least about 110% or up to about 110%, or at least 120%, up to 120%, at least about 120% or up to about 120%.The normalized expression level of at least one stress-related marker can be sustained for at least one day or about one day, at least two days or about two days, at least three days or about three days, at least four days or about four days, at least five days or about five days, at least six days or about six days, at least seven days or about seven days, at least eight days or about eight days, at least nine days or about nine days, at least ten days or about ten days, at least eleven days or about eleven days, at least twelve days or about twelve days, at least thirteen days or about thirteen days, at least fourteen days or about fourteen days, at least three weeks or about three weeks, at least four weeks or about four weeks, at least two months or about two months, at least six months or about six months, or at least one year or about one year. The expression level of at least one stress-related marker in lesion cells after modification of the target gene in the lesion cells as described herein may be at least 5% or at least about 5%, at least 10% or at least about 10%, at least 15% or at least about 15%, at least 20% or at least about 20%, at least 25% or at least about 25%, at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 80% or at least about 80%, or at least 90% or at least about 90% compared to control lesion cells (e.g., those without the system of this disclosure).

[0129] In cells (e.g., diseased cells such as FSHD muscle cells), the systems provided herein or their use can enhance the survival of such cells in vivo. The enhanced survival may be at least 5% or at least about 5%, at least 10% or at least about 10%, at least 15% or at least about 15%, at least 20% or at least about 20%, at least 25% or at least about 25%, at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 80% or at least about 80%, or at least 90% or at least about 90% compared to control cells without the systems of this disclosure (e.g., control diseased cells). Enhanced survival can be measured, for example, by the duration of survival in vivo. Enhanced survival can also be measured at least 1 day or at least about 1 day, at least 2 days or at least about 2 days, at least 3 days or at least about 3 days, at least 4 days or at least about 4 days, at least 5 days or at least about 5 days, at least 6 days or at least about 6 days, at least 7 days or at least about 7 days, at least 8 days or at least about 8 days, at least 9 days or at least about 9 days, at least 10 days or at least about 10 days, at least 11 days or at least about 11 days, at least 12 days or at least about 12 days, at least 13 days or at least about 13 days, at least 14 days or at least about 14 days, at least 3 weeks or at least about 3 weeks, at least 4 weeks or at least about 4 weeks, at least 2 months or at least about 2 months, at least 6 months or at least about 6 months, or at least 1 year or at least about 1 year.Cells can be brought into contact with the system of this disclosure before (for example, before, during, or after) administration to a subject requiring the system of this disclosure (this can, for example, modulate a target gene). Alternatively, the system of the present invention can be administered to a subject requiring the system to bring it into contact with in vivo cells to modulate a target gene or other gene of interest (for example, a downstream gene of the target gene or a cell type-specific gene) in the in vivo cells.

[0130] The systems provided herein or their use in cells (e.g., diseased cells such as FSHD muscle cells) may have only minimal effect (e.g., substantially no effect) on one or more signs of health in a subject. For example, after administering the systems disclosed herein to a subject, one or more signs of health can be measured to assess whether the target gene has been modified in the subject's cells. Examples of health conditions include, but are not limited to, food intake (e.g., grams per subject per day), body weight (e.g., gram-weight), body dimensions (e.g., length, height, or circumference in meters), body mass index (e.g., grams per square centimeter), body fat (e.g., milligrams of fat per gram of body weight), organ weight (e.g., grams of liver, kidney, spleen, small intestine, etc., per gram of body weight), clinical chemistry values, tissue function or integrity (e.g., structure confirmed by biopsy and / or histological examination), bodily injury, pain, grooming, etc. Clinical chemistry values ​​can be determined by measuring one or more markers from body fluids (e.g., blood). Examples of clinical chemistry markers include electrolytes (e.g., sodium, potassium, chloride, bicarbonate), kidney markers (e.g., creatinine, blood urea nitrogen), and liver function markers (e.g., albumin, globulin, albumin / globulin ratio, bilirubin, aspartate transaminase (AST), alanine transaminase (ALT), gamma-glutamine transpeptidase (GGT), alkaline phosphatase (ALP)). Cardiac markers (e.g., H-FABP, troponin, myoglobin, CK-MB, B-type natriuretic peptide (BNP)), minerals (e.g., calcium, magnesium, phosphate, potassium), blood disorder markers (e.g., iron, transferrin, TIBC, vitamin B12, vitamin D, folic acid), and others (e.g., glucose, C-reactive protein, glycated hemoglobin, uric acid, arterial blood gas, adrenocorticotropic hormone (ACTH), neuron-specific enolase (NSE), fecal occult blood test (FOBT), etc.) are included, but are not limited to these.For example, a clinical chemistry marker panel can be used, which includes one or more members consisting of sodium, potassium, chloride, bicarbonate, blood urea nitrogen (BUN), creatinine, glucose, calcium, total protein, albumin, alkaline phosphatase (ALP), alanine aminotransferase (ALT), aspartate aminotransferase (AST), and bilirubin. The signs of the health condition in subjects using the systems or methods provided herein are at least 80%, up to 80%, at least about 80%, or at least about 80%, at least 85%, up to 85%, at least about 85%, or at least about 85%, at least 90%, up to 90%, at least about 90%, or at least about 90%, at least 91%, up to 91%, at least about 91%, or at least about 91%, at least 92%, up to 92%, at least about 92%, or at least about 92%, at least 93%, up to 93%, at least about 93%, or at least about 93%, at least 94%, up to 94%, at least about 94%, or at least about 94%, at least 95%, up to 95%, less It may be at least approximately 95%, or at least approximately 95%, at least 96%, up to 96%, at least approximately 96%, or at least approximately 96%, at least 97%, up to 97%, at least approximately 97%, or at least approximately 97%, at least 98%, up to 98%, at least approximately 98%, or at least approximately 98%, at least 99%, up to 99%, at least approximately 99%, or at least approximately 99%, at least 100%, up to 100%, at least approximately 100%, or at least approximately 100%, at least 105%, up to 105%, at least approximately 105%, or at least approximately 105%, at least 110%, up to 110%, at least approximately 110%, or at least approximately 110%, or at least 120%, up to 120%, at least approximately 120%, or at least approximately 120%.The minimum effect on the health condition following the aforementioned treatment may last (and / or be measured) for at least one day or about one day, at least two days or about two days, at least three days or about three days, at least four days or about four days, at least five days or about five days, at least six days or about six days, at least seven days or about seven days, at least eight days or about eight days, at least nine days or about nine days, at least ten days or about ten days, at least eleven days or about eleven days, at least twelve days or about twelve days, at least thirteen days or about thirteen days, at least fourteen days or about fourteen days, at least three weeks or about three weeks, or at least four weeks or about four weeks.

[0131] The systems, compositions, and methods disclosed herein can be used to treat, suppress, or improve diseases in the subject (e.g., muscular dystrophy and facioscapulohumeral muscular dystrophy (FSHD)).

[0132] Heterogeneous polypeptides As disclosed herein, heterologous polypeptides disclosed herein can be configured to specifically bind to a target polynucleotide sequence, either alone or in combination with one or more co-acting substances such as heterologous polynucleotides (e.g., guide nucleic acids), thereby regulating the expression level and / or epigenetic level of a target gene (e.g., a D4Z4 repeat array) in target cells. The target polynucleotide sequence may be present in the target gene (e.g., within the target gene). Alternatively, the target polynucleotide sequence may be adjacent to the target gene. For example, the target polynucleotide sequence may be adjacent to the terminal (e.g., the 5' or 3' terminal) of the target gene. The target polynucleotide sequence is at least 5 nucleic acid bases or at least approximately 5 nucleic acid bases, at least 10 nucleic acid bases or at least approximately 10 nucleic acid bases, at least 20 nucleic acid bases or at least approximately 20 nucleic acid bases, at least 30 nucleic acid bases or at least approximately 30 nucleic acid bases, at least 40 nucleic acid bases or at least approximately 40 nucleic acid bases, at least 50 nucleic acid bases or at least approximately 50 nucleic acid bases, at least 100 nucleic acid bases or at least approximately 100 nucleic acid bases, at least 150 nucleic acid bases or at least approximately 150 nucleic acid bases, at least 200 nucleic acid bases or at least approximately 200 nucleic acid bases, at least 250 nucleic acid bases or at least approximately 250 nucleic acid bases. The bases may be separated by at least 300 nucleic acid bases or at least about 300 nucleic acid bases, at least 400 nucleic acid bases or at least about 400 nucleic acid bases, at least 500 nucleic acid bases or at least about 500 nucleic acid bases, at least 1,000 nucleic acid bases or at least about 1,000 nucleic acid bases, at least 1,500 nucleic acid bases or at least about 1,500 nucleic acid bases, at least 2,000 nucleic acid bases or at least about 2,000 nucleic acid bases, at least 3,000 nucleic acid bases or at least about 3,000 nucleic acid bases, at least 4,000 nucleic acid bases or at least about 4,000 nucleic acid bases, or at least 5,000 nucleic acid bases or at least about 5,000 nucleic acid salts.The target polynucleotide sequence is a maximum of 5,000 nucleic acid bases or approximately 5,000 nucleic acid bases from the end of the target gene, a maximum of 4,000 nucleic acid bases or approximately 4,000 nucleic acid bases, a maximum of 3,000 nucleic acid bases or approximately 3,000 nucleic acid bases, a maximum of 2,000 nucleic acid bases or approximately 2,000 nucleic acid bases, a maximum of 1,500 nucleic acid bases or approximately 1,500 nucleic acid bases, a maximum of 1,000 nucleic acid bases or approximately 1,000 nucleic acid bases, a maximum of 500 nucleic acid bases or approximately 500 nucleic acid bases, a maximum of 400 nucleic acid bases or approximately 400 nucleic acid bases, It is acceptable for them to be separated by a maximum of 300 nucleic acid bases or approximately 300 nucleic acid bases, a maximum of 200 nucleic acid bases or approximately 200 nucleic acid bases, a maximum of 150 nucleic acid bases or approximately 150 nucleic acid bases, a maximum of 100 nucleic acid bases or approximately 100 nucleic acid bases, a maximum of 50 nucleic acid bases or approximately 50 nucleic acid bases, a maximum of 40 nucleic acid bases or approximately 40 nucleic acid bases, a maximum of 30 nucleic acid bases or approximately 30 nucleic acid bases, a maximum of 20 nucleic acid bases or approximately 20 nucleic acid bases, a maximum of 10 nucleic acid bases or approximately 10 nucleic acid bases, or a maximum of 5 nucleic acid bases or approximately 5 nucleic acid bases.

[0133] While we do not wish to be bound by any particular theory, if the target polynucleotide sequence is not present within the target gene, the target polynucleotide sequence may interact with at least a portion of the target gene (e.g., the promoter sequence of the target gene) (e.g., via direct or indirect binding), and therefore, by binding to or targeting the target polynucleotide sequence at least via heterologous polypeptides (e.g., via complexes comprising heterologous polypeptides and heterologous polynucleotides as disclosed herein), at least a portion of the target gene (e.g., the promoter sequence) can be targeted to modulate the expression level and / or epigenetic level of the target gene within the cell.

[0134] In some cases, the target polynucleotide sequence may comprise multiple target polynucleotide sequences. These multiple target polynucleotide sequences may be located within the target gene. In another embodiment, as disclosed herein, the multiple target polynucleotide sequences may be located outside the target gene but adjacent to it. In yet another embodiment, the multiple target polynucleotide sequences may comprise at least one target polynucleotide sequence located within the target gene (e.g., within the D4Z4 repeat domain) and at least one other target polynucleotide sequence located outside the target gene but adjacent to it. In such cases, targeting both the at least one target polynucleotide sequence and the at least one additional target polynucleotide sequence may yield a more potent effect (e.g., at least 0.1x, at least 0.5x, at least 1x, at least 2x, at least 3x, at least 4x, at least 5x, at least 10x, at least 15x, at least 20x, or more) than targeting either one of these target polynucleotide sequences.

[0135] The heterologous polypeptides disclosed herein may include one or more heterologous gene effectors (e.g., gene effectors that are heterologous to cells containing the gene effectors and / or other components of the complex disclosed herein). The heterologous gene effectors may include domains or candidate domains that can regulate the expression of a target gene (e.g., an endogenous target gene), for example, by activating, repressing, upregulating, downregulating, or stabilizing the expression level or activity level of the target gene. The heterologous gene effectors may be heterologous to other components present in the complex, for example, guide regions (e.g., nucleases and / or guide nucleic acids disclosed herein). In some cases, the heterologous gene effectors may be heterologous to the host cell into which they are introduced.

[0136] The heterogeneous effector may be, or may include, a sequence derived from a suitable source, such as an amino acid sequence derived from a human protein, viral protein, or other protein disclosed herein. The heterogeneous effector may be, or may include, a sequence derived from a protein mainly localized in the nucleus, such as a member of the human nuclear proteome. The heterogeneous effector may be, or may include, one or more native amino acid residues. The heterogeneous effector may be, or may include, one or more synthetic amino acid residues.

[0137] Heterogenetic effectors may be sequences derived from mammalian proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from human proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from viral proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from non-human primate proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from non-human mammalian proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from non-rodent mammalian proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from plant proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from pig proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from rabbit proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from dog proteins, and may contain such sequences. Heterogenetic effectors may be sequences derived from bird proteins, and may contain such sequences. The heterogeneous effector may be a sequence derived from reptile proteins, or may contain such sequences. The heterogeneous effector may be a sequence derived from bacterial proteins, or may contain such sequences. The heterogeneous effector may be a sequence derived from archaeal proteins, or may contain such sequences.

[0138] For example, the amino acid sequences of the heterologous gene effectors disclosed herein do not have to be derived from bacterial proteins, nor do they need to be derived from bacterial proteins (for example, they may be derived from archaeal proteins). Although we do not wish to be bound by any theory, subjects in need of treatment may be treated with a composition comprising a heterologous gene effector derived from a non-bacterial protein, such that (i) bacterial stimulation is not induced in the subject, and / or (ii) a bacterial immune response is not induced in the subject.

[0139] The heterogeneous actuator portion may contain a nuclease (e.g., an endonuclease). For example, the nuclease may be a CRISPR / Cas protein. The nuclease may have a length less than the threshold length. This threshold length may be up to 1,000 amino acids or approximately 1,000 amino acids, up to 950 amino acids or approximately 950 amino acids, up to 900 amino acids or approximately 900 amino acids, up to 850 amino acids or approximately 850 amino acids, up to 800 amino acids or approximately 800 amino acids, up to 750 amino acids or approximately 750 amino acids, up to 700 amino acids or approximately 700 amino acids, or up to 650 amino acids. The length may be up to approximately 650 amino acids, up to 600 amino acids or up to approximately 600 amino acids, up to 550 amino acids or up to approximately 550 amino acids, up to 500 amino acids or up to approximately 500 amino acids, up to 450 amino acids or up to approximately 450 amino acids, up to 400 amino acids or up to approximately 400 amino acids, up to 350 amino acids or up to approximately 350 amino acids, or up to 300 amino acids or up to approximately 300 amino acids.The threshold length may be at least 300 amino acid lengths or at least about 300 amino acid lengths, at least 350 amino acid lengths or at least about 350 amino acid lengths, at least 400 amino acid lengths or at least about 400 amino acid lengths, at least 450 amino acid lengths or at least about 450 amino acid lengths, at least 500 amino acid lengths or at least about 500 amino acid lengths, at least 550 amino acid lengths or at least about 550 amino acid lengths, at least 600 amino acid lengths or at least about 600 amino acid lengths, at least 650 amino acid lengths or at least about 650 amino acid lengths, at least 700 amino acid lengths or at least about 700 amino acid lengths, at least 750 amino acid lengths or at least about 750 amino acid lengths, at least 800 amino acid lengths or at least about 800 amino acid lengths, at least 850 amino acid lengths or at least about 850 amino acid lengths, at least 900 amino acid lengths or at least about 900 amino acid lengths, at least 950 amino acid lengths or at least about 950 amino acid lengths, or at least 1,000 amino acid lengths or at least about 1,000 amino acid lengths.

[0140] While we do not wish to be bound by any particular theory, using a nuclease smaller than the threshold size may offer one or more advantages over using a control nuclease larger than the threshold size. When using a size-limited delivery medium (e.g., when there are limitations on the physical size of the nuclease that can be encapsulated, or when there are limitations on the size of the expression cassette, such as a viral genome), it may be possible to ensure sufficient space for one or more co-acting agents, such as one or more gene regulators (e.g., transcription regulators), and / or one or more heterologous polynucleotides (e.g., one or more guide nucleic acid molecules) (or ensure sufficient space within the expression cassette). In addition to this, or in another aspect, using a nuclease smaller than the threshold size may induce a more potent effect on the regulation of the expression level and / or epigenetic level of a target gene compared to the effect of a control nuclease larger than the threshold size on the regulation of the expression level and / or epigenetic level of the target gene.

[0141] In some examples, the degree of regulation (e.g., increase or decrease) of the expression level and / or epigenetic level of a target gene by the nucleases disclosed herein (e.g., nucleases having a size less than or equal to the threshold size) is at least 0.1 times, up to 0.1 times, at least about 0.1 times or up to about 0.1 times, at least 0.5 times, up to 0.5 times, at least about 0.5 times or up to about 0.5 times, at least 1 time, up to 1x, at least approximately 1x or up to approximately 1x, at least 2x, up to 2x, at least approximately 2x or up to approximately 2x, at least 3x, up to 3x, at least approximately 3x or up to approximately 3x, at least 4x, up to 4x, at least approximately 4x or up to approximately 4x, at least 5x, up to 5x, at least approximately 5x or up to approximately 5x, at least 6x, up to 6x, at least approximately 6x or up to approximately 6x, at least 7x, up to 7x, at least approximately 7x or up to approximately 7x, at least 8x, up to 8x, at least approximately 8x or up to Approximately 8 times, at least 9 times, up to 9 times, at least approximately 9 times or up to approximately 9 times, at least 10 times, up to 10 times, at least approximately 10 times or up to approximately 10 times, at least 15 times, up to 15 times, at least approximately 15 times or up to approximately 15 times, at least 20 times, up to 20 times, at least approximately 20 times or up to approximately 20 times, at least 25 times, up to 25 times, at least approximately 25 times or up to approximately 25 times, at least 30 times, up to 30 times, at least approximately 30 times or up to approximately 30 times, at least 40 times, up to 40 times, at least approximately 40 times Alternatively, it may be up to approximately 40 times, at least 50 times, up to 50 times, at least approximately 50 times, or up to approximately 50 times, at least 60 times, up to 60 times, at least approximately 60 times, or up to approximately 60 times, at least 70 times, up to 70 times, at least approximately 70 times, or up to approximately 70 times, at least 80 times, up to 80 times, at least approximately 80 times, or up to approximately 80 times, at least 90 times, up to 90 times, at least approximately 90 times, or up to approximately 90 times, or at least 100 times, up to 100 times, at least approximately 100 times, or up to approximately 100 times.

[0142] In some examples, the degree of regulation (e.g., increase or decrease) of the expression level and / or epigenetic level of a target gene by the nucleases disclosed herein (e.g., nucleases having a size less than or equal to the threshold size) is at least 0.1 times, up to 0.1 times, at least about 0.1 times or up to about 0.1 times, at least 0.5 times, up to 0.5 times, at least about 0.5 times or up to about 0.5 times, at least 1 time, up to 1x, at least approximately 1x or up to approximately 1x, at least 2x, up to 2x, at least approximately 2x or up to approximately 2x, at least 3x, up to 3x, at least approximately 3x or up to approximately 3x, at least 4x, up to 4x, at least approximately 4x or up to approximately 4x, at least 5x, up to 5x, at least approximately 5x or up to approximately 5x, at least 6x, up to 6x, at least approximately 6x or up to approximately 6x, at least 7x, up to 7x, at least approximately 7x or up to approximately 7x, at least 8x, up to 8x, at least approximately 8x or up to approximately 8 times, at least 9 times, up to 9 times, at least about 9 times or up to about 9 times, at least 10 times, up to 10 times, at least about 10 times or up to about 10 times, at least 15 times, up to 15 times, at least about 15 times or up to about 15 times, at least 20 times, up to 20 times, at least about 20 times or up to about 20 times, at least 25 times, up to 25 times, at least about 25 times or up to about 25 times, at least 30 times, up to 30 times, at least about 30 times or up to about 30 times, at least 40 times, up to 40 times, at least about 40 times or It can last up to approximately 40 times, at least 50 times, up to 50 times, at least approximately 50 times, or up to approximately 50 times, at least 60 times, up to 60 times, at least approximately 60 times, or up to approximately 60 times, at least 70 times, up to 70 times, at least approximately 70 times, or up to approximately 70 times, at least 80 times, up to 80 times, at least approximately 80 times, or up to approximately 80 times, at least 90 times, up to 90 times, at least approximately 90 times, or up to approximately 90 times, or at least 100 times, up to 100 times, at least approximately 100 times, or up to approximately 100 times longer.

[0143] In some cases, the degree of regulation (e.g., increase or decrease) of the expression level and / or epigenetic level of a target gene by the nucleases disclosed herein (e.g., nucleases having a size less than or equal to the threshold size) is greater than the degree of regulation by the control nuclease (e.g., a control nuclease having a size greater than the threshold size) by at least one cell division, up to one cell division, at least about one cell division or up to about one cell division, at least two cell divisions, up to two cell divisions, at least about two cell divisions. Or up to approximately 2 cell divisions, at least 3 cell divisions, up to 3 cell divisions, at least approximately 3 cell divisions, or up to approximately 3 cell divisions, at least 4 cell divisions, up to 4 cell divisions, at least approximately 4 cell divisions, or up to approximately 4 cell divisions, at least 5 cell divisions, up to 5 cell divisions, at least approximately 5 cell divisions, or up to approximately 5 cell divisions, at least 6 cell divisions, up to 6 cell divisions, at least approximately 6 cell divisions, or up to approximately 6 cell divisions, at least 7 cell divisions, up to 7 cell divisions Cell division, at least approximately 7 cell divisions or up to approximately 7 cell divisions, at least 8 cell divisions, up to 8 cell divisions, at least approximately 8 cell divisions or up to approximately 8 cell divisions, at least 9 cell divisions, up to 9 cell divisions, at least approximately 9 cell divisions or up to approximately 9 cell divisions, at least 10 cell divisions, up to 10 cell divisions, at least approximately 10 cell divisions or up to approximately 10 cell divisions, at least 11 cell divisions, up to 11 cell divisions, at least approximately 11 cell divisions or up to approximately 11 16 cell divisions, at least 12 cell divisions, up to 12 cell divisions, at least approximately 12 cell divisions or up to approximately 12 cell divisions, at least 13 cell divisions, up to 13 cell divisions, at least approximately 13 cell divisions or up to approximately 13 cell divisions, at least 14 cell divisions, up to 14 cell divisions, at least approximately 14 cell divisions or up to approximately 14 cell divisions, at least 15 cell divisions, up to 15 cell divisions, at least approximately 15 cell divisions or up to approximately 15 cell divisions, at least 16 cell divisions,Up to 16 cell divisions, at least approximately 16 cell divisions or up to approximately 16 cell divisions, at least 17 cell divisions, up to 17 cell divisions, at least approximately 17 cell divisions or up to approximately 17 cell divisions, at least 18 cell divisions, up to 18 cell divisions, at least approximately 18 cell divisions or up to approximately 18 cell divisions, at least 19 cell divisions, up to 19 cell divisions, at least approximately 19 cell divisions or up to approximately 19 cell divisions, at least 20 cell divisions, up to 20 cell divisions, at least approximately 20 cell divisions or up to approximately 20 cells Cell division can be sustained through at least 25 cell divisions, up to 25 cell divisions, at least approximately 25 cell divisions or up to approximately 25 cell divisions, at least 30 cell divisions, up to 30 cell divisions, at least approximately 30 cell divisions or up to approximately 30 cell divisions, at least 40 cell divisions, up to 40 cell divisions, at least approximately 40 cell divisions or up to approximately 40 cell divisions, at least 50 cell divisions, up to 50 cell divisions, at least approximately 50 cell divisions or up to approximately 50 cell divisions, or at least approximately 100 cell divisions. As disclosed herein, cell division may also be characterized by one parent cell dividing into two daughter cells having substantially the same genetic material as the parent cell.

[0144] The heterogeneous effector may be a sequence derived from a chromatin regulator (CR), or may contain one. The chromatin regulator contains a functional domain derived from various classes of histone modifying enzymes or DNA modifying enzymes (e.g., DNMT, HAT, HMT, etc.). In some embodiments, the heterogeneous effector is a DNMT containing DNMT-A or DNMT-L. In some embodiments, "DNMT-L" is DNMT3L or contains DNMT3L. In some embodiments, "DNMT-A" is DNMT3A or contains DNMT3A. In some embodiments, the heterogeneous effector is a DNMT containing DNMT3A or DNMT3L.

[0145] A heterogeneous gene effector may contain two or more domains derived from chromatin regulators, which may be arranged in tandem, for example, at the C-terminus or N-terminus of a polypeptide sequence, or within the polypeptide sequence, or they may be arranged separately.

[0146] In some embodiments, heterogeneous effectors promote heterochromatin formation. Examples of proteins that can promote heterochromatin formation include, but are not limited to, HP1α, HP1β, KAP1, KRAB, SUV39H1, and G9a.

[0147] In some embodiments, the heterogeneous effector modulates histones via methylation. In some embodiments, the heterogeneous effector modulates histones via acetylation. In some embodiments, the heterogeneous effector modulates histones via phosphorylation. In some embodiments, the heterogeneous effector modulates histones via ADP-ribosylation. In some embodiments, the heterogeneous effector modulates histones via glycosylation. In some embodiments, the heterogeneous effector modulates histones via SUMOylation. In some embodiments, the heterogeneous effector modulates histones via ubiquitination. In some embodiments, the heterogeneous effector modulates histones, for example, by remodeling of histone structure via an ATP hydrolysis-dependent process.

[0148] In some embodiments, heterogeneous effectors facilitate the spatial arrangement of proteins on or near target polynucleotides, such as transcriptional repressors, transcription factors, and histones. In some embodiments, heterogeneous effectors are useful, for example, for manipulating the spatial and temporal organization of genomic DNA and RNA components in the nucleus and / or cytoplasm for the control of diverse cellular functions.

[0149] In some embodiments, the heterogeneous effector is derived from a family of histone acetyltransferases. Examples of histone acetyltransferases include, but are not limited to, the GNAT subfamily, MYST subfamily, p300 / CBP subfamily, HAT1 subfamily, GCN5, PCAF, Tip60, MOZ, MORF, MOF, HBO1, p300, CBP, HAT1, ATF-2, SRC1, and TAFII250.

[0150] In some embodiments, heterogeneous effectors are derived from histone lysine methyltransferases. Examples of histone lysine methyltransferases include the EZH subfamily, non-SET subfamily, other SET subfamily, PRDM subfamily, SET1 subfamily, SET2 subfamily, SUV39 subfamily, SYMD subfamily, ASH1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, NSD2, NSD3, PRDM1, PRDM10, PRDM11, PRDM12, PRDM13, and P Examples include, but are not limited to, RDM14, PRDM15, PRDM16, PRDM2, PRDM4, PRDM5, PRDM6, PRDM7, PRDM8, PRDM9, SET1, SET1L, SET2L, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETDB1, SETDB2, SETMAR, SUV39H1, SUV39H2, SUV420H1, SUV420H2, SYMD1, SYMD2, SYMD3, SYMD4, and SYMD5.

[0151] In some embodiments, the heterogeneous effector is derived from a component of the chromatin remodeling complex. In some embodiments, the heterogeneous effector is a component of the BAF complex, for example, derived from actin, ARIDA / B, BAF155, BAF170, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRG1 / BRM, INI1, or SS18.

[0152] In some embodiments, the heterogeneous effector is derived from a component of the PBAF complex, such as actin, ARID2, BAF155, BAF170, BAF180, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRD7, BRG1, or INI1.

[0153] In some embodiments, the heterogeneous gene effector is derived from a component of the ISWI family of chromatin remodeling complexes, such as the ACF subfamily, RSF subfamily, CERF subfamily, CHRAC subfamily, NURF subfamily, NoRC subfamily, WICH subfamily, b-WICH subfamily, ACF1, ATPase, BPTF, CECR2, CHRAC15, CHRAC17, CSB, DEK, MYBBP1A, NM1, RBAP46 / 48, RHII / Gua, RSF1, SAP155, SNF2H, SNF2H / L, SNF2L, TIP5, or WSTF.

[0154] In some embodiments, the heterogeneous effector is derived from components of the CHD family complex, such as the NuRD complex, NuRD-like complex, or CHD complex. In some embodiments, the heterogeneous effector is derived from CHD1 / 2 / 6 / 7 / 8 / 9, CHD3 / 4, CHD5, GATAD2 A / B, GATAD2 B, HDAC1, HDAC2, HDAC2, MBD2 / 3, MTA1 / 2 / 3, MTA3, RBAP46, or RBAP46 / 48.

[0155] In some embodiments, the heterogeneous effector is derived from components of the INO80 family complex, such as the INO80 complex, Tip60 / p400 complex, SRCAP complex, AMIDA, ARP6, BAF53, BAF53, BAF53A, BRD8, DMAP1, DMAP1, EPC1 / 2, FLJ11730, GAS41, GAS41, IES2, IES6, ING3, INO80, INO80E, MCRS1, MRG15, MRGBP, MRGX, NFRKB, p400, RUVBL1 / 2, RUVBL1 / 2, RUVBL1 / 2, SRCAP, Tip60, TRRAP, UCH37, YL-1, YL-1, YY1, or ZnF-HIT1.

[0156] Heterogenetic effectors may be sequences derived from transcription factors (TRs), or may contain such sequences. TR gene effectors contain transcriptional regulatory domains derived from various transcription factor families (e.g., KRAB, p65, MED, GTF, etc.).

[0157] Heterogenetic effectors may contain transcriptional activator domains. Heterogenetic effectors may also contain tandem transcriptional activation domains, which may be located, for example, at the C-terminus or N-terminus of a polypeptide sequence, or within the polypeptide sequence.

[0158] Examples of transcriptional activation domains include, but are not limited to, GAL4, the herpes simplex activation domain VP16, VP64 (a tetramer of the herpes simplex virus activation domain VP16), the p65 subunit of NF-KB, and the R trans-activator (Rta) of Epstein-Barr virus. In some embodiments, such transcriptional activation domains are used as controls in the methods of the present disclosure. In some embodiments, such transcriptional activation domains are used as one heterogeneous effector in a complex comprising at least one additional heterogeneous effector (e.g., a different effector).

[0159] Heterogenetic effectors may contain transcriptional repressor domains. Heterogenetic effectors may contain two or more transcriptional repressor domains, which may be arranged in tandem, for example, at the C-terminus or N-terminus of a polypeptide sequence, or within the polypeptide sequence, or separately.

[0160] Examples of transcriptional repressor domains include, but are not limited to, the KRAB (Kruppel-associated box) domain of Koxl, the Mad-mSIN3 interaction domain (SID), and the repressor domain (ERD) of ERF. In some embodiments, such transcriptional repressor domains are used as controls in the methods of the present disclosure. In some embodiments, such transcriptional repressor domains are used as one heterogeneous effector in a complex comprising at least one additional heterogeneous effector (e.g., a different effector).

[0161] In some embodiments, the heterologous gene effector is derived from a gene product that is a transcription factor.

[0162] In some embodiments, the heterogeneous effector is derived from gene products that are transcription factors of hematopoietic stem cells. Examples of hematopoietic stem cell transcription factors include AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1α / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, and NF. Examples include, but are not limited to, ATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activators, STAT inhibitors, STAT3, STAT4, STAT5a, STAT6, and TSC22.

[0163] In some embodiments, the heterogeneous effector is derived from a gene product that is a transcription factor of a mesenchymal stem cell. Examples of mesenchymal stem cell transcription factors include, but are not limited to, DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, myocaldin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT activators, STAT inhibitors, STAT1, STAT3, TBX18, Twist-1, and Twist-2.

[0164] In some embodiments, the heterogeneous effector is derived from gene products that are transcription factors of embryonic stem cells. Examples of embryonic stem cell transcription factors include brachiuri, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3α / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, and NFκB / IκB activation. Examples of substances include, but are not limited to, NFκB / IκB inhibitors, NFκB1, NFκB2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activators, STAT inhibitors, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, and ZNF281.

[0165] In some embodiments, the heterogeneous effector is derived from a gene product that is a transcription factor of an induced pluripotent stem cell (iPSC). Examples of iPSC transcription factors include, but are not limited to, KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, and TBX18.

[0166] In some embodiments, the heterologous gene effector is derived from a gene product that is a transcription factor of an epithelial stem cell. Examples of epithelial stem cell transcription factors include ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4α / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT activators, STAT inhibitors, STAT3, SUZ12, TCF-3 / E2A, and TCF7 / TCF One example is, but it is not limited to, this.

[0167] In some embodiments, the heterogeneous effector is derived from gene products that are transcription factors of cancer stem cells. Examples of cancer stem cell transcription factors include androgen R / NR3C4, AP-2γ, β-catenin, β-catenin inhibitors, brachiuri, CREB, ERα / NR3A1, ERβ / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI-2, GLI-3, HIF-1α / HIF1A, HIF-2α / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, Examples include, but are not limited to, MCM7, MITF, c-Myc, Nanog, NFκB / IκB activators, NFκB / IκB inhibitors, NFκB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activators, STAT inhibitors, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, and ZEB1.

[0168] In some embodiments, the heterogeneous effector is derived from a gene product that is a cancer-related transcription factor. Examples of cancer-related transcription factors include ASCL1 / Mash1, ASCL2 / Mash2, ATF1, ATF2, ATF4, BLIMP1 / PRDM1, CDX2, CDX4, DLX5, DNMT1, E2F-1, EGR1, ELF3, Ets-1, FosB / G0S3, FoxC1, FoxC2, FoxF1, GADD153, GATA-2, HMGA2, HMGB1 / HMG-1, HNF-3α / FoxA1, HNF-6 / ONECUT1, HSF1, ID1, ID2, JunD, KLF10, KLF12, KLF17, and LMO2. Examples include, but are not limited to, MEF2C, MYCL1 / L-Myc, NFkB2, Oct-1, p63 / TP73L, Pax3, PITX2, Prox1, RAP80, Rex-1 / ZFP42, RUNX1 / CBFA2, RUNX3 / CBFA3, SALL4, SCL / Tal1, Sirtuin2 / SIRT2, Smad3, Smad4, Smad5, SOX11, STAT5a / b, STAT5a, STAT5b, TCF7 / TCF1, TORC1, TORC2, TRIM32, TRPS1, and TSC22.

[0169] In some embodiments, heterogeneous effectors are derived from gene products that are transcription factors of immune cells. Examples of immune cell transcription factors include, but are not limited to, AP-1, Bcl6, E2A, EBF, Eomes, FoxP3, GATA3, Id2, Ikaros, IRF, IRF1, IRF2, IRF3, IRF3, IRF7, NFAT, NFkB, Pax5, PLZF, PU.1, ROR-γ-T, STAT, STAT1, STAT2, STAT3, STAT4, STAT5, STAT5A, STAT5B, STAT6, T-bet, TCF7, and ThPOK.

[0170] In some embodiments, the heterogeneous effector is derived from a gene product that is an RNA polymerase-related protein. In some embodiments, the heterogeneous effector is derived from a transcription factor having a basic domain. In some embodiments, the heterogeneous effector is derived from a transcription factor having a zinc-coordinated DNA-binding domain. In some embodiments, the heterogeneous effector is derived from a transcription factor having a helix-turn-helix domain. In some embodiments, the heterogeneous effector is derived from a transcription factor having an α-helix DNA-binding domain. In some embodiments, the heterogeneous effector is derived from a transcription factor having an α-helix exposed by a β-structure. In some embodiments, the heterogeneous effector is derived from a transcription factor having an immunoglobulin fold. In some embodiments, the heterogeneous effector is derived from a transcription factor having a β-hairpin exposed by an α / β scaffold. In some embodiments, the heterogeneous effector is derived from a transcription factor having a β-sheet that binds to DNA. In some embodiments, the heterogeneous effector is derived from a transcription factor having a β-barrel DNA-binding domain.

[0171] In some embodiments, the heterologous gene effector is derived from a gene product that is a nuclear receptor, such as a nuclear hormone receptor. Examples of nuclear hormone receptors include NR0B1, NR0B2, NR1A1, NR1A2, NR1B1, NR1B2, NR1B3, NR1C1, NR1C2, NR1C3, NR1D1, NR1D2, NR1F1, NR1F2, NR1F3, NR1H4, NR1H5, NR1H3, NR1H2, NR1I1, NR1I2, NR1I3, NR2A1, NR2A2, NR2B1, and NR2B2. Examples include, but are not limited to, those coded by NR2B3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3A1, NR3A2, NR3B1, NR3B2, NR3B3, NR3C4, NR3C1, NR3C2, NR3C3, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, or NR6A1.

[0172] In some embodiments, the heterogeneous effector is derived from a gene product involved in nucleosome assembly. In some embodiments, the heterogeneous effector is derived from a gene product involved in DNA metabolism. In some embodiments, the heterogeneous effector is derived from a gene product involved in nucleotide metabolism. In some embodiments, the heterogeneous effector is derived from a gene product involved in ribosome biosynthesis. In some embodiments, the heterogeneous effector is derived from a gene product involved in protein folding. In some embodiments, the heterogeneous effector is derived from a gene product involved in translation. In some embodiments, the heterogeneous effector is derived from a gene product involved in signal transduction. In some embodiments, the heterogeneous effector is derived from a gene product involved in protein degradation. In some embodiments, the heterogeneous effector is derived from a gene product involved in the negative regulation of endopeptidase activity.

[0173] In some embodiments, the terms "heterologous gene effector" and "gene regulator" are used herein in the same sense and may include a polypeptide sequence having at least 50% or at least about 50%, at least 55% or at least about 55%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 91% or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or 100% sequence identity with any of the amino acid sequences of the heterologous gene effectors shown in Table 3.

[0174]

Table 1

[0175] In some embodiments, the terms “heterogeneic effector” and “gene regulator” are used interchangeably herein and may include polypeptide sequences exhibiting at least 50% or at least about 50%, at least 55% or at least about 55%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 91% or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or 100% sequence identity with Sequence Number 727 shown below. RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEPGKESGSVGGGSGGSSEQLAQFRSLDGMAAIPALDPEAEPSMDVILVGSSELSSSVSPG TGRDLIAYEVKANQRNIEDICICCGSLQVHTQHPLFEGGICAPCKDKFLDALFLYDDDGYQSYCSICCSGETLLICGNPDCTRCYCFECVDSLVGPGTSGKVHAMSNWVCYLCLPSSRS GLLQRRRKWRSQLKAFYDRESENPLEMFETVPVWRRQPVRVLSLFEDIKKELTSLGFLESGSDPGQLKHVVDVTDTVRKDVEEWGPFDLVYGATPPLGHTCDRPPSWYLFQFHRLLQYARPKPGSPRPFFWMFVDNLVLNKEDLDVASRFLEMEPVTIPDVHGGSLQNAVRVWSNIPAIRSRHWALVSEEELSLLAQNKQSSKLAAKWPTKLVKNCFLPLREYFKYFST (Sequence ID 727) In some embodiments, the heterologous polypeptides disclosed herein comprise a Cas12f-KRAB-DNMT3L modulator. In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a heterologous gene effector comprising KRAB and DNMT3L. In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a heterologous gene effector comprising SEQ ID NO: 727 or a sequence having 80%, 85%, 90%, 95%, 97%, 98% or 99% identity to SEQ ID NO: 727, about 80%, about 85%, about 90%, about 95%, about 97%, about 98% or about 99% identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% identity, or about 100% identity, or a percentage identity within a range bounded by any two of these numerical values (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.).

[0176] The heterologous polynucleotides disclosed herein may comprise one or more guide moieties (e.g., one or more guide nucleic acid molecules) that direct a heterologous gene effector to a target gene (e.g., an endogenous target gene) or a regulatory sequence of a target gene. The guide moiety can be capable of recognizing and specifically binding to the target gene or its regulatory sequence. The guide moiety can be configured to form a complex with the heterologous polypeptide (e.g., a guide nucleic acid that forms a complex with a nuclease such as a CRISPR / Cas protein), and this complex can be configured to exhibit specific binding to a target polypeptide sequence and modify the expression level and / or epigenetic modification level of the target gene as disclosed herein.

[0177] The guide portion may include a guide nucleic acid. The guide portion may include a nuclease and a guide nucleic acid as disclosed herein. The guide portion may include a nuclease or a portion thereof, such as an endonuclease such as a heterologous endonuclease. The nuclease may be, for example, a DNA nuclease and / or an RNA nuclease, a modified nuclease lacking nuclease activity, or a modified nuclease with reduced nuclease activity compared to a wild-type nuclease, a derivative thereof, a variant thereof, or a fragment thereof. In some embodiments, the guide portion has minimal nuclease activity.

[0178] Suitable nucleases, or fragments or derivatives thereof, can be used in the guide portion. Suitable nucleases include, but are not limited to, CRISPR-related (Cas) proteins or Cas nucleases such as type I CRISPR-related (Cas) polypeptide, type II CRISPR-related (Cas) polypeptide, type III CRISPR-related (Cas) polypeptide, type IV CRISPR-related (Cas) polypeptide, type V CRISPR-related (Cas) polypeptide, and type VI CRISPR-related (Cas) polypeptide; zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); meganucleases; RNA-binding proteins (RBPs); CRISPR-related RNA-binding proteins; recombinases; flippases; transposases; Argonaut (Ago) proteins (e.g., prokaryotic Argonaut (pAgo), archaeal Argonaut (aAgo), or eukaryotic Argonaut (eAgo)); derivatives thereof; or variants thereof or functional fragments thereof.

[0179] In some embodiments, the guide portion includes a DNA nuclease, such as a recombinant DNA nuclease lacking nuclease activity (e.g., a programmable DNA nuclease or a targetable DNA nuclease). In some embodiments, the guide portion includes a nuclease-null type DNA-binding protein derived from a DNA nuclease that cannot induce activation or repression of the transcription of a target DNA sequence unless it is complexed with one or more heterogeneous effectors of this disclosure. In some embodiments, the guide portion includes a nuclease-null type DNA-binding protein derived from a DNA nuclease that can induce activation or repression of the transcription of a target DNA sequence (e.g., activation or repression of the transcription of a target DNA sequence can be altered or enhanced by the presence of a heterogeneous effector of this disclosure).

[0180] In some embodiments, the guide portion comprises an RNA nuclease, such as a recombinant RNA nuclease (e.g., a programmable RNA nuclease or a targetable RNA nuclease). In some embodiments, the guide portion comprises a nuclease-null type RNA-binding protein derived from an RNA nuclease that cannot induce activation or repression of the transcription of a target RNA sequence unless it is complexed with one or more heterogeneous effectors of this disclosure. In some embodiments, the guide portion comprises a nuclease-null type RNA-binding protein derived from an RNA nuclease that can induce activation or repression of the transcription of a target RNA sequence (e.g., activation or repression of the transcription of a target RNA sequence can be altered or enhanced by the presence of a heterogeneous effector of this disclosure).

[0181] In some embodiments, the guide portion includes a nucleic acid-guided targeting system. In some embodiments, the guide portion includes a DNA-guided targeting system. In some embodiments, the guide portion includes an RNA-guided targeting system. The guide portion may include, for example, a guide nucleic acid sequence that promotes the specific binding of a CRISPR-Cas system (e.g., its nuclease-deficient form, e.g., dCas9) to a target gene (e.g., an endogenous target gene) or a regulatory sequence of a target gene, and this guide nucleic acid sequence can be utilized. Binding specificity can be determined by using a guide nucleic acid such as a single guide RNA (sgRNA) or a portion thereof. In some embodiments, by using various types of sgRNA, the compositions and methods of the present disclosure can be used for various target genes (e.g., endogenous target genes) or regulatory sequences of target genes (e.g., targeting various target genes or their regulatory sequences).

[0182] Prokaryotic CRISPR-Cas (Clustered regularly interspaced short palindromic repeats-CRISPR associated) systems, such as class II CRISPR-Cas systems like Cas9 and Cpfl, can be adapted in the compositions and methods of this disclosure as tools for regulating gene expression, epigenome editing, and chromatin circularization. Nuclease-inactive Cas (dCas) proteins, when complexed with heterologous gene effectors, can regulate the expression of target genes adjacent to the dCas binding site (e.g., endogenous target genes).

[0183] In some embodiments, the guide portion includes a CRISPR-related (Cas) protein or Cas nuclease that functions in a non-natural CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-related) system. This system can provide acquired immunity against foreign DNA in bacteria.

[0184] In a variety of organisms, including various mammals, animals, plants, microorganisms, and yeasts, the CRISPR / Cas system (e.g., modified CRISPR / Cas systems and / or unmodified CRISPR / Cas systems) can be used as a genome engineering tool or modified to induce specific binding of recombinant proteins to target loci, as disclosed herein. The CRISPR / Cas system may include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for the purpose of targeted regulation of gene expression and / or gene activity or nucleic acid binding. RNA-induced Cas proteins (e.g., Cas nucleases, e.g., Cas9 nucleases) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner. If the Cas protein has nuclease activity, it can cleave DNA.

[0185] In some cases, Cas proteins can be mutated and / or modified to obtain nuclease-deficient proteins or proteins with reduced nuclease activity compared to wild-type Cas proteins. Nuclease-deficient proteins may retain the ability to bind to DNA, but may lack or have reduced nucleic acid cleavage activity.

[0186] In some embodiments, the guide portion includes a Cas protein that forms a complex with a guide nucleic acid, such as a guide RNA or a portion thereof. In some embodiments, the guide portion includes a Cas protein that forms a complex with a single guide nucleic acid, such as a single guide RNA (sgRNA). In some embodiments, the guide portion includes an RNA-binding protein (RBP) that can form a complex with a guide RNA (e.g., sgRNA) and may form a complex with a guide nucleic acid. In some embodiments, the guide portion includes a nuclease-null type DNA-binding protein derived from a DNA nuclease that can induce activation or repression of transcription of a target DNA sequence. In some embodiments, the guide portion includes a nuclease-null type RNA-binding protein derived from RNA.

[0187] In some embodiments, the guide nucleic acids used in the compositions and methods of the present disclosure may have, for example, a nucleotide length of at least 1 nucleotide, at least 2 nucleotides, at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, or more.

[0188] In some embodiments, the guide nucleic acids used in the compositions and methods of the present disclosure are at most 10 nucleotides long, at most 9 nucleotides long, at most 8 nucleotides long, at most 7 nucleotides long, at most 6 nucleotides long, at most 5 nucleotides long, at most 4 nucleotides long, at most 3 nucleotides long, at most 2 nucleotides long, or at most 1 nucleotide long.

[0189] The guide nucleic acid may be guide RNA or a part thereof.

[0190] Any suitable CRISPR / Cas system can be used. The term "CRISPR / Cas system" may refer to systems of various names. A CRISPR / Cas system may be a type I, type II, type III, type IV, type V, or type VI CRISPR / Cas system, or any other suitable CRISPR / Cas system. The CRISPR / Cas systems used herein may be class 1 CRISPR / Cas systems, class 2 CRISPR / Cas systems, or any other appropriately classified CRISPR / Cas system. The distinction between class 1 and class 2 may be based on the gene encoding the effector module. Class 1 CRISPR / Cas systems generally have a multi-subunit crRNA-effector complex, while class 2 CRISPR / Cas systems generally have a single protein such as Cas9, Cpfl, C2c1, C2c2, C2c3, or a crRNA-effector complex. Class 1 CRISPR / Cas systems can be regulated using complexes of multiple Cas proteins. A Class 1 CRISPR / Cas system may include, for example, type I (e.g., type I, type IA, type IB, type IC, type ID, type IE, type IF, or type IU), type III (e.g., type III, type IIIA, type IIIB, type IIIC, or type IIID), and type IV (e.g., type IV, type IVA, or type IVB) CRISPR / Cas. A Class 2 CRISPR / Cas system can be regulated using a single large Cas protein. A Class 2 CRISPR / Cas system may include, for example, type II (e.g., type II, type IIA, or type IIB) and type V CRISPR / Cas. CRISPR systems may be complementary to each other and / or lend out functional units in the trans position to facilitate targeting of CRISPR loci.

[0191] If the guide portion includes a Cas protein or a derivative thereof, the Cas protein or derivative may be a class 1 Cas protein or a class 2 Cas protein. The Cas protein may be a type I Cas protein, a type II Cas protein, a type III Cas protein, a type IV Cas protein, a type V Cas protein, or a type VI Cas protein. The Cas protein may contain one or more domains. Examples of these domains include, but are not limited to, a guide nucleic acid recognition domain and / or a guide nucleic acid binding domain, a nuclease domain (e.g., a DNase domain, an RNase domain, RuvC, or HNH), a DNA binding domain, an RNA binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. The guide nucleic acid recognition domain and / or a guide nucleic acid binding domain may interact with the guide nucleic acid. The nuclease domain may contain catalytic activity for cleaving nucleic acids. The nuclease domain may lack catalytic activity to prevent nucleic acid cleavage. The Cas protein may be a chimeric Cas protein or a fragment thereof fused to another protein or polypeptide. A Cas protein may be a chimera of various Cas proteins, for example, containing domains derived from various Cas proteins.

[0192] Examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e(CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b , Cas8c, Cas9(Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel(CasA), Cse2(CasB), Cse3(CasE), Cse4(CasC), C Examples include, but are not limited to, scl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4, Cul966, Cas13a, Cas13b, Cas13c, Cas13d, Cas13X, and Cas13Y, as well as their homologs or variants thereof.

[0193] In some cases, the Cas proteins disclosed herein do not necessarily have to be Cas9 or Cas12a. The Cas proteins disclosed herein may be smaller in size compared to Cas9 or Cas12a. The Cas proteins disclosed herein may be derived from Un1Cas12f1. For example, the Cas protein disclosed herein may contain an amino acid sequence having at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or substantially about 100% identity with the polypeptide sequence of SEQ ID NO: 43. In another example, the Cas protein disclosed herein may contain an amino acid sequence having at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or substantially about 100% identity with the polypeptide sequence of SEQ ID NO: 44.As disclosed herein, Sequence ID No. 728 codes for, but is not limited to, an example of a Cas12f variant suitable for use in the systems, compositions, and methods of this disclosure. In some embodiments, the Cas12f variants disclosed herein may include an amino acid sequence having at least 50% or about 50%, at least 60% or about 60%, at least 70% or about 70%, at least 75% or about 75%, at least 80% or about 80%, at least 85% or about 85%, at least 90% or about 90%, at least 95% or about 95%, at least 96% or about 96%, at least 97% or about 97%, at least 98% or about 98%, at least 99% or at least about 99%, or substantially about 100% identity with the polypeptide sequence of Sequence ID No. 728. JPEG2026510894000004.jpg216153 In some embodiments, the heterologous polypeptides disclosed herein include a Cas12f-KRAB-DNMT3L modulator. In some embodiments, the Cas12f-KRAB-DNMT3L modulator includes a nuclease comprising Cas12f or a variant thereof. In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a nuclease having the amino acid sequence of SEQ ID NO: 44, or a nuclease having a sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with SEQ ID NO: 44, about 80%, about 85%, about 90%, about 95%, about 97%, about 98%, or about 99% identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity, or about 100% identity, or a percentage of identity within a range of any two of these numbers as upper and lower limits (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.). In some embodiments, the Cas12f-KRAB-DNMT3L modulator comprises a nuclease having the amino acid sequence of SEQ ID NO: 728, or a nuclease having a sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity with SEQ ID NO: 728, about 80%, about 85%, about 90%, about 95%, about 97%, about 98%, or about 99% identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% identity, or about 100% identity, or a percentage of identity within a range of any two of these numbers as upper and lower limits (e.g., 80-100%, 85-97%, 90-95%, 85-95%, etc.).

[0194] In some cases, the Cas protein provided herein may not be the Cas14 protein. In some cases, the dead Cas protein (dCas) provided herein may not be the dead Cas protein.

[0195] The Cas protein or its fragments or derivatives may be derived from any suitable organism. Examples of such organisms include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus species, Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromogenes, Streptomyces viridochromogenes, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas species, and Crocosphaera. watsonii, Cyanothece genus, Microcystis erginosa, Pseudomonas aeruginosa, Synechococcus genus, Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegordia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter genus, Nitrosococcus halophilus, Nitrosococcus watsoni, PseudoalteromonasExamples include, but are not limited to, haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc, Arthrospira maxima, Arthrospira pratensis, Arthrospira, Ringbia, Microcoleus chthonoplastes, Oschilatoria, Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella nobicida. In some embodiments, the organism is Streptococcus pyogenes. In some embodiments, the organism is Staphylococcus aureus. In some embodiments, the organism is Streptococcus thermophilus.

[0196] Cas proteins may originate from various bacterial species, including atypical Veillonella, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria inocure, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, and Oenococcus. kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegordia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma obipneumoniae, Mycoplasma canis, Mycoplasma sinobie, Eubacterium rectore, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Ackermansia muciniphylla, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractor salsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp.Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Satellitella parvula, Proteobacteria, Legionella pneumophila, Parvasterella excrementihominis, Wolinella succinogenes, and Francisella novicida, but are not limited thereto.

[0197] The Cas protein used herein may be wild-type Cas protein or modified Cas protein. The Cas protein may be an active variant, inactive variant, or fragment of wild-type Cas protein or modified Cas protein. Compared to wild-type Cas protein, the Cas protein may contain amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof. The Cas protein may be a polypeptide having at least 5% or about 5%, at least 10% or about 10%, at least 20% or about 20%, at least 30% or about 30%, at least 40% or about 40%, at least 50% or about 50%, at least 60% or about 60%, at least 70% or about 70%, at least 80% or about 80%, at least 90% or about 90%, at least 91% or about 91%, at least 92% or about 92%, at least 93% or about 93%, at least 94% or about 94%, at least 95% or about 95%, at least 96% or about 96%, at least 97% or about 97%, at least 98% or about 98%, at least 99% or about 99%, or 100% sequence identity or similarity with the wild-type Cas protein. The Cas protein may be a polypeptide having up to 5% or approximately 5%, up to 10% or approximately 10%, up to 20% or approximately 20%, up to 30% or approximately 30%, up to 40% or approximately 40%, up to 50% or approximately 50%, up to 60% or approximately 60%, up to 70% or approximately 70%, up to 80% or approximately 80%, up to 90% or approximately 90%, or up to 100% or approximately 100% sequence identity and / or similarity with the wild-type exemplary Cas protein.The variant or fragment may have at least 5% or about 5%, at least 10% or about 10%, at least 20% or about 20%, at least 30% or about 30%, at least 40% or about 40%, at least 50% or about 50%, at least 60% or about 60%, at least 70% or about 70%, at least 80% or about 80%, at least 90% or about 90%, at least 91% or about 91%, at least 92% or about 92%, at least 93% or about 93%, at least 94% or about 94%, at least 95% or about 95%, at least 96% or about 96%, at least 97% or about 97%, at least 98% or about 98%, at least 99% or about 99%, or 100% sequence identity or similarity with the wild-type Cas protein or the modified Cas protein or a portion thereof. The variant or fragment may target a nucleic acid locus that forms a complex with the guide nucleic acid, but may lack nucleic acid cleavage activity.

[0198] The Cas protein may contain one or more nuclease domains, such as a DNase domain. For example, the Cas9 protein may contain a RuvC-like nuclease domain and / or an HNH-like nuclease domain. In the nuclease-active form of Cas9, the RuvC domain and the HNH domain can cleave different strands of double-stranded DNA, thereby cleaving the double strand of DNA. The Cas protein may contain only one nuclease domain (for example, Cpfl contains a RuvC domain but lacks an HNH domain). In some embodiments, no nuclease domains are present. In some embodiments, nuclease domains are present but inactive, have reduced activity, or have minimal activity. In some embodiments, nuclease domains are present and active.

[0199] One or more nuclease domains (e.g., RuvC or HNH) of a Cas protein can be deleted or mutated to render it non-functional or reduce its nuclease activity. For example, in a Cas protein containing at least two nuclease domains (e.g., Cas9), if one of these nuclease domains is deleted or mutated, the resulting Cas protein (known as nickase) can produce single-strand breaks in the CRISPR RNA (crRNA) recognition sequence of double-stranded DNA, but not double-strand breaks. Such a nickase can cleave either the complementary or non-complementary strand, but not both. If all of the nuclease domains of a Cas protein (e.g., both the RuvC and HNH nuclease domains of the Cas9 protein; the RuvC nuclease domain of the Cpfl protein) are deleted or mutated, the resulting Cas protein may have reduced or lost the ability to cleave both strands of double-stranded DNA. An example of a mutation that can convert the Cas9 protein to nickase is the D10A mutation in the RuvC domain of Cas9 from S. pyogenes (a mutation from aspartic acid to alanine at the 10th amino acid position of Cas9). Cas9 can also be converted to nickase by the H939A mutation (a mutation from histidine to alanine at the 839th amino acid position) or the H840A mutation (a mutation from histidine to alanine at the 840th amino acid position) in the HNH domain of Cas9 from S. pyogenes. Examples of mutations that can convert the Cas9 protein to dead Cas9 include the D10A mutation in the RuvC domain of Cas9 from S. pyogenes (a mutation from aspartic acid to alanine at the 10th amino acid position of Cas9), and the H939A mutation (a mutation from histidine to alanine at the 839th amino acid position) or the H840A mutation (a mutation from histidine to alanine at the 840th amino acid position) in the HNH domain.

[0200] A dead Cas protein, which is a nuclease (for example, one derived from a Cas protein such as Un1Cas12f1), may contain one or more mutations compared to the wild type of the protein. These mutations may result in nucleic acid cleavage activity being 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less compared to the nucleic acid cleavage activity of one or more of the nucleic acid cleavage domains of the wild type Cas protein. These mutations may result in one or more of the nucleic acid cleavage domains retaining the ability to cleave the complementary strand of the target nucleic acid while reducing their ability to cleave the non-complementary strand. These mutations may result in one or more of the nucleic acid cleavage domains losing the ability to cleave both the complementary and non-complementary strands of the target nucleic acid. The residues that induce mutations in the nuclease domain may correspond to one or more catalytic residues of the nuclease. For example, by mutating residues of a typical wild-type S. pyogenes Cas9 polypeptide, such as Asp10, His840, Asn854, and Asn856, one or more nucleic acid cleavage domains (e.g., nuclease domains) can be inactivated. The residues that induce mutations in the nuclease domain of the Cas protein may correspond to the Asp10, His840, Asn854, and Asn856 residues of a wild-type Cas9 polypeptide derived from S. pyogenes, as determined by sequence alignment and / or structural alignment.

[0201] For example, the D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 residues (or corresponding mutations in the Cas protein) can be mutated, but are not limited to these. For instance, D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A can be mutated. Mutations other than alanine substitutions may also be preferred.

[0202] The D10A mutation, when combined with one or more of the H840A, N854A, and N856A mutations, can produce a Cas9 protein substantially lacking DNA cleavage activity (e.g., dead Cas9 protein). The H840A mutation, when combined with one or more of the D10A, N854A, and N856A mutations, can produce a site-directed polypeptide substantially lacking DNA cleavage activity. The N854A mutation, when combined with one or more of the H840A, D10A, and N856A mutations, can produce a site-directed polypeptide substantially lacking DNA cleavage activity. The N856A mutation, when combined with one or more of the H840A, N854A, and D10A mutations, can produce a site-directed polypeptide substantially lacking DNA cleavage activity.

[0203] In some embodiments, the Cas protein is a class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified Cas9 protein, or a Cas9 protein-derived Cas protein. For example, a Cas9 protein lacking cleavage activity. In some embodiments, the Cas9 protein is a Cas9 protein derived from S. pyogenes (e.g., SwissProt accession number Q99ZW2). In some embodiments, the Cas9 protein is a Cas9 protein derived from S. aureus (e.g., SwissProt accession number J7RUA5). In some embodiments, the Cas9 protein is a modified Cas9 protein derived from S. pyogenes or S. Aureus. In some embodiments, the Cas9 protein is a Cas9 protein-derived Cas9 protein from S. pyogenes or S. Aureus. For example, a Cas9 protein derived from S. pyogenes or S. Aureus lacking cleavage activity.

[0204] In some embodiments, Cas9 may typically mean a polypeptide having at least 5% or at least about 5%, at least 10% or at least about 10%, at least 20% or at least about 20%, at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 80% or at least about 80%, at least 90% or at least about 90%, or about 100% sequence identity and / or similarity with an exemplary wild-type Cas9 polypeptide (e.g., Cas9 derived from S. pyogenes). In some embodiments, Cas9 may mean a polypeptide having up to about 5%, up to about 10%, up to about 20%, up to about 30%, up to about 40%, up to about 50%, up to about 60%, up to about 70%, up to about 80%, up to about 90%, or about 100% sequence identity and / or similarity with a wild-type Cas9 polypeptide (e.g., Cas9 polypeptide from S. pyogenes). The Cas9 protein may mean a wild-type Cas9 protein or a modified Cas9 protein, which may contain amino acid changes such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0205] The Cas protein has at least 5% or at least about 5%, at least 10% or at least about 10%, at least 20% or at least about 20%, at least 30% or at least about 30%, at least 40% or at least about 40%, at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 80% or at least about 80%, and at least It may also contain amino acid sequences having 90% or at least about 90%, at least 91% or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or 100% sequence identity or sequence similarity.

[0206] Cas proteins or their variants or derivatives can be modified by the compositions and methods of this disclosure, for example, as part of the complexes disclosed herein, to enhance the regulation of gene expression. Cas proteins can be modified to increase or decrease their affinity for binding to nucleic acids, increase or decrease their binding specificity to nucleic acids, increase or decrease their enzymatic activity, and / or increase or decrease their binding to other factors such as heterodimerization domains or oligomerization domains and ligand induction. Cas proteins can also be modified to alter other activities or properties (e.g., stability). For example, one or more nuclease domains of a Cas protein can be modified, deleted or inactivated, or a Cas protein can be cleaved to remove domains that are not essential for the desired function of this protein or complex. Furthermore, Cas proteins can also be modified to modulate (e.g., increase or decrease) the activity of the Cas protein for regulating gene expression by the complexes of this disclosure, including heterogeneous gene effectors.

[0207] For example, Cas proteins can be ligated to heterogeneous gene effectors (e.g., epigenetic modification domains, transcriptional activation domains, and / or transcriptional repression domains) (e.g., by fusion, covalent bond, or non-covalent bond). Cas proteins can be ligated to oligomerization domains or dimerization domains (e.g., heterodimerization domains) disclosed herein (e.g., by fusion, covalent bond, or non-covalent bond). Cas proteins can be ligated to heterogeneous polypeptides that increase or decrease stability (e.g., by fusion, covalent bond, or non-covalent bond). Cas proteins can be ligated to sequences that can promote the degradation of Cas proteins or complexes containing Cas proteins (e.g., by fusion, covalent bond, or non-covalent bond). Examples of sequences that can promote the degradation of Cas proteins or complexes containing Cas proteins include degron, e.g., inducible degron (e.g., inducible auxin).

[0208] The Cas protein can be linked to any number of suitable partners (e.g., by fusion, covalent bonding, or non-covalent bonding), for example, to at least one partner, at least two partners, at least three partners, at least four partners, at least five partners, at least six partners, at least seven partners, or at least eight partners. In some embodiments, the Cas protein of this disclosure is linked to up to two partners, up to three partners, up to four partners, up to five partners, up to six partners, up to seven partners, up to eight partners, or up to ten partners (e.g., by fusion, covalent bonding, or non-covalent bonding). In some embodiments, the Cas protein of this disclosure is linked to 1 to 5 partners, 1 to 4 partners, 1 to 3 partners, 1 to 2 partners, 2 to 5 partners, 2 to 4 partners, 2 to 3 partners, 3 to 5 partners, 3 to 4 partners, or 4 to 5 partners (e.g., by fusion, covalent linkage, or non-covalent linkage). In some embodiments, the Cas protein of this disclosure is linked to 1 partner (e.g., by fusion, covalent linkage, or non-covalent linkage). In some embodiments, the Cas protein of this disclosure is linked to 2 partners (e.g., by fusion, covalent linkage, or non-covalent linkage). In some embodiments, the Cas protein of this disclosure is linked to 3 partners (e.g., by fusion, covalent linkage, or non-covalent linkage). In some embodiments, the Cas protein of this disclosure is linked to four partners (for example, by fusion, covalent bonding, or non-covalent bonding).In some embodiments, the Cas protein of this disclosure is linked to five partners (e.g., by fusion, covalent bonding, or non-covalent bonding). In some embodiments, the Cas protein of this disclosure is linked to six partners (e.g., by fusion, covalent bonding, or non-covalent bonding).

[0209] The Cas protein may be a fusion protein. The fused domain or heterologous polypeptide may be located at the N-terminus, C-terminus, or within the sequence of the Cas protein.

[0210] The Cas protein may be provided in any form. For example, the Cas protein may be provided as a protein, for example, as a standalone Cas protein, or as a ribonucleoprotein in complex with a guide nucleic acid. The Cas protein may be provided as a complex, for example, in complex with a guide nucleic acid and / or one or more heterologous gene effectors of this disclosure. Alternatively, the Cas protein may be provided as a nucleic acid encoding the Cas protein, for example, as RNA (e.g., messenger RNA (mRNA)) or DNA. The nucleic acid encoding the Cas protein may be codon-optimized for efficient translation into the protein in a particular cell or organism.

[0211] Nucleic acids encoding the Cas protein, or fragments or derivatives thereof, can be stably incorporated into the cellular genome. The Cas protein-encoding nucleic acids can be operably ligated to promoters, such as constitutively active promoters or induction-activated promoters in cells. The Cas protein-encoding nucleic acids can be operably ligated to promoters in expression constructs. Expression constructs may include nucleic acid constructs capable of inducing the expression of a target gene or other nucleic acid sequence (e.g., the Cas gene), and such target nucleic acid sequences can be transferred into target cells.

[0212] In some embodiments, the Cas protein or its variant or derivative is a nuclease-inactive dead Cas (dCas) protein. The dead Cas protein may be a protein lacking nucleic acid cleavage activity.

[0213] The Cas protein may include modified versions of the wild-type Cas protein. Modified versions of the wild-type Cas protein may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the Cas protein. For example, the nucleic acid cleavage activity of a modified Cas protein may be 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the wild-type Cas protein (e.g., Cas9 from S. pyogenes). Alternatively, a modified Cas protein may have substantially no nucleic acid cleavage activity. If the Cas protein is a modified version that has substantially no nucleic acid cleavage activity, such a modified version may be referred to herein as an enzymatically inactive "inactive" and / or "dead" (abbreviated as "d"). Dead Cas proteins (e.g., dCas and dCas9) can bind to a target polynucleotide, but do not necessarily cleave the target polynucleotide, or cleave it only minimally. In some embodiments, the dead Cas protein is the dead Cas9 protein.

[0214] The dCas9 polypeptide can activate or repress the transcription of a target gene (e.g., an endogenous target gene) by associating with a single guide RNA (sgRNA) and, for example, by cooperating with a heterologous gene effector disclosed herein. The sgRNA can be introduced into cells expressing the Cas or guide component of this disclosure. In some cases, such cells may contain one or more sgRNAs targeting the same target gene (e.g., an endogenous target gene) or the same regulatory sequence of the target gene. In other cases, the sgRNAs target multiple different nucleic acids within the cell (e.g., multiple different target genes, multiple different regulatory sequences of the target gene, or multiple different sequences within the same target gene or the same regulatory sequence of the target gene).

[0215] "Enzymatically inactive" means a nuclease that can sequence-specifically bind to a nucleic acid sequence in a polynucleotide, but does not cleave the target polynucleotide, or cleaves the target polynucleotide at a significantly lower frequency. The enzymatically inactive guide region may contain an enzymatically inactive domain (e.g., an enzymatically inactive nuclease domain). "Enzymatically inactive" may mean no activity. "Enzymatically inactive" may mean substantially no activity. "Enzymatically inactive" may mean essentially no activity. "Enzymatically inactive" may mean activity of 1% or less, 2% or less, 3% or less, 4% or less, 5% or less, 6% or less, 7% or less, 8% or less, 9% or less, or 10% or less compared to the activity of a comparable wild-type (e.g., nucleic acid cleavage activity or wild-type Cas9 activity).

[0216] In some embodiments, the guide portion does not include a nucleic acid-guided targeting system. For example, the guide portion may include a protein that binds to a target gene (e.g., an endogenous target gene) or a regulatory sequence of a target gene, based on the structural features of a protein such as a specific nuclease disclosed herein.

[0217] In some embodiments, the guide portion comprises a zinc finger nuclease (ZFN), or a variant, fragment, or derivative thereof. A ZFN means a fusion protein of a cleavage domain (e.g., a Fokl cleavage domain) and at least one zinc finger motif (e.g., at least two, at least three, at least four, or at least five zinc finger motifs), where the zinc finger motifs can bind to polynucleotides such as DNA or RNA. In some embodiments, the ZFN is used in the targeting portion of the disclosure to bind to a polynucleotide (e.g., a target gene or its regulatory sequence), but the ZFN does not cleave this polynucleotide, or substantially does not cleave this polynucleotide, for example, a nuclease-inactive dead ZFN. A ZFN or a variant, fragment, or derivative thereof can fuse or associate with one or more heterogeneous effectors to form a complex of the disclosure.

[0218] Two zinc finger neurons (ZFNs) can exert nuclease activity and induce polynucleotide cleavage by forming a heterodimer at a specific location on a polynucleotide in a particular direction and at a specific distance. For example, ZFNs can induce double-strand breaks in DNA by binding to DNA. To cleave DNA by forming a dimer from two cleavage domains, the two ZFNs can attach their respective C-terminuses to the double-stranded DNA strand at a certain distance apart. In some cases, a linker sequence may be required between the zinc finger domain and the cleavage domain to separate each binding site by 5, 6, or 7 base pairs from the 5' end. In some cases, the cleavage domain is fused to the C-terminus of each zinc finger domain.

[0219] In some embodiments, the cleavage domain of the guide portion containing the ZFN includes a modified wild-type cleavage domain. The modified cleavage domain may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the cleavage domain. For example, the nucleic acid cleavage activity of the modified cleavage domain may be 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the corresponding wild-type cleavage domain. Alternatively, the modified cleavage domain may have substantially no nucleic acid cleavage activity. In some embodiments, the modified cleavage domain is enzymatically inactive.

[0220] In some embodiments, the guide portion includes a "TALEN," i.e., a "TAL effector nuclease," or a variant, fragment, or derivative thereof. A TALEN generally refers to a recombinant transcription activator-like effector nuclease containing a central domain consisting of tandem DNA-binding repeats and a cleavage domain. A TALEN can be constructed by fusing a TAL effector DNA-binding domain to a DNA cleavage domain. In some cases, the tandem DNA-binding repeat is 33–35 amino acids long and contains two hypervariable amino acid residues at positions 12 and 13, these hypervariable amino acid residues capable of recognizing at least one specific DNA base pair. Transcription activator-like effector (TALE) proteins can be fused to nucleases such as wild-type or mutant Fok1 endonucleases or the catalytic domain of Fok1. In some embodiments, a TALEN is used in the targeting portion of the disclosure to bind to a polynucleotide (e.g., a target gene or its regulatory sequence), but the TALEN does not cleave the polynucleotide, or substantially does not cleave it, and is, for example, a nuclease-inactive dead TALEN. A TALEN or its variant, fragment, or derivative can fuse or associate with one or more heterogeneous effectors to form a complex of the disclosure.

[0221] In some embodiments, TALENs are recombinant to reduce their nuclease activity. In some embodiments, the nuclease domain of a TALEN includes a modified wild-type nuclease domain. The modified nuclease domain may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the nuclease domain. For example, the nucleic acid cleavage activity of a modified nuclease domain may be 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of a wild-type nuclease domain. Alternatively, a modified nuclease domain may have substantially no nucleic acid cleavage activity. In some embodiments, the modified nuclease domain is enzymatically inactive. TALENs or their variants, fragments, or derivatives can fuse or associate with one or more heterologous effectors to form the complexes of this disclosure.

[0222] For use in TALENs, several mutations have been introduced into Fok1, for example, to improve cleavage specificity or cleavage activity. Such TALENs can be recombined to bind to a desired DNA sequence. Using TALENs, genetic modification can be performed (for example, by editing nucleic acid sequences) by creating double-strand breaks in a target DNA sequence and inducing non-homologous end joining or homologous recombination repair.

[0223] TALE or its variants, fragments, or derivatives can fuse or associate with one or more heterogeneous effectors to form the complexes of the present disclosure. In some embodiments, the transcription activator-like effector (TALE) protein is fused to a heterogeneous effector and does not contain a nuclease. In some embodiments, the TALEN does not cleave polynucleotides or substantially does not cleave polynucleotides, for example, it is a nuclease-inactive dead TALE. TALE or its variants, fragments, or derivatives can fusion or associate with one or more heterogeneous effectors to form the complexes of the present disclosure.

[0224] In some embodiments, the complex of a transcription activator-like effector (TALE) protein and a heterogeneous effector is designed to function as a transcriptional activator. In some embodiments, the complex of a transcription activator-like effector (TALE) protein and a heterogeneous effector is designed to function as a transcriptional repressor. For example, the DNA-binding domain of a transcription activator-like effector (TALE) protein can be fused (e.g., ligated) to one or more heterogeneous effectors containing a transcriptional activation domain, or to one or more heterogeneous effectors containing a transcriptional repression domain.

[0225] In some embodiments, the guide portion includes a meganuclease. A meganuclease generally refers to an endonuclease or homing endonuclease that cleaves at rare sites that may have extremely high sequence specificity. A meganuclease can recognize DNA target sites of at least 12 base pairs in length, for example, DNA target sites of 12–40 base pairs, 12–50 base pairs, or 12–60 base pairs. A meganuclease may be a modular DNA-binding nuclease, such as a fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA-binding domain, or at least one protein that identifies a nucleic acid target sequence. This DNA-binding domain may contain at least one motif that recognizes single-stranded or double-stranded DNA. The nuclease-active form of the meganuclease can form double-strand breaks. In some embodiments, a meganuclease is used in the targeting portion of the disclosure to bind to a polynucleotide (e.g., a target gene or its regulatory sequence), but the meganuclease does not cleave the polynucleotide or substantially does not cleave it, for example, a nuclease-inactive dead meganuclease. A meganuclease or its variant, fragment, or derivative can fuse or associate with one or more heterogeneous effectors to form a complex of the disclosure.

[0226] Meganucleases may be monomers or dimers. In some embodiments, the meganucleases are natural or wild-type (found in nature), and in other embodiments, the meganucleases are unnatural, artificial, recombinant, synthesized, rationally designed, or man-made. In some embodiments, the meganucleases of the present disclosure include I-CreI meganuclease, I-CeuI meganuclease, I-Msol meganuclease, or I-SceI meganuclease, its variants, its derivatives, or functional fragments thereof.

[0227] In some embodiments, the nuclease domain of the meganuclease includes a modified version of the wild-type nuclease domain. The modified nuclease domain may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce or eliminate the nucleic acid cleavage activity of the nuclease domain. For example, the nucleic acid cleavage activity of the modified nuclease domain may be 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the wild-type nuclease domain. Alternatively, the modified nuclease domain may have substantially no nucleic acid cleavage activity. In some embodiments, the modified nuclease domain is enzymatically inactive. In some embodiments, the meganuclease can bind to DNA but cannot cleave it. In some embodiments, a nuclease-inactive meganuclease can fuse with or associate with one or more heterologous gene effectors to form the complex of the present disclosure.

[0228] In some embodiments, the guiding portion of the Disclosure can modulate the expression and / or activity of a target gene (e.g., an endogenous target gene). In some embodiments, the guiding portion of the Disclosure can edit the sequence of a nucleic acid (e.g., a gene and / or gene product). Nuclease-active Cas proteins can edit nucleic acid sequences by introducing double-strand or single-strand breaks at a target polynucleotide.

[0229] In some embodiments, a guide region containing a nuclease can introduce double-strand breaks into target polynucleotides such as DNA. The introduction of double-strand breaks into DNA allows for DNA repair, enabling genetic modification (e.g., nucleic acid editing). In some embodiments, the nuclease induces site-specific single-strand DNA breaks, or nicks, leading to homologous recombination repair.

[0230] Introducing double-strand breaks into DNA allows for DNA repair, enabling genetic modification (e.g., nucleic acid editing). DNA break repair can occur through non-homologous end joining (NHEJ) or homologous recombination repair (HDR). Homologous recombination repair can provide a donor DNA repair template or template polynucleotide containing homologous arms flanking the target DNA site.

[0231] In some embodiments, the guide portion or complex containing the nuclease does not form double-strand breaks in the target polynucleotide, such as DNA.

[0232] In some embodiments, one or more complexes or systems comprising heterologous polypeptides and heterologous polynucleotides are disclosed. In some cases, the complexes of the disclosure may include heterologous gene effectors and guide regions, for example, guide nucleic acids and / or nucleases, the nuclease being, for example, an endonuclease lacking or substantially lacking cleavage activity.

[0233] The complexes of this disclosure may be useful for promoting the regulation of the expression, epigenetic modification, or activity level of a target gene by, for example, positioning one or more heterogeneous effectors near a target gene (e.g., an endogenous target gene) or a regulatory sequence of a target gene.

[0234] In some embodiments, the complex of the disclosure binds to DNA, for example, genomic DNA. In some embodiments, the complex of the disclosure binds to RNA, for example, mRNA, microRNA, siRNA, or non-coding RNA. In some embodiments, the complex of the disclosure binds to both DNA and RNA.

[0235] In some embodiments, the complexes of the present disclosure can modulate (e.g., increase or decrease) the expression and / or activity of a target gene (e.g., an endogenous target gene) by physically interfering with a polynucleotide sequence (e.g., promoters, enhancers, repressors, operators, silencers, insulators, cis-regulators, trans-regulators, epigenetic modification sites (e.g., DNA methylation sites), coding sequences).

[0236] In some embodiments, the complex of the present disclosure can modulate (e.g., increase or decrease) the expression and / or activity of a target gene (e.g., an endogenous target gene) by recruiting additional factors that are effective in repressing or enhancing the expression of the target gene.

[0237] In some embodiments, the complexes of the Disclosure are used to introduce epigenetic modifications to a target gene (e.g., an endogenous target gene) or a regulatory sequence of a target gene (e.g., a promoter, enhancer, silencer, insulator, cis-regulator, trans-regulator, or epigenetic modification site (e.g., a DNA methylation site)). In some embodiments, the complexes of the Disclosure are used to generate a three-dimensional structure, a stereometrically associated domain, or a genomic boundary comprising a target gene or its regulatory sequence (e.g., a gene distal or proximal to the target gene).

[0238] In some embodiments, the complex or system of the disclosure includes a heterogeneous effector and a guide portion. In some embodiments, the complex or system of the disclosure includes one heterogeneous effector and one guide portion. In some embodiments, the complex or system of the disclosure includes two heterogeneous effectors and one guide portion. In some embodiments, the complex or system of the disclosure includes three or more heterogeneous effectors and one guide portion.

[0239] In some embodiments, the complex or system of the Disclosure comprises a heterogeneous effector and a guide nucleic acid. In some embodiments, the complex or system of the Disclosure comprises one heterogeneous effector and one guide nucleic acid. In some embodiments, the complex or system of the Disclosure comprises two heterogeneous effectors and one guide nucleic acid. In some embodiments, the complex or system of the Disclosure comprises three or more heterogeneous effectors and one guide nucleic acid.

[0240] Two components present in the complex or system of the Disclosure may, for example, be linked by covalent bonds if they are present in a fusion protein, crosslinked by, for example, treatment with a crosslinking agent, or linked by a peptide linker or a non-peptide linker, as disclosed herein.

[0241] In some embodiments, the two components present in the complex or system of the present disclosure are parts of a single fusion protein. These components can be linked by a linker, such as a peptide linker or a non-peptide linker.

[0242] In some embodiments, a guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is linked to a heterogeneous effector via a linker. In some embodiments, the guide portion or part thereof is further linked to a second heterogeneous effector via a second linker, which is the same as or different from the first linker. In some embodiments, the guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is fused to the heterogeneous effector without the use of a linker.

[0243] In some embodiments, a guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is linked to an oligomerization domain or dimerization domain (e.g., a heterodimerization domain) via a linker. In some embodiments, the guide portion or part thereof is further linked to a second oligomerization domain or dimerization domain (e.g., a heterodimerization domain) via a second linker which is the same as or different from the first linker. In some embodiments, the guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is fused to a second oligomerization domain or dimerization domain (e.g., a heterodimerization domain) without the use of a linker.

[0244] In some embodiments, a heterogeneous effector is linked to a second heterogeneous effector via a linker. In some embodiments, the heterogeneous effector is further linked to a third heterogeneous effector via a second linker that is the same as or different from the first linker. In some embodiments, the heterogeneous effector is fused to a second heterogeneous effector without a linker.

[0245] In some embodiments, a heterogeneous effector is linked to an oligomerized or dimerized domain (e.g., a heterodimerized domain) via a linker. In some embodiments, the heterogeneous effector is further linked to a second oligomerized or dimerized domain (e.g., a heterodimerized domain) via a second linker that is the same as or different from the first linker. In some embodiments, the heterogeneous effector is fused to a second oligomerized or dimerized domain (e.g., a heterodimerized domain) without the use of a linker.

[0246] Any suitable linker can be used. Flexible linkers may have sequences containing regions consisting of glycine and serine residues. The small size of glycine and serine residues can impart flexibility and mobility to linked functional domains. By incorporating serine or threonine, the linker's stability in aqueous solution can be maintained by forming hydrogen bonds with water molecules, thereby suppressing undesirable interactions between the linker and protein moieties. Flexible linkers may also contain additional amino acids such as threonine or alanine to maintain flexibility, and polar amino acids such as lysine or glutamine to improve solubility. Rigid linkers may, for example, have an α-helix structure. Rigid linkers with an α-helix structure can function as spacers between protein domains.

[0247] The length of the linker sequence is, for example, 1 amino acid residue length, 2 amino acid residue length, 3 amino acid residue length, 4 amino acid residue length, 5 amino acid residue length, 6 amino acid residue length, 7 amino acid residue length, 8 amino acid residue length, 9 amino acid residue length, 10 amino acid residue length, 11 amino acid residue length, 12 amino acid residue length, 13 amino acid residue length, 14 amino acid residue length, 15 amino acid residue length, 16 amino acid residue length, 17 amino acid residue length, 18 amino acid residue length, 19 amino acid residue length, 20 amino acid residue length, 21 amino acid residue length, 22 amino acid residue length, 23 amino acid residue length, 24 amino acid residue length, 25 amino acid residue length, 266 The amino acid residue length may be 27 amino acid residues, 28 amino acid residues, 29 amino acid residues, 30 amino acid residues, 31 amino acid residues, 32 amino acid residues, 33 amino acid residues, 34 amino acid residues, 35 amino acid residues, 36 amino acid residues, 37 amino acid residues, 38 amino acid residues, 39 amino acid residues, 40 amino acid residues, 41 amino acid residues, 42 amino acid residues, 43 amino acid residues, 44 amino acid residues, 45 amino acid residues, 46 amino acid residues, 47 amino acid residues, 48 ​​amino acid residues, 49 amino acid residues, or 50 amino acid residues.

[0248] In some embodiments, the linker sequence is at least 1 amino acid, at least 2 amino acids, at least 3 amino acids, at least 5 amino acids, at least 7 amino acids, at least 9 amino acids, at least 11 amino acids, at least 13 amino acids, at least 15 amino acids, or at least 20 amino acids. In some embodiments, the linker sequence is up to 5 amino acids, up to 7 amino acids, up to 9 amino acids, up to 11 amino acids, up to 13 amino acids, up to 15 amino acids, up to 20 amino acids, up to 25 amino acids, up to 30 amino acids, up to 40 amino acids, or up to 50 amino acids.

[0249] In some embodiments, non-peptide linkers are used. Non-peptide linkers may be, for example, chemical linkers. Chemical linkers can link two parts of the complex or system of the Disclosure. Each chemical linker of the Disclosure may be alkylene, alkenylene, alkynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, and these linkers may be substituted. In some embodiments, the chemical linkers of the Disclosure may be esters, ethers, amides, thioethers, or polyethylene glycol (PEG). In some embodiments, the linker can reverse the order of the amino acid sequence in the compound, for example, so that the amino acid sequence linked by the linker is linked at the beginning rather than at the end. Examples of such linkers include, but are not limited to, diesters of dicarboxylic acids, such as oxalyl diesters, malonyl diesters, succinyl diesters, glutaryl diesters, adipyl diesters, pimethyl diesters, fumaryl diesters, maleyl diesters, phthalyl diesters, isophthalyl diesters, and terephthalyl diesters. Examples of such linkers include, but are not limited to, diamides of dicarboxylic acids, such as oxalyl diamides, malonyl diamides, succinyl diamides, glutaryl diamides, adipyl diamides, pimethyl diamides, fumaryl diamides, maleyl diamides, phthalyl diamides, isophthalyl diamides, and terephthalyl diamides. Examples of such linkers include, but are not limited to, diaminodiamide linkers, such as ethylenediamine, 1,2-di(methylamino)ethane, 1,3-diaminopropane, 1,3-di(methylamino)propane, 1,4-di(methylamino)butane, 1,5-di(methylamino)pentane, 1,6-di(methylamino)hexane, and piperazine.Examples of substituents that may be introduced into the linker include, but are not limited to, hydroxyl groups, sulfhydryl groups, halogens, amino groups, nitro groups, nitroso groups, cyano groups, azide groups, sulfoxide groups, sulfone groups, sulfonamide groups, carboxyl groups, carboxyaldehyde groups, imine groups, alkyl groups, haloalkyl groups, alkenyl groups, haloalkenyl groups, alkynyl groups, haloalkynyl groups, alkoxy groups, aryl groups, aryloxy groups, aralkyl groups, arylalkoxy groups, heterocyclyl groups, acyl groups, acyloxy groups, carbamate groups, amide groups, ureido groups, epoxy groups, and ester groups.

[0250] Two components present in the complex or system of this disclosure can be linked, for example, by ionic bonds, hydrogen bonds, or non-covalent bonds through interactions, such as oligomerized domains or dimerized domains disclosed herein.

[0251] In some embodiments, a guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is linked to a heterogeneous effector via a non-covalent bond. In some embodiments, the guide portion or part thereof is further linked to a second heterogeneous effector via a non-covalent bond. In some embodiments, the guide portion or part thereof is linked to a first heterogeneous effector via a covalent bond (e.g., as a fusion protein, e.g., as a linker-mediated fusion protein) and further linked to a second heterogeneous effector via a non-covalent bond.

[0252] In some embodiments, a guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is linked to an oligomerized domain or dimerized domain (e.g., a heterodimerized domain) via a non-covalent bond. In some embodiments, the guide portion or part thereof is further linked to a second oligomerized domain or dimerized domain (e.g., a heterodimerized domain) via a non-covalent bond. In some embodiments, a guide portion or part thereof (e.g., a nuclease, e.g., dCas9) is fused to a first oligomerized domain or dimerized domain (e.g., a heterodimerized domain) via a covalent bond (e.g., fusion may also be via a linker) and further linked to a second oligomerized domain or dimerized domain (e.g., a heterodimerized domain) via a non-covalent bond.

[0253] In some embodiments, a first component of the guide portion (e.g., a guide nucleic acid) is linked to a second component of the guide portion (e.g., a nuclease) via a non-covalent bond. In some embodiments, a first component of the guide portion (e.g., a guide nucleic acid) is linked to a second component of the guide portion (e.g., a nuclease) via a covalent bond.

[0254] A combination of covalent and non-covalent bonding can be used in the complex or system of the present disclosure. For example, one or more heterogeneous effectors can be fused to a guide region via non-covalent bonding, and one or more oligomeric domains can be attached to a component of the complex or system of the present disclosure (e.g., a nuclease) via covalent bonding.

[0255] In some embodiments, a polypeptide that increases or decreases stability is fused to or associated with a component of the complex or system of the Disclosure (e.g., a guide portion or a heterogeneous effector). This fused polypeptide may be located at the N-terminus, C-terminus, or inside the fusion protein.

[0256] In some embodiments, one or more components of the complex or system of the present disclosure are fused to a domain that induces a desired subcellular localization, such as a protein or nuclear localization signal that targets the inner nuclear membrane, outer nuclear membrane, Cajal body, nuclear plaque, nuclear pore complex, PML body, nucleolus, P granule, GW body, stress granule, spongy body, endoplasmic reticulum, mitochondria, etc.

[0257] In some embodiments, the complex or system of the Disclosure comprises a first protein linked to a first oligomerization domain (e.g., a dimerization domain) and a second protein linked to a second oligomerization domain (e.g., a dimerization domain). In some embodiments, the oligomerization domain or dimerization domain may include a peptide interaction domain and may include, for example, a system utilizing sgRNA2.0, SAM, SunTag, RAB, FLAG-biotin, or an induceable oligomerization system (e.g., a dimerization system) disclosed herein.

[0258] delivery One or more genes encoding heterologous polypeptides (e.g., heterologous gene effectors) and additional molecules operably linked thereto (e.g., heterologous polynucleotides, e.g., one or more guide nucleic acid molecules) can be incorporated into the genome of a cell in which the abnormal expression of a target gene is to be modified. Alternatively, the one or more genes may not be incorporated into the cell's genome, or may not need to be incorporated into the cell's genome. The one or more genes may be a single gene (e.g., a single expression cassette) or multiple genes (e.g., multiple expression cassettes). The one or more genes may be heterologous to the cell.

[0259] Heterogenetic polypeptides (e.g., heterogeneous gene effectors) and additional molecules operably linked thereto (e.g., heterogeneous polynucleotides, e.g., one or more guide nucleic acid molecules) may be introduced into cells by various methods, such as viral or nonviral delivery methods (e.g., delivery and expression). Viral vector delivery systems may include DNA viruses or RNA viruses, and these viruses may have an episomal genome or an integrated genome after being delivered to cells. Nonviral vector delivery systems may include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery medium such as liposomes.

[0260] The one or more genes described herein may further comprise one or more promoters to control the expression of the systems of this disclosure. The promoters disclosed herein may be active in eukaryotic cells, mammalian cells, non-human mammalian cells, or human cells. The promoters may be inductive promoters or constitutively active promoters. In addition, or in other embodiments, the promoters may be tissue-specific promoters or cell-specific promoters. Examples of suitable eukaryotic promoters (i.e., promoters functional in eukaryotic cells) include, but are not limited to, the cytomegalovirus (CMV) early promoter, the herpes simplex virus (HSV) thymidine kinase promoter, the SV40 early and late promoters, promoters derived from retroviral long-chain terminal repeat sequences (LTRs), the human elongation factor 1 promoter (EF1), a hybrid construct comprising a cytomegalovirus (CMV) enhancer fused to a chicken β-actin promoter (CAG), the mouse stem cell virus promoter (MSCV), the phosphoglycerate kinase 1 promoter (PGK), and the mouse metallothionein-I promoter. The promoters may also be fungal promoters. The promoter may be a plant promoter. Databases of plant promoters are publicly known (e.g., PlantProm). The expression vector may further include a ribosome binding site and a transcription terminator for translation initiation. The expression vector may further include a suitable sequence for amplification of expression. In some cases, the promoters disclosed herein may be promoters specific to any of the tissues provided herein, or promoters specific to any of the cell types provided herein.

[0261] The size of a single gene provided herein (for example, a single gene encoding the system of this disclosure) is at least about 2.5 kb or less, at least about 2.6 kb or less, at least about 2.7 kb or less, at least about 2.8 kb or less, at least about 2.9 kb or less, at least about 3.0 kb or less, at least about 3.1 kb or less. , at least approximately 3.2kb or less, at least approximately 3.3kb or less, at least approximately 3.4kb or less, at least approximately 3.5kb or less, at least approximately 3.6kb or less, at least approximately 3.7kb or less, at least approximately 3.8kb or less, at least approximately 3.9kb or less, at least approximately 4.0kb or less, less than At least approximately 4.1kb or less, at least approximately 4.2kb or less, at least approximately 4.3kb or less, at least approximately 4.4kb or less, at least approximately 4.5kb or less, at least approximately 4.6kb or less, at least approximately 4.7kb or less, at least approximately 4.8kb or less, at least approximately 4.9kb or less, at least It may be approximately 5.0kb or less, at least approximately 5.5kb or less, at least approximately 6.0kb or less, at least approximately 6.5kb or less, at least approximately 7.0kb or less, at least approximately 7.5kb or less, at least approximately 8.0kb or less, at least approximately 9.0kb or less, or at least approximately 10kb or less.In some cases, the size of a single gene provided herein is approximately 3kb to 5kb, approximately 3kb to 4.8kb, approximately 3kb to 4.6kb, approximately 3kb to 4.4kb, approximately 3kb to 4.2kb, approximately 3kb to 4.0kb, approximately 3kb to 3.5kb, approximately 3.5kb to 5kb, approximately 3.5kb to 4.8kb, approximately 3.5kb to 4.6kb, and approximately 3.5kb. It may be approximately 4.4kb, 3.5kb to 4.2kb, 3.5kb to 4kb, 4kb to 5kb, 4kb to 4.9kb, 4kb to 4.8kb, 4kb to 4.7kb, 4kb to 4.6kb, 4kb to 4.5kb, 4kb to 4.4kb, 4kb to 4.3kb, 4kb to 4.2kb, or 4kb to 4.1kb.

[0262] By using systems employing RNA or DNA viruses, specific cells can be targeted, and a viral payload can be delivered to the nucleus of these cells. Viral vectors can be brought into contact with cells in vitro, and the resulting modified cells can be administered to subjects such as humans (ex vivo). Alternatively, viral vectors can be administered directly to subjects (in vivo). Examples of viral systems include retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex virus vectors for gene transfer. Gene transfer methods using retroviruses, lentiviruses, or adeno-associated viruses can induce integration into the host genome, resulting in the long-term expression of the inserted transgene.

[0263] In some embodiments, the compositions and systems provided herein are delivered to a subject using a viral vector. In some cases, the viral vector is an adeno-associated virus (AAV) vector. The term "AAV" is an abbreviation for adeno-associated virus and may be used to refer to adeno-associated virus itself or to those derived from adeno-associated virus. Unless otherwise specified, the term encompasses all serotypes, subtypes, native and recombinant forms. The abbreviation "rAAV" means recombinant adeno-associated virus and is also called a recombinant AAV vector (i.e., "rAAV vector"). The term "AAV" also includes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, rh10, and their hybrid forms, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, and sheep AAV. The genomic sequences of various serotypes of AAV, as well as their native terminal repeat (TR) sequences, Rep proteins, and capsid subunits, are known in the art. Such sequences can also be found in academic literature and in public databases such as GenBank. In this specification, “rAAV vector” means an AAV vector containing a polynucleotide sequence derived from something other than AAV (i.e., a polynucleotide heterologous to AAV), such a polynucleotide sequence is usually the target sequence for cell transformation. Generally, this heterologous polynucleotide is positioned adjacent to at least one AAV terminal inverted repeat (ITR), and is usually positioned between two AAV terminal inverted repeats (ITRs). The term “rAAV vector” encompasses both rAAV vector particles and rAAV vector plasmids. rAAV vectors may be single-stranded (ssAAV) or self-complementary (scAAV). The terms "AAV virus," "AAV virus particle," or "rAAV vector particle" refer to a viral particle composed of at least one AAV capsid protein and a polynucleotide rAAV vector enclosed in the capsid.When AAV particles contain heterologous polynucleotides (i.e., polynucleotides other than the wild-type AAV genome, e.g., transgenes delivered to mammalian cells), these AAV particles are generally called “rAAV vector particles” or simply “rAAV vectors.” Therefore, since rAAV particles contain vectors, the production of rAAV particles is always accompanied by the production of rAAV vectors. In some cases, AAV vectors are selected based on the directivity of the viral vector. In some embodiments, a tissue-directed AAV vector (e.g., AAV2 for muscle tissue) may be used to deliver the polynucleotides encoding the compositions or systems provided herein to the target tissue.

[0264] By using systems employing RNA or DNA viruses, it is possible to target specific cells in a living organism and deliver a viral payload to the nucleus of these cells. Viral vectors can be administered directly (in vivo), or they can be brought into contact with cells in vitro, and the resulting modified cells can be administered to subjects such as humans (ex vivo). Examples of viral systems include retroviral vectors, lentiviral vectors, adenovirus vectors, adeno-associated virus vectors, and herpes simplex virus vectors for gene transfer. Gene transfer methods using retroviruses, lentiviruses, or adeno-associated viruses can induce integration into the host genome, resulting in the long-term expression of the inserted transgene. High transduction efficiency can be obtained in many types of cells and target tissues.

[0265] The targeting properties of retroviruses can be modified by incorporating foreign envelope proteins, thereby expanding the range of target cell populations that can be targeted. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and can achieve high viral titers. The choice of retroviral gene transfer system may depend on the type of target tissue. Retroviral vectors may contain cis-acting long-term repeat sequences and can package foreign sequences up to 6-10 kb in length. This minimal cis-acting LTR may be sufficient for vector replication and packaging and can be used to integrate therapeutic genes into target cells, allowing for the permanent expression of the transgene. Examples of retroviral vectors include those based on mouse leukemia virus (MuLV), gibbon leukemia virus (GaLV), simian immunodeficiency virus (SIV), or human immunodeficiency virus (HIV), and combinations thereof.

[0266] Adenovirus-based systems can also be used. Adenovirus-based systems can induce transient expression of transgenes. Adenovirus-based vectors can achieve high transduction efficiency in cells, and cell division may not be required. Furthermore, high titers and high expression levels can be obtained by using adenovirus-based vectors. By transduction of target nucleic acids into cells using adeno-associated virus ("AAV") vectors, for example, nucleic acid and peptide production in vitro, or gene therapy treatment in vivo or ex vivo can be performed.

[0267] By utilizing packaging cells, viral particles capable of infecting host cells can be formed. Examples of such cells include 293 cells (for example, adenovirus packaging) and Psi2 or PA317 cells (for example, retrovirus packaging). Viral vectors can be produced by creating cell lines capable of packaging nucleic acid vectors into viral particles. The vector may contain the minimum viral sequence required for packaging and subsequent integration into the host. The vector may also contain other viral sequences substituted with expression cassettes encoding the polynucleotides to be expressed. Lost viral function can be supplied from the packaging cell line to the trans. For example, an AAV vector may contain an ITR sequence derived from the AAV genome, which is required for packaging and integration into the host genome. Viral DNA can be packaged into cell lines that lack the ITR sequence but may contain helper plasmids encoding other AAV genes (i.e., rep and cap). The cell lines can also be infected with adenovirus as a helper. Helper viruses can promote the replication of AAV vectors and the expression of AAV genes from helper plasmids. Adenovirus contamination can be suppressed, for example, by heat treatment, in which adenoviruses are more susceptible than AAV.

[0268] One or more vectors described herein can be transfected into host cells transiently or non-transiently. Transfection into cells can be performed spontaneously in a subject. Cells can be collected from a subject or cells derived from a subject can be used for transfection. Cells (e.g., cell lines) can also be obtained from cells collected from a subject. In some embodiments, cells transfected with one or more vectors described herein are used to establish novel cell lines containing one or more vector-derived sequences. In some embodiments, cells are modified via the activation of actuator portions such as CRISPR complexes by transiently transfecting cells with the compositions of this disclosure (e.g., transient transfection of one or more vectors, or transient transfection of RNA, etc.), and these modified cells are used to establish novel cell lines containing cells that include the modification but do not contain other exogenous sequences.

[0269] Any suitable vector compatible with the host cell can be used in the method of this disclosure. Examples of vectors for eukaryotic host cells include pXT1, pSG5 (Stratagene) TM ), pSVK3, pBPV, pMSG and pSVLSV40 (Pharmacia TM ) are some examples, but are not limited to these.

[0270] In some embodiments, additional components of the compositions disclosed herein may include additives. Examples of additives include, but are not limited to, solvents, dispersions, diluents or other liquid media; dispersants or suspending agents; surfactants; isotonic agents; thickeners or emulsifiers; preservatives, lipidoids, liposomes, lipid nanoparticles, polymers, Lipoflex, core-shell nanoparticles, peptides, proteins, hyaluronidases, nanoparticle mimetic agents, inert diluents, buffers, lubricants and oils, and combinations thereof. In some examples, the compositions disclosed herein may include (i) heterologous polypeptides or heterologous genes encoding them, and / or (ii) one or more additives in amounts that can improve the stability of cells or modified cells.

[0271] In some embodiments, the Disclosure provides a kit comprising such compositions and instructions thereof, which instruct (i) to contact cells with the compositions (e.g., in vitro, ex vivo, or in vivo), or (ii) to administer cells containing any one of the compositions disclosed herein to a subject. The subject may be a subject having a condition such as a genetic disorder, or a subject suspected of having a condition such as a genetic disorder.

[0272] In some embodiments, the compositions disclosed herein can be administered to a subject by oral administration, intraperitoneal administration, intravenous administration, intra-arterial administration, transdermal administration, intramuscular administration, administration via liposomes, local delivery by catheter or stent, subcutaneous administration, intrafat administration, or intrasacral administration. In specific embodiments, the compositions and systems provided herein (e.g., including polynucleotides encoded in AAV vectors, etc.) can be administered to a subject by intravenous administration.

[0273] Examples of viral vectors that can be used for the delivery of heterologous polypeptides and / or heterologous polynucleotides (or one or more genes encoding them) include, but are not limited to, retroviral vectors, lentiviral vectors, adenovirus vectors, poxvirus vectors, herpesvirus vectors, or adeno-associated virus (AAV) vectors. Examples of AAV vectors include AAV1, AAV10, AAV106.1 / hu.37, AAV11, AAV114.3 / hu.40, AAV12, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV 128.1 / hu.43, AAV128.3 / hu.44, AAV130.4 / hu.48, AAV145.1 / hu.53, AAV145.5 / hu.54, AAV145.6 / hu.55, AAV16.12 / hu. 11, AAV16.3, AAV16.8 / hu. 10, AAV161.10 / hu.60, AAV161.6 / hu.61, AAV1-7 / rh.48, AAV1-8 / rh.49, AAV2, AAV2.5T, AAV2 -15 / rh.62, AAV223.1, AAV223.2, AAV223.4, AAV223.5, AAV223.6, AAV223.7, AAV2-3 / rh.61, A AV24.1, AAV2-4 / rh.50, AAV2-5 / rh.51, AAV27.3, AAV29.3 / bb.1, AAV29.5 / bb.2, AAV2G9, AAV- 2-pre-miRNA-101, AAV3, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-11 / rh.53, AAV3-3, AAV33.12 / hu. 17, AAV33.4 / hu. 15, AAV33.8 / hu. 16, AAV3-9 / rh.52, AAV3a, AAV3b, AAV4, AAV4-19 / rh.55, AAV42.12、AAV42-10、AAV42-11、AAV42-12、AAV42-13、AAV42-15、AAV42-1b、AAV42-2、AAV42-3a、AAV42- 3b、AAV42-4、AAV42-5a、AAV42-5b、AAV42-6b、AAV42-8、AAV42-aa、AAV43-1、AAV43-12、AAV43-20、 AAV43-21、AAV43-23、AAV43-25、AAV43-5、AAV4-4、AAV44.1、AAV44.2、AAV44.5、AAV46.2 / hu.28、AAV46.6 / hu.29、AAV4-8 / r11.64、AAV4-8 / rh.64、AAV4-9 / rh.54、AAV5、AAV52.1 / hu.20、AAV52 / hu. 19、AAV5-22 / rh.58、AAV5-3 / rh.57、AAV54.1 / hu.21、AAV54.2 / hu.22、AAV54.4R / hu.27、AAV54.5 / hu.23、AAV54.7 / hu.24、AAV58.2 / hu.25、AAV6、AAV6.1、AAV6.1.2、AAV6.2、AAV7、AAV7.2、AAV7.3 / hu.7、 AAV8、AAV-8b、AAV-8h、AAV9、AAV9.11、AAV9.13、AAV9.16、AAV9.24、AAV9.45、AAV9.47、AAV9.61、AAV9 .68、AAV9.84、AAV9.9、AAVA3.3、AAVA3.4、AAVA3.5、AAVA3.7、AAV-b、AAVC1、AAVC2、AAVC5、AAVCh.5、AAV AVCh.5R1、AAVcy.2、AAVcy.3、AAVcy.4、AAVcy.5、AAVCy.5R1、AAVCy.5R2、AAVCy.5R3、AAVCy.5R4、AAV cy.6、AAV-DJ、AAV-DJ8、AAVF3、AAVF5、AAV-h、AAVH-1 / hu.1、AAVH2、AAVH-5 / hu.3、AAVH6、AAVhE1.1、AAVH AVhER1.14、AAVhEr1.16、AAVhEr1.18、AAVhEr1.23、AAVhEr1.35、AAVhEr1.36、AAVhEr1.5、AAVhEr1.7 、AAVhEr1.8、AAVhEr2.16、AAVhEr2.29、AAVhEr2.30、AAVhEr2.31、AAVhEr2.36、AAVhEr2.4、AAVhEr3.1、AAVhu.1、AAVhu.10、AAVhu.11、AAVhu.11、AAVhu.12、AAVhu.13、AAVhu.14 / 9、AAVhu.15、AAVhu.16、AAVhu.17、AAVhu.18、AAVhu.19、AAVhu.2、AAVhu.20、AAVhu.21、AAVhu.22、AAVhu.23.2、AAVhu.24、AAVhu.25、AAVhu.27、AAVhu.28、AAVhu.29、AAVhu.29R、AAVhu.3、AAVhu.31、AAVhu.32、AAVhu.34、AAVhu.35、AAVhu.37、AAVhu.39、AAVhu.4、AAVhu.40、AAVhu.41、AAVhu.42、AAVhu.43、AAVhu.44、AAVhu.44R1、AAVhu.44R2、AAVhu.44R3、AAVhu.45、AAVhu.46、AAVhu.47、AAVhu.48、AAVhu.48R1、AAVhu.48R2、AAVhu.48R3、AAVhu.49、AAVhu.5、AAVhu.51、AAVhu.52、AAVhu.53、AAVhu.54、AAVhu.55、AAVhu.56、AAVhu.57、AAVhu.58、AAVhu.6、AAVhu.60、AAVhu.61、AAVhu.63、AAVhu.64、AAVhu.66、AAVhu.67、AAVhu.7、AAVhu.8、AAVhu.9、AAVhu.t 19、AAVLG-10 / rh.40、AAVLG-4 / rh.38、AAVLG-9 / hu.39、AAVLG-9 / hu.39、AAV-LKO1、AAV-LK02、AAVLK03、AAV-LK03、AAV-LK04、AAV-LK05、AAV-LKO6、AAV-LK07、AAV-LK08、AAV-LK09、AAV-LK10、AAV-LK11、AAV-LK12、AAV-LK13、AAV-LK14、AAV-LK15、AAV-LK17、AAV-LK18、AAV-LK19、AAVN721-8 / rh.43、AAV-PAEC、AAV-PAEC11、AAV-PAEC12、AAV-PAEC2、AAV-PAEC4、AAV-PAEC6、AAV-PAEC7、AAV-PAEC8、AAVpi.1、AAVpi.2、AAVpi.3、AAVrh.10、AAVrh.12、AAVrh.13、AAVrh.13R、AAVrh.14、AAVrh.17、AAVrh.18、AAVrh.19、AAVrh.2、AAVrh.20、AAVrh.21、AAVrh.22、AAVrh.23、AAVrh.24、AAVrh.25、AAVrh.2R、AAVrh.31、AAVrh.32、AAVrh.33、AAVrh. h.34、AAVrh.35、AAVrh.36、AAVrh.37、AAVrh.37R2、AAVrh.38、AAVrh.39、AAVrh.40、AAVrh.43、AAVrh.44、AAVrh.45、AAVrh.46、AAVrh.47、AAVrh.48、AAVrh.48、AAVrh.48.1 、AAVrh.48.1.2、AAVrh.48.2、AAVrh.49、AAVrh.50、AAVrh.51、AAVrh.52、AAVrh.53、AAVrh.54、AAVrh.55、AAVrh.56、AAVrh.57、AAVrh.58、AAVrh.59、AAVrh.60、AAVrh.61、AAVrh. AVrh.62、AAVrh.64、AAVrh.64R1、AAVrh.64R2、AAVrh.65、AAVrh.67、AAVrh.68、AAVrh.69、AAVrh.70、AAVrh.72、AAVrh.73、AAVrh.74、AAVrh.8、AAVrh.8R、AAVrh.8R A586R mutation、AAVrh8R R533A mutation、BAAV、BNP61 AAV、BNP62 AAV、BNP63 AAV、ウシAAV、ヤギAAV、JapaneseAAV 10、true type AAV(ttAAV)、UPENN AAV 10、AAV-LK16、AAV Shuffle 100-1、AAV Shuffle 100-2 100-3、AAV Shuffle 100-7、AAV Shuffle 10-2、AAV Shuffle 10-6、AAV Shuffle 10-8、AAV SM 100-10、AAV SM 100-3、AAV SM 10-1、AAV SM 10-2、or AAV SM 10-8 are mentioned but not limited to these. For example、AAVrh.Using 74 as a viral vector, polynucleotide sequences encoding heterogeneous polypeptides and heterogeneous polynucleotides (e.g., a Cas protein-gene effector fusion protein and one or more guide nucleic acid molecules) can be delivered.

[0274] Non-viral methods of nucleic acid delivery include lipofection, nucleofection, microinjection, microparticle guns, virosomes, liposomes, immunoliposomes, polycationic-nucleic acid complexes, lipid-nucleic acid complexes, lipid nanoparticles (LNPs), naked DNA, artificial virions, and drug-enhanced DNA uptake. Cationic and neutral lipids suitable for efficiently introducing polynucleotides can also be used through receptor-recognizing lipofection.

[0275] The compositions disclosed herein (or one or more genes encoding a portion of the compositions disclosed herein), e.g., heterogeneous effector and / or guide nucleic acid molecules, can be administered by any suitable route of administration, including, but not limited to, parenteral administration routes (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intraventricular, intra-articular, intraperitoneal, or intracranial), intranasal, buccal, sublingual, oral, and rectal administration routes. In some cases, the pharmaceutical compositions disclosed herein may be formulated for parenteral administration (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intraventricular, intra-articular, intraperitoneal, or intracranial).

[0276] The compositions disclosed herein (e.g., pharmaceutical compositions) may be suitable for administration to humans. Furthermore, such compositions may be suitable for administration to other animals, such as non-human animals, such as non-human mammals. It is well known that pharmaceutical compositions suitable for administration to humans can be modified to be suitable for administration to various animals, and those skilled in the art and proficient in veterinary pharmacology can design and / or carry out such modifications by simply performing general experiments as necessary. Subjects to which the pharmaceutical compositions of this disclosure are intended include, but are not limited to, humans and / or other primates; mammals, including commercially available mammals such as cattle, pigs, horses, sheep, cats, dogs, mice and / or rats; and birds, including commercially available birds such as poultry, chickens, ducks, geese and / or turkeys.

[0277] target genes This disclosure provides compositions, methods, and systems for regulating the expression of target genes (e.g., endogenous target genes). For example, this specification discloses a complex or system comprising a guide portion and one or more heterogeneous gene effectors that can increase or decrease the activity or expression level of a target gene.

[0278] In some embodiments, the target gene or its regulatory sequence is endogenous to the subject and, for example, present in the subject's genome. In some embodiments, the target gene or its regulatory sequence is not part of a recombinant reporter system.

[0279] In some embodiments, the target gene is exogenous to the host target, and the gene or exogenous gene that targets a pathogen is expressed as a result of a therapeutic intervention such as gene therapy and / or cell therapy. In some embodiments, the target gene is an exogenous reporter gene. In some embodiments, the target gene is an exogenous synthetic gene.

[0280] In some embodiments, the target gene (e.g., an endogenous target gene) is a gene that is overexpressed or underexpressed in a disease or condition. In some embodiments, the target gene is a gene that is overexpressed or underexpressed in a hereditary genetic disorder.

[0281] In some embodiments, the target gene (e.g., endogenous target gene) is a gene that is overexpressed or underexpressed in autoimmune diseases. In some embodiments, the target gene is acute disseminated encephalomyelitis, acute motor axonal neuropathy, Addison's disease, painful steatosis, adult-onset Still's disease, alopecia areata, ankylosing spondylitis, anti-glomerular basement membrane nephritis, anti-neutrophil cytoplasmic antibody-associated vasculitis, anti-N-methyl-D-aspartate receptor encephalitis, antiphospholipid syndrome, anti-synthetic enzyme syndrome, aplastic anemia, autoimmune angioedema, autoimmune encephalitis, autoimmune intestinal disease, autoimmune hemolytic anemia, autoimmune hepatitis, autoimmune inner ear disease, autoimmune Autoimmune lymphoproliferative syndrome, autoimmune neutropenia, autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune polyendocrine syndrome, autoimmune polyendocrine syndrome type 2, autoimmune polyendocrine syndrome type 3, autoimmune progesterone dermatitis, autoimmune retinopathy, autoimmune thrombocytopenic purpura, autoimmune thyroiditis, autoimmune urticaria, autoimmune uveitis, Barlow concentric sclerosis, Behçet's disease, Vickerstaff's encephalitis, bullous pemphigoid, celiac disease Chronic fatigue syndrome, chronic inflammatory demyelinating polyneuropathy, Churg-Strauss syndrome, bullous pemphigoid, Cogan syndrome, cold agglutinin disease, complex regional pain syndrome, Crest syndrome, Crohn's disease, herpetiform dermatitis, dermatomyositis, type 1 diabetes, lupus discoid, endometriosis, enthesitis, enthesitis-associated arthritis, eosinophilic esophagitis, eosinophilic fasciitis, acquired epidermolysis bullosa, erythema nodosum, essential mixed cryoglobulinemia, Evans syndrome, Felty syndrome, fibromyalgia, gastritis, bullous pemphigoid of pregnancy, giant cell arteritis, Goodpasture syndrome, Graves' disease, Graves' ophthalmopathy, Guillain-Barré syndrome, Hashimoto's encephalopathy, Hashimoto's thyroiditis, Henoch-Schönlein purpura, hidradenitis suppurativa, idiopathic dilated cardiomyopathy, idiopathic inflammatory demyelinating disease, IgA nephropathy, IgG4-related systemic disease, inclusion body myositis, inflammatory bowel disease (IBD), intermediate uveitis, interstitial cystitis, juvenile arthritis, Kawasaki disease, Lambert-Eaton myasthenic syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosing, woody conjunctivitis, linear IgA disease, lupus nephritis, lupus vasculitis, Lyme disease, Meniere's disease, microscopic colitis, microscopic polyangiitis, mixed connective tissue disease, Mollen's ulcer, morbid scleroderma,Mucha-Habermann disease, multiple sclerosis, myasthenia gravis, myocarditis, myositis, neuromyelitis optica, neuromyotonia, ocular clonus-myoclonus syndrome, optic neuritis, Ord's thyroiditis, relapsing rheumatoid arthritis, cerebellar degeneration associated with malignant tumors, Parry-Romberg syndrome, Personage-Turner syndrome, streptococcal-associated childhood autoimmune neuropsychiatric disorders, pemphigus vulgaris, pernicious anemia, acute lichenoid rash, POEMS syndrome, polyarteritis nodosa, polymyalgia rheumatica, polymyositis, post-myocardial infarction syndrome, post-pericardiotomy syndrome, primary biliary cirrhosis, primary immunodeficiency, primary sclerosing cholangitis, progressive inflammatory neuropathy, This gene is overexpressed or underexpressed in psoriasis, psoriatic arthritis, pure red cell aplasia, pyoderma gangrenosum, Raynaud's phenomenon, reactive arthritis, relapsing polychondritis, restless legs syndrome, retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, rheumatic vasculitis, sarcoidosis, Schnitzler syndrome, scleroderma, Sjögren's syndrome, generalized rigidity syndrome, subacute bacterial endocarditis, Susac syndrome, Sydenham's chorea, sympathetic ophthalmitis, systemic lupus erythematosus, systemic scleroderma, thrombocytopenia, Tolosa-Hunt syndrome, transverse myelitis, ulcerative colitis, undifferentiated connective tissue disease, urticaria, urticarial vasculitis, vasculitis, or vitiligo.

[0282] In some embodiments, the target gene (e.g., endogenous target gene) is a gene that is overexpressed or underexpressed in cancer, and the cancers include, for example, acute leukemia, astrocytoma, biliary tract cancer (cholangiocarcinoma), bone cancer, breast cancer, brainstem glioma, bronchioloalveolar cell lung cancer, adrenal cancer, anal cancer, bladder cancer, endocrine cancer, esophageal cancer, head and neck cancer, kidney cancer, parathyroid cancer, penile cancer, pleural / peritoneal cancer, salivary gland cancer, small intestine cancer, thyroid cancer, ureteral cancer, urethral cancer, cervical cancer, endometrial cancer, fallopian tube cancer, renal pelvis cancer, vaginal cancer, vulvar cancer, cervix cancer, chronic leukemia, colorectal cancer, colorectal cancer, cutaneous melanoma, ependymoma, epithelioid tumor, and Ewing's cancer. Examples include sarcomas, gastric cancer, glioblastoma, glioblastoma multiforme, glioma, hematological malignancies, hepatocellular carcinoma (liver carcinoma), hepatoma, Hodgkin's disease, intraocular melanoma, Kaposi's sarcoma, lung cancer, lymphoma, medulloblastoma, malignant melanoma, meningioma, mesothelioma, multiple myeloma, muscle cancer, central nervous system (CNS) neoplasms, neurocellular carcinoma, small cell lung cancer, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pediatric malignancies, pituitary adenoma, prostate cancer, rectal cancer, renal cell carcinoma, soft tissue sarcoma, schwannoma, skin cancer, spinal axial tumor, squamous cell carcinoma, gastric cancer, synovial sarcoma, testicular cancer, uterine cancer, and tumors, as well as their metastases and refractory forms, and combinations thereof.

[0283] In some embodiments, the target gene (e.g., endogenous target gene) is a gene associated with differentiation, such as SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30, CD50, AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA -1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1α / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, N FIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activator, S TAT inhibitor, STAT3, STAT4, STAT5a, STAT6, TSC22, DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MY F-5, myocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, S OX2, SOX9, SOX11, STAT activator, STAT inhibitor, STAT1, STAT3, TBX18, Twist-1, Twist-2, Brachiuri, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GAT A-3, GBX2, Goosecoid, HES-1, HNF-3α / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor,NFkB1, NFkB2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activator, STAT inhibitor, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, ZNF281, KLF2, KLF4, c-Maf , c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, TBX18, ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless , HNF-4α / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX 1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT Activator, STAT Inhibitor, STAT3, SUZ12, TCF-3 / E2A, TCF7 / TCF1, Antilogin R / NR3C4, AP-2γ, β-Catenin, β-Catenin Inhibitor, Brakiuri, CREB, ERα / NR3A1, ERβ / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GL I-2, GLI-3, HIF-1α / HIF1A, HIF-2α / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, またはZEB1である。 、

[0284] In some embodiments, the expression level and / or epigenetic level (e.g., methylation level) of a target gene can be modified (e.g., upregulated or downregulated) downstream genes (e.g., one or more downstream genes) of the target gene by regulating the expression level and / or epigenetic level (e.g., methylation level) of the target gene in target cells (e.g., muscle cells). In some cases, the target gene may be encoded by a D4Z4 repeat array (for example, if the target gene is DUX4), and therefore, downstream genes whose expression is modified (e.g., downregulated) include, but are not limited to, ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, DEFB103, ZFN217, RNASEL, EIF2AK2, BMP2, SP1, P21, MYC, MURF1, ATROGIN1, CRYM, PRAMEF1, RFPL2, KHDC1, SPRYD5, TPRX1, HSPA2, FGFR3, SLC2A14, ID2, PVRL3, SFRS2B, THOC4, ZNHIT6, DBR1, TFIP11, FBXO33, USP29, TRIM23, SLC34A2, CSAG3 and / or PNMA6B.

[0285] In some embodiments, the regulation of the expression level and / or epigenetic level (e.g., methylation level) of a target gene in target cells (e.g., muscle cells) can exert an effect on apoptosis in target cells (e.g., muscle cells). In some cases, such regulation of a target gene can reduce stress in target cells. For example, regulation of a target gene (e.g., DUX4) can downregulate one or more stress-related markers in target cells. Examples of one or more stress-related markers include ACTH, glucocorticoid receptor, CRHR-1 / 2, POMC, prolactin, arginine vasopressin receptor V1a, superoxide dismutase 1, superoxide dismutase 2, peroxiredoxin 3, CCR5, iNOS, eNOS, heme oxygenase 2, cyclooxygenase 2, HSP27, HSP40, HSP60, HSP70, HSP70i, HSP90, HSP110, GRP78 / Examples of stress-related markers disclosed herein include, but are not limited to, BIP, AIF, Annexin II, Annexin IV, Caspase 1, Caspase 2, Caspase 3, Caspase 6, Cytokeratin, E-Cadherin, Annexin V, Caspase 5, Caspase 7, Caspase 8, Caspase 9, Caspase 10, BAD, BAX, BAK, BCL2, BID, PARP-1, NOXA, PUMA, RIPK3, RIPK1, FADD, APAF1, DFF40, DFF45, and ROCK. One or more stress-related markers disclosed herein may also be markers of apoptosis. While we do not wish to be bound by any theory, higher expression levels of one or more stress-related markers described herein in cells may indicate higher levels of apoptosis in those cells.

[0286] In some embodiments, the heterologous gene effector is derived from a gene product that is a transcription factor of hematopoietic stem cells. In some embodiments, the target gene is a transcription factor of mesenchymal stem cells. In some embodiments, the target gene is a transcription factor of embryonic stem cells. In some embodiments, the target gene is a transcription factor of induced pluripotent stem cells (iPSCs). In some embodiments, the target gene is a transcription factor of epithelial stem cells. In some embodiments, the target gene is a transcription factor of cancer stem cells.

[0287] In some embodiments, the target gene is an aging-related gene. In some embodiments, the target gene is an aging-related protein. In some embodiments, the target gene is a drug target.

[0288] In some embodiments, the target gene (e.g., endogenous target gene) is a cancer-related gene. Examples of cancer-related genes include A1CF, ABI1, ABL1, ABL2, ACKR3, ACSL3, ACSL6, ACVR1, ACVR2A, AFDN, AFF1, AFF3, AFF4, AKAP9, AKT1, AKT2, AKT3, ALDH2, ALK, AMER1, ANK1, APC, APOBEC3B, AR, ARAF, ARHGAP26, ARHGAP5, ARHGEF10, ARHGEF10L, ARHGEF12, ARID1A, ARID1B, ARID2, ARNT, ASPSCR1, ASXL1, ASXL2, ATF1, ATIC, ATM, ATP1A1, ATP2B3, ATR, ATRX, AXIN1, AXIN2, B2M, BAP1, BARD1, BAX, BAZ1A, BCL10, BCL11A, BCL11B, BCL2, BCL2L12, BCL3, BCL6, BC L7A, BCL9, BCL9L, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BIRC6, BLM, BMP5, BMPR1A, BRAF, BRCA1, BRCA2, BRD3, BRD4, BRIP1, BTG1, BTK, BUB1B, C15or f65, CACNA1D, CALR, CAMTA1, CANT1, CARD11, CARS, CASP3, CASP8, CASP9, CBFA2T3, CBFB, CBL, CBLB, CBLC, CCDC6, CCNB1IP1, CCNC, CCND1, CCND2 , CCND3, CCNE1, CCR4, CCR7, CD209, CD274, CD28, CD74, CD79A, CD79B, CDC73, CDH1, CDH10, CDH11, CDH17, CDK12, CDK4, CDK6, CDKN1A, CDKN1B, CDK N2A, CDKN2C, CDX2, CEBPA, CEP89, CHCHD7, CHD2, CHD4, CHEK2, CHIC2, CHST11, CIC, CIITA, CLIP1, CLP1, CLTC, CLTCL1, CNBD1, CNBP, CNOT3, CNTNA P2, CNTRL, COL1A1, COL2A1, COL3A1, COX6C, CPEB3, CREB1, CREB3L1, CREB3L2, CREBBP, CRLF2, CRNKL1, CRTC1, CRTC3, CSF1R, CSF3R, CSMD3, CTCF,CTNNA2, CTNNB1, CTNND1, CTNND2, CUL3, CUX1, CXCR4, CYLD, CYP2C8, CYSLTR 2. DAXX, DCAF12L2, DCC, DCTN1, DDB2, DDIT3, DDR2, DDX10, DDX3X, DDX5, DDX 6. DEK, DGCR8, DICER1, DNAJB1, DNM2, DNMT3A, DROSHA, DUX4L1(また, Dux4), E BF1, ECT2L, EED, EGFR, EIF1AX, EIF3E, EIF4A2, ELF3, ELF4, ELK4, ELL, ELN, E ML4, EP300, EPAS1, EPHA3, EPHA7, EPS15, ERBB2, ERBB3, ERBB4, ERC1, ERCC2 ERCC3, ERCC4, ERCC5, ERG, ESR1, ETNK1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT 1, EXT2, EZH2, EZR, FAM131B, FAM135B, FAM47C, FANCA, FANCC, FANCD2, FANC E, FANCF, FANCG, FAS, FAT1, FAT3, FAT4, FBLN2, FBXO11, FBXW7, FCGR2B, FCRL 4, FEN1, FES, FEV, FGFR1, FGFR1OP, FGFR2, FGFR3, FGFR4, FH, FHIT, FIP1L1 FKBP9, FLCN, FLI1, FLNA, FLT3, FLT4, FNBP1, FOXA1, FOXL2, FOXO1, FOXO3, FO XO4, FOXP1, FOXR1, FSTL3, FUBP1, FUS, GAS7, GATA1, GATA2, GATA3, GLI1, GM PS, GNA11, GNAQ, GNAS, GOLGA5, GOPC, GPC3, GPC5, GPHN, GRIN2A, GRM3, H3F3A H3F3B, HERPUD1, HEY1, HIF1A, HIP1, HIST1H3B, HIST1H4I, HLA-A, HLF, HMG A1, HMGA2, HMGN2P46, HNF1A, HNRNPA2B1, HOOK3, HOXA11, HOXA13, HOXA9, HOX C11, HOXC13, HOXD11, HOXD13, HRAS, HSP90AA1, HSP90AB1, ID3, IDH1, IDH2. IGF2BP2, IGH, IGK, IGL, IKBKB, IKZF1, IL2, IL21R, IL6ST, IL7R, IRF4, IRS4.ISX, ITGAV, ITK, JAK1, JAK2, JAK3, JAZF1, JUN, KAT6A, KAT6B, KAT7, KCNJ5 KDM5A, KDM5C, KDM6A, KDR, KDSR, KEAP1, KIAA1549, KIF5B, KIT, KLF4, KLF6 KLK2, KMT2A, KMT2C, KMT2D, KNL1, KNSTRN, KRAS, KTN1, LARP4B, LASP1, LATS 1. LATS2, LCK, LCP1, LEF1, LEPROTL1, LHFPL6, LIFR, LMNA, LMO1, LMO2, LPP, L RIG3, LRP1B, LSM14A, LYL1, LZTR1, MACC1, MAF, MAFB, MALAT1, MALT1, MAML2 MAP2K1, MAP2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MAX, MB21D2, MDM2, MDM4 MDS2, MECOM, MED12, MEN1, MET, MGMT, MITF, MLF1, MLH1, MLLT1, MLLT10, ML LT11, MLLT3, MLLT6, MN1, MNX1, MPL, MRTFA, MSH2, MSH6, MSI2, MSN, MTCP1, MT OR, MUC1, MUC16, MUC4, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, MYH11, MYH9, MY O5A, MYOD1, N4BP2, NAB2, NACA, NBEA, NBN, NCKIPSD, NCOA1, NCOA2, NCOA4, N COR1, NCOR2, NDRG1, NF1, NF2, NFATC2, NFE2L2, NFIB, NFKB2, NFKBIE, NIN KX2-1, NONO, NOTCH1, NOTCH2, NPM1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT 5C2, NTHL1, NTRK1, NTRK3, NUMA1, NUP214, NUP98, NUTM1, NUTM2B, NUTM2D, O LIG2, OMD, P2RY8, PABPC1, PAFAH1B2, PALB2, PATZ1, PAX3, PAX5, PAX7, PAX8. PBRM1, PBX1, PCBP1, PCM1, PDCD1LG2, PDE4DIP, PDGFB, PDGFRA, PDGFRB, PER 1, PHF6, PHOX2B, PICALM, PIK3CA, PIK3CB, PIK3R1, PIM1, PLAG1, PLCG1, PML.PMS1、PMS2、POLD1、POLE、POLG、POLQ、POT1、POU2AF1、POU5F1、PPARG、PPFIB P1、PPM1D、PPP2R1A、PPP6C、PRCC、PRDM1、PRDM16、PRDM2、PREX2、PRF1、PRKAC A、PRKAR1A、PRKCB、PRPF40B、PRRX1、PSIP1、PTCH1、PTEN、PTK6、PTPN11、PTP N13、PTPN6、PTPRB、PTPRC、PTPRD、PTPRK、PTPRT、PWWP2A、QKI、RABEP1、RAC1、 RAD17、RAD21、RAD51B、RAF1、RALGDS、RANBP2、RAP1GDS1、RARA、RB1、RBM10、 RBM15、RECQL4、REL、RET、RFWD3、RGPD3、RGS7、RHOA、RHOH、RMI2、RNF213、RNF 43, ROBO2, ROS1, RPL10, RPL22, RPL5, RNP1, RSP2, RSP3, RUNX1, RUNX1T1, S100A7, SALL4, SBDS, SDC4, SDHA, SDHAF2, SDHB, SDHC, SDHD, 44444, 44445, 4 4448、SET、SETBP1、SETD1B、SETD2、SETDB1、SF3B1、SFPQ、SFRP4、SGK1、SH2B 3、SH3GL1、SHTN1、SIRPA、SIX1、SIX2、SKI、SLC34A2、SLC45A3、SMAD2、SMAD3、 SMAD4、SMARCA4、SMARCB1、SMARCD1、SMARCE1、SMC1A、SMO、SND1、SNX29、SOC S1、SOX2、SOX21、SPECC1、SPEN、SPOP、SRC、SRGAP3、SRSF2、SRSF3、SS18、SS18 L1、SSX1、SSX2、SSX4、STAG1、STAG2、STAT3、STAT5B、STAT6、STIL、STK11、ST RN、SUFU、SUZ12、SYK、TAF15、TAL1、TAL2、TBL1XR1、TBX3、TCEA1、TCF12、TCF3 、TCF7L2、TCL1A、TEC、TENT5C、TERT、Tet1、Tet2、TFE3、TFEB、TFG、TFPT、TFR C、TGFBR2、THRAP3、TLX1、TLX3、TMEM127、TMPRSS2、TNC、TNFAIP3、TNFRSF14、Examples include, but are not limited to, TNFRSF17, TOP1, TP53, TP63, TPM3, TPM4, TPR, TRA, TRAF7, TRB, TRD, TRIM24, TRIM27, TRIM33, TRIP11, TRRAP, TSC1, TSC2, TSHR, U2AF1, UBR5, USP44, USP6, USP8, VAV1, VHL, VTI1A, WAS, WDCP, WIF1, WNK2, WRN, WT1, WWTR1, XPA, XPC, XPO1, YWHAE, ZBTB16, ZCCHC8, ZEB1, ZFHX3, ZMYM2, ZMYM3, ZNF331, ZNF384, ZNF429, ZNF479, ZNF521, ZNRF3, and ZRSR2.

[0289] In some embodiments, the target gene provided herein may be Dux4. Target polynucleotide sequences that can be targeted (e.g., bound) by the system of this disclosure (e.g., Cas or dCas protein and / or guide nucleic acid molecule) may be present in or adjacent to the D4Z4 repeat array. For example, a guide nucleic acid molecule (e.g., guide RNA) may include (i) a scaffold sequence configured to form a complex with a Cas protein, and (ii) a spacer sequence that exhibits specific binding to the target polynucleotide sequence. Any suitable scaffold sequence can be used to form a complex with a Cas protein. Examples of suitable scaffold sequences are disclosed, for example, in International Publication WO2023 / 168242 (this document is incorporated herein by reference in its entirety). In another example, heterogeneous nucleic acids (e.g., antisense oligonucleotides, small interfering RNAs, ribozymes, etc.) that can function in the absence of Cas / dCas proteins may exhibit specific binding to the target polynucleotide sequence. In another example, non-Cas / dCas proteins (e.g., ZFNs and Talen) can exhibit specific binding to target polynucleotide sequences.

[0290] In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may be at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, and at least 91% of the target polynucleotide sequence or guide nucleic acid molecule spacer sequence may be at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, and at least 91%. Alternatively, it may include a polynucleotide sequence exhibiting at least approximately 91%, at least 92%, at least approximately 92%, at least 93%, at least approximately 93%, at least 94%, at least approximately 94%, at least 95%, at least approximately 95%, at least 96%, at least approximately 96%, at least 97%, at least approximately 97%, at least 98%, at least approximately 98%, at least 99%, or at least approximately 99%, or substantially approximately 100% sequence identity (or the spacer sequence is encoded by such a polynucleotide sequence).In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule is at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, and at least 91% of the target polynucleotide sequence or guide nucleic acid molecule spacer sequence. Alternatively, it may contain a polynucleotide sequence exhibiting at least approximately 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or substantially about 100% sequence identity (or the spacer sequence is encoded by such a polynucleotide sequence).In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule is at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, and at least 91%. It may contain a polynucleotide sequence exhibiting sequence identity of % or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or substantially about 100% (or the spacer sequence is encoded by such a polynucleotide sequence).

[0291] In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule can be selected to reduce Dux4 expression by at least 50% or at least about 50% of the threshold level in muscle cells compared to control cells. In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may be one or more member polynucleotide sequences selected from the group consisting of SEQ ID NOs: 800, 867, 836, 851, 854, 853, 812, 874, 841, 824, 829, 840, 848, 860, 825, 828, 826, 864, 830, 833, and 865 and their complementary sequences, and at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 7 It may contain polynucleotide sequences exhibiting 5%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 91% or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or 100% sequence identity (or the spacer sequence is encoded by such polynucleotide sequences).

[0292] In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule can be selected to reduce the expression of the target gene (e.g., Dux4) in muscle cells (e.g., FSHD muscle cells) by a threshold level (e.g., by a predetermined threshold level) compared to control cells (e.g., control cells that do not have the system of this disclosure for suppressing the target gene). This threshold level may be at least 40% or at least about 40%, at least 50% or at least about 50%, at least 55% or at least about 55%, at least 60% or at least about 60%, at least 65% or at least about 65%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, at least 90% or at least about 90%, at least 95% or at least about 95%, at least 99% or at least about 99%, or substantially about 100%. For example, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to reduce Dux4 expression in muscle cells by at least 50% or at least approximately 50% compared to control cells.In some examples, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule is at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%, at least 85% or at least about 85%, and less than Both may contain polynucleotide sequences exhibiting 90% or at least about 90%, at least 91% or at least about 91%, at least 92% or at least about 92%, at least 93% or at least about 93%, at least 94% or at least about 94%, at least 95% or at least about 95%, at least 96% or at least about 96%, at least 97% or at least about 97%, at least 98% or at least about 98%, at least 99% or at least about 99%, or substantially about 100% sequence identity (or the spacer sequence is encoded by such polynucleotide sequences).

[0293] In some cases, the target polynucleotide sequence or spacer sequence of the guide nucleic acid molecule can be selected to reduce the expression of the target gene (e.g., Dux4) in muscle cells (e.g., FSHD muscle cells) while exhibiting off-target effects below a threshold level (e.g., a given threshold level) compared to control cells (e.g., control cells that do not have the system of this disclosure for suppressing the target gene). Off-target effects can be measured in vitro, ex vivo, in vivo, or in silico. For example, in an in silico off-target predictive analysis, off-target sites may be identified if similar genomic sequences with an edit distance less than a given, i.e., less than the “edit distance” (e.g., less than the sum of identified mismatches, deletions, and insertions) (e.g., edit distance less than 3) are found in (i) the target cells of interest (e.g., muscle cells such as adult skeletal muscle cells), (ii) in an inactive or “quiescent” chromatin region of the target cell, and / or (iii) not in a promoter region or exon region of the target cell's chromosome (e.g., a region that is neither intergeneric nor intronic). Based on the identification of such off-target sites, off-target activity levels (e.g., predicted off-target activity levels) can be measured (e.g., using one or more methods described, for example, Muhammad Naeem et al., Cells, 9(7), 16008 (2020) (this document is incorporated herein by reference in its entirety)). In some examples, the spacer sequences of the target polynucleotide sequence or guide nucleic acid molecule can be selected to have approximately 5 or fewer, approximately 4 or fewer, approximately 3 or fewer, approximately 2 or fewer, or approximately 1 or fewer off-target sites, or they can be selected so that off-target sites are substantially unidentified.In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule can be selected to exhibit off-target activity levels of approximately 30% or less, approximately 25% or less, approximately 20% or less, approximately 15% or less, approximately 14% or less, approximately 13% or less, approximately 12% or less, approximately 11% or less, approximately 10% or less, approximately 9% or less, approximately 8% or less, approximately 7% or less, approximately 6% or less, approximately 5% or less, approximately 4% or less, approximately 3% or less, approximately 2% or less, approximately 1% or less, or substantially approximately 0%. In some examples, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule is at least 50% or at least about 50%, at least 60% or at least about 60%, at least 70% or at least about 70%, at least 75% or at least about 75%, at least 80% or at least about 80%,...

Claims

1. A system for regulating the abnormal expression of target genes in muscle cells, A heterologous polypeptide containing a nuclease, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule exhibits specific binding to target polynucleotide sequences present in or adjacent to the D4Z4 repeat array within the muscle cell. Once the complex is formed, it can bind to the target polynucleotide sequence and modify the expression level and / or methylation level of the target gene within the muscle cell. The target gene is located within the D4Z4 repeat array. (1) The guide nucleic acid molecule includes a spacer sequence exhibiting the specific binding, and the spacer sequence includes a polynucleotide sequence exhibiting at least about 80%, at least about 90%, or at least about 95% sequence identity with one or more member polynucleotide sequences listed in Table 4, or is encoded by such polynucleotide sequence, and the spacer sequence may also include a polynucleotide sequence exhibiting at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851, or is encoded by such polynucleotide sequence. (2) The system is configured to reduce the expression level of the target gene in the muscle cells by at least about 80% compared to control cells. (3) The system is configured to modify the expression level and / or methylation level of the target gene in the muscle cell while exhibiting off-target effects below a predetermined threshold in the muscle cell, (4) The system is configured to continuously regulate the expression level and / or methylation level of downstream genes of the target gene for at least about 5 days, (5) The system is configured to normalize the apoptosis level of the muscle cells to a level similar to that of healthy muscle cells. (6) The system is configured to improve the survival of the muscle cells in vitro or in vivo, (7) The system is configured to have minimal effect on the expression profile of at least one muscle cell-specific gene different from the target gene in the muscle cell, and / or (8) The system is configured to have minimal effect on at least one sign of a health condition in the subject treated by the system. system.

2. The system according to claim 1, wherein, after the complex is formed, the modified expression level and / or methylation level of the target gene in the muscle cells persists for at least about two days.

3. The system according to claim 2, wherein the modified expression level and / or methylation level of the target gene persists for at least about 3 days, at least about 4 days, at least about 5 days, at least about 6 days, at least about 1 week, at least about 2 weeks, at least about 2 weeks, at least about 4 weeks, or at least about 2 months.

4. The system according to claim 2 or 3, wherein the modified expression level and / or methylation level of the target gene persists for at least about 17 days.

5. The system according to claim 2 or 3, wherein the modified expression level and / or methylation level of the target gene persists for at least about 18 days.

6. The system according to any one of claims 1 to 5, wherein the muscle cells are present in a subject having facioscapulohumeral muscular dystrophy (FSHD) or a subject suspected of having facioscapulohumeral muscular dystrophy (FSHD).

7. The system according to any one of claims 1 to 6, wherein the target gene is Dux4.

8. The system according to any one of claims 1 to 7, wherein the length of the nuclease is approximately 800 amino acids or less.

9. The system according to any one of claims 1 to 8, wherein the length of the nuclease is approximately 750 amino acids or less.

10. The system according to any one of claims 1 to 9, wherein the nuclease is Un1Cas12f1 or a modified variant thereof.

11. The system according to any one of claims 1 to 10, wherein the nuclease comprises an amino acid sequence having at least about 80%, at least about 90%, at least about 95%, or at least about 99% identity with the polypeptide sequence of SEQ ID NO:

43.

12. The system according to any one of claims 1 to 10, wherein the nuclease comprises an amino acid sequence having at least about 80%, at least about 90%, at least about 95%, or at least about 99% identity with the polypeptide sequence of SEQ ID NO:

44.

13. The system according to any one of claims 1 to 12, wherein the heterogeneous polypeptide further comprises a transcription regulator.

14. The system according to claim 13, wherein the transcription regulator comprises at least one methyltransferase.

15. The system according to claim 14, wherein the transcription factor comprises at least one DNA methyltransferase (DNMT).

16. The system according to claim 15, wherein the transcription factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).

17. The system according to claim 14, wherein the transcription factor comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof.

18. The system according to claim 14, wherein the transcription factor comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof.

19. The system according to claim 14, wherein the transcription factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant thereof.

20. The system according to claim 13, wherein the transcription factor comprises DNMT-L (or DNMT3L) or KRAB or a variant thereof.

21. The system according to claim 13, wherein the transcription factor comprises KRAB or a variant thereof.

22. The system according to claim 13, wherein the transcription factor comprises a plurality of different transcription factors.

23. The system according to any one or more of claims 1 to 22, wherein modification of the expression level and / or methylation level of the target gene causes downregulation of a downstream gene of the target gene, and the downstream gene includes one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

24. The system according to any one of claims 1 to 23, wherein modification of the expression level and / or methylation level of the target gene causes downregulation of an apoptosis marker in the muscle cells.

25. The system according to claim 24, wherein the apoptosis marker includes caspase 3.

26. The system according to claim 1, wherein the complex causes modification of the expression level of the target gene within the muscle gene.

27. The system according to claim 26, wherein the modification of the expression level causes downregulation of the target gene.

28. The system according to any one of claims 1 to 27, wherein the complex causes modification of the methylation level of the target gene within the muscle gene.

29. The system according to claim 28, wherein the modification of the methylation level causes downregulation of the target gene.

30. The system according to any one of claims 1 to 29, wherein the nuclease is an inactive nuclease.

31. A composition comprising the system described in any one of the preceding claims.

32. A viral vector comprising the system described in any one of the preceding claims.

33. The viral vector according to claim 32, comprising adeno-associated virus (AAV), retrovirus, lentivirus, poxvirus, or adenovirus.

34. The viral vector according to claim 33, wherein the AAV comprises AAV serotype RH74 AAV.

35. A method for regulating the abnormal expression of a target gene in muscle cells, (a) A step of contacting muscle cells with a complex or system comprising (i) a heterologous polypeptide containing a nuclease and (ii) a guide nucleic acid molecule that exhibits specific binding to a target polynucleotide sequence present in or adjacent to a D4Z4 repeat array within muscle cells, and (b) A step of attaching a target gene to the complex or system after the contact step to modify the expression level and / or methylation level of the target gene in the muscle cells. Includes, The target gene is located within the D4Z4 repeat array. (1) The guide nucleic acid molecule includes a spacer sequence exhibiting the specific binding, and the spacer sequence includes a polynucleotide sequence exhibiting at least about 80%, at least about 90%, or at least about 95% sequence identity with one or more member polynucleotide sequences listed in Table 4, or is encoded by such polynucleotide sequence, and the spacer sequence may also include a polynucleotide sequence exhibiting at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequence of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851, or is encoded by such polynucleotide sequence. (2) The complex or system is configured to reduce the expression level of the target gene in the muscle cells by at least about 80% compared to control cells. (3) The complex or system is configured to modify the expression level and / or methylation level of the target gene in the muscle cell while exhibiting off-target effects below a predetermined threshold in the muscle cell, (4) The complex or system is configured to continuously regulate the expression level and / or methylation level of downstream genes of the target gene for at least about 5 days, (5) The complex or system is configured to normalize the apoptosis level of the muscle cells to a level equivalent to that of healthy muscle cells. (6) The complex or system is configured to improve the survival of the muscle cells in vitro or in vivo, (7) The complex or system is configured to have minimal effect on the expression profile of at least one muscle cell-specific gene different from the target gene in the muscle cell, and / or (8) The complex or system is configured to have minimal effect on at least one sign of a health condition in the subject treated by the complex or system. method.

36. The method according to claim 35, wherein, after the complex is formed, the modified expression level and / or methylation level of the target gene in the muscle cells persists for at least about two days.

37. The method according to claim 35 or 36, wherein the modified expression level and / or methylation level of the target gene persists for at least about 3 days, at least about 4 days, at least about 5 days, at least about 6 days, at least about 1 week, at least about 2 weeks, at least about 2 weeks, at least about 4 weeks, or at least about 2 months.

38. The method according to claim 36, wherein the modified expression level and / or methylation level of the target gene is maintained for at least about 17 days.

39. The method according to claim 36, wherein the modified expression level and / or methylation level of the target gene is maintained for at least about 18 days.

40. The method according to any one of claims 35 to 39, wherein the contact step comprises injecting a composition comprising the complex into a subject who requires regulation of abnormal expression of a target gene in the muscle cells, the subject being a subject having facioscapulohumeral muscular dystrophy (FSHD) or a subject suspected of having facioscapulohumeral muscular dystrophy (FSHD).

41. The method according to any one of claims 35 to 40, wherein the target gene is Dux4.

42. The method according to any one of claims 35 to 41, wherein the length of the nuclease is about 800 amino acids or less.

43. The method according to claim 42, wherein the length of the nuclease is approximately 750 amino acids or less.

44. The method according to any one of claims 35 to 43, wherein the nuclease is Un1Cas12f1 or a modified variant thereof.

45. The method according to any one of claims 35 to 44, wherein the nuclease comprises an amino acid sequence having at least about 80%, at least about 90%, at least about 95%, or at least about 99% identity with the polypeptide sequence of SEQ ID NO:

43.

46. The method according to any one of claims 35 to 45, wherein the nuclease comprises an amino acid sequence having at least about 80%, at least about 90%, at least about 95%, or at least about 99% identity with the polypeptide sequence of SEQ ID NO:

44.

47. The method according to any one of claims 35 to 46, wherein the heterologous polypeptide further comprises a transcription regulator.

48. The method according to claim 47, wherein the transcription regulator comprises at least one methyltransferase.

49. The method according to claim 48, wherein the transcription factor comprises at least one DNA methyltransferase (DNMT).

50. The method according to claim 49, wherein the transcription factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).

51. The method according to claim 49, wherein the transcription regulator comprises (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof.

52. The method according to claim 49, wherein the transcription factor comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant thereof.

53. The method according to claim 49, wherein the transcription factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant thereof.

54. The method according to claim 47, wherein the transcription factor comprises DNMT-L (or DNMT3L) or KRAB or a variant thereof.

55. The method according to claim 47, wherein the transcription regulator comprises KRAB or a variant thereof.

56. The method according to claim 47, wherein the transcription factor comprises a plurality of different transcription factors.

57. The method according to any one of claims 35 to 56, wherein modification of the expression level and / or methylation level of the target gene results in downregulation of a downstream gene of the target gene, the downstream gene comprising one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

58. The method according to any one of claims 35 to 57, wherein the expression level and / or methylation level of the target gene are modified, thereby causing downregulation of an apoptosis marker in the muscle cells.

59. The method according to claim 58, wherein the apoptosis marker includes caspase 3.

60. The method according to any one of claims 35 to 59, wherein the complex causes modification of the expression level of the target gene within the muscle gene.

61. The method according to any one of claims 35 to 60, wherein the modification of the expression level causes downregulation of the target gene.

62. The method according to any one of claims 35 to 61, wherein the complex causes modification of the methylation level of the target gene within the muscle gene.

63. The method according to claim 62, wherein the modification of the methylation level causes downregulation of the target gene.

64. The method according to any one of claims 35 to 63, wherein the nuclease is an inactive nuclease.

65. A system, composition, or vector according to any one of claims 1 to 30, a composition according to claim 31, or a vector according to any one of claims 32 to 34, for use as a pharmaceutical, wherein the pharmaceutical is, for example, a pharmaceutical for suppressing, improving or treating cancer, autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

66. A system according to any one of claims 1 to 30, a composition according to claim 31, or a vector according to any one of claims 32 to 34 for use in intracellular gene editing in vitro or in vivo.

67. A method for providing cells with a system for regulating the abnormal expression of a target gene, comprising the step of introducing into cells the system according to any one of claims 1 to 30, the composition according to claim 31, or the vector according to any one of claims 32 to 34, wherein the cells are cells from a subject such as a human having cancer, an autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).

68. The aforementioned system A heterogeneous polypeptide containing a nuclease having an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. A system, composition, or viral vector according to any one of the preceding claims.

69. The aforementioned system A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterogeneous gene effector contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. A system, composition, or viral vector according to any one of the preceding claims.

70. The aforementioned complex or system A heterogeneous polypeptide containing a nuclease having an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of sequence numbers 43, 44, and 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The method according to any one of the prior claims.

71. The aforementioned complex or system A heterologous polypeptide containing a nuclease operably linked to a heterologous gene effector, A guide nucleic acid molecule configured to form a complex with the aforementioned heterologous polypeptide, Includes, The nuclease contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with SEQ ID NO: 728, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The heterogeneous gene effector contains an amino acid sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with sequence number 727, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The guide nucleic acid molecule includes a spacer sequence encoded by a polynucleotide sequence having 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity with any one of SEQ ID NOs: 800, 867, 836, 851, 874, 841, 830, 833, and 865, approximately 80%, approximately 85%, approximately 90%, approximately 95%, approximately 97%, approximately 98%, or approximately 99% sequence identity, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, or at least 99% sequence identity, or approximately 100% sequence identity. The method according to any one of the prior claims.