Systems and methods for modulating aberrant gene expression - Patents.com

JP2024523394A5Pending Publication Date: 2025-06-13EPICRISPR BIOTECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023577861
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-17
Filing Date
2022-06-16
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Aberrant gene expression, particularly in muscle cells, leads to diseases such as facioscapulohumeral muscular dystrophy (FSHD), and transient modifications are insufficient to treat or cure these conditions.

Method used

A system using a heterologous polypeptide with a nuclease and guide nucleic acid molecule targets the D4Z4 repeat array in muscle cells to modify the expression and methylation levels of target genes, maintaining altered levels for extended periods, typically up to several days or weeks.

Benefits of technology

The system effectively modulates aberrant gene expression in muscle cells, reducing the expression of genes like DUX4 and downstream targets, thereby potentially treating or ameliorating conditions like FSHD by sustaining altered expression and methylation levels for extended durations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present disclosure provides systems, compositions, methods for modulating abnormal expression of a target gene in a cell (e.g., a muscle cell) to treat or ameliorate a disease or condition of a subject (e.g., a muscular dystrophy, such as facioscapulohumeral muscular dystrophy (FSHD)). In some embodiments of any of the systems disclosed herein, upon formation of the complex, the altered expression level and / or methylation level of the target gene in the muscle cell persists for at least about 2 days. In some embodiments of any of the systems disclosed herein, the altered expression level and / or methylation level of the target gene persists for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 63 / 211,791, filed June 17, 2021, which is incorporated by reference herein in its entirety. [Background technology]

[0002] background The abnormal expression of one or more genes can lead to disease or condition.In some cases, the abnormal expression of embryonic transcription factor in muscle cells of a subject can lead to muscular dystrophy.For example, the abnormal expression of transcription factor in muscle cells (for example, the abnormal expression of DUX4 in skeletal muscle cells) can lead to facioscapulohumeral muscular dystrophy (FSHD). Summary of the Invention [Means for solving the problem]

[0003] overview Transiently modifying the abnormal expression of a target gene in a cell may not be sufficient to treat or cure a disease manifested by the abnormal expression of the target gene. Thus, there remains a substantial need for systems and methods for modifying the abnormal expression of a target gene and maintaining the modified expression level of the target gene over a long period of time.

[0004] In one aspect, the disclosure provides a system for regulating aberrant expression of a target gene in a muscle cell, the system comprising: a heterologous polypeptide comprising a nuclease having a length less than or equal to about 900 amino acids; and a guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, the guide nucleic acid molecule exhibiting specific binding to a target polynucleotide sequence at or adjacent to a D4Z4 repeat array in the muscle cell, wherein upon complex formation, the complex is capable of binding to the target polynucleotide sequence resulting in alteration of expression levels and / or methylation levels of the target gene in the muscle cell, the target gene being present within the D4Z4 repeat array.

[0005] In some embodiments of any of the systems disclosed herein, once the complex is formed, the altered expression level and / or methylation level of the target gene in the muscle cell persists for at least about 2 days. In some embodiments of any of the systems disclosed herein, the altered expression level and / or methylation level of the target gene persists for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the systems disclosed herein, the altered expression level and / or methylation level of the target gene persists for at least about 17 days. In some embodiments of any of the systems disclosed herein, the altered expression level and / or methylation level of the target gene persists for at least about 18 days.

[0006] In some embodiments of any of the systems disclosed herein, the muscle cells are present in a subject having or suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the systems disclosed herein, the target gene is Dux4.

[0007] In some embodiments of any of the systems disclosed herein, the nuclease has a length less than or equal to about 800 amino acids. In some embodiments of any of the systems disclosed herein, the nuclease has a length less than or equal to about 750 amino acids.

[0008] In some embodiments of any of the systems disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.

[0009] In some embodiments of any of the systems disclosed herein, the heterologous polypeptide further comprises a transcription regulator. In some embodiments of any of the systems disclosed herein, the transcription regulator comprises at least one methyltransferase. In some embodiments of any of the systems disclosed herein, the transcription regulator comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the systems disclosed herein, the transcription regulator comprises DNMT-A or DNMT-L. In some embodiments of any of the systems disclosed herein, the transcription regulator comprises (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcription regulator comprises (i) DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcription regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises DNMT-L or KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulator comprises a plurality of different transcriptional regulators.

[0010] In some embodiments of any of the systems disclosed herein, alteration of the expression level and / or methylation level of the target gene results in downregulation of a gene downstream of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0011] In some embodiments of any of the systems disclosed herein, altering the expression and / or methylation levels of the target genes results in downregulation of apoptotic markers in muscle cells. In some embodiments of any of the systems disclosed herein, the apoptotic markers include caspase 3.

[0012] In some embodiments of any of the systems disclosed herein, the complex results in an alteration of the expression level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, the alteration of the expression level results in downregulation of the target gene.

[0013] In some embodiments of any of the systems disclosed herein, the complex results in an alteration of the methylation level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, the alteration of the methylation level results in downregulation of the target gene.

[0014] In some embodiments of any of the systems disclosed herein, the nuclease is a non-activated nuclease.

[0015] In another aspect, the present disclosure provides a composition comprising any of the systems disclosed herein.

[0016] In another aspect, the present disclosure provides a viral vector comprising any of the systems or compositions disclosed herein.

[0017] In some embodiments of any of the viral vectors disclosed herein, the viral vector comprises an adeno-associated virus (AAV), a retrovirus, a lentivirus, a poxvirus, or an adenovirus.In some embodiments of any of the viral vectors disclosed herein, the AAV comprises the AAV serotype RH74 AAV.

[0018] In another aspect, the disclosure provides a method for regulating aberrant expression of a target gene in a muscle cell, comprising: (a) contacting the muscle cell with a complex comprising (i) a heterologous polypeptide comprising a nuclease having a length of less than or equal to about 900 amino acids, and (ii) a guide nucleic acid molecule that exhibits specific binding to a target polynucleotide sequence at or adjacent to a D4Z4 repeat array in the muscle cell; and (b) upon contacting, binding of the complex to the target gene results in alteration of expression levels and / or methylation levels of the target gene in the muscle cell, wherein the target gene is present within the D4Z4 repeat array.

[0019] In some embodiments of any of the methods disclosed herein, once the complex is formed, the altered expression level and / or methylation level of the target gene in the muscle cell is sustained for at least about 2 days. In some embodiments of any of the methods disclosed herein, the altered expression level and / or methylation level of the target gene is sustained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the methods disclosed herein, the altered expression level and / or methylation level of the target gene is sustained for at least about 17 days. In some embodiments of any of the methods disclosed herein, the altered expression level and / or methylation level of the target gene is sustained for at least about 18 days.

[0020] In some embodiments of any of the methods disclosed herein, the contacting comprises injecting a composition comprising the complex into a subject in need thereof, the subject having or suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the methods disclosed herein, the target gene is Dux4.

[0021] In some embodiments of any of the methods disclosed herein, the nuclease has a length less than or equal to about 800 amino acids. In some embodiments of any of the methods disclosed herein, the nuclease has a length less than or equal to about 750 amino acids.

[0022] In some embodiments of any of the methods disclosed herein, the nuclease is Un1Cas12f1 or an engineered variant thereof.

[0023] In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43. In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.

[0024] In some embodiments of any of the methods disclosed herein, the heterologous polypeptide further comprises a transcription regulator. In some embodiments of any of the methods disclosed herein, the transcription regulator comprises at least one methyltransferase. In some embodiments of any of the methods disclosed herein, the transcription regulator comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the methods disclosed herein, the transcription regulator comprises DNMT-A or DNMT-L. In some embodiments of any of the methods disclosed herein, the transcription regulator comprises (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcription regulator comprises (i) DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcription regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises DNMT-L or KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulator comprises a plurality of different transcriptional regulators.

[0025] In some embodiments of any of the methods disclosed herein, alteration of the expression level and / or methylation level of the target gene results in downregulation of a gene downstream of the target gene, wherein the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.

[0026] In some embodiments of any of the methods disclosed herein, altering the expression and / or methylation levels of the target gene results in downregulation of an apoptotic marker in the muscle cells. In some embodiments of any of the methods disclosed herein, the apoptotic marker comprises caspase 3.

[0027] In some embodiments of any of the methods disclosed herein, the complex results in altered expression levels of the target gene in the muscle cell.

[0028] In some embodiments of any of the methods disclosed herein, altering the expression level results in downregulation of the target gene.

[0029] In some embodiments of any of the methods disclosed herein, the complex results in an alteration of the methylation level of a target gene in a muscle gene. In some embodiments of any of the methods disclosed herein, the alteration of the methylation level results in downregulation of the target gene.

[0030] In some embodiments of any of the methods disclosed herein, the nuclease is a non-activated nuclease.

[0031] In another aspect, the disclosure provides a system for regulating aberrant expression of a target gene in a muscle cell, the system comprising a heterologous actuator moiety bound to a gene regulator, the heterologous actuator moiety capable of forming a complex with the target gene in the muscle cell, the gene regulator capable of altering the expression level and / or methylation level of the target gene in the muscle cell, wherein upon complex formation, the altered expression level and / or methylation level of the target gene in the muscle cell persists for at least about 2 days.

[0032] In some embodiments of any of the systems disclosed herein, the sustained altered expression and / or methylation level of the target gene is characterized by maintaining at least about 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% of the altered expression and / or methylation level of the target gene.

[0033] In some embodiments of any of the systems disclosed herein, the altered expression and / or methylation levels of the target gene persist for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, or 2 months.

[0034] In some embodiments of any of the systems disclosed herein, the altered expression level of the target gene is a decrease in the expression level of the target gene.

[0035] In some embodiments of any of the systems disclosed herein, the altered methylation level of the target gene is an increase in the degree of methylation of the target gene.

[0036] In some embodiments of any of the systems disclosed herein, the gene regulator comprises an epigenetic regulator. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises a chromatin modifier. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises at least one methyltransferase. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises DNMT-A or DNMT-L. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the epigenetic regulator comprises KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the gene regulator comprises a plurality of different gene regulators.

[0037] In some embodiments of any of the systems disclosed herein, the system further comprises a guide nucleic acid molecule that targets the heterologous actuator moiety to a target gene and enables it to form a complex.

[0038] In some embodiments of any of the systems disclosed herein, the heterologous actuator portion is capable of forming a complex with a first portion of a muscle regulatory gene, and the system further comprises an additional heterologous actuator portion bound to an additional gene regulatory factor, the additional heterologous actuator portion being capable of forming a complex with a second portion of the muscle regulatory gene.

[0039] In some embodiments of any of the systems disclosed herein, the system further comprises an additional guide nucleic acid molecule capable of directing the additional heterologous actuator moiety to a second portion of the muscle regulatory gene.

[0040] In some embodiments of any of the systems disclosed herein, the target gene is a transcription factor.

[0041] In some embodiments of any of the systems disclosed herein, the target gene is within a D4Z4 repeat array. In some embodiments of any of the systems disclosed herein, the target gene encodes DUX4.

[0042] In some embodiments of any of the systems disclosed herein, the target gene is not C9orf72.

[0043] In some embodiments of any of the systems disclosed herein, the muscle cells are skeletal muscle cells.

[0044] In some embodiments of any of the systems disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises an endonuclease. In some embodiments of any of the systems disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises a CRISPR-Cas protein. In some embodiments of any of the systems disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises a dCas protein. In some embodiments of any of the systems disclosed herein, the guide nucleic acid molecule or the additional guide nucleic acid molecule comprises a guide RNA molecule.

[0045] In another aspect, the disclosure provides a method for regulating aberrant expression of a target gene in a muscle cell, the method comprising: (a) contacting the muscle cell with a heterologous actuator moiety bound to a gene regulator, wherein the heterologous actuator moiety is capable of forming a complex with the target gene in the muscle cell, and the gene regulator is capable of altering the expression level and / or methylation level of the target gene in the muscle cell; and (b) upon formation of the complex, sustaining the altered expression level and / or methylation level of the target gene in the muscle cell for at least about 2 days.

[0046] In some embodiments of any of the methods disclosed herein, the sustained altered expression and / or methylation level of the target gene is characterized by maintaining at least about 70%, 75%, 80%, 85%, 90%, 95%, 99%, or 100% of the altered expression and / or methylation level of the target gene.

[0047] In some embodiments of any of the methods disclosed herein, the altered expression and / or methylation levels of the target gene persist for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, or 2 months.

[0048] In some embodiments of any of the methods disclosed herein, the altered expression level of the target gene is a decrease in the expression level of the target gene.

[0049] In some embodiments of any of the methods disclosed herein, the altered methylation level of the target gene is an increase in the degree of methylation of the target gene.

[0050] In some embodiments of any of the methods disclosed herein, the gene regulator comprises an epigenetic regulator. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises a chromatin modifier. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises at least one methyltransferase. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises DNMT-A or DNMT-L. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the epigenetic regulator comprises KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the gene regulator comprises a plurality of different gene regulators.

[0051] In some embodiments of any of the methods disclosed herein, the method further comprises contacting the muscle cell with a guide nucleic acid molecule capable of directing the heterologous actuator moiety to a target gene to form a complex.

[0052] In some embodiments of any of the methods disclosed herein, the heterologous actuator portion is capable of forming a complex with a first portion of a muscle regulatory gene, and the method further comprises contacting the muscle cell with an additional heterologous actuator portion bound to an additional gene regulatory factor, the additional heterologous actuator portion being capable of forming a complex with a second portion of the muscle regulatory gene.

[0053] In some embodiments of any of the methods disclosed herein, the method further comprises contacting the muscle cell with an additional guide nucleic acid molecule capable of directing the additional heterologous actuator portion to a second portion of the muscle regulatory gene.

[0054] In some embodiments of any of the methods disclosed herein, the target gene is a transcription factor.

[0055] In some embodiments of any of the methods disclosed herein, the target gene is in D4Z4 repeat array.In some embodiments of any of the methods disclosed herein, the target gene encodes DUX4.In some embodiments of any of the methods disclosed herein, the target gene is not C9orf72.

[0056] In some embodiments of any of the methods disclosed herein, the muscle cells are skeletal muscle cells.

[0057] In some embodiments of any of the methods disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises an endonuclease. In some embodiments of any of the methods disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises a CRISPR-Cas protein. In some embodiments of any of the methods disclosed herein, the heterologous actuator portion or the additional heterologous actuator portion comprises a dCas protein. In some embodiments of any of the methods disclosed herein, the guide nucleic acid molecule or the additional guide nucleic acid molecule comprises a guide RNA molecule.

[0058] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which shows and describes only exemplary embodiments of the present disclosure. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description should be regarded as illustrative in nature and not restrictive.

[0059] INCORPORATION BY REFERENCE All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference herein.

[0060] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure can be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings, in which: [Brief description of the drawings]

[0061] [Figure 1] Figure 1 provides the different target polynucleotide sequences (e.g., rank #1 to rank #91) between two CpG islands within the D4Z4 repeat array encoding DUX4.

[0062] [Diagram 2] Figure 2 provides regulation of DUX4 expression in a target cell population (e.g., lymphoblasts) by a heterologous actuator moiety bound to a gene regulator (e.g., dCas-KRAB-DNMT3A-DNMT3L) complexed with various guide RNA molecule target polynucleotide sequences (e.g., rank #1 to rank #91) within the D4Z4 repeat array encoding DUX4.

[0063] [Figure 3-1]Figure 3A shows gene expression of DUX4 and DUX4 target genes in immortalized patient-derived human FSHD skeletal myoblasts (SkM) (12ABIC / 12A and 15ABIC / 15A). Gene expression of DUX4 and DUX4 target genes is measured in 12ABIC and 15ABIC undifferentiated cells, 12ABIC and 15ABIC cells after 2 days of differentiation, and 12ABIC and 15ABIC cells after 7 days of differentiation. Each shade of grey on the graph indicates gene expression of a different gene corresponding to the legend on the right. Figure 3B shows the percentage of apoptotic cells in FSHD myoblasts 12ABIC and 15ABIC (right column) compared to their healthy sibling control myoblasts 12UBIC and 15VBIC (left column) after 2 days of differentiation, respectively. White dots on the left image represent apoptotic cells. The graph on the right shows the percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cell cultures after 2 days of differentiation shown in the left image. DAPI staining is used to stain the nuclei. Figure 3C shows the percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. The percentage of apoptotic cells is measured on days 0, 1, 2, and 7 of differentiation. Figure 3D shows the expression of MYHC in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. Myosin heavy chain (MYHC) is a marker for muscle cell differentiation. White dots indicate the expression of MYHC. Figure 3E shows the expression levels of MYOG, MYH2, and MYMK in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells after 7 days of differentiation. MYOG is a myogenic regulator that regulates skeletal muscle differentiation, and MyoMaker (MYMK) is a marker for muscle cell differentiation. DAPI staining is used to stain nuclei. Expression levels of 12ABIC and 15ABIC cells are measured on days 2 and 7 of differentiation. 12A UD and 15A UD: undifferentiated proliferating control myoblasts. Dark grey bars indicate MYOG expression levels, light grey bars indicate MYH2 expression levels, and grey bars indicate MYMK expression levels. [Figure 3-2] Same as above. [Figure 3-3] Same as above. [Diagram 3-4] Same as above. [Figure 3-5] Same as above.

[0064] [Figure 4] Figure 4 depicts the design of multiple gRNAs related to D4Z4 repeat region. Multiple DUX4 targeting gRNAs are designed to extend beyond the DZ4Z repeat region. The positions of DZ4Z repeat region and DUX4 gene relative to each other are shown at the bottom of Figure 4. The newly designed gRNA is shown at the top of Figure 4.

[0065] [Diagram 5] Figure 5 shows the Cas12f effector-modulator vector design. The expression of Cas12f variants, KRAB domain, and DNMT3L domain is under the control of muscle-specific promoter CK8e. The expression of sgRNA spacer sequence with scaffold driven by RNA polymerase III is under the control of human U6g promoter. The vector further contains modified WPRE and polyadenylation regulatory sequence.

[0066] [Figure 6-1]Figure 6A shows the relative expression level of DUX4 in 12ABIC FSHD myoblasts stably expressing Cas12f-KRAB effector-modulator after nucleofection of 78 gRNAs into 12ABIC myoblasts. After nucleofection, cells are cultured in differentiation conditions for 7 days, and then gene expression of DUX4 is measured. The 78 gRNAs tested are listed on the x-axis, and the y-axis represents the relative fold expression of DUX4. The expression level of DUX4 was normalized by the expression of the control gene HPRT1. Figure 6B shows the relative expression level of DUX4 in 12ABIC FSHD myoblasts stably expressing Cas12f-KRAB effector-modulator after nucleofection of 78 gRNAs into 12ABIC myoblasts. After nucleofection, cells are cultured in differentiation conditions for 7 days, and then gene expression of DUX4 and the DUX4 target gene, MBD3L2, is measured. [Figure 6-2] Same as above.

[0067] [Figure 7-1]Figure 7A shows the suppression of DUX4 and DUX4 target genes, DBET / DUX4, MBD3L2, and TRIM48 in immortalized patient-derived FSHD myoblasts transfected with six gRNAs and Cas12f effector-modulator. Cas12f effector-modulator expresses Cas12f variant, KRAB domain, and DNMT-KLa domain. One of the six sgRNAs is a control sgRNA (Empty / trcr) that did not target the DZ4Z repeat region. The expression level of MYOG is measured in cells to assay whether the differentiation potential of DUX4 sgRNA-transfected cells is similar to that of control sgRNA-transfected myoblasts. The expression levels of DUX4, DUX4 target genes, and MYOG are measured 17 days after transfection. Figure 7B shows the suppression of DUX4 and DUX4 target genes, DBET / DUX4, MBD3L2, and TRIM48 in immortalized patient-derived FSHD myoblasts transfected with six gRNAs and Cas12f effector-modulator. Cas12f effector-modulator expresses Cas12f variants, KRAB domain, and DNMT-KLb domain. One of the six sgRNAs is a control sgRNA (empty) that did not target the DZ4Z repeat region. The expression level of MYOG is measured in cells to assay whether the differentiation potential of DUX4 sgRNA-transfected cells is similar to that of control sgRNA-transfected myoblasts. The expression levels of DUX4, DUX4 target genes, and MYOG are measured 18 days after transfection. [Figure 7-2] Same as above.

[0068] [Figure 8-1]Figures 8A and 8B show the apoptosis level of FSHD patient-derived myoblasts transfected with Cas12f effector-modulator and DUX4-targeting gRNA. The percentage of apoptosis-positive cells is measured after 2 days of differentiation after transfection. The image in Figure 8A shows the proportion of apoptotic cells in control 12UBIC cells and 12ABIC cells transfected with Cas12f effector-modulator and DUX4-targeting gRNA. The white dots represent apoptotic cells. The graph in Figure 8B shows the percentage of apoptotic cells measured in the left image, as well as the percentage of apoptotic cells in 12ABIC cells transfected with either DUX4-targeting gRNA or control gRNA that does not target DUX4. DAPI staining is used to stain nuclei. [Figure 8-2] Same as above.

[0069] [Figure 9] Figure 9 shows the workflow of ex vivo FSHD model. The ex vivo model is cultured with immortalized healthy sibling control cells and FSHD skeletal myoblasts, and then the cells are engineered into 3D tissue. The 3D tissue is treated with either control AAV or AAV with Cas12f effector-modulator vector. The 3D tissue is then tested for mechanical force, tetanic force, and fatigue phenotype differences, in addition to measuring 3D tissue morphology and gene expression profile.

[0070] [Figure 10]Figure 10 shows the workflow of the in vivo xenograft model. The in vivo model begins with treating mouse legs with irradiation and TA muscle cardiotoxin to prepare for transplantation of human myoblasts into the mouse legs. After transplantation, the mice are treated with either control AAV or AAV carrying Cas12f effector-modulator vectors. At the designated time points, the mice are euthanized and xenografts and tissue samples are collected for analysis. The collected xenografts are fixed, sectioned, and stained with hematoxylin and eosin. The remaining tissues are used for gene expression assays and to determine AAV tropism within the mice. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0071] Detailed Description Aberrant expression of one or more genes may result in a disease or condition. Aberrant expression can be characterized by abnormally low expression levels of the gene(s). Alternatively, aberrant expression can be characterized by abnormally high expression levels of the gene(s). In some cases, the gene(s) can be genetically modified (e.g., via the action of an endonuclease, e.g., CRISPR-Cas enzyme) to reverse the abnormal expression (e.g., treating Duchenne muscular dystrophy (DMD)). Alternatively, aberrant expression can be transiently modified without genetically modifying such gene(s) of interest, for example, by targeting the gene(s) with a gene effector (e.g., an inactivated CRISPR-Cas enzyme bound to the gene effector). Transiently modifying the abnormal expression of a target gene in a cell may not be sufficient to treat or cure a disease manifested by the abnormal expression of the target gene. Thus, in some embodiments, the present disclosure provides systems and methods for modifying the aberrant expression of a target gene such that the modified expression level of the target gene can be sustained over an extended period of time.

[0072] Altering abnormal expression of target genes

[0073] The present disclosure provides compositions, systems, and methods for regulating abnormal expression of a target gene in a cell (e.g., a muscle cell). For example, the target gene may be present in a D4Z4 repeat array. The target gene may encode at least a portion of DUX4. The compositions, systems, and methods disclosed herein can utilize at least one heterologous polypeptide (e.g., a heterologous actuator moiety, optionally together with a heterologous polynucleotide such as a guide nucleic acid molecule) to modify the expression level and / or epigenetic modification level (e.g., methylation level) of a target gene. For example, the compositions, systems, and methods disclosed herein can utilize a heterologous actuator moiety operably linked (e.g., covalently or non-covalently linked) to a heterologous gene effector or regulator (e.g., a gene actuator, a gene repressor, etc.) to modify the expression level and / or epigenetic modification level of a target gene.

[0074] In some examples, the cell can be a muscle cell. The muscle cells disclosed herein can be any category of muscle cell at any developmental stage. The muscle cell can include undifferentiated muscle cells (e.g., mononuclear cells, e.g., muscle stem cells, muscle satellite cells, myoblasts, etc.). Alternatively or in addition, the muscle cell can include differentiated muscle cells (e.g., multinucleated muscle cells, e.g., myotubes). The muscle cell can be a skeletal muscle cell, a cardiac muscle cell, or a smooth muscle cell. For example, the skeletal muscle cell can be a primary myoblast (e.g., an immortalized primary myoblast cell line). In some examples, the cell can be a non-muscle cell, e.g., a lymphoblast.

[0075] In some examples, the target gene may be present on chromosome 4 of the cell disclosed herein. In some examples, the target gene may be present on chromosome 10 of the cell, for example, on the distal portion of the q (long) arm of chromosome 10 of the cell.

[0076] Prior to modification of the target gene as disclosed herein, abnormal expression of the target gene can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500% or more higher than that of a control cell (e.g., a healthy cell of a healthy subject). Abnormal expression can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1% or less higher than that of a control cell (e.g., a healthy cell of a healthy subject).

[0077] Prior to modification of a target gene as disclosed herein, abnormal expression of the target gene can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99% or more lower than that of a control cell (e.g., a healthy cell of a healthy subject). Abnormal expression can be characterized by an expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or less lower than that of a control cell (e.g., a healthy cell of a healthy subject).

[0078] Prior to modification of the target gene as disclosed herein, abnormal expression of the target gene can be characterized by a duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500% or more longer than that of a control cell (e.g., a healthy cell of a healthy subject). Abnormal expression can be characterized by a duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1% or less longer than that of a control cell (e.g., a healthy cell of a healthy subject).

[0079] Prior to modification of a target gene as disclosed herein, abnormal expression of the target gene can be characterized by a duration of expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99% or more shorter than that of a control cell (e.g., a healthy cell of a healthy subject). Abnormal expression can be characterized by a duration of expression level and / or epigenetic modification level (e.g., methylation level) of the target gene that is at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or less shorter than that of a control cell (e.g., a healthy cell of a healthy subject).

[0080] After modification of the target gene as disclosed herein, the modification of the abnormal expression of the target gene can be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500% or more compared to a control (e.g., without modification). The modification of the abnormal expression of the target gene can be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1% or less compared to a control.

[0081] After modification of a target gene as disclosed herein, the modification of the abnormal expression of the target gene can be characterized by a decrease in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99% or more compared to a control (e.g., without modification). The modification of the abnormal expression of the target gene can be characterized by a decrease in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or less.

[0082] After modification of the target gene as disclosed herein, the modification of the abnormal expression of the target gene can be characterized by an increase in the expression level and / or duration of the epigenetic modification level (e.g., methylation level) of the target gene of at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control (e.g., without modification). The modification of the abnormal expression of the target gene can be characterized by an increase in the expression level and / or duration of the epigenetic modification level (e.g., methylation level) of the target gene of at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less compared to a control.

[0083] After modification of a target gene as disclosed herein, the modification of the abnormal expression of the target gene can be characterized by a decrease in the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more, compared to a control (e.g., without modification). The modification of the abnormal expression of the target gene can be characterized by a decrease in the duration of the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene of at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.

[0084] Following modification of a target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., an aberrantly expressed target gene) may persist for at least about 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 7 days, 2 weeks, 3 weeks, 4 weeks, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 12 months, 2 years, 3 years, 4 years, 5 years, or longer. Following modification of a target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., an aberrantly expressed target gene) may persist for at most about 5 years, 4 years, 3 years, 2 years, 12 months, 11 months, 10 months, 9 months, 8 months, 7 months, 6 months, 5 months, 4 months, 3 months, 2 months, 4 weeks, 3 weeks, 2 weeks, 7 days, 6 days, 5 days, 4 days, 3 days, 2 days, 1 day, or less.

[0085] Following modification of a target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., an aberrantly expressed target gene) may persist for at least about 1 cell division, at least about 2 cell divisions, at least about 3 cell divisions, at least about 4 cell divisions, at least about 5 cell divisions, at least about 6 cell divisions, at least about 7 cell divisions, at least about 8 cell divisions, at least about 9 cell divisions, at least about 10 cell divisions, at least about 15 cell divisions, at least about 20 cell divisions, at least about 25 cell divisions, at least about 30 cell divisions, at least about 40 cell divisions, at least about 50 cell divisions, or at least about 100 cell divisions. Following modification of a target gene as disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., an aberrantly expressed target gene) may persist for at most about 100 cell divisions, at most about 50 cell divisions, at most about 40 cell divisions, at most about 30 cell divisions, at most about 25 cell divisions, at most about 20 cell divisions, at most about 15 cell divisions, at most about 10 cell divisions, at most about 9 cell divisions, at most about 8 cell divisions, at most about 7 cell divisions, at most about 6 cell divisions, at most about 5 cell divisions, at most about 4 cell divisions, at most about 3 cell divisions, at most about 2 cell divisions, or at most about 1 cell division.

[0086] As disclosed herein, non-limiting examples of epigenetic modification may include methylation, acetylation, phosphorylation, ADP-ribosylation, glycosylation, sumoylation, ubiquitination, modification of histone structure (e.g., via ATP hydrolysis-dependent process).For example, epigenetic modification may result in altered methylation level of one or more target genes.

[0087] As disclosed herein, a persistent altered expression level and / or epigenetic modification level (e.g., methylation level) of a target gene can be characterized by maintaining at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the altered expression level and / or methylation level of the target gene. A persistent altered expression level and / or epigenetic modification level (e.g., methylation level) of a target gene can be characterized by maintaining at most about 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% of the altered expression level and / or methylation level of the target gene.

[0088] The systems, compositions, and methods disclosed herein can be used to treat or ameliorate a disease in a subject (eg, a muscular dystrophy, such as facioscapulohumeral muscular dystrophy (FSHD)).

[0089] Heterologous Polypeptides

[0090] The heterologous polypeptides disclosed herein, alone or in combination with one or more co-agents, such as heterologous polynucleotides (e.g., guide nucleic acids), can be configured to specifically bind to a target polynucleotide sequence as disclosed herein to modulate the expression level and / or epigenetic level of a target gene (e.g., D4Z4 repeat array) in a target cell. The target polynucleotide sequence can be present in (e.g., within) the target gene. Alternatively, the target polynucleotide sequence can be adjacent to the target gene. For example, the target polynucleotide sequence can be adjacent to an end (e.g., the 5' end or 3' end) of the target gene. The target polynucleotide sequence can be at least about 5 nucleobases, at least about 10 nucleobases, at least about 20 nucleobases, at least about 30 nucleobases, at least about 40 nucleobases, at least about 50 nucleobases, at least about 100 nucleobases, at least about 150 nucleobases, at least about 200 nucleobases, at least about 250 nucleobases, at least about 300 nucleobases, at least about 400 nucleobases, at least about 500 nucleobases, at least about 1,000 nucleobases, at least about 1,500 nucleobases, at least about 2,000 nucleobases, at least about 3,000 nucleobases, at least about 4,000 nucleobases, or at least about 5,000 nucleobases away from the end of the target gene. The target polynucleotide sequence can be at most about 5,000 nucleobases, at most about 4,000 nucleobases, at most about 3,000 nucleobases, at most about 2,000 nucleobases, at most about 1,500 nucleobases, at most about 1,000 nucleobases, at most about 500 nucleobases, at most about 400 nucleobases, at most about 300 nucleobases, at most about 200 nucleobases, at most about 150 nucleobases, at most about 100 nucleobases, at most about 50 nucleobases, at most about 40 nucleobases, at most about 30 nucleobases, at most about 20 nucleobases, at most about 10 nucleobases, or at most about 5 nucleobases away from the end of the target gene.

[0091] Without wishing to be bound by theory, if the target polynucleotide sequence is not present within a target gene, the target polynucleotide sequence may interact (e.g., via direct or indirect binding) with at least a portion of the target gene (e.g., the promoter sequence of the target gene), such that binding or targeting of the target polynucleotide sequence by at least a heterologous polypeptide (e.g., by a complex comprising a heterologous polypeptide and a heterologous polynucleotide disclosed herein) can target at least a portion of the target gene (e.g., the promoter sequence) and result in modulation of the expression levels and / or epigenetic levels of the target gene in a cell.

[0092] In some examples, the target polynucleotide sequence may comprise multiple target polynucleotide sequences. The multiple target polynucleotide sequences may be present within the target gene. Alternatively, the multiple target polynucleotide sequences may be adjacent to but outside the target gene disclosed herein. In yet another alternative, the multiple target polynucleotide sequences may comprise at least one target polynucleotide sequence within the target gene (e.g., within the D4Z4 repeat domain) and at least one additional target polynucleotide sequence adjacent but outside the target gene. In such a case, targeting both the at least one target polynucleotide sequence and the at least one additional target polynucleotide sequence may produce a greater effect (e.g., a greater degree of modulation of the expression and / or epigenetic level of the target gene) (e.g., at least 0.1-fold, at least 0.5-fold, at least 1-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold, or more) compared to targeting only one of the at least one target polynucleotide sequence and the at least one additional target polynucleotide sequence.

[0093] The heterologous polypeptides disclosed herein may include one or more heterologous gene effectors (e.g., gene effectors heterologous to the cell containing the gene effector and / or another component in the complex of the present disclosure). Heterologous gene effectors may include domains that can modulate or are candidates for modulating the expression of a target gene (e.g., a target endogenous gene), for example, activating, suppressing, up-regulating, down-regulating, or stabilizing the expression or activity level of the gene. Heterologous gene effectors may be heterologous with respect to another component present in the complex, such as a guide moiety (e.g., a nuclease and / or a guide nucleic acid disclosed herein). In some examples, heterologous gene effectors may be heterologous with respect to the host cell into which they are introduced.

[0094] Heterologous effector may be or contain any suitable origin, such as human protein, viral protein, or other protein disclosed herein. Heterologous effector may be or contain a sequence from a protein that is mainly localized in the nucleus, such as a member of the human nuclear proteome. Heterologous effector may be or contain one or more natural amino acid residues. Heterologous effector may be or contain one or more synthetic amino acid residues.

[0095] The heterologous gene effector may be or comprise a mammalian protein. The heterologous gene effector may be or comprise a human protein. The heterologous gene effector may be or comprise a viral protein. The heterologous gene effector may be or comprise a non-human primate protein. The heterologous gene effector may be or comprise a non-human mammalian protein. The heterologous gene effector may be or comprise a non-rodent mammalian protein. The heterologous gene effector may be or comprise a plant protein. The heterologous gene effector may be or comprise a porcine protein. The heterologous gene effector may be or comprise a rabbit protein. The heterologous gene effector may be or comprise a canine protein. The heterologous gene effector may be or comprise an avian protein. The heterologous gene effector may be or comprise a reptilian protein. The heterologous genetic effector may be or may comprise a bacterial protein.The heterologous genetic effector may be or may comprise an archaeal protein.

[0096] For example, the amino acid sequences of the heterologous genetic effectors disclosed herein may not, and need not, be derived from bacterial proteins (e.g., may be derived from Archaea). Without wishing to be bound by theory, a subject in need thereof may be treated with a composition comprising a heterologous genetic effector derived from a non-bacterial protein such that the composition (i) does not induce bacterial irritation in the subject and / or (ii) does not elicit a bacterial immune response in the subject.

[0097] The heterologous actuator portion may include a nuclease (e.g., an endonuclease). For example, the nuclease may be a CRISPR / Cas protein. The nuclease may have a length that is less than a threshold length. The threshold length may be at most about 1,000 amino acids, at most about 950 amino acids, at most about 900 amino acids, at most about 850 amino acids, at most about 800 amino acids, at most about 750 amino acids, at most about 700 amino acids, at most about 650 amino acids, at most about 600 amino acids, at most about 550 amino acids, at most about 500 amino acids, at most about 450 amino acids, at most about 400 amino acids, at most about 350 amino acids, or at most about 300 amino acids. The threshold length can be at least about 300 amino acids, at least about 350 amino acids, at least about 400 amino acids, at least about 450 amino acids, at least about 500 amino acids, at least about 550 amino acids, at least about 600 amino acids, at least about 650 amino acids, at least about 700 amino acids, at least about 750 amino acids, at least about 800 amino acids, at least about 850 amino acids, at least about 900 amino acids, at least about 950 amino acids, or at least about 1000 amino acids.

[0098] Without wishing to be bound by theory, using a nuclease size that is less than the threshold length may have one or more advantages over using a control nuclease with a size greater than the threshold length.When using a delivery vehicle with a limited size (e.g., limited physical size to capture the nuclease, or limited expression cassette size, e.g., viral genome), it can leave enough room (or enough space in the expression cassette) for one or more co-agents, such as one or more gene regulators (e.g., transcription factors) and / or one or more heterologous polynucleotides (e.g., one or more guide nucleic acid molecules).Alternatively or in addition, using a nuclease with a size less than or equal to the threshold size can induce a greater effect on the modulation of the expression level and / or epigenetic level of the target gene compared to the effect of a control nuclease with a size greater than the threshold size on the modulation of the expression level and / or epigenetic level of the target gene.

[0099] In some examples, the degree of modulation (e.g., increase or decrease) of expression levels and / or epigenetic levels of a target gene by a nuclease disclosed herein (e.g., having a size less than or equal to the threshold size) is at least or at most about 0.1-fold, at least or at most about 0.5-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3-fold, at least or at most about 4-fold, at least or at most about 5-fold greater than the degree by a control nuclease (e.g., having a size greater than the threshold size). , at least or up to about 6 times greater, at least or up to about 7 times greater, at least or up to about 8 times greater, at least or up to about 9 times greater, at least or up to about 10 times greater, at least or up to about 15 times greater, at least or up to about 20 times greater, at least or up to about 25 times greater, at least or up to about 30 times greater, at least or up to about 40 times greater, at least or up to about 50 times greater, at least or up to about 60 times greater, at least or up to about 70 times greater, at least or up to about 80 times greater, at least or up to about 90 times greater, or at least or up to about 100 times greater.

[0100] In some examples, the degree of modulation (e.g., increase or decrease) of expression levels and / or epigenetic levels of a target gene by a nuclease disclosed herein (e.g., having a size less than or equal to the threshold size) is at least or at most about 0.1-fold, at least or at most about 0.5-fold, at least or at most about 1-fold, at least or at most about 2-fold, at least or at most about 3-fold, at least or at most about 4-fold, at least or at most about 5-fold greater than the degree by a control nuclease (e.g., having a size greater than the threshold size). , at least or up to about 6 times, at least or up to about 7 times, at least or up to about 8 times, at least or up to about 9 times, at least or up to about 10 times, at least or up to about 15 times, at least or up to about 20 times, at least or up to about 25 times, at least or up to about 30 times, at least or up to about 40 times, at least or up to about 50 times, at least or up to about 60 times, at least or up to about 70 times, at least or up to about 80 times, at least or up to about 90 times, or at least or up to about 100 times longer.

[0101] In some examples, the degree of modulation (e.g., increase or decrease) of expression levels and / or epigenetic levels of a target gene by a nuclease disclosed herein (e.g., having a size less than or equal to the threshold size) is at least or up to about 1 cell division, at least or up to about 2 cell divisions, at least or up to about 3 cell divisions, at least or up to about 4 cell divisions, at least or up to about 5 cell divisions, at least or up to about 6 cell divisions, at least or up to about 7 cell divisions, at least or up to about 8 cell divisions, at least or up to about 9 cell divisions, at least or up to about 10 cell divisions, at least or up to about 12 cell divisions, at least or up to about 14 cell divisions, at least or up to about 16 cell divisions, at least or up to about 18 cell divisions, at least or up to about 19 cell divisions, at least or up to about 20 cell divisions, at least or up to about 21 cell divisions, at least or up to about 22 cell divisions, at least or up to about 23 cell divisions, at least or up to about 24 cell divisions, at least or up to about 25 cell divisions, at least or up to about 26 cell divisions, at least or up to about 27 cell divisions, at least or up to about 28 cell divisions, at least or up to about 29 cell divisions, at least or up to about 30 cell divisions, at least or up to about 31 cell divisions, at least or up to about 32 cell divisions, at least or up to about 33 cell divisions, at least or up to about 34 cell divisions, at least or up to about 35 cell divisions, at least or up to about 36 cell divisions, at least or up to about 37 cell divisions, at least or up to about 38 cell divisions, at least or up to about 39 cell divisions, at least or up to about The cell division may last for at least about 10 cell divisions, at least or up to about 11 cell divisions, at least or up to about 12 cell divisions, at least or up to about 13 cell divisions, at least or up to about 14 cell divisions, at least or up to about 15 cell divisions, at least or up to about 16 cell divisions, at least or up to about 17 cell divisions, at least or up to about 18 cell divisions, at least or up to about 19 cell divisions, at least or up to about 20 cell divisions, at least or up to about 25 cell divisions, at least or up to about 30 cell divisions, at least or up to about 40 cell divisions, at least or up to about 50 cell divisions, or at least about 100 cell divisions. As disclosed herein, cell division may be characterized by the division of a parent cell into two daughter cells having substantially the same genetic material as the parent cell.

[0102] The heterologous gene effector may be or may include the sequence of a chromatin regulator (CR). Chromatin regulators include functional domains of various classes of histones and DNA modifying enzymes (e.g., DNMTs, HATs, HMTs, etc.).

[0103] Heterologous gene effectors can contain two or more domains from a chromatin regulator located, for example, in tandem or separately at the C-terminus, N-terminus, or within the polypeptide sequence.

[0104] In some embodiments, heterologous gene effectors that facilitate heterochromatin formation. Non-limiting examples of proteins that can facilitate heterochromatin formation include HP1α, HP1β, KAP1, KRAB, SUV39H1, and G9a.

[0105] In some embodiments, the heterologous genetic effector modulates histones through methylation. In some embodiments, the heterologous genetic effector modulates histones through acetylation. In some embodiments, the heterologous genetic effector modulates histones through phosphorylation. In some embodiments, the heterologous genetic effector modulates histones through ADP-ribosylation. In some embodiments, the heterologous genetic effector modulates histones through glycosylation. In some embodiments, the heterologous genetic effector modulates histones through sumoylation. In some embodiments, the heterologous genetic effector modulates histones through ubiquitination. In some embodiments, the heterologous genetic effector modulates histones by remodeling histone structure, for example, via an ATP hydrolysis-dependent process.

[0106] In some embodiments, heterologous genetic effectors facilitate the spatial positioning of proteins at or near a target polynucleotide, e.g., transcriptional repressors, transcription factors, histones, etc. In some embodiments, heterologous genetic effectors are useful for manipulating the spatiotemporal organization of genomic DNA and RNA components in the nucleus and / or cytoplasm, e.g., to regulate diverse cellular functions.

[0107] In some embodiments, the heterologous gene effector is derived from a histone acetyltransferase, non-limiting examples of which include the GNAT subfamily, the MYST subfamily, the p300 / CBP subfamily, the HAT1 subfamily, GCN5, PCAF, Tip60, MOZ, MORF, MOF, HBO1, p300, CBP, HAT1, ATF-2, SRC1, and TAFII250.

[0108] In some embodiments, the heterologous gene effector is derived from a histone lysine methyltransferase. Non-limiting examples of histone lysine methyltransferase include EZH subfamily, non-SET subfamily, other SET subfamily, PRDM subfamily, SET1 subfamily, SET2 subfamily, SUV39 subfamily, SYMD subfamily, ASH1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, NSD2, NSD3, PRDM1, PRDM10, PRDM11, PRDM12, These include PRDM13, PRDM14, PRDM15, PRDM16, PRDM2, PRDM4, PRDM5, PRDM6, PRDM7, PRDM8, PRDM9, SET1, SET1L, SET2L, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETDB1, SETDB2, SETMAR, SUV39H1, SUV39H2, SUV420H1, SUV420H2, SYMD1, SYMD2, SYMD3, SYMD4, and SYMD5.

[0109] In some embodiments, the heterologous effector is derived from a component of a chromatin remodeling complex. In some embodiments, the heterologous effector is a component of BAF, such as actin, ARIDA / B, BAF155, BAF170, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRG1 / BRM, INI1, or SS18.

[0110] In some embodiments, the heterologous effector is derived from a component of PBAF, such as actin, ARID2, BAF155, BAF170, BAF180, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRD7, BRG1, or INI1.

[0111] In some embodiments, the heterologous gene effector is derived from a component of the ISWI family chromatin remodeling complex, e.g., ACF subfamily, RSF subfamily, CERF subfamily, CHRAC subfamily, NURF subfamily, NoRC subfamily, WICH subfamily, b-WICH subfamily, ACF1, ATPase, BPTF, CECR2, CHRAC15, CHRAC17, CSB, DEK, MYBBP1A, NM1, RBAP46 / 48, RHII / Gua, RSF1, SAP155, SNF2H, SNF2H / L, SNF2L, TIP5, or WSTF.

[0112] In some embodiments, the heterologous gene effector is derived from a member of the CHD family complex, such as NuRD complex, NuRD-like complex, or CHD complex. In some embodiments, the heterologous gene effector is derived from CHD1 / 2 / 6 / 7 / 8 / 9, CHD3 / 4, CHD5, GATAD2 A / B, GATAD2 B, HDAC1, HDAC2, HDAC2, MBD2 / 3, MTA1 / 2 / 3, MTA3, or RBAP46, RBAP46 / 48.

[0113] In some embodiments, the heterologous gene effector is derived from a component of the INO80 family complex, such as the INO80 complex, Tip60 / p400 complex, SRCAP complex, AMIDA, ARP6, BAF53, BAF53, BAF53A, BRD8, DMAP1, DMAP1, EPC1 / 2, FLJ11730, GAS41, GAS41, IES2, IES6, ING3, INO80, INO80E, MCRS1, MRG15, MRGBP, MRGX, NFRKB, p400, RUVBL1 / 2, RUVBL1 / 2, RUVBL1 / 2, SRCAP, Tip60, TRRAP, UCH37, YL-1, YL-1, YY1, or ZnF-HIT1.

[0114] Heterologous gene effectors can be or include sequences derived from transcriptional regulators (TRs). TR gene effectors contain transcriptional regulatory domains from various families of transcription factors (e.g., KRAB, p65, MED, GTF, etc.).

[0115] Heterologous gene effectors can include a transcription activation domain. Heterologous gene effectors can include two or more tandem transcription activation domains, for example, located at the C-terminus, N-terminus, or within the polypeptide sequence.

[0116] Non-limiting examples of transcription activation domains include GAL4, herpes simplex activation domain VP16, VP64 (tetramer of herpes simplex activation domain VP16), NF-KB p65 subunit, Epstein-Barr virus R transactivator (Rta). In some embodiments, such transcription activation domains are used as controls in the methods of the present disclosure. In some embodiments, such transcription activation domains are used as one heterologous gene effector in a complex that includes at least one additional heterologous gene effector (e.g., a different effector).

[0117] The heterologous genetic effector may include a transcriptional repressor domain. The heterologous genetic effector may include two or more transcriptional repressor domains located, for example, in tandem or separately at the C-terminus, N-terminus, or within the polypeptide sequence.

[0118] Non-limiting examples of transcriptional repressor domains include the KRAB (Kruppel-associated box) domain of Koxl, the Mad mSIN3 interaction domain (SID), and the ERF repressor domain (ERD).In some embodiments, such transcriptional repressor domains are used as controls in the methods of the present disclosure.In some embodiments, such transcriptional repressor domains are used as one heterologous gene effector in a complex that includes at least one additional heterologous gene effector (e.g., a different effector).

[0119] In some embodiments, the heterologous gene effector is derived from a gene product that is a transcription factor.

[0120] In some embodiments, the heterologous gene effector is derived from a gene product that is a hematopoietic stem cell transcription factor. Non-limiting examples of hematopoietic stem cell transcription factors include AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1 alpha / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB , c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activators, STAT inhibitors, STAT3, STAT4, STAT5a, STAT6, and TSC22.

[0121] In some embodiments, the heterologous gene effector is derived from a gene product that is a mesenchymal stem cell transcription factor, non-limiting examples of which include DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, Myocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT activator, STAT inhibitor, STAT1, STAT3, TBX18, Twist-1, and Twist-2.

[0122] In some embodiments, the heterologous gene effector is derived from a gene product that is an embryonic stem cell transcription factor. Non-limiting examples of embryonic stem cell transcription factors include Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3alpha / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NFkB2, Oc These include t-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activators, STAT inhibitors, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, and ZNF281.

[0123] In some embodiments, the heterologous gene effector is derived from a gene product that is an induced pluripotent stem cell (iPSC) transcription factor. Non-limiting examples of iPSC transcription factors include KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, and TBX18.

[0124] In some embodiments, the heterologous gene effector is derived from a gene product that is an epithelial stem cell transcription factor. Non-limiting examples of epithelial stem cell transcription factors include ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4alpha / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p5 3, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT activators, STAT inhibitors, STAT3, SUZ12, TCF-3 / E2A, and TCF7 / TCF1.

[0125] In some embodiments, the heterologous gene effector is derived from a gene product that is a cancer stem cell transcription factor. Non-limiting examples of cancer stem cell transcription factors include androgen R / NR3C4, AP-2 gamma, beta-catenin, beta-catenin inhibitor, brachyury, CREB, ER alpha / NR3A1, ER beta / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI-2, GLI-3, HIF-1 alpha / HIF1A, HIF-2 alpha / EPAS1, HMGA1B, c-Jun, JunB, KL These include F4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB activators, NFkB / IkB inhibitors, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activators, STAT inhibitors, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, and ZEB1.

[0126] In some embodiments, the heterologous gene effector is derived from a gene product that is a cancer-associated transcription factor. Non-limiting examples of cancer-associated transcription factors include ASCL1 / Mash1, ASCL2 / Mash2, ATF1, ATF2, ATF4, BLIMP1 / PRDM1, CDX2, CDX4, DLX5, DNMT1, E2F-1, EGR1, ELF3, Ets-1, FosB / G0S3, FoxC1, FoxC2, FoxF1, GADD153, GATA-2, HMGA2, HMGB1 / HMG-1, HNF-3alpha / FoxA1, HNF-6 / ONECUT1, HSF1, ID1, ID2, JunD, KLF10, KLF12, These include KLF17, LMO2, MEF2C, MYCL1 / L-Myc, NFkB2, Oct-1, p63 / TP73L, Pax3, PITX2, Prox1, RAP80, Rex-1 / ZFP42, RUNX1 / CBFA2, RUNX3 / CBFA3, SALL4, SCL / Tal1, Sirtuin 2 / SIRT2, Smad3, Smad4, Smad5, SOX11, STAT5a / b, STAT5a, STAT5b, TCF7 / TCF1, TORC1, TORC2, TRIM32, TRPS1, and TSC22.

[0127] In some embodiments, the heterologous gene effector is derived from a gene product that is an immune cell transcription factor. Non-limiting examples of immune cell transcription factors include AP-1, Bcl6, E2A, EBF, Eomes, FoxP3, GATA3, Id2, Ikaros, IRF, IRF1, IRF2, IRF3, IRF3, IRF7, NFAT, NFkB, Pax5, PLZF, PU.1, ROR-gamma-T, STAT, STAT1, STAT2, STAT3, STAT4, STAT5, STAT5A, STAT5B, STAT6, T-bet, TCF7, and ThPOK.

[0128] In some embodiments, the heterologous effector is derived from a gene product that is an RNA polymerase associated protein. In some embodiments, the heterologous effector is derived from a transcription factor with a basic domain. In some embodiments, the heterologous effector is derived from a transcription factor with a zinc-coordinating DNA binding domain. In some embodiments, the heterologous effector is derived from a transcription factor with a helix-turn-helix domain. In some embodiments, the heterologous effector is derived from a transcription factor with an alpha-helical DNA binding domain. In some embodiments, the heterologous effector is derived from a transcription factor with an alpha-helix exposed by a beta structure. In some embodiments, the heterologous effector is derived from a transcription factor with an immunoglobulin fold. In some embodiments, the heterologous effector is derived from a transcription factor with a beta hairpin exposed by an alpha / beta scaffold structure. In some embodiments, the heterologous effector is derived from a transcription factor with a beta sheet that binds to DNA. In some embodiments, the heterologous effector is derived from a transcription factor with a beta-barrel DNA binding domain.

[0129] In some embodiments, the heterologous gene effector is derived from a gene product that is a nuclear receptor, such as a nuclear hormone receptor. Non-limiting examples of nuclear hormone receptors include NR0B1, NR0B2, NR1A1, NR1A2, NR1B1, NR1B2, NR1B3, NR1C1, NR1C2, NR1C3, NR1D1, NR1D2, NR1F1, NR1F2, NR1F3, NR1H4, NR1H5, NR1H3, NR1H2, NR1I1, NR1I2, NR1I3, NR2A1, NR2A2, NR2B 1, NR2B2, NR2B3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3A1, NR3A2, NR3B1, NR3B2, NR3B3, NR3C4, NR3C1, NR3C2, NR3C3, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, and NR6A1.

[0130] In some embodiments, the heterologous gene effector is derived from a gene product involved in nucleosome assembly. In some embodiments, the heterologous gene effector is derived from a gene product involved in DNA metabolism. In some embodiments, the heterologous gene effector is derived from a gene product involved in nucleotide metabolism. In some embodiments, the heterologous gene effector is derived from a gene product involved in ribosome biogenesis. In some embodiments, the heterologous gene effector is derived from a gene product involved in protein folding. In some embodiments, the heterologous gene effector is derived from a gene product involved in translation. In some embodiments, the heterologous gene effector is derived from a gene product involved in signal transduction. In some embodiments, the heterologous gene effector is derived from a gene product involved in protein degradation. In some embodiments, the heterologous gene effector is derived from a gene product involved in negative regulation of endopeptidase activity.

[0131] In some embodiments, a heterologous gene effector or gene regulator, as used interchangeably herein, may comprise a polypeptide sequence that exhibits at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or substantially about 100% sequence identity to any of the heterologous gene effector amino acid sequences provided in Table 3. Table 3. Heterologous gene effector amino acid sequences [Table 3-1] [Table 3-2]

[0132] The heterologous polynucleotides disclosed herein may include one or more guide moieties (e.g., one or more guide nucleic acid molecules) to direct a heterologous gene effector to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence. The guide moiety can confer the ability to recognize and specifically bind to a target gene or a target gene regulatory sequence. The guide moiety can be configured to form a complex with a heterologous polypeptide (e.g., a guide nucleic acid complexed with a nuclease, e.g., a CRISPR / Cas protein), and the complex can be configured to exhibit specific binding to a target polypeptide sequence disclosed herein to modify the expression level and / or epigenetic modification level of the target gene.

[0133] The guide portion may comprise a guide nucleic acid. The guide portion may comprise a nuclease and a guide nucleic acid as disclosed herein. The guide portion may comprise a nuclease or a part thereof, such as an endonuclease, such as a heterologous endonuclease. The nuclease may be, for example, a DNA nuclease and / or an RNA nuclease, a modified nuclease that is deficient in nuclease or has reduced nuclease activity compared to a wild-type nuclease, a derivative thereof, a variant thereof, or a fragment thereof. In some embodiments, the guide portion has minimal nuclease activity.

[0134] Any suitable nuclease, fragment or derivative thereof can be used in the guide portion.Suitable nucleases include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases, including type I CRISPR-associated (Cas) polypeptides, type II CRISPR-associated (Cas) polypeptides, type III CRISPR-associated (Cas) polypeptides, type IV CRISPR-associated (Cas) polypeptides, type V CRISPR-associated (Cas) polypeptides, and type VI CRISPR-associated (Cas) polypeptides; zinc finger nucleases (ZFN); transcription activator-like effector nucleases (TALEN); meganucleases; RNA-binding proteins (RBPs); CRISPR-associated RNA-binding proteins; recombinases; flippases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaeal Argonaute (aAgo), and eukaryotic Argonaute (eAgo)); any derivative thereof; any variant thereof and any fragment thereof.

[0135] In some embodiments, the guide portion comprises a DNA nuclease, such as an engineered (e.g., programmable or targetable) DNA nuclease that is nuclease-deficient. In some embodiments, the guide portion comprises a nuclease-null DNA binding protein derived from a DNA nuclease that does not induce transcriptional activation or repression of a target DNA sequence unless present in a complex with one or more heterologous genetic effectors of the present disclosure. In some embodiments, the guide portion comprises a nuclease-null DNA binding protein derived from a DNA nuclease that can induce transcriptional activation or repression of a target DNA sequence (e.g., can be altered or enhanced by the presence of a heterologous genetic effector of the present disclosure).

[0136] In some embodiments, the guide portion comprises an RNA nuclease, such as an engineered (e.g., programmable or targetable) RNA nuclease. In some embodiments, the guide portion comprises a nuclease null RNA binding protein derived from an RNA nuclease that does not induce transcriptional activation or repression of a target RNA sequence unless present in a complex with one or more heterologous genetic effectors of the present disclosure. In some embodiments, the guide portion comprises a nuclease null RNA binding protein derived from an RNA nuclease that can induce transcriptional activation or repression of a target RNA sequence (e.g., can be altered or enhanced by the presence of a heterologous genetic effector of the present disclosure).

[0137] In some embodiments, the guide portion comprises a nucleic acid guided targeting system. In some embodiments, the guide portion comprises a DNA guided targeting system. In some embodiments, the guide portion comprises an RNA guided targeting system. The guide portion can comprise and utilize a guide nucleic acid sequence that facilitates specific binding of, for example, a CRISPR-Cas system (e.g., a nuclease-deficient version thereof, e.g., dCas9) to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence. The binding specificity can be determined by using a guide nucleic acid, for example, a single guide RNA (sgRNA) or a portion thereof. In some embodiments, the compositions and methods of the present disclosure can be used (e.g., targeted) with different target genes (e.g., target endogenous genes) or target gene regulatory sequences by using different sgRNAs.

[0138] Prokaryotic CRISPR-Cas (Clustered regularly interspaced short palindromic repeats-CRISPR-related) systems, such as class II CRISPR-Cas systems, such as Cas9 and Cpfl, can be repurposed as tools to regulate gene expression, epigenome editing, and chromatin looping in the compositions and methods of the present disclosure. Nuclease-deactivated Cas (dCas) proteins complexed with heterologous gene effectors can allow regulation of expression of target genes (e.g., target endogenous genes) adjacent to the site where dCas is bound.

[0139] In some embodiments, the guide portion comprises a CRISPR-associated (Cas) protein or Cas nuclease that functions in a non-naturally occurring CRISPR (clustered regularly interspaced short palindromic repeats) / Cas (CRISPR-associated) system. In bacteria, this system can provide adaptive immunity against foreign DNA.

[0140] In a wide variety of organisms, including various mammals, animals, plants, microorganisms, and yeasts, CRISPR / Cas systems (e.g., modified and / or unmodified) can be utilized as genome engineering tools, or can be modified to target specific binding of engineered proteins as disclosed herein to target loci. CRISPR / Cas systems can include guide nucleic acids, such as guide RNAs (gRNAs), complexed with Cas proteins for targeted regulation of gene expression and / or activity or nucleic acid binding. RNA-guided Cas proteins (e.g., Cas nucleases, e.g., Cas9 nucleases) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner. If Cas proteins possess nuclease activity, they can cleave DNA.

[0141] In some examples, Cas proteins are mutated and / or modified to produce nuclease-deficient proteins or proteins with reduced nuclease activity compared to wild-type Cas proteins. Nuclease-deficient proteins can retain DNA binding ability but may lack or have reduced nucleic acid cleavage activity.

[0142] In some embodiments, the guide portion comprises a Cas protein complexed with a guide nucleic acid, e.g., a guide RNA or a portion thereof. In some embodiments, the guide portion comprises a Cas protein complexed with a single guide nucleic acid, e.g., a single guide RNA (sgRNA). In some embodiments, the guide portion comprises an RNA binding protein (RBP) that optionally complexes with a guide nucleic acid, e.g., a guide RNA (e.g., sgRNA), that can complex with a Cas protein. In some embodiments, the guide portion comprises a nuclease null DNA binding protein derived from a DNA nuclease that can induce activation or repression of transcription of a target DNA sequence. In some embodiments, the guide portion comprises a nuclease null RNA binding protein derived from an RNA.

[0143] In some embodiments, a guide nucleic acid used in the compositions and methods of the disclosure can be, for example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more nucleotides.

[0144] In some embodiments, a guide nucleic acid used in the compositions and methods of the disclosure is at most at most 10, at most 9, at most 8, at most 7, at most 6, at most 5, at most 4, at most 3, at most 2, or at most 1 nucleotide.

[0145] The guide nucleic acid can be a guide RNA or a part thereof.

[0146] Any suitable CRISPR / Cas system can be used. CRISPR / Cas systems can be referred to using a variety of naming systems. CRISPR / Cas systems can be type I, type II, type III, type IV, type V, type VI systems, or any other suitable CRISPR / Cas system. CRISPR / Cas systems, as used herein, can be class 1, class 2, or any other suitable classified CRISPR / Cas system. The determination of class 1 or class 2 can be based on genes encoding effector modules. Class 1 systems generally have a multi-subunit crRNA-effector complex, while class 2 systems generally have a single protein, such as Cas9, Cpfl, C2c1, C2c2, C2c3, or crRNA-effector complex. Class 1 CRISPR / Cas systems can use a complex of multiple Cas proteins to effect regulation. Class 1 CRISPR / Cas systems can include, for example, type I (e.g., type I, IA, IB, IC, ID, IE, IF, IU), type III (e.g., type III, IIIA, IIIB, IIIC, IIID), and type IV (e.g., type IV, IVA, IVB) CRISPR / Cas types. Class 2 CRISPR / Cas systems can use a single large Cas protein to provide regulation. Class 2 CRISPR / Cas systems can include, for example, type II (e.g., type II, IIA, IIB) and type V CRISPR / Cas types. CRISPR systems can be complementary to each other and / or functional units can be in trans to facilitate CRISPR gene localization.

[0147] When the guide portion includes a Cas protein or a derivative thereof, the Cas protein or a derivative thereof can be a class 1 or class 2 Cas protein. The Cas protein can be a type I, type II, type III, type IV, type V Cas protein, or type VI Cas protein. The Cas protein can include one or more domains. Non-limiting examples of domains include guide nucleic acid recognition and / or binding domains, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domains, RNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. The guide nucleic acid recognition and / or binding domains can interact with the guide nucleic acid. The nuclease domains can include catalytic activity for nucleic acid cleavage. The nuclease domains can lack catalytic activity to prevent nucleic acid cleavage. The Cas protein can be a chimeric Cas protein or a fragment thereof fused to other proteins or polypeptides. The Cas protein can be a chimera of various Cas proteins, including, for example, domains from different Cas proteins.

[0148] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Casl, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, CaslOd, Cas10, CaslOd, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (Csx12), and the like. asE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4, Cul966, Cas13a, Cas13b, Cas13c, Cas13d, Cas13X, Cas13Y, and homologs or modified versions thereof.

[0149] In some examples, the Cas protein disclosed herein may not be, and need not be, Cas9 or Cas12a. The Cas protein disclosed herein may have a smaller size compared to Cas9 or Cas12a. The Cas protein disclosed herein may be derived from Un1Cas12f1 (or Cas14a1). For example, the Cas protein disclosed herein may comprise an amino acid sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO:43. In another example, the Cas protein disclosed herein may comprise an amino acid sequence that is at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO: 44. As disclosed herein, SEQ ID NO: 43 encodes the polypeptide sequence of Un1Cas12f1 (or Cas14a1). As disclosed herein, SEQ ID NO: 44 encodes an engineered variant of Un1Cas12f1 with reduced nuclease activity.

[0150] SEQ ID NO: 43 (Un1Cas12f1) [ka]

[0151] SEQ ID NO: 44 (Deactivated nuclease variant of Un1Cas12f1) [ka] [ka]

[0152] Cas protein or its derivative can be derived from any suitable organism. For example, Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, Alicyclobacillus acidocaldarius, Bacillus pseudomycoides、Bacillus selenitireducens、Exiguobacterium sibiricum、Lactobacillus delbrueckii、Lactobacillus salivarius、Microscilla marina、Burkholderiales bacterium、Polaromonas nap hthalenivorans、Polaromonas sp.、Crocosphaera watsonii、Cyanothece sp.、Microcystis aeruginosa、Pseudomonas aeruginosa、Synechococcus sp.、Acetohalobium arabaticum、Ammonifex degensii、Caldicellulosiruptor becscii、Candidatus Desulforudis、Clostridium botulinum、Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, the organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, the organism is Staphylococcus aureus (S. aureus). In some embodiments, the organism is Streptococcus thermophilus (S. thermophilus).

[0153] Cas proteins include, but are not limited to, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp.Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, Proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida.

[0154] The Cas protein as used herein may be a wild type or modified version of a Cas protein. The Cas protein may be an active variant, an inactive variant, or a fragment of a wild type or modified Cas protein. The Cas protein may include amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof, compared to a wild type version of the Cas protein. The Cas protein may be a polypeptide having at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity or similarity to a wild type Cas protein. A Cas protein can be a polypeptide having at most about 5%, at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 70%, at most about 80%, at most about 90%, or at most about 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas protein. A variant or fragment can comprise at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or portion thereof. The variant or fragment can be targeted to a nucleic acid locus in a complex with a guide nucleic acid, but lacks nucleic acid cleavage activity.

[0155] Cas proteins may contain one or more nuclease domains, such as DNase domains. For example, Cas9 proteins may contain RuvC-like nuclease domains and / or HNH-like 20 nuclease domains. The nuclease active forms of Cas9, RuvC and HNH domains, each cleave different strands of double-stranded DNA to create double-stranded breaks in DNA. Cas proteins may contain only one nuclease domain (e.g., Cpfl contains a RuvC domain but lacks an HNH domain). In some embodiments, the nuclease domain is not present. In some embodiments, the nuclease domain is present but inactive or has reduced or minimal activity. In some embodiments, the nuclease domain is present and active.

[0156] One or more nuclease domains of a Cas protein (e.g., RuvC, HNH) may be deleted or mutated so that they are no longer functional or contain reduced nuclease activity. For example, in a Cas protein that contains at least two nuclease domains (e.g., Cas9), if one of the nuclease domains is deleted or mutated, the resulting Cas protein is known as a nickase and can generate a single-stranded break at the CRISPR RNA (crRNA) recognition sequence in double-stranded DNA, but cannot generate a double-stranded break. Such a nickase can cleave a complementary or non-complementary strand, but not both. If all of the nuclease domains of a Cas protein (e.g., both RuvC and HNH nuclease domains in Cas9 protein; RuvC nuclease domain in Cpfl protein) are deleted or mutated, the resulting Cas protein may have reduced or no ability to cleave both strands of double-stranded DNA. An example of a mutation that can convert the Cas9 protein into a nickase is the D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain of S. pyogenes Cas9. H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of S. pyogenes Cas9 can convert Cas9 into a nickase. An example of a mutation that can convert the Cas9 protein into a dead Cas9 is the D10A (aspartate to alanine at position 10 of Cas9) mutation in the RuvC domain and H939A (histidine to alanine at amino acid position 839) or H840A (histidine to alanine at amino acid position 840) in the HNH domain of S. pyogenes Cas9.

[0157] A nuclease-dead Cas protein (e.g., from any Cas protein, e.g., Un1Cas12f1) may contain one or more mutations compared to the wild-type version of the protein. The mutations may result in 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of nucleic acid cleavage activity in one or more of the multiple nucleic acid cleavage domains of the wild-type Cas protein. The mutations may result in one or more of the multiple nucleic acid cleavage domains retaining the ability to cleave a complementary strand of a target nucleic acid, but having a reduced ability to cleave a non-complementary strand of the target nucleic acid. The mutations may result in one or more of the multiple nucleic acid cleavage domains lacking the ability to cleave complementary and non-complementary strands of a target nucleic acid. The residues to be mutated in the nuclease domain can correspond to one or more catalytic residues of nuclease.For example, residues such as Asp10, His840, Asn854, and Asn856 in the wild-type exemplary S.pyogenes Cas9 polypeptide can be mutated to inactivate one or more of multiple nucleic acid cleavage domains (e.g., nuclease domains).The residues to be mutated in the nuclease domain of Cas protein can correspond to residues Asp10, His840, Asn854, and Asn856 in the wild-type S.pyogenes Cas9 polypeptide, for example, as determined by sequence and / or structural alignment.

[0158] As non-limiting examples, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 (or corresponding mutations in any of the Cas proteins) can be mutated, such as D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. Mutations other than alanine substitutions may be suitable.

[0159] The D10A mutation can be combined with one or more of the H840A, N854A, or N856A mutations to produce a Cas9 protein that is substantially devoid of DNA cleavage activity (e.g., a dead Cas9 protein). The H840A mutation can be combined with one or more of the D10A, N854A, or N856A mutations to produce a site-directed polypeptide that is substantially devoid of DNA cleavage activity. The N854A mutation can be combined with one or more of the H840A, D1OA, or N856A mutations to produce a site-directed polypeptide that is substantially devoid of DNA cleavage activity. The N856A mutation can be combined with one or more of the H840A, N854A, or D10A mutations to produce a site-directed polypeptide that is substantially devoid of DNA cleavage activity.

[0160] In some embodiments, the Cas protein is a class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, or is derived from a Cas9 protein. For example, a Cas9 protein that lacks cleavage activity. In some embodiments, the Cas9 protein is a S. pyogenes Cas9 protein (e.g., SwissProt Accession No. Q99ZW2). In some embodiments, the Cas9 protein is a S. aureus Cas9 (e.g., SwissProt Accession No. J7RUA5). In some embodiments, the Cas9 protein is a modified version of a S. pyogenes or S. Aureus Cas9 protein. In some embodiments, the Cas9 protein is derived from a S. pyogenes or S. Aureus Cas9 protein. For example, a S. pyogenes or S. Aureus Cas9 protein that lacks cleavage activity.

[0161] In some embodiments, Cas9 may generally refer to a polypeptide having at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, or about 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas9 polypeptide (e.g., Cas9 of S. pyogenes). In some embodiments, Cas9 may refer to a polypeptide having no more than about 5%, no more than about 10%, no more than about 20%, no more than about 30%, no more than about 40%, no more than about 50%, no more than about 60%, no more than about 70%, no more than about 80%, no more than about 90%, or no more than about 100% sequence identity and / or sequence similarity to a wild-type Cas9 polypeptide (e.g., from S. pyogenes). Cas9 may refer to a wild-type or modified form of the Cas9 protein, which may include amino acid alterations, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.

[0162] The Cas protein can comprise an amino acid sequence having at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity or similarity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein.

[0163] Cas proteins, variants or derivatives thereof, for example as part of the complexes disclosed herein, can be modified to enhance the regulation of gene expression by the compositions and methods of the present disclosure. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, enzymatic activity, and / or binding to other factors, such as heterodimerization or oligomerization domains, and to induce ligands. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for the desired function of the protein or complex. Cas proteins can be modified to modulate (e.g., enhance or reduce) the activity of the Cas protein to regulate gene expression by the complexes of the present disclosure that include heterologous gene effectors.

[0164] For example, the Cas protein can be coupled (e.g., fused, covalently or non-covalently linked) to a heterologous genetic effector (e.g., an epigenetic modification domain, a transcriptional activation domain, and / or a transcriptional repressor domain). The Cas protein can be coupled (e.g., fused, covalently or non-covalently linked) to an oligomerization or dimerization domain (e.g., a heterodimerization domain) disclosed herein. The Cas protein can be coupled (e.g., fused, covalently or non-covalently linked) to a heterologous polypeptide that provides increased or decreased stability. The Cas protein can be coupled (e.g., fused, covalently or non-covalently linked) to a sequence that can facilitate degradation of the Cas protein or a complex containing the Cas protein, such as a degron, e.g., an inducible degron (e.g., auxin-inducible).

[0165] A Cas protein can be conjugated (e.g., fused, covalently or non-covalently linked) to any suitable number of partners, such as at least 1, at least 2, at least 3, at least 4, or at least 5, at least 6, at least 7, or at least 8 partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, or at most 10 partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to 1-5, 1-4, 1-3, 1-2, 2-5, 2-4, 2-3, 3-5, 3-4, or 4-5 partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to one partner. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to two partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to three partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to four partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to five partners. In some embodiments, a Cas protein of the disclosure is conjugated (e.g., fused, covalently or non-covalently linked) to six partners.

[0166] The Cas protein may be a fusion protein. The fused domain or heterologous polypeptide may be located at the N-terminus, C-terminus, or internally of the Cas protein.

[0167] The Cas protein may be provided in any form. For example, the Cas protein may be provided in the form of a protein, e.g., the Cas protein alone or in a complex with a guide nucleic acid as a ribonucleoprotein. The Cas protein may be provided in a complex, e.g., in a complex with a guide nucleic acid and / or one or more heterologous gene effectors of the present disclosure. The Cas protein may be provided in the form of a nucleic acid, e.g., RNA (e.g., messenger RNA (mRNA)), or DNA, that encodes the Cas protein. The nucleic acid encoding the Cas protein may be codon-optimized for efficient translation into a protein in a particular cell or organism.

[0168] The nucleic acid encoding the Cas protein, its fragment or derivative can be stably integrated into the genome of the cell. The nucleic acid encoding the Cas protein can be operably linked to a promoter, for example, a promoter that is constitutively or inducibly active in the cell. The nucleic acid encoding the Cas protein can be operably linked to a promoter in an expression construct. The expression construct can include any nucleic acid construct that can direct the expression of a gene or other nucleic acid sequence of interest (e.g., Cas gene) and can transfer such a nucleic acid sequence of interest into a target cell.

[0169] In some embodiments, the Cas protein, its variant or derivative is a nuclease-dead Cas (dCas) protein. A dead Cas protein can be a protein that lacks nucleic acid cleavage activity.

[0170] The Cas protein may include modified forms of wild-type Cas proteins. Modified forms of wild-type Cas proteins may include amino acid changes (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the Cas protein. For example, modified forms of Cas proteins may have 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of wild-type Cas proteins (e.g., Cas9 of S.pyogenes). Modified forms of Cas proteins may not have substantial nucleic acid cleavage activity. When a Cas protein is a modified form that does not have substantial nucleic acid cleavage activity, it may be referred to as enzymatically inactive, "deactivated," and / or "dead" (abbreviated by "d"). Dead Cas proteins (e.g., dCas, dCas9) may bind to target polynucleotides but may not cleave or barely cleave the target polynucleotide. In some embodiments, the dead Cas protein is a dead Cas9 protein.

[0171] The dCas9 polypeptide can associate with a single guide RNA (sgRNA) to activate or repress transcription of a target gene (e.g., a target endogenous gene), for example in combination with a heterologous gene effector(s) disclosed herein. The sgRNA can be introduced into a cell expressing a Cas or guide moiety component of the present disclosure. In some examples, such a cell can contain one or more different sgRNAs that target the same target gene (e.g., a target endogenous gene) or target gene regulatory sequence. In other examples, the sgRNAs target different nucleic acids in the cell (e.g., different target genes, different target gene regulatory sequences, or different sequences within the same target gene or target gene regulatory sequence).

[0172] Enzymatically inactive may refer to a nuclease that binds to a nucleic acid sequence in a polynucleotide in a sequence-specific manner but does not cleave the target polynucleotide or cleaves it at a substantially reduced frequency. An enzymatically inactive guide portion may include an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive may refer to no activity. Enzymatically inactive may refer to substantially no activity. Enzymatically inactive may refer to essentially no activity. Enzymatically inactive may refer to 1% or less, 2% or less, 3% or less, 4% or less, 5% or less, 6% or less, 7% or less, 8% or less, 9% or less, or 10% or less activity compared to a comparable wild-type activity (e.g., nucleic acid cleavage activity, wild-type Cas9 activity).

[0173] In some embodiments, the guide portion does not contain a nucleic acid guide targeting system. For example, the guide portion may include a protein that binds to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence based on a protein structural feature, such as a certain nuclease disclosed herein.

[0174] In some embodiments, the guide portion comprises a zinc finger nuclease (ZFN), or a variant, fragment, or derivative thereof. ZFN may refer to a fusion between a cleavage domain, such as the cleavage domain of Fokl, and at least one zinc finger motif (e.g., at least 2, at least 3, at least 4, or at least 5 zinc finger motifs) that can bind to polynucleotides such as DNA and RNA. In some embodiments, ZFN is used in the targeting portion of the present disclosure to bind to a polynucleotide (e.g., a target gene or a target gene regulatory sequence), but the ZFN does not cleave or does not substantially cleave the polynucleotide, e.g., a nuclease-dead ZFN. ZFN or a variant, fragment, or derivative thereof can be fused to or associated with one of more heterologous gene effectors to form the complex of the present disclosure.

[0175] Heterodimerization of two individual ZFNs at a certain position in a polynucleotide in a specific orientation and spacing can result in cleavage of the polynucleotide in nuclease-active ZFNs. For example, when ZFNs bind to DNA, they can induce double-stranded breaks in the DNA. To allow the two cleavage domains to dimerize and cleave the DNA, two individual ZFNs can bind to opposite strands of DNA at their C-termini at a certain distance apart. In some cases, the linker sequence between the zinc finger domain and the cleavage domain may require that the 5' ends of each binding site are about 5-7 base pairs apart. In some cases, the cleavage domain is fused to the C-terminus of each zinc finger domain.

[0176] In some embodiments, the cleavage domain of the guide moiety comprising ZFN comprises a modified version of the wild-type cleavage domain. The modified version of the cleavage domain may comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid cleavage activity of the cleavage domain. For example, the modified version of the cleavage domain may have 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the corresponding wild-type cleavage domain. The modified version of the cleavage domain may not have substantial nucleic acid cleavage activity. In some embodiments, the cleavage domain is enzymatically inactive.

[0177] In some embodiments, the guide moiety comprises a "TALEN" or "TAL-effector nuclease", or a variant, fragment, or derivative thereof. TALEN generally refers to an engineered transcription activator-like effector nuclease that contains a central domain and a cleavage domain of a DNA-binding tandem repeat. TALENs can be produced by fusing a TAL effector DNA-binding domain to a DNA-cleavage domain. In some examples, the DNA-binding tandem repeat comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. A transcription activator-like effector (TALE) protein can be fused to a nuclease, such as a wild-type or mutated Fok1 endonuclease, or the catalytic domain of Fok1. In some embodiments, a TALEN is used in the targeting moiety of the present disclosure to bind to a polynucleotide (e.g., a target gene or a target gene regulatory sequence), but the TALEN does not cleave or does not substantially cleave the polynucleotide, e.g., a nuclease-dead TALEN. TALENs, or variants, fragments, or derivatives thereof, can be fused to or associated with one or more heterologous gene effectors to form a complex of the present disclosure.

[0178] In some embodiments, the TALEN is engineered for reduced nuclease activity. In some embodiments, the nuclease domain of the TALEN comprises a modified version of the wild-type nuclease domain. The modified version of the nuclease domain may comprise an amino acid change (e.g., deletion, insertion, or substitution) that reduces the nucleic acid cleavage activity of the nuclease domain. For example, the modified version of the nuclease domain may have 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified version of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is enzymatically inactive. The TALEN, or its variant, fragment, or derivative, may be fused or associated with one or more heterologous gene effectors to form the complex of the present disclosure.

[0179] For example, some mutations of Fok1 have been made for its use in TALEN, improving cleavage specificity or activity.Such TALEN can be engineered to bind to any desired DNA sequence.TALEN can be used to create double-strand breaks in target DNA sequence, and then undergo NHEJ or HDR to generate genetic modification (for example, nucleic acid sequence editing).

[0180] A TALE, or a variant, fragment, or derivative thereof, can be fused to or associated with one or more heterologous gene effectors to form the complex of the present disclosure. In some embodiments, a transcription activator-like effector (TALE) protein is fused to the heterologous gene effector and does not contain a nuclease. In some embodiments, the TALEN is a TALE that does not cleave or does not substantially cleave a polynucleotide, e.g., a nuclease-dead TALE. A TALE, or a variant, fragment, or derivative thereof, can be fused to or associated with one or more heterologous gene effectors to form the complex of the present disclosure.

[0181] In some embodiments, the complex of the transcription activator-like effector (TALE) protein and the heterologous gene effector is designed to function as a transcription activator. In some embodiments, the complex of the transcription activator-like effector (TALE) protein and the heterologous gene effector is designed to function as a transcription repressor. For example, the DNA binding domain of the transcription activator-like effector (TALE) protein can be fused (e.g., linked) to one or more heterologous gene effectors that contain a transcription activation domain or one or more heterologous gene effectors that contain a transcription repression domain.

[0182] In some embodiments, the guide moiety comprises a meganuclease. Meganuclease generally refers to a rare-cutting endonuclease, or a homing endonuclease, that is highly sequence specific. Meganucleases can recognize DNA target sites ranging in length from at least 12 base pairs in length, for example, 12-40 base pairs, 12-50 base pairs, or 12-60 base pairs in length. Meganucleases can be modular DNA-binding nucleases, for example, any fusion protein that includes at least one catalytic domain of an endonuclease and at least one DNA-binding domain or protein that specifies a nucleic acid target sequence. The DNA-binding domain can contain at least one motif that recognizes single-stranded or double-stranded DNA. Nuclease-active meganucleases can generate double-stranded breaks. In some embodiments, meganucleases are used in the targeting moieties of the present disclosure to bind to polynucleotides (e.g., target genes or target gene regulatory sequences), but the meganuclease does not cleave or does not substantially cleave the polynucleotide, e.g., a nuclease-dead meganuclease. Meganucleases, or variants, fragments, or derivatives thereof, can be fused or associated with one or more heterologous genetic effectors to form the complexes of the present disclosure.

[0183] Meganucleases can be monomeric or dimeric. In some embodiments, meganucleases are naturally occurring (found in nature) or wild type, while in other examples, meganucleases are non-natural, artificial, engineered, synthetic, rationally designed, or man-made. In some embodiments, meganucleases of the present disclosure include I-CreI meganuclease, I-CeuI meganuclease, I-Msol meganuclease, I-SceI meganuclease, variants thereof, derivatives thereof, and fragments thereof.

[0184] In some embodiments, the nuclease domain of the meganuclease comprises a modified version of the wild-type nuclease domain. The modified version of the nuclease domain may comprise amino acid changes (e.g., deletions, insertions, or substitutions) that reduce or eliminate the nucleic acid cleavage activity of the nuclease domain. For example, the modified version of the nuclease domain may have 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified version of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is enzymatically inactive. In some embodiments, the meganuclease can bind to DNA but cannot cleave DNA. In some embodiments, the nuclease-inactive meganuclease is fused or associated with one or more heterologous gene effectors to generate the complex of the present disclosure.

[0185] In some embodiments, the guide moiety can regulate the expression and / or activity of a target gene (e.g., a target endogenous gene). In some embodiments, the guide moiety can edit the sequence of a nucleic acid (e.g., a gene and / or a gene product). A nuclease-active Cas protein can edit a nucleic acid sequence by generating a double-stranded or single-stranded break in a target polynucleotide.

[0186] In some embodiments, the guide portion containing nuclease can generate double-strand breaks in target polynucleotide, such as DNA. The double-strand breaks in DNA can result in DNA break repair, which allows the introduction of genetic modification(s) (e.g., nucleic acid editing). In some embodiments, the nuclease induces site-specific single-strand DNA breaks or nicks, thus resulting in HDR.

[0187] Double-strand breaks in DNA can result in DNA break repair, which allows the introduction of genetic modification(s) (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor DNA repair template or template polynucleotide can be provided that contains homologous arms adjacent to the site of target DNA.

[0188] In some embodiments, a complex comprising a guide moiety or a nuclease does not generate a double-stranded break in a target polynucleotide, e.g., DNA.

[0189] III. Complex

[0190] In some embodiments, one or more complexes are disclosed herein that comprise heterologous polypeptides and heterologous polynucleotides.In some examples, the complexes can comprise heterologous gene effectors and guide moieties, such as guide nucleic acids and / or nucleases, such as endonucleases that lack or substantially lack cleavage activity.

[0191] The complexes of the present disclosure may be useful, for example, to bring one or more heterologous gene effectors into close proximity to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence, thereby facilitating modulation of expression, epigenetic modification, or activity levels of the target gene.

[0192] In some embodiments, the complex of the present disclosure binds to DNA, for example genomic DNA.In some embodiments, the complex of the present disclosure binds to RNA, for example mRNA, microRNA, siRNA or non-coding RNA.In some embodiments, the complex of the present disclosure binds to DNA and RNA.

[0193] In some embodiments, the complex can modulate (e.g., increase or decrease) the expression and / or activity of a target gene (e.g., a target endogenous gene) by physical disruption of a polynucleotide sequence (e.g., a promoter, enhancer, repressor, operator, or silencer, insulator, cis-regulatory element, trans-regulatory element, site of epigenetic modification (e.g., DNA methylation), coding sequence).

[0194] In some embodiments, the complex can modulate (e.g., increase or decrease) the expression and / or activity of a target gene (e.g., a target endogenous gene) by recruiting additional factors effective to suppress or enhance expression of the target gene.

[0195] In some embodiments, the complexes of the present disclosure are used to introduce epigenetic modifications into a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence (e.g., a promoter, enhancer, silencer, insulator, cis-regulatory element, trans-regulatory element, site of epigenetic modification (e.g., DNA methylation)). In some embodiments, the complexes of the present disclosure are used to generate three-dimensional structures, topologically related domains, or genomic boundaries that include a target gene or a target gene regulatory sequence (e.g., a gene distal or proximal to a target gene).

[0196] In some embodiments, the complex comprises a heterologous gene effector and a guide portion. In some embodiments, the complex comprises one heterologous gene effector and one guide portion. In some embodiments, the complex comprises two heterologous gene effectors and one guide portion. In some embodiments, the complex comprises three or more heterologous gene effectors and one guide portion.

[0197] In some embodiments, the complex comprises a heterologous genetic effector and a guide nucleic acid. In some embodiments, the complex comprises one heterologous genetic effector and one guide nucleic acid. In some embodiments, the complex comprises two heterologous genetic effectors and one guide nucleic acid. In some embodiments, the complex comprises three or more heterologous genetic effectors and one guide nucleic acid.

[0198] The two components present in a complex, e.g., a fusion protein, can be covalently linked or crosslinked, e.g., treated with a crosslinking agent, or linked by a peptide or non-peptide linker as disclosed herein.

[0199] In some embodiments, the two components present in the complex are part of the same fusion protein. The components can be optionally linked by a linker, such as a peptide or non-peptide linker.

[0200] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is linked to the heterologous gene effector by a linker. In some embodiments, the guide portion or a portion thereof is further linked to a second heterologous gene effector by the same or a different second linker. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is fused to the heterologous gene effector without a linker.

[0201] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is linked to the oligomerization domain or dimerization (e.g., heterodimerization) domain by a linker. In some embodiments, the guide portion or a portion thereof is further linked to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by the same or a different second linker. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is fused to the second oligomerization domain or dimerization (e.g., heterodimerization) domain without a linker.

[0202] In some embodiments, the heterologous gene effector is linked to the second heterologous gene effector by a linker.In some embodiments, the heterologous gene effector is further linked to a third heterologous gene effector by the same or different second linker.In some embodiments, the heterologous gene effector is fused to the second heterologous gene effector without a linker.

[0203] In some embodiments, the heterologous effector is linked to the oligomerization domain or dimerization (e.g., heterodimerization) domain by a linker. In some embodiments, the heterologous effector is further linked to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by the same or different second linker. In some embodiments, the heterologous effector is fused to the second oligomerization domain or dimerization (e.g., heterodimerization) domain without a linker.

[0204] Any suitable linker can be used. Flexible linkers can have sequences that contain stretches of glycine and serine residues. The small size of glycine and serine residues provides flexibility and allows mobility of the connected functional domains. Incorporating serine and threonine can maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thereby reducing unfavorable interactions between the linker and the protein moiety. Flexible linkers can also contain additional amino acids such as threonine and alanine to maintain flexibility, as well as polar amino acids such as lysine and glutamine to improve solubility. Rigid linkers can have, for example, an alpha-helical structure. Alpha-helical rigid linkers can act as spacers between protein domains.

[0205] The linker sequence can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues in length.

[0206] In some embodiments, the linker is at least 1, at least 2, at least 3, at least 5, at least 7, at least 9, at least 11, at least 13, at least 15, or at least 20 amino acids. In some embodiments, the linker is at most 5, at most 7, at most 9, at most 11, at most 13, at most 15, at most 20, at most 25, at most 30, at most 40, or at most 50 amino acids.

[0207] In some embodiments, a non-peptide linker is used. The non-peptide linker can be, for example, a chemical linker. The two parts of the complex of the present disclosure can be connected by a chemical linker. Each chemical linker of the present disclosure can be an alkylene, alkenylene, alkynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, any of which is optionally substituted. In some embodiments, the chemical linker of the present disclosure can be an ester, ether, amide, thioether, or polyethylene glycol (PEG). In some embodiments, the linker can reverse the order of the amino acid sequence in the compound, for example, so that the amino acid sequence linked by the linker is head-to-head, rather than head-to-tail. Non-limiting examples of such linkers include diesters of dicarboxylic acids, such as oxalyl diesters, malonyl diesters, succinyl diesters, glutaryl diesters, adipyl diesters, p-methyl diesters, fumaryl diesters, maleyl diesters, phthalyl diesters, isophthalyl diesters, and terephthalyl diesters. Non-limiting examples of such linkers include diamides of dicarboxylic acids, such as oxalyl diamide, malonyl diamide, succinyl diamide, glutaryl diamide, adipyl diamide, dimethyl diamide, fumaryl diamide, maleyl diamide, phthalyl diamide, isophthalyl diamide, and terephthalyl diamide. Non-limiting examples of such linkers include diamides of diamino linkers, such as ethylenediamine, 1,2-di(methylamino)ethane, 1,3-diaminopropane, 1,3-di(methylamino)propane, 1,4-di(methylamino)butane, 1,5-di(methylamino)pentane, 1,6-di(methylamino)hexane, and pipyrizine.Non-limiting examples of optional substituents include hydroxyl groups, sulfhydryl groups, halogens, amino groups, nitro groups, nitroso groups, cyano groups, azide groups, sulfoxide groups, sulfone groups, sulfonamide groups, carboxyl groups, carboxaldehyde groups, imine groups, alkyl groups, haloalkyl groups, alkenyl groups, haloalkenyl groups, alkynyl groups, haloalkynyl groups, alkoxy groups, aryl groups, aryloxy groups, aralkyl groups, arylalkoxy groups, heterocyclyl groups, acyl groups, acyloxy groups, carbamate groups, amide groups, ureido groups, epoxy groups, and ester groups.

[0208] The two components present in the complex can be non-covalently linked, for example, by ionic bonds, hydrogen bonds, interactions mediated by oligomerization or dimerization domains disclosed herein, and the like.

[0209] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is attached to the heterologous genetic effector by a non-covalent bond. In some embodiments, the guide portion or a portion thereof is further attached to a second heterologous genetic effector by a non-covalent bond. In some embodiments, the guide portion or a portion thereof is attached to a first heterologous genetic effector by a covalent bond (e.g., a fusion protein, optionally with a linker), and the guide portion or a portion thereof is further attached to a second heterologous genetic effector by a non-covalent bond.

[0210] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is attached to the oligomerization domain or dimerization (e.g., heterodimerization) domain by a non-covalent bond. In some embodiments, the guide portion or a portion thereof is further attached to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by a non-covalent bond. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease, e.g., dCas9) is fused to a first oligomerization domain or dimerization (e.g., heterodimerization) domain by a covalent bond (e.g., fused via a linker, if desired) and attached to a second oligomerization domain or dimerization (e.g., heterodimerization) domain by a non-covalent bond.

[0211] In some embodiments, the first component of the guide portion (e.g., the guide nucleic acid) is non-covalently linked to the second component of the guide portion (e.g., the nuclease). In some embodiments, the first component of the guide portion (e.g., the guide nucleic acid) is covalently linked to the second component of the guide portion (e.g., the nuclease).

[0212] Any combination of covalent and non-covalent attachments can be used in the complexes of the present disclosure, for example, one or more heterologous genetic effectors can be non-covalently fused to a guide moiety and one or more oligomerization domains can be covalently attached to a component of the complex (e.g., a nuclease).

[0213] In some embodiments, the polypeptide that provides increased or decreased stability is fused or otherwise associated with a component of the complex of the present disclosure, such as a guide moiety or a heterologous gene effector. The fusion polypeptide can be located at the N-terminus, C-terminus, or internal to the fusion protein.

[0214] In some embodiments, one or more components of the complexes of the present disclosure are fused to a domain the directs desirable sub-cellular localization, such as a nuclear localization signal or a protein for targeting to the inner nuclear membrane, the outer nuclear membrane, Cajal bodies, nuclear speckles, nuclear pore complexes, PML bodies, nucleoli, P granules, GW bodies, stress granules, sponge bodies, endoplasmic reticulum, mitochondria, etc.

[0215] In some embodiments, the complex of the present disclosure comprises a first protein linked to a first oligomerization (e.g., dimerization) domain and a second protein linked to a second oligomerization (e.g., dimerization) domain. In some embodiments, the oligomerization or dimerization domain may comprise a peptide interaction domain, such as sgRNA2.0, SAM, SunTag, RAB, FLAG-biotin, or a system that utilizes an inducible oligomerization (e.g., dimerization) system disclosed herein.

[0216] delivery

[0217] The gene or genes encoding any heterologous polypeptides (e.g., heterologous gene effectors) and any additional molecules operably linked thereto (e.g., heterologous polynucleotides, e.g., one or more guide nucleic acid molecules), as disclosed herein, may be integrated into the genome of the cell in which the aberrant expression of the target gene is to be modified. Alternatively, the gene or genes may not, and do not necessarily have to, be integrated into the genome of the cell.

[0218] Any heterologous polypeptide (e.g., heterologous gene effector) and any additional molecules operably linked thereto (e.g., heterologous polynucleotides, e.g., one or more guide nucleic acid molecules) can be introduced (e.g., delivered, expressed, etc.) into cells by a variety of methods, including viral and non-viral delivery methods. Viral vector delivery systems can include DNA and RNA viruses, and can have either episomal or integrated genomes after delivery to a cell. Non-viral vector delivery systems can include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acid, and nucleic acid complexed with a delivery vehicle, such as a liposome.

[0219] RNA or DNA virus-based systems can be used to target specific cells and transport the viral payload to the nucleus of the cell. Viral vectors can be used to treat cells in vitro and the modified cells can be administered as needed (ex vivo). Alternatively, viral vectors can be administered directly to subjects (in vivo). Viral-based systems can include retrovirus, lentivirus, adenovirus, adeno-associated virus, and herpes simplex virus vectors for gene transfer. Integration into the host genome can occur by retrovirus, lentivirus, and adeno-associated virus gene transfer methods, which can result in long-term expression of the inserted transgene.

[0220] Non-limiting examples of viral vectors that can be utilized to deliver heterologous polypeptides and / or heterologous polynucleotides (or one or more genes encoding same) include, but are not limited to, retroviral vectors, lentiviral vectors, adenoviral vectors, poxvirus vectors, herpesvirus vectors, adeno-associated virus (AAV) vectors. Non-limiting examples of AAV vectors include AAV1, AAV10, AAV106.1 / hu.37, AAV11, AAV114.3 / hu.40, AAV12, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV128.1 / hu.43, AAV128.3 / hu.44, AAV130.4 / hu.48, AAV145.1 / hu.53, AAV145.5 / hu.54, AAV145.6 / hu.55, AAV16.12 / hu. 11, AAV16.3, AAV16.8 / hu. 10, AAV161.10 / hu.60, AAV161.6 / hu.61, AAV1-7 / rh.48, AAV1-8 / rh.49, AAV2, AAV2.5T, AAV2 -15 / rh.62, AAV223.1, AAV223.2, AAV223.4, AAV223.5, AAV223.6, AAV223.7, AAV2-3 / rh.61, A AV24.1, AAV2-4 / rh.50, AAV2-5 / rh.51, AAV27.3, AAV29.3 / bb.1, AAV29.5 / bb.2, AAV2G9, AAV- 2-pre-miRNA-101, AAV3, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-11 / rh.53, AAV3-3, AAV33.12 / hu. 17, AAV33.4 / hu. 15, AAV33.8 / hu. 16, AAV3-9 / rh.52, AAV3a, AAV3b, AAV4, AAV4-19 / rh.55, AAV42.12、AAV42-10、AAV42-11、AAV42-12、AAV42-13、AAV42-15、AAV42-1b、AAV42-2、AAV42-3a、AAV42- 3b、AAV42-4、AAV42-5a、AAV42-5b、AAV42-6b、AAV42-8、AAV42-aa、AAV43-1、AAV43-12、AAV43-20、 AAV43-21、AAV43-23、AAV43-25、AAV43-5、AAV4-4、AAV44.1、AAV44.2、AAV44.5、AAV46.2 / hu.28、AAV46.6 / hu.29、AAV4-8 / r11.64、AAV4-8 / rh.64、AAV4-9 / rh.54、AAV5、AAV52.1 / hu.20、AAV52 / hu. 19、AAV5-22 / rh.58、AAV5-3 / rh.57、AAV54.1 / hu.21、AAV54.2 / hu.22、AAV54.4R / hu.27、AAV54.5 / hu.23、AAV54.7 / hu.24、AAV58.2 / hu.25、AAV6、AAV6.1、AAV6.1.2、AAV6.2、AAV7、AAV7.2、AAV7.3 / hu.7、 AAV8、AAV-8b、AAV-8h、AAV9、AAV9.11、AAV9.13、AAV9.16、AAV9.24、AAV9.45、AAV9.47、AAV9.61、AAV9 .68、AAV9.84、AAV9.9、AAVA3.3、AAVA3.4、AAVA3.5、AAVA3.7、AAV-b、AAVC1、AAVC2、AAVC5、AAVCh.5、AAVCh. AVCh.5R1、AAVcy.2、AAVcy.3、AAVcy.4、AAVcy.5、AAVCy.5R1、AAVCy.5R2、AAVCy.5R3、AAVCy.5R4、AAV cy.6、AAV-DJ、AAV-DJ8、AAVF3、AAVF5、AAV-h、AAVH-1 / hu.1、AAVH2、AAVH-5 / hu.3、AAVH6、AAVhE1.1、A AVhER1.14、AAVhEr1.16、AAVhEr1.18、AAVhEr1.23、AAVhEr1.35、AAVhEr1.36、AAVhEr1.5、AAVhEr1.7 、AAVhEr1.8、AAVhEr2.16、AAVhEr2.29、AAVhEr2.30、AAVhEr2.31、AAVhEr2.36、AAVhEr2.4、AAVhEr3.1、AAVhu.1、AAVhu.10、AAVhu.11、AAVhu.11、AAVhu.12、AAVhu.13、AAVhu.14 / 9、AAVhu.15、AAVhu.16、AAVhu.17、AAVhu.18、AAVhu.19、AAVhu.2、AAVhu.20、AAVhu.21、AAVhu.22、AAVhu.23.2、AAVhu.24、AAVhu.25、AAVhu.27、AAVhu.28、AAVhu.29、AAVhu.29R、AAVhu.3、AAVhu.31、AAVhu.32、AAVhu.34、AAVhu.35、AAVhu.37、AAVhu.39、AAVhu.4、AAVhu.40、AAVhu.41、AAVhu.42、AAVhu.43、AAVhu.44、AAVhu.44R1、AAVhu.44R2、AAVhu.44R3、AAVhu.45、AAVhu.46、AAVhu.47、AAVhu.48、AAVhu.48R1、AAVhu.48R2、AAVhu.48R3、AAVhu.49、AAVhu.5、AAVhu.51、AAVhu.52、AAVhu.53、AAVhu.54、AAVhu.55、AAVhu.56、AAVhu.57、AAVhu.58、AAVhu.6、AAVhu.60、AAVhu.61、AAVhu.63、AAVhu.64、AAVhu.66、AAVhu.67、AAVhu.7、AAVhu.8、AAVhu.9、AAVhu.t 19、AAVLG-10 / rh.40、AAVLG-4 / rh.38、AAVLG-9 / hu.39、AAVLG-9 / hu.39、AAV-LKO1、AAV-LK02、AAVLK03、AAV-LK03、AAV-LK04、AAV-LK05、AAV-LKO6、AAV-LK07、AAV-LK08、AAV-LK09、AAV-LK10、AAV-LK11、AAV-LK12、AAV-LK13、AAV-LK14、AAV-LK15、AAV-LK17、AAV-LK18、AAV-LK19、AAVN721-8 / rh.43、AAV-PAEC、AAV-PAEC11、AAV-PAEC12、AAV-PAEC2、AAV-PAEC4、AAV-PAEC6、AAV-PAEC7、AAV-PAEC8、AAVpi.1、AAVpi.2、AAVpi.3、AAVrh.10、AAVrh.12、AAVrh.13、AAVrh.13R, AAVrh.14, AAVrh.17, AAVrh.18, AAVrh.19, AAVrh.2, AAVrh.20, AAVrh.21, AAVrh .22, AAVrh.23, AAVrh.24, AAVrh.25, AAVrh.2R, AAVrh.31, AAVrh.32, AAVrh.33, AAVr h.34, AAVrh.35, AAVrh.36, AAVrh.37, AAVrh.37R2, AAVrh.38, AAVrh.39, AAVrh.40, A AVrh.43, AAVrh.44, AAVrh.45, AAVrh.46, AAVrh.47, AAVrh.48, AAVrh.48, AAVrh.48.1 , AAVrh.48.1.2, AAVrh.48.2, AAVrh.49, AAVrh.50, AAVrh.51, AAVrh.52, AAVrh.53, A AVrh.54, AAVrh.55, AAVrh.56, AAVrh.57, AAVrh.58, AAVrh.59, AAVrh.60, AAVrh.61, A AVrh.62, AAVrh.64, AAVrh.64R1, AAVrh.64R2, AAVrh.65, AAVrh.67, AAVrh.68, AAVrh .69, AAVrh.70, AAVrh.72, AAVrh.73, AAVrh.74, AAVrh.8, AAVrh.8R, AAVrh8R, AAVrh8R Examples of AAV include A586R mutant, AAVrh8R R533A mutant, BAAV, BNP61 AAV, BNP62 AAV, BNP63 AAV, bovine AAV, caprine AAV, Japanese AAV 10, true type AAV (ttAAV), UPENN AAV 10, AAV-LK16, AAAV, AAV shuffle 100-1, AAV shuffle 100-2, AAV shuffle 100-3, AAV shuffle 100-7, AAV shuffle 10-2, AAV shuffle 10-6, AAV shuffle 10-8, AAV SM 100-10, AAV SM 100-3, AAV SM 10-1, AAV SM 10-2, and AAV SM 10-8. For example, AAVrh.74 can be used as a viral vector to deliver polynucleotide sequences encoding heterologous polypeptides and heterologous polynucleotides (e.g., Cas protein-gene effector fusions and one or more guide nucleic acid molecules).

[0221] Non-viral delivery methods of nucleic acids may include lipofection, nucleofection, microinjection, biolistics, virosomes, liposomes, immunoliposomes, polycation or lipid:nucleic acid conjugates, lipid nanoparticles (LNPs), naked DNA, artificial virions, and agent-enhanced uptake of DNA. Cationic and neutral lipids that are suitable for efficient receptor-recognition lipofection of polynucleotides may be used.

[0222] Any of the compositions disclosed herein (or one or more genes encoding any portion of the compositions), such as the heterologous gene effector(s) and / or guide nucleic acid molecule(s), can be administered by any suitable route of administration, including, but not limited to, parenteral (e.g., intravenous, intratumor, subcutaneous, intramuscular, intracerebral, intraventricular, intraarticular, intraperitoneal, or intracranial), intranasal, buccal, sublingual, oral, or rectal routes of administration. In some examples, the pharmaceutical composition is formulated for parenteral (e.g., intravenous, intratumor, subcutaneous, intramuscular, intracerebral, intraventricular, intraarticular, intraperitoneal, or intracranial) administration.

[0223] Target Gene

[0224] The present disclosure provides compositions, methods and systems for modulating the expression of target genes (e.g., target endogenous genes).For example, disclosed herein is a complex that comprises a guide portion and one or more heterologous gene effectors that can increase or decrease the activity or expression level of target genes.

[0225] In some embodiments, the target gene or its regulatory sequence is endogenous to the subject, e.g., present in the subject's genome. In some embodiments, the target gene or its regulatory sequence is not part of an engineered reporter system.

[0226] In some embodiments, the target gene is exogenous to the host subject, for example, a pathogen target gene or an exogenous gene expressed as a result of therapeutic intervention, for example, gene therapy and / or cell therapy. In some embodiments, the target gene is an exogenous reporter gene. In some embodiments, the target gene is an exogenous synthetic gene.

[0227] In some embodiments, the target gene (e.g., the target endogenous gene) is a gene that is overexpressed or underexpressed in a disease or condition. In some embodiments, the target gene is a gene that is overexpressed or underexpressed in an inherited genetic disease.

[0228] In some embodiments, the target gene (e.g., the target endogenous gene) is a gene that is overexpressed or underexpressed in an autoimmune disease. In some embodiments, the target gene is a gene that is overexpressed or underexpressed in an autoimmune disease, such as acute disseminated encephalomyelitis, acute motor axonal neuropathy, Addison's disease, adiposity, adult-onset Still's disease, alopecia areata, ankylosing spondylitis, anti-glomerular basement membrane nephritis, antineutrophil cytoplasmic antibody-associated vasculitis, anti-N-methyl-D-aspartate receptor encephalitis, antiphospholipid syndrome, antisynthetase syndrome, aplastic anemia, autoimmune angioedema, autoimmune encephalitis, autoimmune enteropathy, autoimmune hemolytic anemia, autoimmune hepatitis, autoimmune inner ear disease, autoimmune lymphoproliferative syndrome, autoimmune neutropenia, autoimmune pulmonary fibrosis ... Immune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune polyendocrine syndrome, autoimmune polyendocrine syndrome type 2, autoimmune polyendocrine syndrome type 3, autoimmune progesterone dermatitis, autoimmune retinopathy, autoimmune thrombocytopenic purpura, autoimmune thyroiditis, autoimmune urticaria, autoimmune uveitis, Baro concentric sclerosis, Behcet's disease, Bickerstaff encephalitis, bullous pemphigoid, celiac disease, chronic fatigue syndrome, chronic inflammatory demyelinating polyneuropathy, Churg-Strauss syndrome, cicatricial pemphigoid, Cogan's syndrome, cold agglutinin disease, complex regional pain syndrome, CREST syndrome, Crohn's disease, dermatitis herpetiformis, dermatomyositis, type 1 diabetes, discoid lupus erythematosus, endometriosis, enthesitis, enthesitis-associated arthritis, eosinophilic esophagitis, eosinophilic fasciitis, epidermolysis bullosa acquisita, erythema nodosum, essential mixed cryoglobulinemia, Evans' syndrome, Felty's syndrome, fibromyalgia, gastritis, pemphigoid of pregnancy, giant cell arteritis, Goodpasture's syndrome, Graves' disease, Graves' ophthalmopathy, Guillain-Barré syndrome, Hashimoto's encephalopathy, Hashimoto's thyroiditis, Henot's syndrome Hochönlein purpura, hidradenitis suppurativa, idiopathic dilated cardiomyopathy, idiopathic inflammatory demyelinating disease, IgA nephropathy, IgG4-related systemic disease, inclusion body myositis, inflammatory bowel disease (IBD), intermedium uveitis, interstitial cystitis, juvenile arthritis, Kawasaki disease, Lambert-Eaton myasthenic syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosus, lignified conjunctivitis, linear IgA disease, lupus nephritis, lupus vasculitis, Lyme disease, Meniere's disease, microscopic colitis, microscopic polyangiitis, mixed connective tissue disease, Mooren's ulcer, scleroderma, Mucha-Habermann's disease, multiple sclerosis,Myasthenia gravis, myocarditis, myositis, neuromyelitis optica, neuromyotonia, opsoclonus-myoclonus syndrome, optic neuritis, Ord's thyroiditis, relapsing rheumatism, paraneoplastic cerebellar degeneration, Parry-Romberg syndrome, Parsonage-Turner syndrome, streptococcal-associated pediatric autoimmune neuropsychiatric disorder, pemphigus vulgaris, pernicious anemia, acute pityriasis lichenoides, POEMS syndrome, polyarteritis nodosa, polymyalgia rheumatica, polymyositis, post-myocardial infarction syndrome, post-pericardiotomy syndrome, primary biliary cirrhosis, primary immunodeficiency syndrome, primary sclerosing cholangitis, progressive inflammatory neuropathy, psoriasis, psoriatic arthritis, red bud syndrome Genes that are overexpressed or underexpressed in aplasia, pyoderma gangrenosum, Raynaud's phenomenon, reactive arthritis, relapsing polychondritis, restless legs syndrome, retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, rheumatic vasculitis, sarcoidosis, Schnitzler's syndrome, scleroderma, Sjogren's syndrome, stiff-person syndrome, subacute bacterial endocarditis, Susac's syndrome, Sydenham's chorea, sympathetic ophthalmia, systemic lupus erythematosus, systemic scleroderma, thrombocytopenia, Tolosa Hunt syndrome, transverse myelitis, ulcerative colitis, undifferentiated connective tissue disease, urticaria, urticarial vasculitis, vasculitis, or vitiligo.

[0229] In some embodiments, the target gene (e.g., target endogenous gene) is a cancer, such as acute leukemia, astrocytoma, bile duct cancer (cholangiocarcinoma), bone cancer, breast cancer, brain stem glioma, bronchioloalveolar cell lung cancer, adrenal gland cancer, cancer of the anal region, bladder cancer, cancer of the endocrine system, esophageal cancer, cancer of the head or neck, kidney cancer, parathyroid cancer, penile cancer, cancer of the pleura / peritoneum, salivary gland cancer, small intestine cancer, thyroid cancer, ureter cancer, urethra cancer, cervical cancer, endometrial cancer, , fallopian tube cancer, renal pelvis cancer, vaginal cancer, vulvar cancer, cervical cancer, chronic leukemia, colon cancer, colorectal cancer, cutaneous melanoma, ependymoma, epidermoid tumor, Ewing's sarcoma, gastric cancer, glioblastoma, glioblastoma multiforme, glioma, hematologic malignancies, hepatocellular (liver) carcinoma, hepatoma, Hodgkin's disease, intraocular melanoma, Kaposi's sarcoma, lung cancer, lymphoma, medulloblastoma, melanoma, meningioma, mesothelioma, multiple myeloma, muscle cancer, central nervous system (CNS) neoplasms, neuronal carcinoma The genes are overexpressed or underexpressed in tumors and their metastases, including small cell lung cancer, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pediatric malignancies, pituitary adenoma, prostate cancer, rectal cancer, renal cell carcinoma, soft tissue sarcoma, Schwannoma, skin cancer, spinal axis tumor, squamous cell carcinoma, gastric cancer, synovial sarcoma, testicular cancer, uterine cancer, or refractory versions or combinations of any of the above cancers.

[0230] In some embodiments, the target gene (e.g., a target endogenous gene) is a differentiation-related gene, such as, for example, SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30, CD50, AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1 alpha / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activators, STAT inhibitors -, STAT3, STAT4, STAT5a, STAT6, TSC22, DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, My ocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT activators, STAT inhibitors, STAT1, STAT3, TBX18, Twist-1, Twist-2, Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3alpha / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB activators, NFkB / IkB inhibitors,NFkB1, NFkB2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activators, STAT inhibitors, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, ZNF281, KLF2, KLF4, c-Maf, c-My c, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, TBX18, ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4alpha / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, SUZ12, TCF-3 / E2A, TCF7 / TCF1, Androgen R / NR3C4, AP-2 gamma, beta-catenin, beta-catenin inhibitor, Brachyury, CREB, ER alpha / NR3A1, ER beta / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI- 2, GLI-3, HIF-1 alpha / HIF1A, HIF-2 alpha / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1, or ZEB1.

[0231] In some embodiments, modulation of the expression level and / or epigenetic level (e.g., methylation level) of a target gene in a target cell (e.g., a muscle cell) may result in modification (e.g., upregulation or downregulation) of a gene downstream of the target gene (e.g., one or more downstream genes). In some examples, the target gene can be encoded by a D4Z4 repeat array (e.g., the target gene is DUX4), and the downstream genes whose expression is then altered (e.g., downregulated) can include, but are not limited to, ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, DEFB103, ZFN217, RNASEL, EIF2AK2, BMP2, SP1 P21, MYC, MURF1, ATROGIN1, CRYM, PRAMEF1, RFPL2, KHDC1, SPRYD5, TPRX1, HSPA2, FGFR3, SLC2A14, ID2, PVRL3, SFRS2B, THOC4, ZNHIT6, DBR1, TFIP11, FBXO33, USP29, TRIM23, SLC34A2, CSAG3, and / or PNMA6B.

[0232] In some embodiments, modulation of the expression level and / or epigenetic level (e.g., methylation level) of a target gene in a target cell can result in apoptosis of the target cell (e.g., muscle cell). In some examples, such modulation of a target gene can reduce stress in the target cell. For example, modulation of a target gene (e.g., DUX4) can result in downregulation of one or more stress-related markers in the target cell. Non-limiting examples of one or more stress-related markers include ACTH, glucocorticoid receptor, CRHR-1 / 2, POMC, prolactin, arginine vasopressin receptor V1a, superoxide dismutase 1, superoxide dismutase 2, peroxiredoxin-3, CCR5, iNOS, eNOS, heme oxygenase-2, cyclooxygenase-2, HSP27, HSP40, HSP60, HSP70, HSP70i, HSP90, HSP1 10, GRP78 / BIP, AIF, annexin II, annexin IV, caspase 1, caspase 2, caspase 3, caspase 6, cytokeratin, E-cadherin, and / or annexin V, caspase 5, caspase 7, caspase 8, caspase 9, caspase 10, BAD, BAX, BAK, BCL2, BID, PARP-1, NOXA, PUMA, RIPK3, RIPK1, FADD, APAF1, DFF40, DFF45, ROCK. One or more stress-related markers disclosed herein can be apoptosis markers.

[0233] In some embodiments, the heterologous gene effector is derived from a gene product that is a hematopoietic stem cell transcription factor. In some embodiments, the target gene is a mesenchymal stem cell transcription factor. In some embodiments, the target gene is an embryonic stem cell transcription factor. In some embodiments, the target gene is an induced pluripotent stem cell (iPSC) transcription factor. In some embodiments, the target gene is an epithelial stem cell transcription factor. In some embodiments, the target gene is a cancer stem cell transcription factor.

[0234] In some embodiments, the target gene is an aging-associated gene. In some embodiments, the target gene is a senescence-associated protein. In some embodiments, the target gene is a drug target.

[0235] In some embodiments, the target gene (e.g., the target endogenous gene) is a cancer-associated gene. Non-limiting examples of cancer-associated genes include A1CF, ABI1, ABL1, ABL2, ACKR3, ACSL3, ACSL6, ACVR1, ACVR2A, AFDN, AFF1, AFF3, AFF4, AKAP9, AKT1, AKT2, AKT3, ALDH2, ALK, AMER1, ANK1, APC, APOBEC3B, AR, ARAF, ARHGAP26, ARHGAP5, ARHGEF10, ARHGEF10L, ARHGEF12, ARID1A, ARID1B, ARID2, ARNT, ASPSCR1, ASXL1, AS XL2, ATF1, ATIC, ATM, ATP1A1, ATP2B3, ATR, ATRX, AXIN1, AXIN2, B2M, BAP1, BARD1, BAX, BAZ1A, BCL10, BCL11A, BCL11B, BCL2, BCL2L12, BCL3, BCL 6, BCL7A, BCL9, BCL9L, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BIRC6, BLM, BMP5, BMPR1A, BRAF, BRCA1, BRCA2, BRD3, BRD4, BRIP1, BTG1, BTK, BUB1B, C1 5orf65, CACNA1D, CALR, CAMTA1, CANT1, CARD11, CARS, CASP3, CASP8, CASP9, CBFA2T3, CBFB, CBL, CBLB, CBLC, CCDC6, CCNB1IP1, CCNC, CCND1, CCN D2, CCND3, CCNE1, CCR4, CCR7, CD209, CD274, CD28, CD74, CD79A, CD79B, CDC73, CDH1, CDH10, CDH11, CDH17, CDK12, CDK4, CDK6, CDKN1A, CDKN1B, CD KN2A, CDKN2C, CDX2, CEBPA, CEP89, CCHHD7, CHD2, CHD4, CHEK2, CHIC2, CHST11, CIC, CIITA, CLIP1, CLP1, CLTC, CLTCL1, CNBD1, CNBP, CNOT3, CNTN AP2, CNTRL, COL1A1, COL2A1, COL3A1, COX6C, CPEB3, CREB1, CREB3L1, CREB3L2, CREBBP, CRLF2, CRNKL1, CRTC1, CRTC3, CSF1R, CSF3R, CSMD3, CTCF,<h2 style=";text-align:left;direction:ltr">CTNNA2、CTNNB1、CTNND1、CTNND2、CUL3、CUX1、CXCR4、CYLD、CYP2C8、CYSLTR 2、DAXX、DCAF12L2、DCC、DCTN1、DDB2、DDIT3、DDR2、DDX10、DDX3X、DDX5、DDX6 、DEK、DGCR8、DICER1、DNAJB1、DNM2、DNMT3A、DROSHA、DUX4L1、EBF1、ECT2L、 EED、EGFR、EIF1AX、EIF3E、EIF4A2、ELF3、ELF4、ELK4、ELL、ELN、EML4、EP300、 EPAS1、EPHA3、EPHA7、EPS15、ERBB2、ERBB3、ERBB4、ERC1、ERCC2、ERCC3、ERC C4、ERCC5、ERG、ESR1、ETNK1、ETV1、ETV4、ETV5、ETV6、EWSR1、EXT1、EXT2、EZH 2、EZR、FAM131B、FAM135B、FAM47C、FANCA、FANCC、FANCD2、FANCE、FANCF、FA NCG、FAS、FAT1、FAT3、FAT4、FBLN2、FBXO11、FBXW7、FCGR2B、FCRL4、FEN1、FES 、FEV、FGFR1、FGFR1OP、FGFR2、FGFR3、FGFR4、FH、FHIT、FIP1L1、FKBP9、FLCN 、FLI1、FLNA、FLT3、FLT4、FNBP1、FOXA1、FOXL2、FOXO1、FOXO3、FOXO4、FOXP1、 FOXR1, FSTL3, FUBP1, FUS, GAS7, GATA1, GATA2, GATA3, GLI1, GMPS, GNA11, GNAQ, GNAS, GOLGA5, GOPC, GPC3, GPC5, GPHN, GRIN2A, GRM3, H3F3A, H3F3B, HER PUD1、HEY1、HIF1A、HIP1、HIST1H3B、HIST1H4I、HLA-A、HLF、HMGA1、HMGA2、H MGN2P46、HNF1A、HNRNPA2B1、HOOK3、HOXA11、HOXA13、HOXA9、HOXC11、HOXC13 、HOXD11、HOXD13、HRAS、HSP90AA1、HSP90AB1、ID3、IDH1、IDH2、IGF2BP2、IG H、IGK、IGL、IKBKB、IKZF1、IL2、IL21R、IL6ST、IL7R、IRF4、IRS4、ISX、ITGAV、ITK、JAK1、JAK2、JAK3、JAZF1、JUN、KAT6A、KAT6B、KAT7、KCNJ5、KDM5A、KDM5 C, KDM6A, KDR, KDSR, KEAP1, KIAA1549, KIF5B, KIT, KLF4, KLF6, KLK2, KMT2A KMT2C, KMT2D, KNL1, KNSTRN, KRAS, KTN1, LARP4B, LASP1, LATS1, LATS2, LC K, LCP1, LEF1, LEPROTL1, LHFPL6, LIFR, LMNA, LMO1, LMO2, LPP, LRIG3, LRP1B LSM14A, LYL1, LZTR1, MACC1, MAF, MAFB, MALAT1, MALT1, MAML2, MAP2K1, MA P2K2, MAP2K4, MAP3K1, MAP3K13, MAPK1, MAX, MB21D2, MDM2, MDM4, MDS2, MECO MED12, MEN1, MET, MGMT, MITF, MLF1, MLH1, MLLT1, MLLT10, MLLT11, MLLT3 MLLT6, MN1, MNX1, MPL, MRTFA, MSH2, MSH6, MSI2, MSN, MTCP1, MTOR, MUC1, MU C16, MUC4, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, MYH11, MYH9, MYO5A, MYOD1 N4BP2, NAB2, NACA, NBEA, NBN, NCKIPSD, NCOA1, NCOA2, NCOA4, NCOR1, NCOR2 NDRG1, NF1, NF2, NFATC2, NFE2L2, NFIB, NFKB2, NFKBIE, NIN, NKX2-1, NONO NOTCH1, NOTCH2, NPM1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1 NTRK1, NTRK3, NUMA1, NUP214, NUP98, NUTM1, NUTM2B, NUTM2D, OLIG2, OMD, P 2RY8, PABPC1, PAFAH1B2, PALB2, PATZ1, PAX3, PAX5, PAX7, PAX8, PBRM1, PBX1 PCBP1, PCM1, PDCD1LG2, PDE4DIP, PDGFB, PDGFRA, PDGFRB, PER1, PHF6, PHO X2B, PICALM, PIK3CA, PIK3CB, PIK3R1, PIM1, PLAG1, PLCG1, PML, PMS1, PMS2.<h2 style=";text-align:left;direction:ltr">POLD1、POLE、POLG、POLQ、POT1、POU2AF1、POU5F1、PPARG、PPFIBP1、PPM1D、P PP2R1A、PPP6C、PRCC、PRDM1、PRDM16、PRDM2、PREX2、PRF1、PRKACA、PRKAR1A 、PRKCB、PRPF40B、PRRX1、PSIP1、PTCH1、PTEN、PTK6、PTPN11、PTPN13、PTPN6 、PTPRB、PTPRC、PTPRD、PTPRK、PTPRT、PWWP2A、QKI、RABEP1、RAC1、RAD17、RAD 21、RAD51B、RAF1、RALGDS、RANBP2、RAP1GDS1、RARA、RB1、RBM10、RBM15、REC QL4、REL、RET、RFWD3、RGPD3、RGS7、RHOA、RHOH、RMI2、RNF213、RNF43、ROBO2、 ROS1, RPL10, RPL22, RPL5, RPN1, RSPO2, RSPO3, RUNX1, RUNX1T1, S100A7, SALL4, SBDS, SDC4, SDHA, SDHAF2, SDHB, SDHC, SDHD, 44444, 44445, 44448, SET SETBP1, SETD1B, SETD2, SETDB1, SF3B1, SFPQ, SFRP4, SGK1, SH2B3, SH3GL1, SHTN1, SIRPA, SIX1, SIX2, SKI, SLC34A2, SLC45A3, SMAD2, SMAD3, SMAD4, SMA RCA4、SMARCB1、SMARCD1、SMARCE1、SMC1A、SMO、SND1、SNX29、SOCS1、SOX2、S OX21、SPECC1、SPEN、SPOP、SRC、SRGAP3、SRSF2、SRSF3、SS18、SS18L1、SSX1、S SX2、SSX4、STAG1、STAG2、STAT3、STAT5B、STAT6、STIL、STK11、STRN、SUFU、S UZ12、SYK、TAF15、TAL1、TAL2、TBL1XR1、TBX3、TCEA1、TCF12、TCF3、TCF7L2、T CL1A、TEC、TENT5C、TERT、Tet1、Tet2、TFE3、TFEB、TFG、TFPT、TFRC、TGFBR2、 THRAP3、TLX1、TLX3、TMEM127、TMPRSS2、TNC、TNFAIP3、TNFRSF14、TNFRSF17、TOP1, TP53, TP63, TPM3, TPM4, TPR, TRA, TRAF7, TRB, TRD, TRIM24, TRIM27, TRIM33, TRIP11, TRRAP, TSC1, TSC2, TSHR, U2AF1, UBR5, USP44, USP6, USP8, VAV1, VHL, VTI1A, WAS, WDCP, WIF1, WNK2, WRN, WT1, WWTR1, XPA, XPC, XPO1, YWHAE, ZBTB16, ZCCHC8, ZEB1, ZFHX3, ZMYM2, ZMYM3, ZNF331, ZNF384, ZNF429, ZNF479, ZNF521, ZNRF3, and ZRSR2.

[0236] In some embodiments, the target gene (e.g., a target endogenous gene) is an immune cell-associated gene, such as a cytokine, a cytokine receptor, a chemokine, a chemokine receptor, a co-inhibitory immune receptor, a co-stimulatory immune receptor, an immune cell transcription factor, etc.

[0237] In some embodiments, the target gene (e.g., target endogenous gene) is a cytokine, e.g., 4-1BBL, APRIL, CD153, CD154, CD178, CD70, G-CSF, GITRL, GM-CSF, IFN-α, IFN-β, IFN-γ, IL-1RA, IL-1α, IL-1β, IL-2, IL-3, IL-4, IL-5, IL-6, IL-7, IL-9, IL- 10, IL-11, IL-12, IL-13, IL-14, IL-15, IL-16, IL-17, IL-18, IL-20, IL-23, LIF, LIGHT, LT-β, ​​M-CSF, MS P, OSM, OX40L, SCF, TALL-1, TGF-β, TGF-β1, TGF-β2, TGF-β3, TNF-α, TNF-β, TRAIL, TRANCE, or TWEAK.

[0238] In some embodiments, the target gene (e.g., a target endogenous gene) is a cytokine receptor, e.g., common gamma chain receptor, common beta chain receptor, interferon receptor, TNF family receptor, TGF-B receptor, Apo3, BCMA, CD114, CD115, CD116, CD117, CD118, CD120, CD120a, CD120b, CD121, CD121a, CD121b, CD122, CD123, CD124, CD126, CD127, CD130, CD131, CD132, CD212, CD213, CD213a1, CD213a13, CD213a2, CD25, CD27, CD30 , CD4, CD40, CD95(Fas), CDw119, CDw121b, CDw125, CDw131, CDw136, CDw137(41BB), CDw210, CDw217, GITR, HVEM, IL-11R, IL-11Ra, IL-14R, IL-15R, IL-15Ra, IL-18R, IL -18Rα, IL-18Rβ, IL-20R, IL-20Rα, IL-20Rβ, IL-9R, LIFR, LTβR, OPG, OSMR, OX40, RANK, TACI, TGF-βR1, TGF-βR2, TGF-βR3, TRAILR1, TRAILR2, TRAILR3, or TRAILR4.

[0239] In some embodiments, the target gene (e.g., target endogenous gene) is a chemokine, e.g., ACT-2, AMAC-a, ATAC, ATAC, BLC, CCL1, CCL11, CCL13, CCL14, CCL15, CCL16, CCL17, CCL18, CCL19, CCL2, CCL20, CCL21, CCL22, CCL23, CCL24, CCL25, CCL26, CCL27, CCL3, CCL4, CCL5, CCL7, CCL 8, CKb-6, CKb-8, CTACK, CX3CL1, CXCL1, CXCL10, CXCL11, CXCL12, CXCL13, CXCL14, CXCL2, CXCL3, CXCL4, CXCL5, CXCL6, CXCL7, CXCL8, CXCL9, DC-CK1, ELC, ENA-78, Eotaxin, Eotaxin-2, Eotaxin-3, Escaine, Exodus-1, Exodus-2, Exodus-3, Fractalkine , GCP-2, GROa, GROb, GROg, HCC-1, HCC-2, HCC-4, I-309, IL-8, ILC, IP-10, I-TAC, LAG-1, LARC, LCC-1, LD78α, LEC, Lkn-1, LMC, lymphotactin, lymphotactin b, MCAF, MCP-1, MCP-2, MCP-3, MCP-4, MDC, MDNCF, MGSA-a, MGSA-b, MGSA-g, M ig, MIP-1d, MIP-1α, MIP-1β, MIP-2a, MIP-2b, MIP-3, MIP-3α, MIP-3β, MIP-4, MIP-4a, MIP-5, MPIF-1, MPIF-2, NAF, NAP-1, NAP-2, oncostatin, PARC, PF4, PPBP, RANTES, SCM-1a, SCM-1b, SDF-1α / β-, SLC, STCP-1, TARC, TECK, XCL1, or XCL2.

[0240] In some embodiments, the target gene (e.g., a target endogenous gene) is a chemokine receptor, such as CCR1, CCR2, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCR10, CX3CR1, CXCR1, CXCR2, CXCR3, CXCR4, CXCR5, XCR1, or XCR1.

[0241] In some embodiments, the target gene (e.g., target endogenous gene) is an activating NK receptor, e.g., CD100 (SEMA4D), CD16 (FcgRIIIA), CD160 (BY55), CD244 (2B4, SLAMF4), CD27, CD94-NKG2C, CD94-NKG2E, CD94-NKG2H, CD96, CRTAM, DAP12, DNAM1 (CD226), KIR2DL4, K IR2DS1, KIR2DS2, KIR2DS3, KIR2DS4, KIR2DS5, KIR3DS1, Ly49, NCR, NKG2D (KLRK1, CD314), NKp30 (NCR3), NKp44 (NCR2), NKp46 (NCR1), NKp80 (KLRF1, CLEC5C), NTB-A (SLAMF6), PSGL1, or SLAMF7 (CRACC, CS1, CD319).

[0242] In some embodiments, the target gene (e.g., target endogenous gene) is an inhibitory NK receptor, e.g., CD161 (NKR-P1A, NK1.1), CD94-NKG2A, CD96, CEACAM1, KIR2DL1, KIR2DL2, KIR2DL3, KIR2DL4, KIR2DL5A, KIR2DL5B, KIR3DL1, KIR3DL2, KIR3DL3, KLRG1, LAIR1, LIR1 (ILT2, LILRB1), Ly49a , Ly49b, NKR-P1A (KLRB1), SIGLEC-10, SIGLEC-11, SIGLEC-14, SIGLEC-16, SIGLEC-3 (CD33), SIGLEC-5 (CD170), SIGLEC -6 (CD327), SIGLEC-7 (CD328), SIGLEC-8, SIGLEC-9 (CD329), SIGLEC-E, SIGLEC-F, SIGLEC-G, SIGLEC-H, or TIGIT.

[0243] In some embodiments, the target gene (e.g., a target endogenous gene) is a co-inhibitory immune receptor, e.g., 2B4, B7-1, BTLA, CD160, CTLA-4, DR6, Fas, LAG3, LAIR1, Ly108, PD-1, PD-L1, PD1H, TIGIT, TIM1, TIM2, or TIM3.

[0244] In some embodiments, the target gene (e.g., a target endogenous gene) is a costimulatory immune receptor, e.g., 2B4, 4-1BB, CD2, CD4, CD8, CD21, CD27, CD28, CD30, CD40, CD84, CD226, CD355, CRACC, DcR3, DR3, GITR, HVEM, ICOS, Ly9, Ly108, LIGHT, LTβR, OX40, SLAM, TIM1, or TIM2.

[0245] In some embodiments, the target gene (e.g., the target endogenous gene) is itself a genetic effector, such as any of the genetic effectors disclosed herein (e.g., a transcription factor disclosed herein).

[0246] In some embodiments, the target gene (e.g., a target endogenous gene) is an immune cell transcription factor, e.g., AP-1, Bcl6, E2A, EBF, Eomes, FoxP3, GATA3, Id2, Ikaros, IRF, IRF1, IRF2, IRF3, IRF3, IRF7, NFAT, NFkB, Pax5, PLZF, PU.1, ROR-gamma-T, STAT, STAT1, STAT2, STAT3, STAT4, STAT5, STAT5A, STAT5B, STAT6, T-bet, TCF7, or ThPOK.

[0247] In some embodiments, the target gene is a kinase, such as a tyrosine kinase or a serine / threonine kinase. In some embodiments, the target gene is a phosphatase, such as a tyrosine phosphatase or a serine / threonine phosphatase.

[0248] In some embodiments, the target gene is a receptor. In some embodiments, the target gene is an ion channel. In some embodiments, the target gene is a GPCR. In some embodiments, the target gene is a receptor tyrosine kinase. In some embodiments, the target gene is a ribosomal protein. In some embodiments, the target gene is a membrane protein. In some embodiments, the target gene is a cytoplasmic protein. In some embodiments, the target gene is a nuclear protein. In some embodiments, the target gene is a mitochondrial protein. In some embodiments, the target gene is a ubiquitin ligase. In some embodiments, the target gene is a methyltransferase. In some embodiments, the target gene is a glycosyltransferase. In some embodiments, the target gene is a hydrolase.

[0249] In some embodiments, CD45 is the target gene used in the compositions and methods of the present disclosure (e.g., gene expression activation screening). In some embodiments, CD45 is not used as a target gene. The compositions and methods disclosed herein for identifying complexes that modulate CD45 expression can be modified and adapted to other target genes (e.g., target endogenous genes), including the genes disclosed herein.

[0250] In some embodiments, CD71 is the target gene used in the compositions and methods of the present disclosure (e.g., gene expression reduction screening). In some embodiments, CD71 is not used as a target gene. The compositions and methods disclosed herein for identifying complexes that modulate CD71 expression can be modified and adapted to other target genes (e.g., target endogenous genes), including the genes disclosed herein.

[0251] cell

[0252] The compositions, methods and systems of the present disclosure can be applied to various types of cells and their populations.For example, the complexes of the present disclosure can be used to induce the expression, epigenetic modification or activity level change of target genes (e.g., target endogenous genes) in specific types of cells or their populations.The methods of the present disclosure can be used to identify complexes that can induce the expression or activity change of target genes (e.g., target endogenous genes) in specific types of cells or their populations.

[0253] In some embodiments, the complex or heterologous gene effector identified by the method of the present disclosure results in a desired change in expression of a target gene (e.g., a target endogenous gene) that is specific to a particular cell type. In some embodiments, the complex or heterologous gene effector identified by the method of the present disclosure results in a desired change in expression of a target gene (e.g., a target endogenous gene) that is applicable to two or more cell types. In some embodiments, the complex or heterologous gene effector identified by the method of the present disclosure results in a desired change in expression of a target gene (e.g., a target endogenous gene) that is applicable to three or more cell types. In some embodiments, the complex or heterologous gene effector identified by the method of the present disclosure results in a desired change in expression of a target gene (e.g., a target endogenous gene) that is applicable to a class of cell types, such as cell types that reside in similar tissues or are derived from the same or similar differentiation lineages, and have overlapping functional roles, such as stem cells, immune cells, T cells, effector T cells, etc. In some embodiments, the complexes or heterologous gene effectors identified by the methods of the present disclosure are broadly applicable to a wide variety of cell types, e.g., when introduced into cells using suitable methods, induce expression levels of the target gene that are above or below a certain threshold in multiple target cell types, resulting in a desired change in expression of a target gene (e.g., a target endogenous gene).

[0254] In some embodiments, the disclosed compositions, complexes, systems, or methods are used to effect changes in expression, epigenetic modification, or activity levels of target genes in primary cells. In some embodiments, the disclosed compositions, complexes, systems, or methods are used to effect changes in expression, epigenetic modification, or activity levels of target genes in cell lines. In some embodiments, the disclosed compositions, complexes, systems, or methods are used to effect changes in expression, epigenetic modification, or activity levels of target genes in immortalized cells.

[0255] In some embodiments, the disclosed compositions, complexes, systems, or methods are used to effect a change in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in a mammalian cell, such as a human cell, a non-human primate cell, a non-rodent mammalian cell, a non-human mammalian cell, a porcine cell, a rabbit cell, a canine cell, etc. In some embodiments, the disclosed compositions, complexes, systems, or methods are used to effect a change in expression, epigenetic modification, or activity levels of a target gene in a plant cell, an avian cell, a reptilian cell, a bacterial cell, or an archaeal cell.

[0256] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modifications, or activity levels of a target gene (e.g., a target endogenous gene) in a human cell.

[0257] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (eg, a target endogenous gene) in a stem cell.

[0258] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in a differentiated cell.

[0259] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in disease-associated cells.

[0260] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in a cancer cell.

[0261] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in a non-cancer cell.

[0262] The disclosed compositions, conjugates, systems, or methods can be used to treat lymphoid cells, such as B cells, T cells (cytotoxic T cells, natural killer T cells, regulatory T cells, helper T cells), natural killer cells, cytokine-induced killer (CIK) cells (see, e.g., US20080241194); myeloid cells, such as granulocytes (basophilic granulocytes, eosinophilic granulocytes, neutrophilic granulocytes / hypersegmented neutrophils), monocytes / macrophages, erythrocytes ... Blood cells, reticulocytes, mast cells, platelets / megakaryocytes, dendritic cells; cells derived from the endocrine system, including thyroid (thyroid epithelial cells, parafollicular cells), parathyroid (parathyroid chief cells, oxyphyllin cells), adrenal (chromaffin cells), and pineal (pineal cells) cells; glial cells (astrocytes, microglia), magnocellular neurosecretory cells, stellate cells, Boettcher cells, and pituitary (gonadotropes, corticotropes, thyrotropes, somatotropes, prolactin receptors, and thyroid gland) cells. cells of the nervous system, including inflammatory bowel disease (IGDs); cells of the respiratory system, including lung cells (type I pneumocytes, type II pneumocytes), Clara cells, goblet cells, and dust cells; cells of the circulatory system, including cardiac muscle cells and pericytes; cells of the digestive system, including stomach (chief cells, parietal cells), goblet cells, Paneth cells, G cells, D cells, ECL cells, I cells, K cells, and S cells; enteroendocrine cells, including enterochromaffin cells, APUD cells, liver cells (e.g., hepatocytes, or Kupffer cells), and cartilage / bone / muscle; osteoblasts Bone cells, including osteocytes, osteoclasts, dental cells (cementoblasts, ameloblasts);chondrocytes, including chondroblasts, chondrocytes;skin cells, including hair cells, keratinocytes, melanocytes (nevus cells);muscle cells, including muscle cells;cells of the urinary system, including podocytes, juxtaglomerular cells, intraglomerular / extrauterine mesangial cells, renal proximal tubule brush border cells, macula densa cells;cells of the reproductive system, including sperm, Sertoli cells, Leydig cells, and oocytes;and adipocytes, fibroblasts, tendon cells, epithelial keratinocytes, epithelial basal cells, fingernail and toenail keratinocytes, nail bed basal cells, hair medulla cells, cortical cells, hair cuticle cells, root sheath cuticle cells, root sheath cells of Huxley's layer, root sheath cells of Henle's layer, outer root sheath cells, hair matrix cells, moist stratified barrier epithelial cells, surface epithelial cells of the stratified squamous epithelium of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra, and vagina, and Basal cells of epithelium, urothelial cells, exocrine epithelial cells, mucous cells of salivary glands, serous cells of salivary glands, von Ebner's gland cells of the tongue, mammary gland cells, lacrimal gland cells, ear canal gland cells, dark cells of eccrine sweat glands, light cells of eccrine sweat glands, apocrine sweat gland cells, Molar gland cells of the eyelid, sebaceous gland cells, Bowman's gland cells of the nose, Brunner's gland cells of the duodenum, seminal vesicle cells, prostatic cells, bulbar urethral gland cells, Bartholin's gland cells, urethral gland cells, endometrial cells, isolated goblet cells of the respiratory and digestive tracts, mucus in the stomach lining cells, gastric gland zymogen cells, gastric acid-secreting cells, pancreatic acinar cells, small intestinal Paneth cells, type II pneumocytes of the lung, pulmonary Clara cells, hormone-secreting cells, anterior pituitary cells, somatotropes, lactotropes, thyrotropes, gonadotropes, corticotropes, intermediate pituitary cells, large cell neurosecretory cells, intestinal and respiratory tract cells, thyroid cells, thyroid epithelial cells, parafollicular cells, parathyroid cells, chief parathyroid cells, acidophilic cells, adrenal cells, chromaffin cells, Leiden cells of the testes Wick cells, theca cells of the ovarian follicle, luteal cells of ruptured follicles, granulosa lutein cells, luteal cells, juxtaglomerular cells, macula densa cells of the kidney, metabolic and storage cells, barrier function cells (e.g., lung, gastrointestinal tract, exocrine glands and urogenital tract), kidney, type I pneumocytes, pancreatic duct cells (central acinar cells), non-striated duct cells (sweat glands, salivary glands, mammary glands, etc.), duct cells (seminal vesicles, prostate, etc.), epithelial cells lining closed internal body cavities, ciliated cells with propulsive function, extracellular matrix secreting cells, contractile cells;Skeletal muscle cells, stem cells, cardiomyocytes, cells of the blood and immune system, erythrocytes, megakaryocytes, monocytes, connective tissue macrophages (various types), epidermal Langerhans cells, osteoclasts, dendritic cells, microglial cells, neutrophil granulocytes, eosinophil granulocytes, basophil granulocytes, mast cells, helper T cells, suppressor T cells, cytotoxic T cells, natural killer T cells, B cells, natural killer cells, reticulocytes, stem cells and committed progenitor cells of the blood and immune system (various types), pluripotent stem cells, totipotent stem cells, induced pluripotent stem cells, adult stem cells, sensory transduction cells, neurons, autonomic nerve cells, The expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) can be altered in sensory organs and peripheral neuronal support cells, central nervous system neurons and glial cells, lens cells, pigment cells, melanocytes, retinal pigment epithelial cells, germ cells, oogonia / egg cells, sperm, spermatocytes, spermatogonia, sperm, nursery cells, follicular cells, Sertoli cells, thymic epithelial cells, stromal cells, interstitial kidney cells, common bone marrow progenitor cells, common lymphoid progenitor cells, and stem cells that are or will be differentiated into any of the cell types disclosed herein;

[0263] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in stem cells, such as isolated stem cells (e.g., ESCs) or artificial stem cells (e.g., iPSCs).

[0264] The disclosed compositions, complexes, systems, or methods can be used to effect changes in expression, epigenetic modification, or activity levels of a target gene (e.g., a target endogenous gene) in hematopoietic stem cells, e.g., hematopoietic stem cells of a subject, e.g., bone marrow or peripheral blood (e.g., mobilized by administration of a mobilized peripheral blood apheresis product, e.g., GCSF, GM-CSF, Mozobil, or a combination thereof).

[0265] In some examples, the pluripotency of stem cells (e.g., ESCs or iPSCs) can be determined in part by assessing the pluripotency characteristics of the cells. Pluripotency characteristics may include, but are not limited to, pluripotent stem cell morphology; unlimited self-renewal capacity; expression of pluripotent stem cell markers, including, but not limited to, SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30 and / or CD50; differentiation capacity into all three somatic cell lineages (ectoderm, mesoderm, and endoderm); ability to form teratomas containing the three somatic cell lineages; and / or (vi) formation of embryoid bodies containing cells from the three somatic cell lineages.

[0266] The disclosed compositions, conjugates, systems, or methods can be used to treat immune cells, such as lymphocytes, T cells, CD4+ T cells, CD8+ T cells, alpha-beta T cells, gamma-delta T cells, regulatory T cells (Tregs), cytotoxic T lymphocytes, Th1 cells, Th2 cells, Th17 cells, Th9 cells, naive T cells, memory T cells, effector T cells, effector-memory T cells (TEMs), central memory T cells (TCMs), resident memory T cells (TRMs), follicular helper T cells (TFHs), natural killer T cells (NKTs), tumor-infiltrating lymphocytes (TILs), natural killer cells (NKs), innate lymphoid cells (ILs), and the like. C) The expression, epigenetic modification, or activity level of a target gene (e.g., a target endogenous gene) can be altered in ILC1 cells, ILC2 cells, ILC3 cells, lymphoid tissue-derived (LTi) cells, B cells, B1 cells, B1a cells, B1b cells, B2 cells, plasma cells, regulatory B cells, memory B cells, marginal zone B cells, follicular B cells, germinal center B cells, antigen presenting cells (APCs), monocytes, macrophages, M1 macrophages, M2 macrophages, tissue-associated macrophages, dendritic cells, plasmacytoid dendritic cells, neutrophils, mast cells, basophils, eosinophils, common myeloid precursors, common lymphoid precursors, or any combination thereof.

[0267] The compositions, complexes, systems, or methods of the present disclosure can be used to effect changes in expression, epigenetic modification, or activity levels of engineered cells used to manufacture biologics, such as antibodies or other protein-based therapeutics. EXAMPLES

[0268] Example 1 Regulation of DUX4 expression in target cell populations A lymphoblast population was used as an example target cell population for modulating DUX4 expression levels by the compositions and methods disclosed herein. A lymphoblast population was contacted with (i) a heterologous actuator moiety linked to a genetic regulator, and (ii) a guide RNA (see Table 1) designed to direct the heterologous actuator moiety to a target polynucleotide sequence between two CpG islands in the D4Z4 repeat array encoding DUX4 in the lymphoblast population. Some guide RNAs were able to enable the heterologous actuator moiety linked to the genetic regulator to form a complex with its respective target polynucleotide sequence in the lymphoblast population, resulting in between about 0.2-fold and about 0.8-fold reduction in DUX4 expression levels (Figure 2). Table 1. Guide RNA molecules used in Figure 2 [Table 1-1] [Table 1-2] Example 2 Validation of in vitro FSHD models

[0269] To develop an in vitro FSHD model, two immortalized patient-derived FSHD skeletal myoblast (SkM) cells, 12ABIC (12A) and 15ABIC (15A), were expanded in full growth medium. After confluence, the full growth medium was changed to differentiation medium conditions for either 2 or 7 days (D2 or D7 in Figure 3A). The experiment included a negative control, which was undifferentiated proliferating myoblasts (UD in Figure 3A). After differentiation, total RNA was extracted from myoblasts and qRT-PCR was performed using TaqMan probes for DUX4 and DUX4 target genes ZSAC4, LEUTZ, MBD3L2, TRIM48, and TRIM43. GAPDH was included as an internal reference control for qRT-PCR measurements, and the double delta Ct method was used to calculate gene fold changes. 12ABIC and 15ABIC cells showed increased DUX4 and DUX4 target gene expression consistent with FSHD symptoms in patients.

[0270] Next, 12ABIC and 15ABIC cells were tested as to whether they also exhibited increased apoptosis consistent with the FSHD phenotype in patients. 12ABIC and 15ABIC cells, as well as two matched healthy control cells 12UBIC and 15VBIC, were grown and then differentiated for 2 days and stained for the apoptosis marker, caspase 3. The assay also included DAPI staining as a positive control. The cells were then imaged and analyzed using a CellXpress PICO imager. As shown in Figure 3B, 12ABIC and 15ABIC cells had increased levels of apoptosis compared to their healthy sibling control myoblasts 12UBIC and 15VBIC after 2 days of differentiation. 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells were grown, differentiated, stained, and imaged and analyzed using a CellXpress PICO imager on day 7 of differentiation. The percent of apoptotic cells for each cell type was plotted at days 0, 1, 2, and 7 of differentiation (FIG. 3C). 12ABIC and 15ABIC cells had higher apoptosis levels at days 2 and 7 compared to their corresponding healthy controls. This increase in apoptosis during differentiation was consistent with the in vivo phenotype of FSHD.

[0271] To ensure that 12ABIC and 15ABIC cells had similar myoblast differentiation compared to their healthy sibling controls, all four cell types were immunostained for myosin heavy chain (MYHC), a late muscle gene used as a marker for myocyte differentiation (Figure 3D). The immunostaining assay also included DAPI staining as a positive control. In addition, the four cell types were assayed via qRT-PCR for expression of MYHC, MYOG, and MyoMaker (MYMK) (Figure 3E). GADPH was included in the qRT-PCR and was an internal control for the qRT-PCR measurements. MYOG is an essential myogenic regulator that regulates skeletal muscle differentiation, and MYMK was a late muscle gene used as a marker for myocyte differentiation. The results of both the immunostaining and qRT-PCR experiments showed that 12ABIC and 15ABIC cells had similar differentiation to their corresponding healthy sibling controls.

[0272] Overall, the results of these validation experiments demonstrated that 12ABIC and 15ABIC cells presented the in vivo phenotype of FSHD myoblasts and were thus suitable in vitro models of FSHD. Example 3 Targeting DUX4 to downregulate expression

[0273] To target DUX4 for downregulation, multiple gRNAs were designed across the entire DZ4Z locus, including the region encoding DBET, a long non-coding RNA at the 5' end of the D4Z4 locus. The DZ4Z locus is known to upregulate DUX4 gene expression upon depletion of repeat units, and DBET lncRNAs have been shown to positively regulate expression of DUX4 from the DZ4Z locus. gRNAs were designed using the ChopChop CRISPR guide design tool. When designing gRNAs, the Hg38 human genome assembly was used with TTTR as the PAM sequence requirement. A map of gRNAs designed to the DZ4Z locus is shown in Figure 4.

[0274] After designing the gRNA, different gRNAs were tested with Cas12f variant constructs bound to KRAB modulator. 12ABIC cells stably expressing Cas12f-KRAB effector-modulator were generated after lentiviral transduction of 12ABIC myoblasts. The design of Cas12f-KRAB effector-modulator vector is shown in Figure 5. The vector contained a muscle-specific promoter (CK8e) to drive the expression of Cas12f variant effector, as well as KRAB and DNMT3L domains. A human U6g promoter was included to drive the expression of sgRNA spacer sequences along with the scaffold driven by RNA polymerase III. The vector further contained modified WPRE and polyadenylation regulatory sequences. Cas12f-KRAB effector-modulator was labeled by mCherry, so after transduction, mCherry+ cells were selected for enrichment. After enrichment, the annealed crRNA:trcrRNA construct for 78 guides was nucleofected in Cas12f-KRAB effector-modulator expressing 12ABIC cells. After myoblast differentiation for 7 days (for 7 cells), cells were assayed for the expression of DUX4 (Figures 6A and 6B) and MDB3L2 (Figure 6B) using Quantigene assay probes. The relative expression of DUX4 was normalized to the expression of the control gene HPRT1. The experiment showed that different gRNAs were able to downregulate the expression of DUX4 and MDB3L2 in cells expressing Cas12f-KRAB effector-modulator. In addition, the downregulation of DUX4 and MDB3L2 by different gRNAs was positively correlated (Figure 6B).

[0275] Six gRNAs from the initial screen were further tested. The six gRNAs were transfected into immortalized patient-derived FSHD myoblasts along with one of two different Cas12f-KRAB effector-modulators. The two different Cas12f-KRAB effector-modulators contained one of two different DNMT3L domains (e.g., DNMT3L-Kla or DNMT3L-Klb). After transfection, cells were differentiated and expression of DUX4, as well as DUX4 target genes, MBD3L2, TRIM48, and MYOG, was assayed using qRT-PCR 17 days (Figure 7A) and 18 days (Figure 7B) after transfection to measure sustained DUX4 suppression. MYOG was included as a positive control to ensure that the differentiation potential of DUX4 sgRNA-transfected cells was similar to control sgRNA-transfected myoblasts. Overall, it was found that Cas12f-KRAB-DNMT3L modulators led to sustained repression of DUX4 and DUX4 target genes.

[0276] In addition to testing the expression levels of DUX4 and DUX4 target genes in patient-derived myoblasts treated with Cas12f-KRAB-DNMT3L modulators, cells were also tested for the apoptosis levels of treated cells. Treated cells were stained for the apoptosis marker caspase 3 after 2 days of differentiation. After staining, apoptosis-positive cells were counted using a high content imager, and the percent positive cells were calculated based on the total number of nuclei stained by DAPI (blue cells). Cells treated with Cas12f-KRAB-DNMT3L modulators showed reduced apoptosis compared to cells transfected with control sgRNA (Figures 8A and 8B). Example 4 Establishment and validation of ex vivo FSHD model

[0277] Immortalized healthy sibling control cells and FSHD skeletal myoblasts were thawed and expanded for ex vivo 3D studies. Skeletal myoblasts were split on a 2D surface and then engineered into 3D Mantarray tissue following established Curi Bio lab protocols described in Fayazi, M., "Passive-Stretch Induced Skeletal Muscle Injury Platform for Duchenne Muscular Dystrophy Modeling," Archives of Physical Medicine and Rehabilitation, volume 103, issue 3, March 2022, page e26, which is incorporated herein by reference in its entirety. Briefly, 3D skeletal myoblast tissues were cultured for 7 days to allow for compaction, which was then cultured for an additional 14 days. Functional measurements were performed 3 times a week during culture to assess contractile force over time and stimulate to assess phenotypic differences in mechanical force, tetanic force, and fatigue (Figure 9). Once the model was established, patient-derived FSHD skeletal myoblasts were used to test the efficacy of Cas12f effector-modulator AAVs targeting control and DZ4Z loci on rescue of 3D tissue morphology, gene expression profile, and mechanical force assessment.

[0278] In some examples, 3D skeletal myoblast tissue treated with a system, composition, or method disclosed herein (e.g., to modulate expression levels, or epigenetic levels, of a gene encoded by a D4Z4 repeat array, such as DUX4), can be characterized by exhibiting (i) enhanced mechanical force (e.g., maximum mechanical force, average mechanical force over a period of time), (ii) enhanced tetanic force (e.g., force indicative of a sustained muscle contraction evoked when motor nerves innervating skeletal muscle release action potentials at a very high rate), and / or (iii) reduced fatigue (e.g., as measured via contraction against a fixed, immovable object (static testing or isometric measurements) or via dynamic muscle contractions at controlled speeds (repetitive contractions or isokinetic assessments)) compared to control 3D skeletal myoblast tissue (e.g., not treated with a system, composition, or method disclosed herein).

[0279] In some examples, 3D skeletal myoblast tissue treated by a system, composition, or method disclosed herein can be characterized by exhibiting a mechanical force that is at least or up to about 1%, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 30%, at least or up to about 40%, at least or up to about 50%, at least or up to about 60%, at least or up to about 70%, at least or up to about 80%, at least or up to about 90%, at least or up to about 95%, at least or up to about 100%, at least or up to about 120%, at least or up to about 150%, at least or up to about 200%, at least or up to about 300%, at least or up to about 400%, or at least or up to about 500% greater than a mechanical force of a control 3D skeletal myoblast tissue.

[0280] In some examples, 3D skeletal myoblast tissue treated by a system, composition, or method disclosed herein can be characterized by exhibiting a tetanic force that is at least or up to about 1%, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 30%, at least or up to about 40%, at least or up to about 50%, at least or up to about 60%, at least or up to about 70%, at least or up to about 80%, at least or up to about 90%, at least or up to about 95%, at least or up to about 100%, at least or up to about 120%, at least or up to about 150%, at least or up to about 200%, at least or up to about 300%, at least or up to about 400%, or at least or up to about 500% greater than the tetanic force of a control 3D skeletal myoblast tissue.

[0281] In some examples, 3D skeletal myoblast tissue treated by a system, composition, or method disclosed herein can be characterized as exhibiting at least or up to about 1%, at least or up to about 5%, at least or up to about 10%, at least or up to about 15%, at least or up to about 20%, at least or up to about 30%, at least or up to about 40%, at least or up to about 50%, at least or up to about 60%, at least or up to about 70%, at least or up to about 80%, at least or up to about 90%, at least or up to about 95%, at least or up to about 99%, or at least or up to about 100% less fatigue than control 3D skeletal myoblast tissue. Example 5 In vivo FSHD model for DUX4 targeting

[0282] On the morning of the designated -7th day, mice can be anesthetized using intraperitoneal administration of 90-200 mg / kg ketamine and 10 mg / kg xylazine. The hind limbs of the mice can be subjected to X-ray irradiation at 25 Gy at 2.2 Gy / min for 11-12 minutes. Six days later, 60 μL of 0.3 mg / kg cardiotoxin can be administered along the length of the TA muscle to promote degradation. One day later, 60 μL of 2E10^6 human myoblasts can be administered along the TA muscle. Isoflurane anesthesia can be used for subsequent cardiotoxin and human myoblast administration. One day after administration of myoblasts, the Cas12f modulator vectors of the previous examples and the control AAVrh74 vector can be administered via retro-orbital venous plexus injection. On days 4 and 21, animals can be euthanized. The TA muscle and other major organs (e.g., heart, lungs, liver) can be harvested. Harvested TA muscles can be sectioned, fixed, and H&E stained. Remaining organs can be processed, total RNA / DNA extracted, and gene expression experiments can be performed using qRT-PCR. Gene expression experiments can measure the expression of DUX4 and DUX4 target genes and determine the level of DUX4 repression. Gene expression experiments can measure the expression of one or more downstream genes of DUX4, such as ZSCAN4, LEUTX, MBD3L2, TRIM48, and / or TRIM43. Gene expression experiments can also examine the enrichment of human myoblasts in mice as well as AAV tropism to specific tissues. The experimental workflow is shown in Figure 10. Table 2. Guide RNA molecules for binding to target polynucleotide sequences to modify the expression or epigenetic levels of genes (e.g., DUX4) encoded by the D4Z4 repeat array in target cells (e.g., muscle cells). [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10] [Table 2-11] [Table 2-12] [Table 2-13] [Table 2-14] [Table 2-15]

[0283] It is understood that different aspects of the present invention can be realized individually, collectively, or in combination with each other. The various aspects of the present invention described herein can be applied to any of the specific applications disclosed herein. The compositions of matter disclosed herein in the composition section of this disclosure can be utilized in the method section, including the use and production methods disclosed herein, and vice versa.

[0284] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention is limited to the specific examples provided herein. Although the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it is understood that all aspects of the present invention are not limited to the specific descriptions, configurations, or relative proportions described herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in the practice of the present invention. It is therefore contemplated that the present invention includes within its scope any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within these claims and their equivalents are covered thereby.

Claims

1. A system for regulating abnormal expression of a target gene in muscle cells, comprising: a heterologous polypeptide comprising a nuclease having a length of less than about 900 amino acids or equal thereto; and a guide nucleic acid molecule that forms a complex with the heterologous polypeptide, the guide nucleic acid molecule showing specific binding to a target polynucleotide sequence in or adjacent to the D4Z4 repeat array in the muscle cells wherein, when the complex is formed, the complex binds to the target polynucleotide sequence, resulting in modification of the expression level and / or methylation level of the target gene in the muscle cells, and the target gene is present within the D4Z4 repeat array.

2. The system according to claim 1, wherein when the complex is formed, the modified expression level and / or methylation level of the target gene in the muscle cells persists for at least about 2 days.

3. (i) The modified expression level and / or methylation level of the target gene persists for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months, or (ii) The modified expression level and / or methylation level of the target gene persists for at least about 17 days, or (iii) The modified expression level and / or methylation level of the target gene persists for at least about 18 days. The system according to claim 2.

4. The system according to claim 1, wherein the muscle cells are present in a subject having or suspected of having facioscapulohumeral muscular dystrophy (FSHD).

5. The system according to claim 1, wherein the target gene is Dux4.

6. (i) The nuclease has a length of less than about 800 amino acids or equal thereto, or (ii) The nuclease has a length of less than about 750 amino acids or equal thereto. The system according to claim 1.

7. (i) The nuclease is Un1Cas12f1 or a modified variant thereof, (ii) The nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 43 or 44, and / or (iii) the nuclease is an inactivated nuclease, The system according to claim 1.

8. The system according to claim 1, wherein the heterologous polypeptide further comprises a transcriptional regulator.

9. The system according to claim 8, wherein the transcriptional regulator comprises at least one methyltransferase.

10. The system according to claim 8, wherein the transcriptional regulator comprises at least one DNA methyltransferase (DNMT), and optionally, the transcriptional regulator comprises DNMT-A or DNMT-L.

11. (i) the transcriptional regulator comprises (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB, or (ii) the transcriptional regulator comprises (i) DNMT-L, and (ii) KRAB or a variant of KRAB, or (iii) the transcriptional regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB, The system according to claim 8.

12. (i) the transcriptional regulator comprises DNMT-L, or KRAB or a variant of KRAB, or (ii) the transcriptional regulator comprises KRAB or a variant of KRAB, or (iii) the transcriptional regulator comprises a plurality of different transcriptional regulators, The system according to claim 8.

13. The modification of the expression level and / or the methylation level of the target gene results in downregulation of a gene downstream of the target gene, and the downstream gene comprises one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43. The system according to claim 1.

14. The modification of the expression level and / or the methylation level of the target gene results in downregulation of an apoptosis marker in the muscle cell, and optionally, the apoptosis marker comprises caspase 3. The system according to claim 1.

15. The complex results in the modification of the expression level of the target gene in the muscle gene, and optionally, the modification of the expression level results in downregulation of the target gene. The system according to claim 1.

16. The system according to claim 1, wherein the complex brings about the modification of the methylation level of the target gene in the muscle gene, and optionally, the modification of the methylation level brings about down-regulation of the target gene.

17. A composition comprising the system according to any one of the preceding claims.

18. A viral vector comprising one or more nucleic acids encoding the system according to any one of claims 1 to 16.

19. The viral vector according to claim 18, comprising adeno-associated virus (AAV), retrovirus, lentivirus, poxvirus, or adenovirus, and optionally, the AAV comprises AAV serotype RH74 AAV.

20. An in vitro or ex vivo method for regulating abnormal expression of a target gene in a muscle cell, comprising: (a) contacting the muscle cell with the complex of the heterologous polypeptide of the system according to any one of claims 1 to 16 and the guide nucleic acid molecule; and (b) upon the contacting, binding of the target gene to the complex results in modification of the expression level and / or the methylation level of the target gene in the muscle cell. A method comprising the steps of:

21. The system according to any one of claims 1 to 16 for use in a method for regulating abnormal gene expression of a target gene in a muscle cell, the method comprising: (a) contacting the muscle cell with the complex of the heterologous polypeptide of the system and the guide nucleic acid molecule; and (b) upon the contacting, binding of the target gene to the complex results in modification of the expression level and / or the methylation level of the target gene in the muscle cell, and optionally, the contacting comprises injecting a composition comprising the complex into a subject in need thereof, the subject having or suspected of having facioscapulohumeral muscular dystrophy (FSHD). A system. ​