Systems and methods for modulating abnormal gene expression
By regulating target gene expression and methylation through a complex of heterologous peptides and guide nucleic acid molecules, the treatment challenge of abnormal target gene expression has been solved. This approach achieves long-term reduction of target gene expression and methylation modification, thereby improving the survival rate and differentiation capacity of muscle cells. It is suitable for the treatment of facioscapulohumeral muscular dystrophy.
Patent Information
- Application Number
- CN202480032796.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-10-24
- Filing Date
- 2024-03-14
- Publication Date
- 2025-12-12
Smart Images

Figure CN121127598A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to U.S. Provisional Patent Application No. 63 / 490678, filed March 16, 2023; U.S. Provisional Patent Application No. 63 / 520253, filed August 17, 2023; and U.S. Provisional Patent Application No. 63 / 592882, filed October 24, 2023, each of which is expressly incorporated herein by reference.
[0003] References to sequence lists
[0004] This application is submitted together with an electronic sequence list in XML format. The sequence list XML is provided as a file named SequenceListing_EPICR022WO.xml, created on March 14, 2024, and is 929,129 bytes in size. The electronic information of this sequence list is incorporated herein by reference in its entirety. Background Technology
[0005] Aberrant expression of one or more genes can lead to disease or condition. In some cases, aberrant expression of germinal transcription factors in a subject's muscle cells can lead to muscular dystrophy. For example, aberrant expression of transcription factors in muscle cells (such as aberrant expression of DUX4 in skeletal muscle cells) can lead to facioscapulohumeral muscular dystrophy (FSHD). Summary of the Invention
[0006] Transiently modified aberrant expression of a target gene in a cell may not be sufficient to treat or cure the disease manifested by the aberrant expression of that target gene. Therefore, there remains a need for systems and methods that modify the aberrant expression of a target gene and maintain the modified expression level of the target gene over a longer period of time.
[0007] In one aspect, this disclosure provides a system for regulating the aberrant expression of a target gene in muscle cells, the system comprising: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide containing a nuclease, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or near a D4Z4 repeat array in muscle cells, wherein, after the complex is formed, the complex is capable of binding the target polynucleotide sequence, thereby achieving modification of the expression level and / or methylation level of the target gene in muscle cells, wherein the target gene is located within a D4Z4 repeat array, and wherein:
[0008] (1) The guide nucleic acid molecule contains a spacer sequence with specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of one or more members in Table 4; optionally, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of SEQ ID NO:800, SEQ ID NO:836, or SEQ ID NO:851; and / or
[0009] (2) Configure the system to reduce the expression level of target genes in muscle cells by at least 80% or at least about 80% compared with control cells; and / or
[0010] (3) The system is configured to modify the expression level and / or methylation level of target genes in muscle cells, while exhibiting off-target effects below a predetermined threshold in muscle cells; and / or
[0011] (4) Configure the system to continuously modulate the expression level and / or methylation level of downstream genes of the target gene for at least approximately 5 days; and / or
[0012] (5) Configure the system to normalize the apoptosis level of muscle cells to a level comparable to that of healthy muscle cells; and / or
[0013] (6) Configure the system to improve the survival rate of muscle cells in vitro or in vivo; and / or
[0014] (7) The system is configured to have minimal impact on the expression profile of at least one muscle cell-specific gene in muscle cells, wherein the muscle cell-specific gene differs from the target gene; and / or
[0015] (8) Configure the system so that after treating the subject with the system, it has minimal impact on at least one health indicator of the subject.
[0016] In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene in muscle cells is maintained for at least about 2 days after complex formation. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least about 17 days. In some embodiments of any of the systems disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least about 18 days.
[0017] In some embodiments of any of the systems disclosed herein, muscle cells are present in subjects who have or are suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the systems disclosed herein, the target gene is Dux4.
[0018] In some embodiments of any of the systems disclosed herein, the nuclease has a length of less than or equal to about 800 amino acids. In some embodiments of any of the systems disclosed herein, the nuclease has a length of less than or equal to about 750 amino acids.
[0019] In some embodiments of any of the systems disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO:43. In some embodiments of any of the systems disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO:44.
[0020] In some embodiments of any of the systems disclosed herein, the heterologous polypeptide further comprises a transcriptional regulatory factor. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises at least one methyltransferase. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises: (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L), and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises (i) DNMT-L and (ii) KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises DNMT-L (or DNMT3L), or KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises KRAB or a variant of KRAB. In some embodiments of any of the systems disclosed herein, the transcriptional regulatory factor comprises a variety of different transcriptional regulatory factors.
[0021] In some embodiments of any of the systems disclosed herein, modifications to the expression level and / or methylation level of the target gene cause downregulation of downstream genes of the target gene, wherein the downstream genes comprise one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.
[0022] In some embodiments of any of the systems disclosed herein, modification of the target gene expression level and / or methylation level leads to downregulation of apoptosis markers in muscle cells. In some embodiments of any of the systems disclosed herein, the apoptosis markers include caspase 3.
[0023] In some embodiments of any of the systems disclosed herein, the complex causes modification of the expression level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, the modification of the expression level causes downregulation of the target gene.
[0024] In some embodiments of any of the systems disclosed herein, the complex causes modification of the methylation level of the target gene in the muscle gene. In some embodiments of any of the systems disclosed herein, modification of the methylation level causes downregulation of the target gene.
[0025] In some implementations of any of the systems disclosed herein, the nuclease is an inactivated nuclease.
[0026] In another aspect, this disclosure provides compositions comprising any of the systems disclosed herein.
[0027] In another respect, this disclosure provides viral vectors comprising any of the systems or compositions disclosed herein.
[0028] In some embodiments of any of the viral vectors disclosed herein, the viral vector includes adeno-associated virus (AAV), retrovirus, lentivirus, poxvirus, or adenovirus. In some embodiments of any of the viral vectors disclosed herein, the AAV includes AAV serotype RH74 AAV.
[0029] On the other hand, this disclosure provides a method for regulating the aberrant expression of a target gene in muscle cells, the method comprising: (a) contacting muscle cells with a complex or system comprising (i) a heterologous polypeptide comprising a nuclease, and (ii) a guide nucleic acid molecule exhibiting specific binding to a target polynucleotide sequence at or adjacent to a D4Z4 repeat array in muscle cells; and (b) upon contact, binding a target gene to the complex or system to modify the expression level and / or methylation level of the target gene in the muscle cells, wherein the target gene is located within a D4Z4 repeat array, and wherein:
[0030] (1) The guide nucleic acid molecule contains a spacer sequence with specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity with the polynucleotide sequences of one or more members in Table 4; optionally, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity with the polynucleotide sequences of SEQ ID NO:800, SEQ ID NO:836, or SEQ ID NO:851; and / or
[0031] (2) The complex or system is configured to reduce the expression level of the target gene in muscle cells by at least or at least about 80% compared with control cells; and / or
[0032] (3) The complex or system is configured to modify the expression level and / or methylation level of target genes in muscle cells, while exhibiting off-target effects below a predetermined threshold in muscle cells; and / or
[0033] (4) Configure the complex or system to continuously modulate the expression level and / or methylation level of downstream genes of the target gene for at least approximately 5 days; and / or
[0034] (5) Configure the complex or system to normalize the apoptosis level of muscle cells to a level comparable to that of healthy muscle cells; and / or
[0035] (6) To formulate a complex or system to improve the survival rate of muscle cells in vitro or in vivo; and / or
[0036] (7) The complex or system is configured to have minimal impact on the expression profile of at least one muscle cell-specific gene in muscle cells, wherein the muscle cell-specific gene is different from the target gene; and / or
[0037] (8) Configure the complex or system so that it has minimal effect on at least one health indicator of the subject after treatment with the system.
[0038] In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene in muscle cells is maintained for at least about 2 days after the formation of the complex or system. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least or at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least or at least about 17 days. In some embodiments of any of the methods disclosed herein, the modified expression level and / or methylation level of the target gene is maintained for at least or at least about 18 days.
[0039] In some embodiments of any of the methods disclosed herein, contact includes injecting a composition comprising a complex or system into a subject in need, wherein the subject has or is suspected of having facioscapulohumeral muscular dystrophy (FSHD). In some embodiments of any of the methods disclosed herein, the target gene is Dux4.
[0040] In some embodiments of any of the methods disclosed herein, the nuclease has a length of 800 amino acids or less. In some embodiments of any of the methods disclosed herein, the nuclease has a length of 750 amino acids or less.
[0041] In some embodiments of any of the methods disclosed herein, the nuclease is Un1Cas12f1 or a modified variant thereof.
[0042] In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO:43. In some embodiments of any of the methods disclosed herein, the nuclease comprises an amino acid sequence that is at least or at least about 80%, at least or at least about 90%, at least or at least about 95%, or at least or at least about 99% identical to the polypeptide sequence of SEQ ID NO:44.
[0043] In some embodiments of any of the methods disclosed herein, the heterologous polypeptide further comprises a transcriptional regulatory factor. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises at least one methyltransferase. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises at least one DNA methyltransferase (DNMT). In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L). In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises: (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L), and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises DNMT-L (or DNMT3L), or KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises KRAB or a variant of KRAB. In some embodiments of any of the methods disclosed herein, the transcriptional regulatory factor comprises a variety of different transcriptional regulatory factors.
[0044] In some embodiments of any of the methods disclosed herein, modification of the expression level and / or methylation level of the target gene causes downregulation of downstream genes of the target gene, wherein the downstream genes comprise one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48, and TRIM43.
[0045] In some embodiments of any of the methods disclosed herein, modification of the target gene expression level and / or methylation level causes downregulation of apoptosis markers in muscle cells. In some embodiments of any of the methods disclosed herein, the apoptosis markers include caspase 3.
[0046] In some embodiments of any of the methods disclosed herein, the complex or system causes modification of the expression level of the target gene in the muscle gene.
[0047] In some implementations of any of the methods disclosed herein, modification of expression levels leads to downregulation of the target gene.
[0048] In some embodiments of any of the methods disclosed herein, the complex or system causes modification of the methylation level of the target gene in the muscle gene. In some embodiments of any of the methods disclosed herein, modification of the methylation level causes downregulation of the target gene.
[0049] In some embodiments of any of the methods disclosed herein, the nuclease is an inactivated nuclease.
[0050] In some embodiments, any of the compositions or carriers disclosed herein are used as medicines, for example for inhibiting, improving or treating cancer, autoimmune diseases or facioscapulohumeral muscular dystrophy (FSHD).
[0051] In some embodiments, any of the compositions or vectors disclosed herein are used for editing genes in cells, either in vitro or in vivo.
[0052] A method for delivering a system to cells that regulates the aberrant expression of a target gene is also provided, the method comprising introducing any of the systems, compositions or vectors disclosed herein into cells, preferably into cells of a subject (e.g., a person with cancer, an autoimmune disease or facioscapulohumeral muscular dystrophy (FSHD)).
[0053] In some embodiments of any of the systems, compositions, or viral vectors disclosed herein, the system comprises: a heterologous polypeptide and a guide nucleic acid molecule, said heterologous polypeptide comprising a nuclease, said nuclease comprising an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, said guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence consisting of a sequence of the same type as SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:828. Any of NO:865 encodes a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity.
[0054] In some embodiments of any of the systems, compositions, or viral vectors disclosed herein, the system comprises: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence composed of sequences corresponding to SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, and SEQ ID NO:827. Either NO:833 or SEQ ID NO:865 encodes a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity.
[0055] In some embodiments of any of the methods disclosed herein, the complex or system comprises: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease comprising an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865.
[0056] In some embodiments of any of the methods disclosed herein, the complex or system comprises: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence composed of sequences corresponding to SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:827, SEQ ID NO:828 ... Either NO:833 or SEQ ID NO:865 encodes a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity.
[0057] Additional aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only exemplary embodiments of the disclosure are shown and described. As will be understood, this disclosure is capable of other and different embodiments, and several details thereof can be modified in various obvious respects without departing from this disclosure. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0058] By incorporating references
[0059] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the same extent as each individual publication, patent or patent application is expressly and individually indicated to be incorporated by reference. Attached Figure Description
[0060] The novel features of this disclosure are particularly set forth in the appended claims. The features and advantages of this disclosure will be better understood by referring to the following detailed description, which illustrates exemplary embodiments in which the principles of the invention are utilized, and the accompanying drawings.
[0061] Figure 1 Different target polynucleotide sequences (e.g., rank #1 to rank #91) between two CpG islands within the D4Z4 repeat array encoding DUX4 are provided.
[0062] Figure 2 Provides heterologous execution via coupling with gene regulators (e.g., dCas-KRAB-DNMT3A-DNMT3L). Apparatus The regulation of DUX4 expression in target cell populations (e.g., lymphoblasts) by a certain actuator moiety, wherein the gene regulator is complexed with multiple guide RNA molecules that target DUX4 within a D4Z4 repeat array containing polynucleotide sequences (e.g., rank #1 to rank #91).
[0063] Figure 3A This study describes the gene expression of DUX4 and its target genes in immortalized patient-derived human FSHD skeletal muscle myoblasts (SkM) (12ABIC / 12A and 15ABIC / 15A). Gene expression of DUX4 and its target genes was measured in the following cell types: undifferentiated 12ABIC and 15ABIC cells, 12ABIC and 15ABIC cells two days after differentiation, and 12ABIC and 15ABIC cells seven days after differentiation. Each grayscale value in the figure represents the gene expression of the different gene corresponding to the legend on the right. Figure 3BThe figures depict the percentage of apoptotic cells in FSHD myoblasts 12ABIC and 15ABIC (right column) two days after differentiation, compared to their healthy siblings (left column) of control myoblasts 12UBIC and 15VBIC. White dots in the left image represent apoptotic cells. The right image illustrates the percentage of apoptotic cells in the 12ABIC, 15ABIC, 12UBIC, and 15VBIC cell cultures two days after differentiation, as shown in the left image. Cell nuclei were stained using DAPI staining. Figure 3C The percentage of apoptotic cells in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells was described 7 days after differentiation. The percentage of apoptotic cells was measured on days 0, 1, 2, and 7 of differentiation. Figure 3D The expression of MYHC in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells 7 days after differentiation was described. Myosin heavy chain (MYHC) is a marker of muscle cell differentiation. White dots indicate MYHC expression. Figure 3E The expression levels of MYOG, MYH2, and MYMK in 12ABIC, 15ABIC, 12UBIC, and 15VBIC cells were described 7 days after differentiation. MYOG is a myogenic regulator of skeletal muscle differentiation, while MyoMaker (MYMK) is a marker of muscle cell differentiation. Cell nuclei were stained using DAPI staining. Expression levels in 12ABIC and 15ABIC cells were measured on days 2 and 7 of differentiation. 12AUD and 15AUD: undifferentiated, proliferating control myoblasts. Dark gray bars depict MYOG expression levels, light gray bars depict MYH2 expression levels, and gray bars depict MYMK expression levels.
[0064] Figure 4 The design of multiple gRNAs associated with the D4Z4 repeat region is illustrated. Multiple gRNAs targeting DUX4 were designed to span the D4Z4 repeat region. The relative positions of the D4Z4 repeat region and the DUX4 gene are shown in [the diagram / illustration]. Figure 4 The bottom is shown. The newly designed gRNA is... Figure 4 It is shown at the top.
[0065] Figure 5 The design of the Cas12f effector-regulator vector is described. Expression of the Cas12f variant, KRAB domain, and DNMT3L domain is controlled by the muscle-specific promoter CK8e. Expression of the scaffolded sgRNA spacer sequence driven by RNA polymerase III is controlled by the human U6g promoter. The vector additionally includes a modified WPRE and a polyadenylation regulatory sequence.
[0066] Figure 6AThis study describes the relative expression level of DUX4 in 12ABIC FSHD myoblasts stably expressing the Cas12f-KRAB effector-regulator after nuclear transfection of 78 gRNAs into 12ABIC myoblasts. DUX4 gene expression was measured 7 days after nuclear transfection and cell culture under differentiation conditions. The 78 tested gRNAs are plotted on the x-axis, and the y-axis represents the relative fold change in DUX4 expression. DUX4 expression levels were normalized using the expression of the control gene HPRT1. Figure 6B This study describes the relative expression level of DUX4 in 12ABICFSHD myoblasts that stably express the Cas12f-KRAB effector-regulator after nuclear transfection of 78 gRNAs into 12ABIC myoblasts. Following nuclear transfection, the gene expression of DUX4 and its target gene MBD3L2 was measured after 7 days of cell culture under differentiation conditions.
[0067] Figure 7A This study describes the repression of DUX4 and its target genes (DBET / DUX4, MBD3L2, and TRIM48) in immortalized FSHD myoblasts derived from patients transfected with six gRNAs and a Cas12f effector-regulator. The Cas12f effector-regulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLa domain. One of the six sgRNAs is a control sgRNA (empty / trcr) that does not target the D4Z4 repeat region. The expression level of MYOG in cells was measured to analyze whether the differentiation capacity of cells transfected with DUX4 sgRNA was similar to that of myoblasts transfected with control sgRNA. The expression levels of DUX4, DUX4 target genes, and MYOG were measured 17 days post-transfection. Figure 7B This study describes the repression of DUX4 and its target genes (DBET / DUX4, MBD3L2, and TRIM48) in immortalized FSHD myoblasts derived from patients transfected with six gRNAs and a Cas12f effector-regulator. The Cas12f effector-regulator expresses a Cas12f variant, a KRAB domain, and a DNMT-KLb domain. One of the six sgRNAs is a control sgRNA (empty) that does not target the D4Z4 repeat region. The expression level of MYOG in cells was measured to analyze whether the differentiation capacity of cells transfected with DUX4 sgRNA was similar to that of myoblasts transfected with the control sgRNA. The expression levels of DUX4, DUX4 target genes, and MYOG were measured 18 days post-transfection.
[0068] Figure 8A and Figure 8BThe apoptosis levels of myoblasts from FSHD patients transfected with Cas12f effector-regulatory and DUX4-targeting gRNAs were described. The percentage of apoptotic-positive cells was determined 2 days after differentiation following transfection. Figure 8A The images depict the proportion of apoptotic cells in control 12UBIC cells and 12ABIC cells transfected with Cas12f effector-regulator and DUX4-targeting gRNA. White dots represent apoptotic cells. Figure 8B The charts in the image depict the percentage of apoptotic cells measured in the left image, as well as the percentage of apoptotic cells in 12ABIC cells transfected with either DUX4-targeting gRNA or a control gRNA that does not target DUX4. Cell nuclei were stained using DAPI staining.
[0069] Figures 9A-9E The effect of the exemplary system on a 3D ex vivo FSHD organoid model is shown. Figure 9A The workflow for an ex vivo FSHD model is described. An ex vivo model culture was used to immortalize healthy sibling control cells and FSHD skeletal muscle myoblasts, which were then engineered into 3D tissue. This 3D tissue was contacted with a control AAV or an AAV having the exemplary system described herein. Subsequently, in addition to determining the morphological and gene expression profiles of the 3D tissue, phenotypic differences in mechanical force, tetanic force, and fatigue were tested. Figure 9B The average active twitching force over time is shown in both 3D organoid tissues treated with GFP control (-) (top) and treated with the exemplary system (bottom). Figure 9C The dominant forces at the endpoint (day 46) are shown in both 3D organoid tissues treated with GFP control (-) (top) and treated with the exemplary system (bottom). Figure 9D The mean tetanic force plotted over time is shown in both 3D organoid tissues treated with GFP control (-) (top) and treated with the exemplary system (bottom). Figure 9E Normalized tetanic force at the endpoint (day 46) is shown in both 3D organoid tissue treated with GFP control (-) (top) and treated with the exemplary system (bottom).
[0070] Figures 10A-10G The effect of the exemplary system on inhibiting DUX4 and DUX4 pathway genes in humanized mice is shown. Figure 10AThe workflow for an in vivo xenograft model is described. An in vivo model was constructed by treating mouse legs with irradiation and TA myocardial toxin to prepare for transplantation of human myoblasts into mouse legs. After transplantation, mice were exposed to either control AAV or AAV with the exemplary system described herein. Mice were euthanized at designated time points, and xenograft and tissue samples were collected for analysis. The collected xenografts were fixed, sectioned, and stained with hematoxylin and eosin. Remaining tissue was used for gene expression analysis and to determine AAV tropism in mice. Figure 10B The mRNA expression of DUX4 in humanized TA muscle is shown. Figure 10C Gene expression of DUX4 pathway genes plotted as a comprehensive score is shown. Figure 10D The biological distribution of an exemplary system is shown. Figure 10E and Figure 10F The DUX4 protein ( was shown) Figure 10E ) and SLC34A2 protein ( Figure 10F Quantitative analysis of staining. Figure 10G The quantification of TUNEL staining is shown.
[0071] Figure 11 The methylation of a target gene (e.g., D4Z4) is described using an exemplary system / method.
[0072] Figure 12 The proposed treatment options for FSHD patients are shown.
[0073] Figure 13A and Figure 13B The unmethylated or methylated states at the target gene / target site are shown. Figure 13A The methylation / unmethylation state of target sites at multiple time points is shown through an exemplary system containing multiple gene repressor factors. Figure 13B The percentage of CpG methylation at the target locus (D4Z4 locus) is shown in myoblasts derived from healthy sibling controls, controlled AAV-treated cells, or cells treated with the exemplary system.
[0074] Figure 14 The screening method and results for high-throughput screening of guide RNA spacer sequences for DUX4 are shown.
[0075] Figure 15 The off-target analysis, based on verification and computer simulation, is shown.
[0076] Figure 16 The virus delivery payload of an exemplary system is described.
[0077] Figure 17The expression of myogenic genes in patient-derived FSHD myoblasts is shown.
[0078] Figure 18 The expression of DUX4 and its downstream genes is shown in patient-derived FSHD myoblasts using an exemplary system.
[0079] Figures 19A-19D Apoptotic cells in patient-derived FSHD myoblasts via an exemplary system are shown. Figure 19A The percentage of apoptotic cells among differentiated patient-derived FSHD myoblasts and patient-derived FSHD myoblasts in contact with the exemplary system is shown. Figure 19B The mRNA expression of DUX4 and exemplary system payloads is shown. Figure 19C Live-cell imaging analysis of caspase 3 / 7 staining is shown. Figure 19D The normalized intensity of the total caspase 3 / 7 signal at the endpoint is shown.
[0080] Figure 20 The effect of the exemplary system on the skeletal muscle of a humanized FSHD mouse model is shown.
[0081] Figure 21 Histology of TA muscle processed using an exemplary system is shown.
[0082] Figure 22 The expression of DUX4 and DUX4 pathway genes is shown in a humanized FSHD mouse model in contact with an exemplary system.
[0083] Figure 23 A 6-month non-GLP toxicology study in immune-active mice is shown.
[0084] Figure 24 Clinical chemical evaluations in systematically treated animals are shown. At each time point, the left column represents control animals, and the right column represents systematically treated animals.
[0085] Figure 25 The histopathological evaluation of the animals is shown 3 months after systemic administration.
[0086] Figure 26 Blood chemistry and hematology data from a non-human primate model in contact with the system are shown.
[0087] Figure 27 This study demonstrates pharmacokinetic studies using a non-human primate model.
[0088] Figure 28 The tissue-specific mRNA expression levels and tropism of candidate guide nucleic acid molecules were shown.
[0089] Figures 29A-29E (together) Figures 10A-10G The effective dose range of the exemplary system is shown to affect the molecular and cellular phenotypes in humanized mice. Figure 29A A schematic diagram of an in vivo xenograft model for FSHD is shown. Figures 29B-29C The qPCR of the exemplary system in TA muscle biopsy obtained from FSHD humanized mice is shown 24 days after intravenous administration of the exemplary system or solvent. Figure 29B ) and qRT-PCR ( Figure 29C )analyze. Figure 29B and Figure 10D Related. Figures 29D-29E The quantitative analysis of the DUX4 cascade in TA muscle biopsies obtained from FSHD humanized mice is shown 24 days after intravenous administration of an exemplary system or solvent. Figure 29D The qRT-PCR analysis of the four DUX4 target genes is shown as a comprehensive score. Figure 29E Quantitative image analysis of SLC34A2 immunohistochemical staining, plotted as H score, is shown.
[0090] Figures 30A-30H The effect of dose escalation of an exemplary system on 3D engineered muscle tissue (EMT) derived from myoblasts in patients with FSHD is shown. Figure 30A A schematic diagram of 3D EMT tissue casting and stimulation model is shown. Figures 30B-30C The maximum active twitching force in tissues transduced via the exemplary system is shown from day 6 to day 46 in the control 3D EMT without transduction. Figure 30B ) and taut force ( Figure 30C ). Figures 30D-30E The muscle contractile force normalized to untransduced 3D EMT is shown at day 46. Maximum active twitching force (VTF) up to an exemplary systemic dose level of 7.5E8 vg / tissue is also shown. Figure 30D ) and taut force ( Figure 30E The normalized force increase is shown. Data are presented as mean ± SEM. Figures 30F-30G This demonstrates a quantitative PCR evaluation of the genome of an exemplary system vector in 3D EMT. Figure 30F ) and qRT-PCR for evaluating mRNA expression ( Figure 30G Data are presented as mean ± SEM. Figure 30H The expression of DUX4 and DUX4 pathway genes in 3D EMT transduced with the exemplary system and in untransduced control tissues is shown. Data are presented as mean ± SEM.
[0091] Figures 31A-31HAn exemplary system is shown as a mechanism of action for gene-targeted therapy for FSHD by directly regulating the epigenetic state of the DUX4 locus and thus alleviating its pathological overexpression. Figure 31A A schematic diagram illustrating an exemplary systemic mechanism of action analysis using on-target methylation analysis in FSHD myoblasts is shown. Figure 31B The mRNA expression of DUX4 and six DUX4 genes (MBD3L2, ZSCAN4, LEUTX, TRIM43, and TRIM48) in patient-derived myoblasts is shown. Figure 31C Primary myoblasts (07ABIC) derived from FSHD patients exposed to the exemplary system showed increased methylation compared to myoblasts treated with (-) control test sample. Figure 31D The mRNA expression of DUX4 and six DUX4 genes (i.e., MBD3L2, ZSCAN4, LEUTX, TRIM43, and TRIM48) in patient-derived myoblasts (07ABIC) exposed to the (-) control and exemplary system, as analyzed by RT-qPCR, is shown. Mean ± SEM (n=3) of experimental replicates is plotted. Figure 31E Myoblasts (07ABIC) in contact with the exemplary system showed a significant reduction in apoptosis. Figure 31F The enzymatic transformation of 100 ng of DNA is shown, followed by targeted methylation analysis of unmethylated and methylated DNA standards using custom primers (Table 17). Methylation % was calculated using QUMA analysis from at least 12 clones each time, and the median (Q2) was plotted. Figure 31F The top subplot of the clone analysis (22 CpGs). Lollipop plots indicating the methylation status of each CpG in each clone were also plotted. Figure 31F (The bottom sub-figure). Each black circle in the figure represents a methylated CpG dinucleotide. Figure 31G The targeted methylation analysis of unmethylated and methylated DNA standards (Sigma-Aldrich) using primer pairs described in Table 17 after enzymatic transformation of 100 ng of DNA is described in next-generation sequencing analysis. Figure 31H The results of downsampling analysis of target methylation in next-generation sequencing are shown. Detailed Implementation
[0092] Although various embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that these embodiments are provided by way of example only. Many variations, modifications, and substitutions will occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0093] When the terms "at least," "greater than," or "greater than or equal to" precede the first value in a series of two or more values, the terms "at least," "greater than," or "greater than or equal to" apply to each value in that series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0094] When the terms "not exceeding," "less than," or "less than or equal to" precede the first value in a series of two or more values, the terms "not exceeding," "less than," or "less than or equal to" apply to each value in that series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0095] Unless the context clearly indicates otherwise, the singular terms “a,” “an,” and “the” include the plural referent. Similarly, unless the context clearly indicates otherwise, the word “or” is intended to include “and.” The abbreviation “eg” is used herein to denote non-restrictive instances. Therefore, the abbreviation “eg” is synonymous with the term “for example.” Numbers provided in the ranges include overlapping ranges and integers within them; for example, the ranges 1-4 and 5-7 include, for example, 1-7, 1-6, 1-5, 2-5, 2-7, 4-7, 1, 2, 3, 4, 5, 6, and 7.
[0096] The terms “about” or “approximately” generally mean within an acceptable margin of error for a particular value as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to convention in the art, “about” may mean within 1 or more standard deviations. Alternatively, “about” may mean a range of no more than 20%, no more than 10%, no more than 5%, or no more than 1% of a given value. Or, particularly for biological systems or processes, the term may mean within an order of magnitude of the value, preferably within 5 times, and more preferably within 2 times. When a particular value is described in this application and claims, unless otherwise stated, the term “about” should be considered to mean within an acceptable margin of error for that particular value.
[0097] The use of alternatives (e.g., "or") should be understood as referring to one, two, or any combination of alternatives. The term "and / or" should be understood as referring to one or two alternatives.
[0098] The term "cell" generally refers to a biological cell. A cell can be the basic structural, functional, and / or biological unit of an organism. Cells can originate from any organism having one or more cells. Some non-limiting examples include: prokaryotic cells, eukaryotic cells, bacterial cells, archaea cells, cells of unicellular eukaryotes, protozoan cells, cells from plants (e.g., cells from crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, squash, hay, potatoes, cotton, hemp, tobacco, flowering plants, conifers, gymnosperms, ferns, lycophytes, hornwort, liverwort, mosses), and algal cells (e.g., *Botryococcus braunii*, *Chlamydomonas reinhardtii*, *Nannochloropsis gaditana*, *Chlorella pyrenoidosa*, *Sargassum patens*). Cells derived from various organisms include: agardh, algae (e.g., giant kelp), fungi (e.g., yeast cells, mushroom cells), animal cells, cells derived from invertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells derived from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), and cells derived from mammals (e.g., pigs, cattle, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.). Sometimes, cells are not derived from natural organisms (e.g., cells can be synthetically manufactured, sometimes referred to as artificial cells).
[0099] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate composition. A nucleotide may comprise a synthetic nucleotide. A nucleotide may comprise a synthetic nucleotide analog. A nucleotide may be a monomeric unit of a nucleic acid sequence, such as deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). The term nucleotide may include ribonucleoside triphosphates (adenosine triphosphate (ATP), uridine triphosphate (UTP), cytidine triphosphate (CTP), guanosine triphosphate (GTP)) and deoxyribonucleoside triphosphates (such as dATP, dCTP, dITP, dUTP, dGTP, dTTP) or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-denitro-dGTP, and 7-denitro-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxyribonucleoside triphosphates (ddNTPs) and derivatives thereof. Exemplary examples of dideoxyribonucleoside triphosphates may include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides can be unlabeled or detectably labeled using known techniques. Labeling can also be implemented using quantum dots. Detectable labels can include, for example, radioactive isotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labeling of nucleotides can include, but is not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2',7'-dimethoxy-4',5'-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides may include: [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP, all available from Perkin Elmer (Foster City, Calif.); and FluoroLink deoxynucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink, all available from Amersham (Arlington Heights, Ill.). Cy5-dUTP; luciferin-15-dATP, luciferin-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, luciferin-12-ddUTP, luciferin-12-UTP, and luciferin-15-2'-dATP, all available from Boehringer Mannheim (Indianapolis, Ind.); and chromosome marker nucleotides, BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, luciferin-12-UTP, and luciferin-12-dUTP, available from Oregon, are all available from Molecular Probes (Eugene, Oreg.). Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides can also be labeled or marked by chemical modifications. Chemically modified mononucleotides can be biotin-dNTPs.Some non-limiting examples of biotinylated dNTPs may include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0100] The terms “polynucleotide,” “oligonucleotide,” or “nucleic acid” used interchangeably in this document generally refer to polymers of nucleotides of any length, whether deoxyribonucleotides, ribonucleotides, or analogs thereof; whether single-stranded, double-stranded, or multi-stranded. Polynucleotides can be exogenous or endogenous to cells. Polynucleotides can exist in cell-free environments. Polynucleotides can be genes or segments thereof. Polynucleotides can be DNA. Polynucleotides can be RNA. Polynucleotides can have any three-dimensional structure and can perform any function, known or unknown. Polynucleotides can contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). Modifications to the nucleotide structure, if present, can be made before or after polymer assembly. Some non-limiting examples of analogues include: 5-bromouracil, peptide nucleic acid, xeno nucleic acid, morpholinos, locked nucleic acid, glycol nucleic acid, threonine, dideoxynucleotide, cordycepin, 7-denitro-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queuosine, and wyosine. Non-restricted examples of polynucleotides include: coding or non-coding regions of genes or gene fragments, loci defined by linkage analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides (including cell-free DNA (cfDNA) and cell-free RNA (cfRNA)), nucleic acid probes, and primers. Nucleotide sequences can be interrupted by non-nucleotide components.
[0101] The term "sequence identity" typically refers to the precise nucleotide-to-nucleotide or amino acid-to-amino acid correspondence between two polynucleotide or polypeptide sequences. Generally, techniques used to determine sequence identity involve determining the nucleotide sequence of the polynucleotide and / or the amino acid sequence it encodes, and comparing these sequences to a second nucleotide or amino acid sequence. Two or more sequences (polynucleotides or amino acids) can be compared by determining their "percentage of identity." Whether it's a nucleic acid sequence or an amino acid sequence, the percentage of identity between two sequences is the number of exact matches between the two aligned sequences divided by the length of the longer sequence, and then multiplied by 100. For example, the percentage of identity can also be determined by comparing sequence information using the Advanced BLAST computer program (including version 2.2.9) available from the National Institutes of Health. This BLAST program is based on the alignment method of Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 87:2264-2268 (1990), and discussed in Altschul et al., J. Mol. Biol., 215:403-410 (1990), Karlin and Altschul, Proc. Natl. Acad. Sci. USA, 90:5873-5877 (1993), and Altschul et al., Nucleic Acids Res., 25:3389-3402 (1997). The program can be used to determine the percentage of identity across the full length of the compared proteins. Default parameters are provided in programs such as blastp to optimize retrieval of short query sequences. The procedure also allows the use of SEG filters to filter out fragments of the query sequence identified by the SEG procedure of Wootton and Federhen, Computers and Chemistry 17:149-163 (1993). The expected degree of sequence identity ranges from approximately 50% to 100%, and integer values between therebetween. Generally, this disclosure covers sequences that have at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 98% sequence identity with any of the sequences provided herein.
[0102] The term "gene" generally refers to a nucleic acid (e.g., DNA, such as genomic DNA and cDNA) involved in encoding RNA transcripts and its corresponding nucleotide sequence. The term used herein to refer to genomic DNA includes the intermediate non-coding region and regulatory regions, and may include the 5' and 3' ends. In some uses, the term includes the transcribed sequence, including the 5' and 3' untranslated regions (5'-UTR and 3'-UTR), exons, and introns. In some genes, the transcribed region will contain an "open reading frame" encoding a polypeptide. In some uses of the term, "gene" contains only the coding sequence necessary to encode a polypeptide (e.g., "open reading frame" or "coding region"). In some cases, the gene does not encode a polypeptide, such as ribosomal RNA genes (rRNA) and transfer RNA (tRNA) genes. In some cases, the term "gene" includes not only the transcribed sequence but also additionally non-transcribed regions, which include upstream and downstream regulatory regions, enhancers, and promoters. For example, a gene may refer to the portion of a gene that is close to or adjacent to the transcription start site (TSS) of that gene. The gene (e.g., a targeting gene as disclosed herein) can be at least, at most, at least about, or at most about 2000 nucleotides away from the TSS of the gene, at least, at most, at least about, or at most about 1800 nucleotides away, at least, at most, at least about, or at most about 1600 nucleotides away, at least, at most, at least about, or at most about 1500 nucleotides away, at least, at most, at least about, or at most about 1400 nucleotides away, at least, at most, at least about, or at most about 1200 nucleotides away, at least, at most, at least about, or at most about 1000 nucleotides away, at least, at most, at least about, or at most about 9 nucleotides away. 00 nucleobases, at least, at most, at least about or at most about 800 nucleobases, at least, at most, at least about or at most about 700 nucleobases, at least, at most, at least about or at most about 600 nucleobases, at least, at most, at least about or at most about 500 nucleobases, at least, at most, at least about or at most about 400 nucleobases, at least, at most, at least about or at most about 300 nucleobases, at least, at most, at least about or at most about 200 nucleobases, at least, at most, at least about or at most about 100 nucleobases, or at least, at most, at least about or at most about 50 nucleobases.
[0103] A gene can refer to an "endogenous gene" or natural gene located in its natural location within an organism's genome. A gene can also refer to an "exogenous gene" or non-natural gene. A non-natural gene can refer to a gene that is not normally found in the host organism but was introduced into the host organism through gene transfer. A non-natural gene can also refer to a gene that is not located in its natural location within an organism's genome. A non-natural gene can also refer to a naturally occurring nucleic acid or polypeptide sequence containing mutations, insertions, and / or deletions (e.g., a non-natural sequence).
[0104] The term "expression" generally refers to one or more processes of transcription from a DNA template into polynucleotides (such as mRNA or other RNA transcripts), and / or the subsequent translation of transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides can be collectively referred to as "gene products." If the polynucleotides are derived from genomic DNA, expression can include splicing of mRNA in eukaryotic cells. In terms of expression, "upregulation" generally refers to an increase in the expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression level in the wild-type state; while "downregulation" generally refers to a decrease in the expression level of a polynucleotide (e.g., RNA such as mRNA) and / or polypeptide sequence relative to its expression level in the wild-type state. Expression of a transfected gene can occur transiently or stably in a cell. During "transient expression," the transfected gene does not transfer to daughter cells during cell division. Because its expression is limited to the transfected cell, gene expression disappears over time. Conversely, when a gene is co-transfected with another gene that confers a selective advantage to the transfected cell, stable expression of the transfected gene can occur. This selective advantage can be resistance to a certain toxin presented to the cell.
[0105] The term "expression profile" typically refers to the quantitative (e.g., abundance) and qualitative expression of one or more genes in a sample (e.g., cells). These genes may be expressed and identified as nucleic acid molecules (e.g., mRNA or other RNA transcripts). Alternatively or supplementally, these genes may be expressed and identified as peptides (e.g., proteins determined by Western blotting). A gene expression profile can be defined as the shape of the gene's expression level over a period of time (e.g., at least, at most, at least about, or at most about 1 hour, at least, at most, at least about, or at most about 2 hours, at least, at most, at least about, or at most about 3 hours, at least, at most, at least about, or at most about 4 hours, at least, at most, at least about, or at most about 5 hours, at least, at most, at least about, or at most about 6 hours, at least, at most, at least about, or at most about 7 hours, at least, at most, at least about, or at most about 8 hours, at least, at most, at least about, or at most about 9 hours, at least, at most, at least about, or at most about 10 hours, at least, at most, at least about, or at most about 11 hours, at least, at most, at least about, or at most about 12 hours, at least, at most, at least about, or at most about 16 hours, at least, at most, at least... Approximately or at most about 18 hours, at least, at most, at least about or at most about 24 hours, at least, at most, at least about or at most about 36 hours, at least, at most, at least about or at most about 48 hours, at least, at most, at least about or at most about 3 days, at least, at most, at least about or at most about 4 days, at least, at most, at least about or at most about 5 days, at least, at most, at least about or at most about 6 days, at least, at most, at least about or at most about 7 days, at least, at most, at least about or at most about 8 days, at least, at most, at least about or at most about 9 days, at least, at most, at least about or at most about 10 days, at least, at most, at least about or at most about 11 days, at least, at most, at least about or at most about 12 days, at least, at most, at least about or at most about 13 days, at least, at most, at least about or at most about 14 days, etc.).Alternatively, a gene expression profile can be defined as the gene's expression level at a time point of interest (e.g., the gene's expression level measured at the following times: at least, at most, at least about, or at most about 1 hour after cells have been treated to induce this expression level; at least, at most, at least about, or at most about 2 hours; at least, at most, at least about, or at most about 3 hours; at least, at most, at least about, or at most about 4 hours; at least, at most, at least about, or at most about 5 hours; at least, at most, at least about, or at most about 6 hours; at least, at most, at least about, or at most about 7 hours; at least, at most, at least about, or at most about 8 hours; at least, at most, at least about, or at most about 9 hours; at least, at most, at least about, or at most about 10 hours; at least, at most, at least about, or at most about 11 hours; at least, at most, at least about, or at most about 12 hours; at least, at most, at least... Approximately or at most about 16 hours, at least, at most, at least about or at most about 18 hours, at least, at most, at least about or at most about 24 hours, at least, at most, at least about or at most about 36 hours, at least, at most, at least about or at most about 48 hours, at least, at most, at least about or at most about 3 days, at least, at most, at least about or at most about 4 days, at least, at most, at least about or at most about 5 days, at least, at most, at least about or at most about 6 days, at least, at most, at least about or at most about 7 days, at least, at most, at least about or at most about 8 days, at least, at most, at least about or at most about 9 days, at least, at most, at least about or at most about 10 days, at least, at most, at least about or at most about 11 days, at least, at most, at least about or at most about 12 days, at least, at most, at least about or at most about 13 days, or at least, at most, at least about or at most about 14 days).
[0106] The terms “peptide,” “polypeptide,” or “protein,” used interchangeably throughout this document, generally refer to polymers of at least two amino acid residues linked by peptide bonds. This terminology does not imply a specific length of polymer, nor is it intended to suggest or distinguish whether the peptide was produced using recombinant technology, chemical synthesis, or enzymatic synthesis, or whether it is naturally occurring. These terms apply to naturally occurring amino acid polymers as well as amino acid polymers containing at least one modified amino acid. In some cases, the polymer may be interrupted by non-amino acid chains. These terms include amino acid chains of any length, including full-length proteins, and proteins with or without secondary and / or tertiary structures (e.g., domains). These terms also cover modified amino acid polymers, such as those modified by disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, oxidation, and any other manipulation, such as conjugation with labeled components. As used herein, the terms “amino acid” and “multiple amino acids” generally refer to natural and non-natural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids can include both natural and non-natural amino acids that have been chemically modified to include groups or chemical moieties not naturally present on the amino acid. Amino acid analogs can refer to amino acid derivatives. The term "amino acid" includes both D-amino acids and L-amino acids.
[0107] Regarding peptides, the terms “derivative,” “variant,” or “fragment” used herein generally refer to peptides that are related to wild-type peptides, for example, by their amino acid sequence, structure (e.g., secondary and / or tertiary structure), activity (e.g., enzyme activity), and / or function. Compared to wild-type peptides, peptide derivatives, variants, and fragments may contain one or more amino acid variations (e.g., mutations, insertions, and deletions), truncations, modifications, or combinations thereof.
[0108] Regarding polypeptide molecules (e.g., proteins), the terms "engineered," "chimeric," or "recombinant" as used herein generally refer to polypeptide molecules that have heterologous or altered amino acid sequences due to the application of genetic engineering techniques to the nucleic acid encoding the polypeptide molecule and the cell or organism expressing the polypeptide molecule. Regarding polynucleotide molecules (e.g., DNA or RNA molecules), the terms "engineered" or "recombinant" as used herein generally refer to polynucleotide molecules that have heterologous or altered nucleic acid sequences due to the application of genetic engineering techniques. Genetic engineering techniques include, but are not limited to: PCR and DNA cloning techniques; transfection, transformation, and other gene transfer techniques; homologous recombination; site-directed mutagenesis; and gene fusion. In some cases, engineered or recombinant polynucleotides (e.g., genomic DNA sequences) can be modified or altered through gene editing.
[0109] The terms "engineered" and "modified" are used interchangeably herein. The terms "engineered cell" and "modified cell" are used interchangeably herein. The terms "engineered feature" and "modified feature" are used interchangeably herein.
[0110] The terms "enhanced expression," "increased expression," or "upregulated expression" generally refer to the production of a moiety of interest (e.g., a polynucleotide or polypeptide) at a level higher than the normal expression level of that moiety of interest in a host strain (e.g., a host cell). This normal expression level can be essentially zero (or null) or above zero. The moiety of interest can comprise an endogenous gene or polypeptide construct of the host strain. Alternatively, the moiety of interest can comprise a heterologous gene or polypeptide construct introduced into or derived from the host strain. For example, a heterologous gene encoding a polypeptide of interest can be knocked into the genome of a host strain to enhance the expression of that polypeptide of interest in the host strain.
[0111] The terms "enhanced activity," "increased activity," or "upregulated activity" generally refer to modifications that increase the activity of the moiety of interest (e.g., a polynucleotide or polypeptide) to a level higher than the normal activity level of that moiety of interest in a host strain (e.g., a host cell). This normal activity level can be essentially zero (or null) or above zero. The moiety of interest may comprise a polypeptide construct of the host strain. The moiety of interest may also comprise a heterologous polypeptide construct derived from or introduced into the host strain. For example, a heterologous gene encoding the polypeptide of interest can be knocked into the genome of a host strain (KI) to enhance the activity of that polypeptide of interest in the host strain.
[0112] The terms "reduced expression," "lowered expression," or "downregulated expression" generally refer to the production of a moiety of interest (e.g., a polynucleotide or polypeptide) at a level below the normal expression level of that moiety of interest in a host strain (e.g., a host cell). This normal expression level is above zero. The moiety of interest may comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest may be knocked out or knocked down in the host strain. In some instances, reduced expression of the moiety of interest may include complete suppression of such expression in the host strain.
[0113] The terms "reduced activity," "decreased activity," or "downregulated activity" generally refer to the modification of a moiety of interest (e.g., a polynucleotide or polypeptide) to a level below the normal activity level of that moiety of interest in a host strain (e.g., a host cell). This normal activity level is above zero. The moiety of interest may comprise an endogenous gene or polypeptide construct of the host strain. In some cases, the moiety of interest may be knocked out or knocked down in the host strain. In some instances, reduced activity of the moiety of interest may include complete inhibition of this activity in the host strain.
[0114] The terms “subject,” “individual,” or “patient,” used interchangeably in this document, generally refer to vertebrates, preferably mammals such as humans. Mammals include, but are not limited to, rodents, apes, humans, livestock, locomotor animals, and pets. It also encompasses tissues, cells, and their progeny from biological entities obtained in vivo or cultured in vitro.
[0115] The terms “treatment” or “treatment” generally refer to a method of obtaining a beneficial or desired outcome, including but not limited to therapeutic and / or preventive benefits. For example, treatment may include administration of a system or cell population disclosed herein. A therapeutic benefit refers to any treatment-related improvement or effect on one or more diseases, conditions, or symptoms. For preventive benefits, the composition may be administered to a subject at risk of developing a particular disease, condition, or symptom, or to a subject who reports one or more physiological symptoms of a disease, even if the disease, condition, or symptom may not yet have manifested.
[0116] The term "effective amount" or "therapeutic effective amount" generally refers to an amount of composition (e.g., a composition comprising a heterologous peptide, a heterologous polynucleotide, and / or modified cells (e.g., modified stem cells)) sufficient to produce the desired activity when administered to a subject in need. In the context of this disclosure, the term "therapeutic effective" generally means an amount of composition sufficient to delay the manifestation, halt the progression, alleviate, or reduce at least one symptom of a disorder treated by the methods of this disclosure.
[0117] The term "muscle cell" as used in this article generally refers to any cell that makes up muscle tissue. Myoblasts, satellite cells, myotubes, and myofibrils are all included in the term "muscle cell." Muscle cell effects can be induced in skeletal muscle, cardiac muscle, and smooth muscle.
[0118] Aberrant expression of one or more genes can lead to disease or symptoms. Aberrant expression may be characterized by abnormally low expression levels of the gene, or abnormally high expression levels. In some cases, the gene can be genetically modified (e.g., by action of a nuclease, such as a CRISPR-Cas enzyme) to reverse the aberrant expression (e.g., for the treatment of Duchenne muscular dystrophy (DMD)). Alternatively, the aberrant expression can be transiently modified without genetic modification of the genes of interest, for example, by targeting the gene with a gene effector (e.g., an inactivated CRISPR-Cas enzyme coupled to the gene effector). Transient modification of the aberrant expression of a target gene in cells may not be sufficient to treat or cure the disease manifested by the aberrant expression of the target gene. Therefore, in some embodiments, this disclosure provides systems and methods for modifying the aberrant expression of a target gene such that the modified expression level of the target gene can be maintained for a long period of time.
[0119] Modification of abnormal expression of target genes
[0120] This disclosure provides compositions, systems, and methods for regulating aberrant expression of target genes in cells (e.g., muscle cells). For example, the target gene may be within a D4Z4 repeat array. The target gene may encode at least a portion of DUX4. The compositions, systems, and methods disclosed herein may utilize at least a heterologous peptide (e.g., a heterologous execution). Apparatus The components optionally contain heterologous polynucleotides, such as guide nucleic acid molecules, to modify the expression level and / or epigenetic modification level (e.g., methylation level) of a target gene. For example, the compositions, systems, and methods disclosed herein can utilize heterologous components operably coupled (e.g., covalently or non-covalently) to heterologous gene effectors or regulatory factors (e.g., gene activators, gene repressors, etc.). Apparatus Partially, to modify the expression level and / or epigenetic modification level of the target gene. In some embodiments, the heterologous peptide contains a heterologous execution. Apparatus Part of the heterogeneous execution Apparatus Partially coupled (e.g., covalently or nonvalently) to a heterologous gene effector or regulator (e.g., gene actuator, gene repressor, etc.). In some embodiments, the heterologous peptide comprises a heterologous effector. Apparatus Part of the heterogeneous execution ApparatusThe product contains a nuclease and a guide nucleic acid molecule configured to form a complex with a heterologous polypeptide. In some embodiments, the guide nucleic acid molecule configured to form a complex with a heterologous polypeptide comprises: (i) a scaffold sequence configured to form a complex with the heterologous polypeptide; and (ii) a spacer sequence that exhibits specific binding to a target polynucleotide sequence, for example, targeting at least a portion of DUX4 as described herein. In some embodiments, the nuclease is fused (directly or indirectly) to a heterologous gene effector or regulator (e.g., a gene actuator, gene repressor, etc.). In some embodiments, the nuclease is fused at its C-terminus to a heterologous gene effector or regulator (e.g., a gene actuator, gene repressor, etc.).
[0121] In some embodiments, the system of this disclosure includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or adjacent to a D4Z4 repeat array in muscle cells; wherein, after the complex is formed, the complex is capable of binding the target polynucleotide sequence to achieve modification of the expression level and / or methylation level of a target gene in muscle cells, wherein the target gene is located within a D4Z4 repeat array. In some embodiments, (1) the guide nucleic acid molecule comprises a spacer sequence having specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity with a polynucleotide sequence of one or more members in Table 4, optionally wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that is identical to SEQ ID NO:800, SEQ ID NO:836, or SEQ ID NO:800. The polynucleotide sequence of NO:851 exhibits at least or at least about 80%, at least or at least about 90%, or at least or at least about 95% sequence identity; and / or (2) the system is configured to reduce the expression level of the target gene in muscle cells by at least or at least about 80% compared to control cells; and / or (3) the system is configured to modify the expression level and / or methylation level of the target gene in muscle cells while exhibiting off-target effects below a predetermined threshold level in muscle cells; and / or (4) the system is configured to continuously regulate the expression level of downstream genes of the target gene. And / or methylation levels for at least approximately 5 days; and / or (5) configuring the system to normalize the apoptosis levels of muscle cells to levels comparable to those of healthy muscle cells; and / or (6) configuring the system to improve the survival rate of muscle cells in vitro or in vivo; and / or (7) configuring the system to exhibit minimal impact on the expression profile of at least one muscle cell-specific gene in muscle cells, wherein the muscle cell-specific gene is different from the target gene; and / or (8) configuring the system to exhibit minimal impact on at least one health indicator of the subject after administration of the system to the subject.In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease comprising an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence consisting of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:828. Any of SEQ ID NO:865 has, has about, or has at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof) comprising an amino acid sequence of any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having a sequence of any of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865.In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease being operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence formed by amino acids having, having, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, and SEQ ID NO:828. Any one of SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833 and SEQ ID NO:865 encodes a polynucleotide sequence that has about, or has at least 80%, 85%, 90%, 95%, 97%, 98%, 99% or about 100% sequence identity. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease being operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence of any one of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence of any one of SEQ ID NOs:15-42 and SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865.In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease being operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence of any one of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease being operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprising an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence composed of sequences corresponding to SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, and SEQ ID NO:827. Either NO:833 or SEQ ID NO:865 encodes a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity.In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease (e.g., Cas12f or a variant thereof), the nuclease being operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865. In any system or complex of this disclosure, the guide nucleic acid molecule may comprise (i) a scaffold sequence configured to be complexed with the nuclease, and (ii) a spacer sequence as described herein. In any of the systems or complexes disclosed herein, in some embodiments, a heterologous gene effector (directly or indirectly) is fused to the C-terminus of a nuclease. In any of the systems or complexes disclosed herein, in some embodiments, a heterologous gene effector (directly or indirectly) is fused to the N-terminus of a nuclease. In any of the systems or complexes disclosed herein, in some embodiments, a heterologous gene effector is fused internally to the nuclease. In some embodiments, the heterologous polypeptide comprises a Cas12f-KRAB-DNMT3L regulator, wherein the nuclease Cas12f (or a variant thereof) as described herein is operatively coupled (e.g., fused at the C-terminus, with or without a linker) to a heterologous gene effector comprising KRAB and DNMT3L as described herein.
[0122] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:800. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:800. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0123] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:867. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:867. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0124] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:836. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:836. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0125] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:851. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:851. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0126] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:874. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:874. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0127] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:841. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:841. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0128] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:830. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:830. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0129] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:833. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:833. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0130] In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:865. In some embodiments, the system includes: a heterologous polypeptide and a guide nucleic acid molecule, the heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises the amino acid sequence of SEQ ID NO:728 and the heterologous gene effector comprises the amino acid sequence of SEQ ID NO:727, the guide nucleic acid molecule being configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence of SEQ ID NO:865. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the C-terminus of the nuclease. In some embodiments, the heterologous gene effector is fused (directly or indirectly) to the N-terminus of the nuclease.
[0131] In some cases, the cell can be a muscle cell. The muscle cells disclosed herein can be any type of muscle cell at any developmental stage. Muscle cells can contain undifferentiated muscle cells (e.g., monocytes, such as muscle stem cells, muscle satellite cells, myoblasts, etc.). Alternatively or supplemented, muscle cells can contain differentiated muscle cells (e.g., multinucleated muscle cells, such as myotubes). Muscle cells can be skeletal muscle cells, cardiomyocytes, or smooth muscle cells. For example, skeletal muscle cells can be primary myoblasts (e.g., immortalized primary myoblast lines). In some cases, the cell can be a non-muscle cell, such as a lymphoblast.
[0132] In some cases, the target gene may be located on chromosome 4 of the cells disclosed herein. In other cases, the target gene may be located on chromosome 10 of the cells, for example, on the distal portion of the q (long) arm of chromosome 10.
[0133] Prior to the target gene modifications disclosed herein, aberrant expression of the target gene may be characterized by the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500% or more higher than the level in control cells (e.g., healthy cells in healthy subjects). Aberrant expression may be characterized by the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1% or less higher than the level in control cells (e.g., healthy cells in healthy subjects).
[0134] Prior to the target gene modifications disclosed herein, aberrant expression of the target gene may be characterized by the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99% or more lower than the level in control cells (e.g., healthy cells in healthy subjects). Aberrant expression may be characterized by the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or less lower than the level in control cells (e.g., healthy cells in healthy subjects).
[0135] Prior to the target gene modifications disclosed herein, aberrant expression of the target gene may be characterized by the following: the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500% or more longer than the duration in control cells (e.g., healthy cells in healthy subjects). Aberrant expression may be characterized by the following: the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1% or less longer than the duration in control cells (e.g., healthy cells in healthy subjects).
[0136] Prior to the target gene modifications disclosed herein, aberrant expression of the target gene may be characterized by the following: the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99% or more shorter than the duration in control cells (e.g., healthy cells in healthy subjects). Aberrant expression may be characterized by the following: the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) being at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1% or less shorter than the duration in control cells (e.g., healthy cells in healthy subjects).
[0137] Following the target gene modifications disclosed herein, the modification leading to aberrant target gene expression may be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control (e.g., no modification). The modification leading to aberrant target gene expression may also be characterized by an increase in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less compared to a control.
[0138] Following the target gene modifications disclosed herein, the modification resulting in aberrant target gene expression may be characterized by a reduction in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more compared to a control (e.g., no modification). The modification resulting in aberrant target gene expression may also be characterized by a reduction in the expression level and / or epigenetic modification level (e.g., methylation level) of the target gene by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.
[0139] Following the target gene modifications disclosed herein, the modification of aberrant target gene expression is characterized by an increase in the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) by at least or at least about 1%, 5%, 10%, 20%, 50%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control (e.g., no modification). The modification of aberrant target gene expression is characterized by an increase in the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) by at most or at most about 500%, 400%, 300%, 200%, 150%, 100%, 50%, 20%, 10%, 5%, 1%, or less compared to a control.
[0140] Following the target gene modifications disclosed herein, the modification of aberrant target gene expression is characterized by a reduction in the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) by at least or at least about 1%, 5%, 10%, 20%, 30%, 40%, 50%, 70%, 99%, or more compared to a control (e.g., no modification). The modification of aberrant target gene expression is characterized by a reduction in the duration of the target gene expression level and / or epigenetic modification level (e.g., methylation level) by at most or at most about 100%, 70%, 50%, 40%, 30%, 20%, 10%, 5%, 1%, or less.
[0141] The system provided herein or its use can, after modification of the target gene, continuously regulate (e.g., reduce) the expression level of at least one downstream gene of the target gene (e.g., downstream genes of Dux4 such as ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43 and / or RFLP2) in vitro, in vitro or in vivo. In some cases, the regulation of the expression level of at least one downstream gene, compared with control cells lacking the system, can be a downregulation of at least, at most, at least about, or at most about 10%, at least, at most, at least about, or at most about 20%, at least, at most, at least about, or at most about 30%, at least, at most, at least about, or at most about 40%, at least, at most, at least about, or at most about 50%, at least, at most, at least about, or at most about 60%, at least, at most, at least about, or at most about 70%, at least, at most, at least about, or at most about 75%, at least, at most, at least about, or at most about 80%, at least, at most, at least about, or at most about 85%, at least, at most, at least about, or at most about 90%, at least, at most, at least about, or at most about 95%, at least, at most, at least about, or at most about 99%, or substantially about 100%.
[0142] Following the target gene modification disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., aberrantly expressed target gene) or any other gene of interest (e.g., downstream genes of the target gene, cell type-specific genes, etc.) can be maintained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, or at least or at least about 11 days. At least or at least about 12 days, at least or at least about 13 days, at least or at least about 2 weeks, at least or at least about 3 weeks, at least or at least about 4 weeks, at least or at least about 2 months, at least or at least about 3 months, at least or at least about 4 months, at least or at least about 5 months, at least or at least about 6 months, at least or at least about 7 months, at least or at least about 8 months, at least or at least about 9 months, at least or at least about 10 months, at least or at least about 11 months, at least or at least about 12 months, at least or at least about 2 years, at least or at least about 3 years, at least or at least about 4 years, at least or at least about 5 years or more. Following the target gene modifications disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., aberrantly expressed target gene) or any other gene of interest (e.g., downstream genes of the target gene, cell type-specific genes, etc.) can be maintained for up to or up to approximately 5 years, 4 years, 3 years, 2 years, 12 months, 11 months, 10 months, 9 months, 8 months, 7 months, 6 months, 5 months, etc. Months, 4 months, 3 months, 2 months, 4 weeks, 3 weeks, 2 weeks, up to or up to 13 days, up to or up to 12 days, up to or up to 11 days, up to or up to 10 days, up to or up to 9 days, up to or up to 8 days, up to or up to 7 days, up to or up to 6 days, up to or up to 5 days, up to or up to 4 days, up to or up to 3 days, up to or up to 2 days, up to or up to 1 day or less.
[0143] Following the target gene modification disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., aberrantly expressed target gene) or any other gene of interest (e.g., downstream genes of the target gene, cell type-specific genes, etc.) can maintain at least or at least about 1 cell division, at least or at least about 2 cell divisions, at least or at least about 3 cell divisions, at least or at least about 4 cell divisions, at least or at least about 5 cell divisions, at least or at least about 6 cell divisions, at least or at least about 7 cell divisions, at least or at least about 8 cell divisions, at least or at least about 9 cell divisions, at least or at least about 10 cell divisions, at least or at least about 15 cell divisions, at least or at least about 20 cell divisions, at least or at least about 25 cell divisions, at least or at least about 30 cell divisions, at least or at least about 40 cell divisions, at least or at least about 50 cell divisions, or at least or at least about 100 cell divisions. Following the target gene modifications disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., an aberrantly expressed target gene) or any other gene of interest (e.g., a downstream gene of the target gene, a cell type-specific gene, etc.) can be maintained for up to or up to about 100 cell divisions, up to or up to about 50 cell divisions, up to or up to about 40 cell divisions, up to or up to about 30 cell divisions, or up to or up to about 25 cell divisions. Cell division, up to or about 20 cell divisions, up to or about 15 cell divisions, up to or about 10 cell divisions, up to or about 9 cell divisions, up to or about 8 cell divisions, up to or about 7 cell divisions, up to or about 6 cell divisions, up to or about 5 cell divisions, up to or about 4 cell divisions, up to or about 3 cell divisions, up to or about 2 cell divisions, or up to or about 1 cell division.
[0144] Following the target gene modification disclosed herein, the modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (e.g., aberrantly expressed target gene) or any other gene of interest (e.g., downstream genes of the target gene, cell type-specific genes, etc.) can be determined at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 2 weeks, at least or at least about 3 weeks, or at least or at least about 4 weeks.
[0145] As disclosed herein, non-limiting examples of epigenetic modifications may include methylation, acetylation, phosphorylation, ADP ribosylation, glycosylation, SUMOylation, ubiquitination, and modifications to histone structures (e.g., via ATP hydrolysis-dependent processes). For example, epigenetic modifications can produce modified methylation levels in one or more target genes or any other gene of interest (e.g., downstream genes of the target gene, cell type-specific genes, etc.).
[0146] As disclosed herein, the persistent modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (or any other gene of interest) may be characterized by maintaining at least or at least about 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of the modified expression level and / or methylation level of the target gene (or any other gene of interest). The persistent modified expression level and / or epigenetic modification level (e.g., methylation level) of the target gene (or any other gene of interest) may be characterized by maintaining at most or at most about 100%, 99%, 98%, 97%, 96%, 95%, 90%, 85%, 80%, 75%, or 70% of the modified expression level and / or methylation level of the target gene (or any other gene of interest).
[0147] The system provided herein or its use may have minimal (e.g., essentially no) impact on the expression profile of at least one cell type-specific gene in target cells that is not a target gene. For example, the system provided herein or its use may have minimal (e.g., comparable) variation in the expression profile of at least one myogenic gene in muscle cells (e.g., during differentiation from myoblasts into myotubes or myofibrils). Non-limiting examples of at least one myogenic gene may include p21, MyoD, Ezh2, Notch1, MyoG, MyHC, MEF2, Smad4, IGF2, Sirt1, myostatin, Myh2, Myh4, Myh1, myomixer, myomaker, and Mrf4. Compared to control cells lacking the system, muscle cells containing this system exhibit the following expression profiles: at least, at most, at least about, or at most about 80%; at least, at most, at least about, or at most about 85%; at least, at most, at least about, or at most about 90%; at least, at most, at least about, or at most about 91%; at least, at most, at least about, or at most about 92%; at least, at most, at least about, or at most about 93%; at least, at most, at least about, or at most about 94%. At least, at most, at least about or at most about 95%, at least, at most, at least about or at most about 96%, at least, at most, at least about or at most about 97%, at least, at most, at least about or at most about 98%, at least, at most, at least about or at most about 99%, at least, at most, at least about or at most about 100%, at least, at most, at least about or at most about 105%, at least, at most, at least about or at most about 110%, or at least, at most, at least about or at most about 120%. Following the target gene modifications described herein, the minimal impact on the expression profile of at least one cell type-specific gene can be maintained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, or at least or at least about 4 weeks. The expression profile of at least one cell type-specific gene can be determined by a variety of methods, such as, but not limited to, reverse transcription polymerase chain reaction (RT-PCR or real-time quantitative RT-PCR) or Western blotting.
[0148] The system provided herein, or its use in cells (e.g., diseased cells, such as FSHD muscle cells), can normalize the level of apoptosis in cells, for example, by regulating the expression level of at least one stress-related marker (e.g., caspase protein) provided herein in cells to a level similar to, close to, or equivalent to the “normal” level in control cells (e.g., healthy muscle cells in a comparable state of differentiation). Following target gene modification in the cells provided herein, the regulated expression level of at least one stress-related marker in the cells, compared to the normal level in control cells, can be at least, at most, at least about, or at most about 80%, at least, at most, at least about, or at most about 85%, at least, at most, at least about, or at most about 90%, at least, at most, at least about, or at most about 91%, at least, at most, at least about, or at most about 92%, at least, at most, at least about, or at most about 93%, at least, at most, at least About or at most about 94%, at least, at most, at least about or at most about 95%, at least, at most, at least about or at most about 96%, at least, at most, at least about or at most about 97%, at least, at most, at least about or at most about 98%, at least, at most, at least about or at most about 99%, at least, at most, at least about or at most about 100%, at least, at most, at least about or at most about 105%, at least, at most, at least about or at most about 110%, or at least, at most, at least about or at most about 120%. The normalized expression level of at least one stress-related biomarker can be maintained for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, at least or at least about 4 weeks, at least or at least about 2 months, at least or at least about 6 months, or at least or at least about 1 year. After modifying the target gene in the cells provided herein in the diseased cells, the expression level of at least one stress-related marker may be at least or at least about 5%, at least or at least about 10%, at least or at least about 15%, at least or at least about 20%, at least or at least about 25%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, or at least or at least about 90% lower than the level in control diseased cells (e.g., lacking the system).
[0149] The system described herein, or its use in cells (e.g., diseased cells, such as FSHD muscle cells), can improve cell survival in vivo. The improved survival rate, compared to control cells lacking this system (e.g., control diseased cells), can be at least or at least about 5%, at least or at least about 10%, at least or at least about 15%, at least or at least about 20%, at least or at least about 25%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, or at least or at least about 90%. The improved survival rate can be determined, for example, by the duration of survival in vivo. After modulation of the target gene in cells via the system, the improved survival rate can be determined at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, at least or at least about 4 weeks, at least or at least about 2 months, at least or at least about 6 months, or at least or at least about 1 year later. Cells can be contacted with the system (e.g., to achieve modulation of the target gene) before, during, or after administration to a subject in need (e.g., transplantation into muscle tissue). Alternatively, the system can be administered to a subject in need to contact cells in vivo, thereby achieving in vivo modulation of the target gene or any other gene of interest in the cells (e.g., downstream genes of the target gene, cell type-specific genes, etc.).
[0150] The system described herein, or its use in cells (e.g., diseased cells, such as FSHD muscle cells), can have minimal (e.g., essentially no) impact on one or more health indicators of a subject. For example, one or more health indicators can be measured after the subject is given the system to achieve modification of a target gene in the subject's cells. Non-limiting examples of health conditions may include food intake (e.g., grams per subject per day), weight (e.g., grams), body dimensions (e.g., length, height, or circumference in meters), body mass index (e.g., grams per square centimeter), body fat (e.g., milligrams of fat per gram of body weight), organ weight (e.g., grams per gram of body weight of liver, kidney, spleen, small intestine, etc.), clinical chemistry, tissue function or integrity (e.g., structures determined by biopsy and / or histology), physical injury, pain, appearance, etc. Clinical chemistry can be determined by measuring one or more markers from bodily fluids (e.g., blood), and non-limiting examples of clinical chemistry markers may include electrolytes (e.g., sodium, potassium, chloride, bicarbonate), renal markers (e.g., creatinine, blood urea nitrogen), liver function markers (e.g., albumin, globulin, albumin / globulin ratio, bilirubin, aspartate aminotransferase (AST), alanine aminotransferase (ALT), gamma-glutamyl transferase (GGT), alkaline phosphatase (ALP)), and cardiac markers. Biomarkers (e.g., H-FABP, troponin, myoglobin, CK-MB, B-type natriuretic peptide (BNP)), minerals (e.g., calcium, magnesium, phosphate, potassium), blood disorder markers (e.g., iron, transferrin, TIBC, vitamin B12, vitamin D, folic acid), and others (e.g., glucose, C-reactive protein, glycated hemoglobin, uric acid, arterial blood gas, adrenocorticotropic hormone (ACTH), neuron-specific enolase (NSE), fecal occult blood test (FOBT), etc.). For example, a group of clinical chemical biomarkers may be used, including one or more members of sodium, potassium, chloride, bicarbonate, blood urea nitrogen (BUN), creatinine, glucose, calcium, total protein, albumin, alkaline phosphatase (ALP), alanine aminotransferase (ALT), aspartate aminotransferase (AST), or bilirubin.Compared to control subjects who did not receive this treatment, subjects receiving the system or method described herein may have health indicators of at least, at most, at least about, or at most about 80%, at least, at most, at least about, or at most about 85%, at least, at most, at least about, or at most about 90%, at least, at most, at least about, or at most about 91%, at least, at most, at least about, or at most about 92%, at least, at most, at least about, or at most about 93%, at least, at most, at least about, or at most about 94%. At least, at most, at least about or at most about 95%, at least, at most, at least about or at most about 96%, at least, at most, at least about or at most about 97%, at least, at most, at least about or at most about 98%, at least, at most, at least about or at most about 99%, at least, at most, at least about or at most about 100%, at least, at most, at least about or at most about 105%, at least, at most, at least about or at most about 110%, or at least, at most, at least about or at most about 120%. The minimum impact on health following treatment may last (and / or may be determined at the following times) for at least or at least about 1 day, at least or at least about 2 days, at least or at least about 3 days, at least or at least about 4 days, at least or at least about 5 days, at least or at least about 6 days, at least or at least about 7 days, at least or at least about 8 days, at least or at least about 9 days, at least or at least about 10 days, at least or at least about 11 days, at least or at least about 12 days, at least or at least about 13 days, at least or at least about 14 days, at least or at least about 3 weeks, or at least or at least about 4 weeks.
[0151] The systems, compositions, and methods disclosed herein can be used to treat, inhibit, or improve a disease in a subject (e.g., muscular dystrophy, such as facioscapulohumeral muscular dystrophy (FSHD)).
[0152] Heteropeptides
[0153] Whether alone or in combination with one or more auxiliaries such as heteropolynucleotides (e.g., guide nucleic acids), the heteropolypeptides disclosed herein can be configured to specifically bind to target polynucleotide sequences, thereby modulating the expression and / or epigenetic levels of target genes (e.g., D4Z4 repeat arrays) in target cells. The target polynucleotide sequence may be located at (e.g., within) the target gene. Alternatively, the target polynucleotide sequence may be adjacent to the target gene. For example, the target polynucleotide sequence may be adjacent to the end of the target gene (e.g., the 5' or 3' end). The target polynucleotide sequence may be at least or at least about 5 nucleotides from the end of the target gene, at least or at least about 10 nucleotides, at least or at least about 20 nucleotides, at least or at least about 30 nucleotides, at least or at least about 40 nucleotides, at least or at least about 50 nucleotides, at least or at least about 100 nucleotides, at least or at least about 150 nucleotides, at least or at least about 200 nucleotides, at least or at least about 250 nucleotides, at least or at least about 300 nucleotides, at least or at least about 400 nucleotides, at least or at least about 500 nucleotides, at least or at least about 1000 nucleotides, at least or at least about 1500 nucleotides, at least or at least about 2000 nucleotides, at least or at least about 3000 nucleotides, at least or at least about 4000 nucleotides, or at least or at least about 5000 nucleotides. The target polynucleotide sequence can be located at most or at most approximately 5000 nucleotides from the end of the target gene, at most or at most approximately 4000 nucleotides, at most or at most approximately 3000 nucleotides, at most or at most approximately 2000 nucleotides, at most or at most approximately 1500 nucleotides, at most or at most approximately 1000 nucleotides, at most or at most approximately 500 nucleotides, at most or at most approximately 400 nucleotides. Up to or about 300 nucleobases, up to or about 200 nucleobases, up to or about 150 nucleobases, up to or about 100 nucleobases, up to or about 50 nucleobases, up to or about 40 nucleobases, up to or about 30 nucleobases, up to or about 20 nucleobases, up to or about 10 nucleobases, or up to or about 5 nucleobases.
[0154] Without being limited by theory, when the target polynucleotide sequence is not within the target gene, the target polynucleotide sequence can interact with at least a portion of the target gene (e.g., the promoter sequence of the target gene) (e.g., through direct or indirect binding), such that the binding or targeting of the target polynucleotide sequence by at least a heterologous polypeptide (e.g., a complex comprising a heterologous polypeptide and a heterologous polynucleotide disclosed herein) can target said at least a portion of the target gene (e.g., the promoter sequence), thereby achieving regulation of the expression level and / or epigenetic level of the target gene in the cell.
[0155] In some cases, the target polynucleotide sequence may comprise multiple target polynucleotide sequences. These multiple target polynucleotide sequences may be within the target gene. Alternatively, as disclosed herein, the multiple target polynucleotide sequences may be outside the target gene but adjacent to it. However, in another alternative, the multiple target polynucleotide sequences may comprise at least one target polynucleotide sequence within the target gene (e.g., within a D4Z4 repeat domain) and at least one additional target polynucleotide sequence outside the target gene but adjacent to it. In this case, targeting both at least one target polynucleotide sequence and at least one additional target polynucleotide sequence may produce a greater effect (e.g., a greater degree of regulation of the expression level and / or epigenetic level of the target gene) (e.g., at least 0.1-fold, at least 0.5-fold, at least 1-fold, at least 2-fold, at least 3-fold, at least 4-fold, at least 5-fold, at least 10-fold, at least 15-fold, at least 20-fold, or more) compared to targeting only one of the at least one target polynucleotide sequence and at least one additional target polynucleotide sequence.
[0156] The heterologous peptides disclosed herein may contain one or more heterologous gene effectors (e.g., gene effectors that are heterologous to cells containing gene effectors and / or another component of the complex disclosed herein). Heterologous gene effectors may contain domains capable of or expected to regulate the expression of target genes (e.g., target endogenous genes), such as activating, repressing, upregulating, downregulating, or stabilizing gene expression or activity levels. Heterologous gene effectors may be heterologous relative to another component present in the complex (e.g., guide portions, such as nucleases and / or guide nucleic acids disclosed herein). In some cases, heterologous gene effectors may be heterologous relative to the host cells in which they are introduced.
[0157] Heterogeneous effectors may be or may contain sequences from any suitable source, such as amino acid sequences from human proteins, viral proteins, or other proteins disclosed herein. Heterogeneous effectors may be or may contain sequences from proteins predominantly localized to the cell nucleus, such as members of the human nuclear proteome. Heterogeneous effectors may be or may contain one or more native amino acid residues. Heterogeneous effectors may be or may contain one or more synthetic amino acid residues.
[0158] Heterogeneous effectors can or may contain sequences from mammalian proteins. Heterogeneous effectors can or may contain sequences from human proteins. Heterogeneous effectors can or may contain sequences from viral proteins. Heterogeneous effectors can or may contain sequences from non-human primate proteins. Heterogeneous effectors can or may contain sequences from non-human mammalian proteins. Heterogeneous effectors can or may contain sequences from non-rodent mammalian proteins. Heterogeneous effectors can or may contain sequences from plant proteins. Heterogeneous effectors can or may contain sequences from pig proteins. Heterogeneous effectors can or may contain sequences from lagomorpha proteins. Heterogeneous effectors can or may contain sequences from canine proteins. Heterogeneous effectors can or may contain sequences from avian proteins. Heterogeneous effectors can or may contain sequences from reptile proteins. Heterogeneous effectors can or may contain sequences from bacterial proteins. Heterogeneous effectors can or may contain sequences from archaea proteins.
[0159] For example, the amino acid sequences of the heterologous gene effectors disclosed herein may not be derived from, and need not be derived from, bacterial proteins (e.g., they may be derived from archaea proteins). Without wishing to be limited by theory, subjects in need may be treated with compositions containing heterologous gene effectors of non-bacterial protein origin, such that the compositions do not (i) induce bacterial stimulation and / or (ii) elicit a bacterial immune response in the subject.
[0160] heterogeneous execution ApparatusThe portion may contain a nuclease (e.g., an endonuclease). For example, the nuclease may be a CRISPR / Cas protein. The length of the nuclease may be less than a threshold length. The threshold length may be up to or up to about 1000 amino acids, up to or up to about 950 amino acids, up to or up to about 900 amino acids, up to or up to about 850 amino acids, up to or up to about 800 amino acids, up to or up to about 750 amino acids, up to or up to about 700 amino acids, up to or up to about 650 amino acids, up to or up to about 600 amino acids, up to or up to about 550 amino acids, up to or up to about 500 amino acids, up to or up to about 450 amino acids, up to or up to about 400 amino acids, up to or up to about 350 amino acids, or up to or up to about 300 amino acids. The threshold length can be at least or at least about 300 amino acids, at least or at least about 350 amino acids, at least or at least about 400 amino acids, at least or at least about 450 amino acids, at least or at least about 500 amino acids, at least or at least about 550 amino acids, at least or at least about 600 amino acids, at least or at least about 650 amino acids, at least or at least about 700 amino acids, at least or at least about 750 amino acids, at least or at least about 800 amino acids, at least or at least about 850 amino acids, at least or at least about 900 amino acids, at least or at least about 950 amino acids, or at least or at least about 1000 amino acids.
[0161] To avoid being limited by theory, using nucleases smaller than the threshold length can offer one or more advantages compared to using control nucleases larger than the threshold length. When using delivery media with limited size (e.g., limited physical size for encapsulating the nuclease or limited expression cassette size, such as a viral genome), sufficient space can be reserved (or sufficient space within the expression cassette) for one or more auxiliaries (such as one or more gene regulators (e.g., transcription regulators)) and / or one or more heteropolynucleotides (e.g., one or more guide nucleic acid molecules)). Alternatively or complementary, using nucleases smaller than or equal to the threshold size can elicit a greater effect on the regulation of target gene expression levels and / or epigenetic levels compared to the effect of control nucleases larger than the threshold size.
[0162] In some instances, the nucleases disclosed herein (e.g., having a size less than or equal to a threshold size) may regulate the expression levels and / or epigenetic levels of target genes (e.g., increase or decrease) by at least, at most, at least about, or at most about 0.1 times, at least, at most, at least about, or at most about 0.5 times, at least, at most, at least about, or at most about 1 time, at least, at most, at least about, or at most about 2 times, at least, at most, at least about, or at most about 3 times, at least, at most, at least about, or at most about 4 times, at least, at most, at least about, or at most about 5 times, at least, at most, at least about, or at most about 6 times, at least, at most, at least about, or at most Approximately 7 times, at least, at most, at least about or at most about 8 times, at least, at most, at least about or at most about 9 times, at least, at most, at least about or at most about 10 times, at least, at most, at least about or at most about 15 times, at least, at most, at least about or at most about 20 times, at least, at most, at least about or at most about 25 times, at least, at most, at least about or at most about 30 times, at least, at most, at least about or at most about 40 times, at least, at most, at least about or at most about 50 times, at least, at most, at least about or at most about 60 times, at least, at most, at least about or at most about 70 times, at least, at most, at least about or at most about 80 times, at least, at most, at least about or at most about 90 times, or at least, at most, at least about or at most about 100 times.
[0163] In some instances, the degree of regulation (e.g., increase or decrease) of the expression level and / or epigenetic level of the nuclease disclosed herein (e.g., having a size less than or equal to a threshold size) of the target gene expression level can last at least, at most, at least about, or at most about 0.1 times longer than the degree of regulation of the control nuclease (e.g., having a size greater than the threshold size), at least, at most, at least about, or at most about 0.5 times, at least, at most, at least about, or at most about 1 time, at least, at most, at least about, or at most about 2 times, at least, at most, at least about, or at most about 3 times, at least, at most, at least about, or at most about 4 times, at least, at most, at least about, or at most about 5 times, at least, at most, at least about, or at most about 6 ... Up to approximately 7 times, at least, at most, at least about or at most 8 times, at least, at most, at least about or at most 9 times, at least, at most, at least about or at most 10 times, at least, at most, at least about or at most 15 times, at least, at most, at least about or at most 20 times, at least, at most, at least about or at most 25 times, at least, at most, at least about or at most 30 times, at least, at most, at least about or at most 40 times, at least, at most, at least about or at most 50 times, at least, at most, at least about or at most 60 times, at least, at most, at least about or at most 70 times, at least, at most, at least about or at most 80 times, at least, at most, at least about or at most 90 times, or at least, at most, at least about or at most 100 times.
[0164] In some instances, the degree of regulation (e.g., increase or decrease) of the expression level and / or epigenetic level of the nucleases disclosed herein (e.g., having a size less than or equal to the threshold size) of the target gene expression level can last longer (or be maintained longer) than the degree of regulation by a control nuclease (e.g., having a size greater than the threshold size) by at least, at most, at least about, or at most about 1 cell division, at least, at most, at least about, or at most about 2 cell divisions, at least, at most, at least about, or at most about 3 cell divisions, at least, at most, at least about, or at most about 4 cell divisions, at least, at most, at least about, or at most about 5 cell divisions, at least, at most, at least about, or at most about 6 cell divisions, at least, at most, at least about, or at most about 7 cell divisions, at least, at most, at least about, or at most about 8 cell divisions, at least, at most, at least about, or at most about 9 cell divisions, at least, at most, at least about, or at most about 10 cell divisions. At least, at most, at least about, or at most about 11 cell divisions; at least, at most, at least about, or at most about 12 cell divisions; at least, at most, at least about, or at most about 13 cell divisions; at least, at most, at least about, or at most about 14 cell divisions; at least, at most, at least about, or at most about 15 cell divisions; at least, at most, at least about, or at most about 16 cell divisions; at least, at most, at least about, or at most about 17 cell divisions; at least, at most... At least approximately or at most approximately 18 cell divisions, at least approximately or at most approximately 19 cell divisions, at least approximately or at most approximately 20 cell divisions, at least approximately or at most approximately 25 cell divisions, at least approximately or at most approximately 30 cell divisions, at least approximately or at most approximately approximately 40 cell divisions, at least approximately or at most approximately 50 cell divisions, or at least approximately 100 cell divisions. As disclosed herein, cell division may be characterized by the division of a parent cell into two daughter cells having substantially the same genetic material as the parent cell.
[0165] The heterologous gene effector can be or may contain sequences from chromatin regulatory factors (CRs). Chromatin regulatory factors include functional domains from various histones and DNA modifying enzymes (e.g., DNMTs, HATs, or HMTs). In some embodiments, the heterologous gene effector is DNMT, which contains DNMT-A or DNMT-L. In some embodiments, "DNMT-L" is DNMT3L or contains DNMT3L. In some embodiments, "DNMT-A" is DNMT3A or contains DNMT3A. In some embodiments, the heterologous gene effector is DNMT, which contains DNMT3A or DNMT3L.
[0166] Heterogeneous gene effectors may contain two or more domains derived from chromatin regulators, for example, located in tandem or separately at the C-terminus, N-terminus, or within the polypeptide sequence.
[0167] In some implementations, heterologous gene effectors promote heterochromatin formation. Non-limiting examples of proteins that can promote heterochromatin formation include HP1α, HP1β, KAP1, KRAB, SUV39H1, or G9a.
[0168] In some embodiments, heterologous gene effectors regulate histones via methylation. In some embodiments, heterologous gene effectors regulate histones via acetylation. In some embodiments, heterologous gene effectors regulate histones via phosphorylation. In some embodiments, heterologous gene effectors regulate histones via ADP-ribosylation. In some embodiments, heterologous gene effectors regulate histones via glycosylation. In some embodiments, heterologous gene effectors regulate histones via SUMOylation. In some embodiments, heterologous gene effectors regulate histones via ubiquitination. In some embodiments, heterologous gene effectors regulate histones by remodeling histone structure, for example, through ATP-dependent processes.
[0169] In some embodiments, heterologous gene effectors facilitate the spatial localization of proteins on or near target polynucleotides, such as transcriptional repressors, transcription factors, or histones. In some embodiments, heterologous gene effectors are used to manipulate the spatiotemporal configuration of genomic DNA and RNA components in the cell nucleus and / or cytoplasm, for example, to regulate different cellular functions.
[0170] In some implementations, the heterologous gene effector is derived from histone acetyltransferases. Non-limiting examples of histone acetyltransferases include the GNAT subfamily, MYST subfamily, p300 / CBP subfamily, HAT1 subfamily, GCN5, PCAF, Tip60, MOZ, MORF, MOF, HBO1, p300, CBP, HAT1, ATF-2, SRC1, or TAFII250.
[0171] In some implementations, the heterologous gene effector originates from histone lysine methyltransferases. Non-limiting examples of histone lysine methyltransferases include the EZH subfamily, Non-SET subfamily, other SET subfamily, PRDM subfamily, SET1 subfamily, SET2 subfamily, SUV39 subfamily, SYMD subfamily, ASH1L, EHMT1, EHMT2, EZH1, EZH2, MLL, MLL2, MLL3, MLL4, MLL5, NSD1, NSD2, NSD3, PRDM1, PRDM10, PRDM11, PRDM12, PRDM13, PRDM14, PR... DM15, PRDM16, PRDM2, PRDM4, PRDM5, PRDM6, PRDM7, PRDM8, PRDM9, SET1, SET1L, SET2L, SETD2, SETD3, SETD4, SETD5, SETD6, SETD7, SETD8, SETDB1, SETDB2, SETMAR, SUV39H1, SUV39H2, SUV420H1, SUV420H2, SYMD1, SYMD2, SYMD3, SYMD4 or SYMD5.
[0172] In some embodiments, the heterologous gene effector is derived from a component of the chromatin remodeling complex. In some embodiments, the heterologous gene effector is a component of BAF, such as actin, ARIDA / B, BAF155, BAF170, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRG1 / BRM, INI1, or SS18.
[0173] In some implementations, the heterologous gene effector is derived from a component of PBAF. For example, actin, ARID2, BAF155, BAF170, BAF180, BAF45 A / B / C / D, BAF53 A / B, BAF57, BAF60 A / B / C, BRD7, BRG1, or INI1.
[0174] In some implementations, the heterologous gene effector is derived from components of the ISWI family chromatin remodeling complex. Examples include the ACF subfamily, RSF subfamily, CERF subfamily, CHRAC subfamily, NURF subfamily, NoRC subfamily, WICH subfamily, b-WICH subfamily, ACF1, ATPase, BPTF, CECR2, CHRAC15, CHRAC17, CSB, DEK, MYBBP1A, NM1, RBAP46 / 48, RHII / Gua, RSF1, SAP155, SNF2H, SNF2H / L, SNF2L, TIP5, or WSTF.
[0175] In some embodiments, the heterologous gene effector is derived from a component of the CHD family complex, such as the NuRD complex, NuRD-like complex, or CHD complex. In some embodiments, the heterologous gene effector is derived from CHD1 / 2 / 6 / 7 / 8 / 9, CHD3 / 4, CHD5, GATAD2 A / B, GATAD2 B, HDAC1, HDAC2, HDAC2, MBD2 / 3, MTA1 / 2 / 3, MTA3, or RBAP46, RBAP46 / 48.
[0176] In some implementations, the heterologous gene effector is derived from components of the INO80 family complex, such as the INO80 complex, Tip60 / p400 complex, SRCAP complex, AMIDA, ARP6, BAF53, BAF53, BAF53A, BRD8, DMAP1, DMAP1, EPC1 / 2, FLJ11730, GAS41, GAS41, IES2, IES6, ING3, INO80, INO80E, MCRS1, MRG15, MRGBP, MRGX, NFRKB, p400, RUVBL1 / 2, RUVBL1 / 2, RUVBL1 / 2, SRCAP, Tip60, TRRAP, UCH37, YL-1, YL-1, YY1, or ZnF-HIT1.
[0177] Heterogeneous gene effectors can be or may contain sequences from transcriptional regulatory factors (TRs). TR gene effectors include transcriptional regulatory domains from various transcription factor families, such as KRAB, p65, MED, or GTFs.
[0178] Heterogeneous gene effectors may contain transcription activator domains. Heterogeneous gene effectors may contain two or more tandem transcription activator domains, for example, located at the C-terminus, N-terminus, or within the polypeptide sequence.
[0179] Non-limiting examples of transcriptional activation domains include GAL4, herpes simplex virus activation domain VP16, VP64 (a tetramer of herpes simplex virus activation domain VP16), the NF-KB p65 subunit, or the EBV R transactivator (Rta). In some embodiments, such a transcriptional activation domain is used as a control in the methods of this disclosure. In some embodiments, such a transcriptional activation domain is used as a heterogeneous effector in a complex containing at least one additional heterogeneous effector (e.g., a different effector).
[0180] Heterogeneous gene effectors may contain transcriptional repression domains. Heterogeneous gene effectors may contain two or more transcriptional repression domains, for example, located in tandem or separately at the C-terminus, N-terminus, or within the polypeptide sequence.
[0181] Non-limiting examples of transcriptional repression domains include the KRAB (Kruppel-associated box) domain of Kox1, the MadmSIN3 interaction domain (SID), or the ERF repression domain (ERD). In some embodiments, such a transcriptional repression domain is used as a control in the methods of this disclosure. In some embodiments, such a transcriptional repression domain is used as a heterogeneous effector in a complex containing at least one additional heterogeneous effector (e.g., a different effector).
[0182] In some implementations, the heterologous gene effector is derived from a gene product, which is a transcription factor.
[0183] In some embodiments, the heterologous gene effector is derived from a gene product, which is a hematopoietic stem cell transcription factor. Non-limiting examples of hematopoietic stem cell transcription factors include AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, Helios, HES-1, HHEX, HIF-1α / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activator, STAT repressor, STAT3, STAT4, STAT5a, STAT6 or TSC22.
[0184] In some embodiments, the heterologous gene effector is derived from a gene product, which is a mesenchymal stem cell transcription factor. Non-limiting examples of mesenchymal stem cell transcription factors include DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, Myocardin, MyoD, Myogenin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT activator, STAT repressor, STAT1, STAT3, TBX18, Twist-1, or Twist-2.
[0185] In some embodiments, the heterologous gene effector is derived from a gene product, which is an embryonic stem cell transcription factor. Non-limiting examples of embryonic stem cell transcription factors include Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3α / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB repressor, NFkB1, NFk B2, Oct-3 / 4, Otx2, p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activator, STAT repressor, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206 or ZNF281.
[0186] In some embodiments, the heterologous gene effector is derived from a gene product, which is an induced pluripotent stem cell (iPSC) transcription factor. Non-limiting examples of iPSC transcription factors include KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, or TBX18.
[0187] In some embodiments, the heterologous gene effector is derived from a gene product, which is an epithelial stem cell transcription factor. Non-limiting examples of epithelial stem cell transcription factors include ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4α / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, and Nrf. 2. p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2, RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT activator, STAT repressor, STAT3, SUZ12, TCF-3 / E2A or TCF7 / TCF1.
[0188] In some embodiments, the heterologous gene effector is derived from a gene product, which is a cancer stem cell transcription factor. Non-limiting examples of cancer stem cell transcription factors include androgens R / NR3C4, AP-2γ, β-catenin, β-catenin repressor, Brachyury, CREB, ERα / NR3A1, ERβ / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, GLI-2, GLI-3, HIF-1α / HIF1A, HIF-2α / EPAS1, HMGA1B, c-Jun, JunB, and KLF4. c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1 or ZEB1.
[0189] In some implementations, the heterologous gene effector is derived from a gene product, which is a cancer-associated transcription factor. Non-limiting examples of cancer-associated transcription factors include ASCL1 / Mash1, ASCL2 / Mash2, ATF1, ATF2, ATF4, BLIMP1 / PRDM1, CDX2, CDX4, DLX5, DNMT1, E2F-1, EGR1, ELF3, Ets-1, FosB / G0S3, FoxC1, FoxC2, FoxF1, GADD153, GATA-2, HMGA2, HMGB1 / HMG-1, and HNF-3α / F oxA1, HNF-6 / ONECUT1, HSF1, ID1, ID2, JunD, KLF10, KLF12, KLF17, LMO2, MEF2C, MYCL1 / L-Myc, NFkB2, Oct-1, p63 / TP73L, Pax3, PITX2, Prox1, RAP80, Rex-1 / ZFP42, RUNX1 / CBFA2, RUNX3 / CBFA3, SALL4, SCL / Tal1, Sirtuin 2 / SIRT2, Smad3, Smad4, Smad5, SOX11, STAT5a / b, STAT5a, STAT5b, TCF7 / TCF1, TORC1, TORC2, TRIM32, TRPS1 or TSC22.
[0190] In some embodiments, the heterologous gene effector is derived from a gene product, which is an immune cell transcription factor. Non-limiting examples of immune cell transcription factors include AP-1, Bcl6, E2A, EBF, Eomes, FoxP3, GATA3, Id2, Ikaros, IRF, IRF1, IRF2, IRF3, IRF7, NFAT, NFkB, Pax5, PLZF, PU.1, ROR-γ-T, STAT, STAT1, STAT2, STAT3, STAT4, STAT5, STAT5A, STAT5B, STAT6, T-bet, TCF7, or ThPOK.
[0191] In some embodiments, the heterologous gene effector originates from a gene product, said gene product being an RNA polymerase-associated protein. In some embodiments, the heterologous gene effector originates from a transcription factor having a basic domain. In some embodiments, the heterologous gene effector originates from a transcription factor having a zinc-coordinated DNA-binding domain. In some embodiments, the heterologous gene effector originates from a transcription factor having a helix-turn-helix domain. In some embodiments, the heterologous gene effector originates from a transcription factor having an α-helix DNA-binding domain. In some embodiments, the heterologous gene effector originates from a transcription factor having an α-helix exposed through a β-structure. In some embodiments, the heterologous gene effector originates from a transcription factor having an immunoglobulin fold. In some embodiments, the heterologous gene effector originates from a transcription factor having a β-hairpin exposed through an α / β scaffold. In some embodiments, the heterologous gene effector originates from a transcription factor having a DNA-binding β-sheet. In some embodiments, the heterologous gene effector originates from a transcription factor having a β-barrel DNA-binding domain.
[0192] In some embodiments, the heterologous gene effector is derived from a gene product, which is a nuclear receptor, such as a nuclear hormone receptor. Non-limiting examples of nuclear hormone receptors include NR0B1, NR0B2, NR1A1, NR1A2, NR1B1, NR1B2, NR1B3, NR1C1, NR1C2, NR1C3, NR1D1, NR1D2, NR1F1, NR1F2, NR1F3, NR1H4, NR1H5, NR1H3, NR1H2, NR1I1, NR1I2, NR1I3, NR2A1, and NR2A2. 2. Those encoded with NR2B1, NR2B2, NR2B3, NR2C1, NR2C2, NR2E1, NR2E3, NR2F1, NR2F2, NR2F6, NR3A1, NR3A2, NR3B1, NR3B2, NR3B3, NR3C4, NR3C1, NR3C2, NR3C3, NR4A1, NR4A2, NR4A3, NR5A1, NR5A2, or NR6A1.
[0193] In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in nucleosome assembly. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in DNA metabolism. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in nucleotide metabolism. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in ribosome biogenesis. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in protein folding. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in translation. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in signal transduction. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in proteolysis. In some embodiments, the heterologous gene effector originates from the gene product, and the gene product participates in the negative regulation of endopeptidase activity.
[0194] In some embodiments, the heterologous gene effectors or gene regulators used interchangeably herein may comprise polypeptide sequences that show at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity with any of the heterologous gene effector amino acid sequences provided in Table 3.
[0195] Table 3. Amino acid sequences of heterologous gene effectors.
[0196] SEQ ID NO: Heterologous gene effector amino acid sequence 15 NNSQGRVTFEDVTVNFTQGEWQRLNPEQRNLYRDVMLENYSNLVSVGQGETTKPDVILRLEQGKEPWLEEEEVLGSGRAEKNGDI 16 SGHPGSWEMNSVAFEDVAVNFTQEEWALLDPSQKNLYRDVMQETFRNLASIGNKGEDQSIEDQYKNSSRNLRHIISHSGNNPYGC 17 AAATLRTPTQGTVTFEDVAVHFSWEEWGLLDEAQRCLYRDVMLENLALLTSLDVHHQKQHLGEKHFRSNVGRALFVKTCTFHVSG 18 TTFKEAMTFKDVAVVFTEEELGLLDLAQRKLYRDVMLENFRNLLSVGHQAFHRDTFHFLREEKIWMMKTAIQREGNSGDKIQTEM 19 VPAETSSSGLLEEQKMMKSQGLVSFKDVAVDFTQEEWQQLDPSQRTLYRDVMLENYSHLVSMGYPVSKPDVISKLEQGEEPWIIK 20 MKSQGLVSFKDVAVDFTQEEWQQLDPSQRTLYRDVMLENYSHLVSMGYPVSKPDVISKLEQGEEPWIIKGDISNWIYPDEYQADG 21 AEGSVMFSDVSIDFSQEEWDCLDPVQRDLYRDVMLENYGNLVSMGLYTPKPQVISLLEQGKEPWMVGRELTRGLCSDLESMCETK 22 AAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGKVLTPHPSILSWARLFLLFL 23 AAAALRDPAQVPVAADLLTDHEEQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEP 24 AAAALRDPAQVPVAADLLTDHEEQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGCWHGAEAEEAPEQIASVG 25 AAAALRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGCWHGAEAEEAPEQIASVGLLSSNIQQHQKQH 26 AAAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGKVLTPHPSILSWARLFLLF 27 YVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFMPAWEVVTSAIPRGSWWVELREV 28 AAAALRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFMPAWEVVTSAIPR 29 AAAALRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGCWHGAEAEEAPEQIASVGLLSSNIQQHQKQHC 30 AAAALRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFMPAWEVVTSAIP 31 AAAALRDPAQVPVAADLLTDHEEGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPF 32 VTCAHLGRRARLPAAQPSACPGTCFSQEERMAAGYLPRWSQELVTFEDVSMDFSQEEWELLEPAQKNLYREVMLENYRNVVSLEA 33 LVTFEDVSMDFSQEEWELLEPAQKNLYREVMLENYRNVVSLEALKNQCTDVGIKEGPLSPAQTSQVTSLSSWTGYLLFQPVASSH 34 KNATIVMSVRREQGSSSGEGSLSFEDVAVGFTREEWQFLDQSQKVLYKEVMLENYINLVSIGYRGTKPDSLFKLEQGEPPGIAEG 35 SSGEGSLSFEDVAVGFTREEWQFLDQSQKVLYKEVMLENYINLVSIGYRGTKPDSLFKLEQGEPPGIAEGAAHSQICPDADFLE 36 GPLQFRDVAIEFSLEEWHCLDTAQRNLYRNVMLENYSNLVFLGITVSKPDLITCLEQGRKPLTMKRNEMIAKPSVSFLQVHSESQ 37 GPLQFRDVAIEFSLEEWHCLDTAQRNLYRNVMLENYSNLVFLGITVSKPDLITCLEQGRKPLTMKRNEMIAKPSVMCSHFAQDLW 38 APPSAPLPAQGPGKARPSRKRGRRPRALKFVDVAVYFSPEEWGCLRPAQRALYRDVMRETYGHLGALGCAGPKPALISWLERNTD 39 QTNTKDWTVTPEHVLPESQSLLTFEEVAMYFSQEEWELLDPTQKALYNDVMQENYETVISLALFVLPKPKVISCLEQGEEPWVQV 40 AAATLRDPAQQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFMPAWEVVTSAIL 41 AAATLRDPAQGYVTFEDVAVYFSQEEWRLLDDAQRLLYRNVMLENFTLLASLGLASSKTHEITQLESWEEPFMPAWEVVTSAILR 42 DSVAFEDVAVNFTQEEWALLDPSQKNLYREVMQETLRNLTSIGKKWNNQYIEDEHQNPRRNLRRLIGERLSESKESHQHGEVLTQ
[0197] In some embodiments, the heterologous gene effectors or gene regulators used interchangeably herein may comprise polypeptide sequences that show at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity with SEQ ID NO:727 shown below:
[0198] RTLVTFKDVFVDFTREEWKLLDTAQQIVYRNVMLENYKNLVSLGYQLTKPDVILRLEKGEEPGKESGSVGGSGGSSEQLAQFRSLDGMAAIPALDPEAEPSMDVILVGSSELSSSVS PGTGRDLIAYEVKANQRNIEDICICCGSLQVHTQHPLFEGGICAPCKDKFLDALFLYDDDGYQSYCSICCSGETLLICGNPDCTRCYCFECVDSLVGPGTSGKVHAMSNWVCYLCLPS SRSGLLQRRRKWRSQLKAFYDRESENPLEMFETVPVWRRQPVRVLSLFEDIKKELTSLGFLESGSDPGQLKHVVDVTDTVRKDVEEWGPFDLVYGATPPLGHTCDRPPSWYLFQFHRL LQYARPKPGSPRPFFWMFVDNLVLNKEDLDVASRFLEMEPVTIPDVHGGSLQNAVRVWSNIPAIRSRHWALVSEEELSLLAQNKQSSKLAAKWPTKLVKNCFLPLREYFKYFST (SEQ ID NO: 727)
[0199] In some embodiments, the heterologous peptide comprises a Cas12f-KRAB-DNMT3L regulator. In some embodiments, the Cas12f-KRAB-DNMT3L regulator comprises a heterologous gene effector comprising KRAB and DNMT3L. In some embodiments, the Cas12f-KRAB-DNMT3L regulator comprises a heterologous gene effector comprising SEQ ID NO:727, or a sequence having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99% identity with SEQ ID NO:727, or about 100% identity, or a percentage identity within a range defined by any two of the foregoing values (e.g., 80%-100%, 85%-97%, 90%-95%, 85%-95%, etc.).
[0200] The heteropolynucleotides disclosed herein may include one or more guide moieties (e.g., one or more guide nucleic acid molecules) to direct heterogene effectors toward target genes (e.g., target endogenous genes) or target gene regulatory sequences. The guide moieties may confer the ability to recognize and specifically bind to target genes or target gene regulatory sequences. The guide moieties may be configured to form complexes with heteropeptides (e.g., guide nucleic acids that form complexes with nucleases such as CRISPR / Cas proteins), and these complexes may be configured to exhibit specific binding to the target peptide sequences disclosed herein, thereby modifying the expression level and / or epigenetic modification level of the target gene.
[0201] The guide portion may contain a guide nucleic acid. The guide portion may contain a nuclease and a guide nucleic acid disclosed herein. The guide portion may contain a nuclease or a portion thereof, such as a nuclease (e.g., a heterologous nuclease). The nuclease may be, for example, a DNA nuclease and / or an RNA nuclease, a modified nuclease that is nuclease-deficient or has reduced nuclease activity compared to a wild-type nuclease, its derivatives, variants, or fragments thereof. In some embodiments, the guide portion has minimal nuclease activity.
[0202] Any suitable nuclease, fragment, or derivative thereof may be used in the wizard section. Suitable nucleases include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases, including type I, type II, type III, type IV, type V, and type VI CRISPR-associated (Cas) peptides; zinc finger nucleases (ZFNs); transcription activator-like effector nucleases (TALENs); meganucleases; RNA-binding proteins (RBPs); CRISPR-associated RNA-binding proteins; recombinases; flipases; transposases; Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaea Argonaute (aAgo), or eukaryotic Argonaute (eAgo)); or any derivative thereof; or any variant thereof or any functional fragment thereof.
[0203] In some embodiments, the guide portion comprises a DNA nuclease, such as a nuclease-deficient engineered (e.g., programmable or targetable) DNA nuclease. In some embodiments, the guide portion comprises a nuclease-inactivated DNA-binding protein derived from a DNA nuclease, said protein not inducing transcriptional activation or repression of the target DNA sequence unless it is present in the complex with one or more heterologous gene effectors of this disclosure. In some embodiments, the guide portion comprises a nuclease-inactivated DNA-binding protein derived from a DNA nuclease, said protein may induce transcriptional activation or repression of the target DNA sequence (e.g., which may be altered or enhanced by the presence of heterologous gene effectors of this disclosure).
[0204] In some embodiments, the guide portion comprises an RNA nuclease, such as an engineered (e.g., programmable or targetable) RNA nuclease. In some embodiments, the guide portion comprises a nuclease-inactivated RNA-binding protein derived from an RNA nuclease, said protein not inducing transcriptional activation or repression of the target RNA sequence unless it is present in the complex with one or more heterologous gene effectors of this disclosure. In some embodiments, the guide portion comprises a nuclease-inactivated RNA-binding protein derived from an RNA nuclease, said protein being capable of inducing transcriptional activation or repression of the target RNA sequence (e.g., which may be altered or enhanced by the presence of heterologous gene effectors of this disclosure).
[0205] In some embodiments, the guide portion comprises a nucleic acid-guided targeting system. In some embodiments, the guide portion comprises a DNA-guided targeting system. In some embodiments, the guide portion comprises an RNA-guided targeting system. The guide portion may comprise and utilize, for example, a guide nucleic acid sequence that facilitates the specific binding of a CRISPR-Cas system (e.g., its nuclease-deficient form, such as dCas9) to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence. Binding specificity can be determined by using a guide nucleic acid (such as a single guide RNA (sgRNA) or a portion thereof). In some embodiments, the use of different sgRNAs allows the compositions and methods of this disclosure to be used with different target genes (e.g., target endogenous genes) or target gene regulatory sequences (e.g., targeting them).
[0206] Prokaryotic CRISPR-Cas (clustered regularly spaced short palindromic repeats-CRISPR-associated) systems, such as class II CRISPR-Cas systems (like Cas9 and Cpfl), can be reused in the compositions and methods of this disclosure as tools for regulating gene expression, epigenome editing, and chromatin circularization. Nuclease-inactivated Cas (dCas) proteins complexed with heterologous gene effectors can allow for the regulation of expression of target genes (e.g., target endogenous genes) adjacent to the dCas binding site.
[0207] In some implementations, the guide portion includes a CRISPR-associated (Cas) protein or Cas nuclease that functions within a non-naturally occurring CRISPR (clustered regularly spaced short palindromic repeats) / Cas (CRISPR-associated) system. In bacteria, this system can provide adaptive immunity against foreign DNA.
[0208] In a wide range of organisms, including various mammals, animals, plants, microorganisms, and yeast, the CRISPR / Cas system (e.g., modified and / or unmodified) can be used as a tool for genome engineering, or can be modified to guide engineered proteins to specifically bind to target loci disclosed herein. The CRISPR / Cas system may contain a guide nucleic acid (e.g., guide RNA (gRNA)) complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid binding. RNA-guided Cas proteins (e.g., Cas nucleases, such as Cas9 nuclease) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner. If the Cas protein possesses nuclease activity, it can cleave DNA.
[0209] In some cases, Cas proteins are mutated and / or modified to produce nuclease-deficient proteins or proteins with reduced nuclease activity relative to wild-type Cas proteins. Nuclease-deficient proteins may retain the ability to bind DNA, but may lack or have reduced nucleic acid cleavage activity.
[0210] In some embodiments, the guide portion comprises a Cas protein that forms a complex with a guide nucleic acid (such as guide RNA or a portion thereof). In some embodiments, the guide portion comprises a Cas protein that forms a complex with a single guide nucleic acid (such as single guide RNA (sgRNA)). In some embodiments, the guide portion comprises an RNA-binding protein (RBP) that optionally complexes with a guide nucleic acid (such as guide RNA (e.g., sgRNA)) capable of forming a complex with the Cas protein. In some embodiments, the guide portion comprises a nuclease-inactivated DNA-binding protein derived from a DNA nuclease, which can induce transcriptional activation or repression of a target DNA sequence. In some embodiments, the guide portion comprises an RNA-inactivated RNA-binding protein derived from an RNA nuclease.
[0211] In some embodiments, the guide nucleic acid used in the compositions and methods of this disclosure may be, for example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10 or more nucleotides.
[0212] In some embodiments, the guide nucleic acids used in the compositions and methods of this disclosure are up to 10, up to 9, up to 8, up to 7, up to 6, up to 5, up to 4, up to 3, up to 2, or up to 1 nucleotide.
[0213] Guide nucleic acids can be guide RNA or a portion thereof.
[0214] Any suitable CRISPR / Cas system can be used. CRISPR / Cas systems can be referred to using various nomenclature systems. A CRISPR / Cas system can be type I, type II, type III, type IV, type V, type VI, or any other suitable CRISPR / Cas system. The CRISPR / Cas system used in this paper can be class 1, class 2, or any other suitable classification of CRISPR / Cas system. Class 1 or class 2 can be determined based on the gene encoding the effector module. Class 1 systems typically have multi-subunit crRNA effector complexes, while class 2 systems typically have a single protein (e.g., Cas9, Cpfl, C2c1, C2c2, C2c3) or crRNA effector complex. Class 1 CRISPR / Cas systems can be regulated using complexes of multiple Cas proteins. Class 1 CRISPR / Cas systems may include, for example, type I (e.g., I, IA, IB, IC, ID, IE, IF, or IU), type III (e.g., III, IIIA, IIIB, IIIC, or IIID), and type IV (e.g., IV, IVA, or IVB) CRISPR / Cas types. Class 2 CRISPR / Cas systems can achieve regulation using a single large Cas protein. Class 2 CRISPR / Cas systems may include, for example, type II (e.g., II, IIA, or IIB) and type V CRISPR / Cas types. CRISPR systems can be complementary to each other and / or can provide trans-functional units to facilitate CRISPR locus targeting.
[0215] When the guide portion contains a Cas protein or a derivative thereof, the Cas protein or derivative thereof may be a class 1 or class 2 Cas protein. The Cas protein may be a type I, II, III, IV, V, or VI Cas protein. The Cas protein may contain one or more domains. Non-limiting examples of domains include guide nucleic acid recognition and / or binding domains, nuclease domains (e.g., DNase or RNase domains, RuvC or HNH), DNA-binding domains, RNA-binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. The guide nucleic acid recognition and / or binding domain may interact with the guide nucleic acid. The nuclease domain may contain catalytic activity for nucleic acid cleavage. The nuclease domain may lack catalytic activity to prevent nucleic acid cleavage. The Cas protein may be a chimeric Cas protein or fragment thereof fused with other proteins or peptides. The Cas protein may be a chimera of multiple Cas proteins, for example, containing domains from different Cas proteins.
[0216] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csn1 or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Cse1 (CasA), Cse2 (CasB), and Cse3 (CasE). Cse4 (CasC), Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cul966, Cas13a, Cas13b, Cas13c, Cas13d, Cas13X or Cas13Y, and their homologs or modified versions.
[0217] In some cases, the Cas protein disclosed herein may not be, and does not need to be, Cas9 or Cas12a. The Cas protein disclosed herein may have a smaller size compared to Cas9 or Cas12a. The Cas protein disclosed herein may be derived from Un1Cas12f1. For example, the Cas protein disclosed herein may contain an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO:43. In another example, the Cas protein disclosed herein may comprise an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO:44. As disclosed herein, SEQ ID NO:43 encodes a polypeptide sequence of Un1Cas12f1. As disclosed herein, SEQ ID NO:44 encodes an engineered variant of Un1Cas12f1 having reduced nuclease activity. As disclosed herein, SEQ ID NO:728 encodes a non-limiting example of a Cas12f variant suitable for use in the systems, compositions, and methods of this disclosure. In some embodiments, the Cas12f variants disclosed herein may contain an amino acid sequence that is at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% identical to the polypeptide sequence of SEQ ID NO:728.
[0218] SEQ ID NO: 43 (Un1Cas12f1)
[0219] 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC
[0220] 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE
[0221] 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSDVCYTRAA
[0222] 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG
[0223] 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI
[0224] 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKIGEKS
[0225] 301 AWMLNLSIDV PKIDKGVDPS IIGGIDVGVK SPLVCAINNA FSRYSISDND
[0226] 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI
[0227] 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN
[0228] 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC
[0229] 501 EKCNFKENAD YNAALNISNP KLKSTKEEP
[0230] SEQ ID NO: 44 (Inactivated nuclease variant of Un1Cas12f1)
[0231] 1 MAKNTITKTL KLRIVRPYNS AEVEKIVADE KNNREKIALE KNKDKVKEAC
[0232] 51 SKHLKVAAYC TTQVERNACL FCKARKLDDK FYQKLRGQFP DAVFWQEISE
[0233] 101 IFRQLQKQAA EIYNQSLIEL YYEIFIKGKG IANASSVEHY LSRVCYRRAA
[0234] 151 ELFKNAAIAS GLRSKIKSNF RLKELKNMKS GLPTTKSDNF PIPLVKQKGG
[0235] 201 QYTGFEISNH NSDFIIKIPF GRWQVKKEID KYRPWEKFDF EQVQKSPKPI
[0236] 251 SLLLSTQRRK RNKGWSKDEG TEAEIKKVMN GDYQTSYIEV KRGSKICEKS
[0237] 301 AWMLNLSIDV PKIDKGVDPS IIGGIAVGVR SPLVCAINNA FSRYSISDND
[0238] 351 LFHFNKKMFA RRRILLKKNR HKRAGHGAKN KLKPITILTE KSERFRKKLI
[0239] 401 ERWACEIADF FIKNKVGTVQ MENLESMKRK EDSYFNIRLR GFWPYAEMQN
[0240] 451 KIEFKLKQYG IEIRKVAPNN TSKTCSKCGH LNNYFNFEYR KKNKFPHFKC
[0241] 501 EKCNFKENAA YNAALNISNP KLKSTKERP
[0242] SEQ ID NO: 728 (Cas12f variant)
[0243] MAKNTITKTLKLRIVRPYNSAEVEKIVADEKERRKQAGGTGELDDKFYQKLRGQFPDAVFWQEISEIFRQLQKQAAEIYNQSLIELYYEIFIKGKGIANASVEHYLSRVCYRRAAELFKNAA IAGLRSKIKSNFRLKELKNMKSGLPTTKSDNFPIPLVKQKGGQYTGFEISNHNSDFIIKIPFGRWQVKKEIDKYRPWEKFDFEQVQKSPKPISLLLSTQRRKRNKGWSKDEGTEAEIKKVMNGD YQTSYIEVKRGSKICEKSAWMLNLSIDVPKIDKGVDPSIIGGIAVGVRSPLVCAINNAFSRYSISDNDLFHFNKKMFARRRILLKKNRHKRAGHGAKNKLKPITILTEKSERFRKKLIERWAC EIADFFIKNKVGTVQMENLESMKRKEDSYFNIRLRGFWPYAEMQNKIEFKLKQYGIEIRKVAPNNTSKTCSKCGHLNNYFNFEYRKKNKFPHFKCEKCNFKENAAYNAALNISNPKLKSTKERP
[0244] In some embodiments, the heterologous polypeptide comprises the Cas12f-KRAB-DNMT3L regulator. In some embodiments, the Cas12f-KRAB-DNMT3L regulator comprises a nuclease comprising Cas12f or a variant thereof. In some embodiments, the Cas12f-KRAB-DNMT3L regulator comprises a nuclease having the amino acid sequence of SEQ ID NO:44, or having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99% identity or about 100% identity, or a percentage identity within a range defined by any two of the foregoing values (e.g., 80%-100%, 85%-97%, 90%-95%, 85%-95%, etc.) with SEQ ID NO:44. In some embodiments, the Cas12f-KRAB-DNMT3L regulator comprises a nuclease having the amino acid sequence of SEQ ID NO:728, or having, having about, or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99% identity or about 100% identity with SEQ ID NO:728, or a percentage identity within a range defined by any two of the foregoing values (e.g., 80%-100%, 85%-97%, 90%-95%, 85%-95%, etc.).
[0245] In some cases, the Cas protein described in this article may not be the Cas14 protein. In some cases, the dead Cas protein (dCas) described in this article may not be the dead Cas protein.
[0246] Cas proteins or fragments or derivatives thereof can originate from any suitable organism. Non-limiting examples include *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Streptococcus* sp., *Staphylococcus aureus*, *Nocardiopsis dassonvillei*, *Streptomyces pristinae spiralis*, *Streptomyces viridochromogenes*, *Streptomyces viridochromogenes*, *Streptomyces viridochromogenes*, *Streptosporangium roseum*, *Streptosporangium roseum*, *AlicyclobacHlus acidocaldarius*, *Bacillus pseudomycoides*, *Bacillus selenitireducens*, *Exiguobacterium sibiricum*, *Lactobacillus delbrueckii*, *Lactobacillus salivarius*, and *Microscilla*. marina, Burkholderiales bacterium, Polaromonas naphthalenivorans, Polaromonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp.Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena The species *Variabilis*, *Nodularia spumigena*, *Nostoc* sp., *Arthrospira maxima*, *Arthrospira platensis*, *Arthrospira* sp., *Lyngbya* sp., *Microcoleus chthonoplastes*, *Oscillatoria* sp., *Petrotoga mobilis*, *Thermosipho africanus*, *Acaryochloris marina*, *Leptotrichia shahii*, or *Francisella novicida*. In some respects, the organism is *S. pyogenes*. In some respects, the organism is *S. aureus*. In some respects, the organism is *S. thermophilus*.
[0247] Cas proteins can originate from a variety of bacterial species, including but not limited to Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, and Mycoplasma gallisepticum. Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, and Lactobacillus coryniformis subsp.Torquens), Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermuscellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinella succinogenes, Campylobacter jejuni subsp.*Jejuni*, *Helicobacter mustelae*, *Bacillus cereus*, *Acidovorax ebreus*, *Clostridium perfringens*, *Parvibaculum lavamentivorans*, *Roseburia intestinalis*, *Neisseria meningitidis*, *Pasteurella multocida subsp. Multocida*, *Sutterella wadsworthensis*, *proteobacterium*, *Legionellapneumophila*, *Parasutterella excrementihominis*, *Wolinella succinogenes*, or *Francisella novicida*.
[0248] The Cas protein used in this article may be a wild-type or modified form of the Cas protein. The Cas protein may be an active variant, inactive variant, or fragment of the wild-type or modified Cas protein. Relative to the wild-type version of the Cas protein, the Cas protein may contain amino acid alterations, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof. The Cas protein may be a polypeptide having at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to the wild-type Cas protein. The Cas protein can be a polypeptide that has at most or at most about 5%, at most or at most about 10%, at most or at most about 20%, at most or at most about 30%, at most or at most about 40%, at most or at most about 50%, at most or at most about 60%, at most or at most about 70%, at most or at most about 80%, at most or at most about 90%, or at most or at most about 100% sequence identity and / or sequence similarity to the wild-type exemplary Cas protein. The variant or fragment may contain at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to wild-type or modified Cas proteins or a portion thereof. The variant or fragment may target a nucleic acid locus that complexes with a guide nucleic acid while lacking nucleic acid cleavage activity.
[0249] Cas proteins may contain one or more nuclease domains, such as DNase domains. For example, the Cas9 protein may contain a RuvC-like nuclease domain and / or an HNH-like 20 nuclease domain. In the nuclease-active form of Cas9, the RuvC and HNH domains can each cleave different strands of double-stranded DNA, thereby creating double-strand breaks in the DNA. Cas proteins may contain only one nuclease domain (e.g., Cpfl contains a RuvC domain but lacks an HNH domain). In some embodiments, the nuclease domain is absent. In some embodiments, the nuclease domain is present but inactive, or has reduced or minimal activity. In some embodiments, the nuclease domain is present and active.
[0250] One or more nuclease domains (e.g., RuvC or HNH) of a Cas protein can be deleted or mutated, rendering them nonfunctional or containing reduced nuclease activity. For example, in a Cas protein containing at least two nuclease domains (e.g., Cas9), if one nuclease domain is deleted or mutated, the resulting Cas protein (called a nickase) can produce single-strand breaks, but not double-strand breaks, at CRISPR RNA (crRNA) recognition sequences within double-stranded DNA. This nickase can cleave either the complementary or non-complementary strand, but may not cleave both simultaneously. If all nuclease domains of a Cas protein (e.g., both the RuvC and HNH nuclease domains in Cas9; the RuvC nuclease domain in Cpfl) are deleted or mutated, the resulting Cas protein may have reduced or no ability to cleave both strands of double-stranded DNA. An example of a mutation that can convert a Cas9 protein into a nickase is the D10A mutation (changing the 10th aspartic acid in Cas9 to alanine) in the RuvC domain of Cas9 from Streptococcus pyogenes. The H939A (histidine at amino acid position 839 is replaced with alanine) or H840A (histidine at amino acid position 840 is replaced with alanine) mutations in the HNH domain of Cas9 from Streptococcus pyogenes can convert Cas9 into a nicking enzyme. Examples of mutations that can convert the Cas9 protein into dead Cas9 are: the D10A mutation (aspartic acid at amino acid position 10 is replaced with alanine) in the RuvC domain of Cas9 from Streptococcus pyogenes, and the H939A (histidine at amino acid position 839 is replaced with alanine) or H840A (histidine at amino acid position 840 is replaced with alanine) mutations in the HNH domain.
[0251] Nuclease-dead Cas proteins (e.g., proteins derived from any Cas protein, such as Un1Cas12f1) may contain one or more mutations relative to the wild-type version of that protein. Mutations may cause the cleavage activity of one or more of the multiple cleavage domains of the wild-type Cas protein to be no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1%. Mutations may cause one or more of the multiple cleavage domains to retain the ability to cleave the complementary strand of the target nucleic acid, but reduce its ability to cleave the non-complementary strand of the target nucleic acid. Mutations may cause one or more of the multiple cleavage domains to retain the ability to cleave the non-complementary strand of the target nucleic acid, but reduce its ability to cleave the complementary strand of the target nucleic acid. Mutations may cause one or more of the multiple cleavage domains to lack the ability to cleave both the complementary and non-complementary strands of the target nucleic acid. The residues to be mutated in the nuclease domain may correspond to one or more catalytic residues of the nuclease. For example, residues such as Asp10, His840, Asn854, and Asn856 in the wild-type exemplary Streptococcus pyogenes Cas9 polypeptide can be mutated to inactivate one or more of multiple nucleic acid cleavage domains (e.g., nuclease domains). The residues to be mutated in the Cas protein nuclease domain may correspond to the Asp10, His840, Asn854, and Asn856 residues in the wild-type Streptococcus pyogenes Cas9 polypeptide, for example, determined by sequence and / or structural alignment.
[0252] As a non-limiting example, residues D10, G12, G17, E762, H840, N854, N863, H982, H983, A984, D986, and / or A987 (or the corresponding mutation of any Cas protein) can be mutated. Examples include D10A, G12A, G17A, E762A, H840A, N854A, N863A, H982A, H983A, A984A, and / or D986A. Mutations other than alanine substitutions are acceptable.
[0253] The D10A mutation can be combined with one or more of the H840A, N854A, or N856A mutations to produce a Cas9 protein that is substantially lacking in DNA cleavage activity (e.g., the dead Cas9 protein). The H840A mutation can be combined with one or more of the D10A, N854A, or N856A mutations to produce a site-directed peptide that is substantially lacking in DNA cleavage activity. The N854A mutation can be combined with one or more of the H840A, D10A, or N856A mutations to produce a site-directed peptide that is substantially lacking in DNA cleavage activity. The N856A mutation can be combined with one or more of the H840A, N854A, or D10A mutations to produce a site-directed peptide that is substantially lacking in DNA cleavage activity.
[0254] In some embodiments, the Cas protein is a type 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, or a Cas9 protein derived from it. For example, a Cas9 protein lacking cleavage activity. In some embodiments, the Cas9 protein is a Cas9 protein derived from *Streptococcus pyogenes* (e.g., SwissProt accession number Q99ZW2). In some embodiments, the Cas9 protein is a Cas9 protein derived from *Staphylococcus aureus* (e.g., SwissProt accession number J7RUA5). In some embodiments, the Cas9 protein is a modified version of a Cas9 protein derived from *Streptococcus pyogenes* or *Staphylococcus aureus*. In some embodiments, the Cas9 protein is derived from a Cas9 protein derived from *Streptococcus pyogenes* or *Staphylococcus aureus*. For example, a *Streptococcus pyogenes* or *Staphylococcus aureus* Cas9 protein lacking cleavage activity.
[0255] In some embodiments, Cas9 can generally refer to a polypeptide having at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, or about 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas9 polypeptide (e.g., from Streptococcus pyogenes). In some embodiments, Cas9 can refer to a polypeptide having at most about 5%, at most about 10%, at most about 20%, at most about 30%, at most about 40%, at most about 50%, at most about 60%, at most about 70%, at most about 80%, at most about 90%, or about 100% sequence identity and / or sequence similarity to a wild-type Cas9 polypeptide (e.g., from Streptococcus pyogenes). Cas9 can refer to wild-type or modified forms of the Cas9 protein, which can include amino acid alterations such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof.
[0256] The Cas protein may contain an amino acid sequence having at least or at least about 5%, at least or at least about 10%, at least or at least about 20%, at least or at least about 30%, at least or at least about 40%, at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 80%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity or sequence similarity to the nuclease domain (e.g., RuvC domain or HNH domain) of the wild-type Cas protein.
[0257] Cas proteins, their variants, or derivatives can be modified to enhance the regulation of gene expression by the compositions and methods of this disclosure, for example, as part of the complexes disclosed herein. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, enzyme activity, and / or binding to other factors, such as heterodimerization or oligomerization domains and inducible ligands. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for the desired function of the protein or complex. Cas proteins can be modified to modulate (e.g., enhance or reduce) the activity of the Cas protein in regulating gene expression through the complexes of this disclosure containing heterogeneous gene effectors.
[0258] For example, the Cas protein can be coupled (e.g., by fusion, covalent coupling, or non-covalent coupling) to heterogene effectors (e.g., epigenetic modification domains, transcriptional activation domains, and / or transcriptional repression domains). The Cas protein can be coupled (e.g., by fusion, covalent coupling, or non-covalent coupling) to oligomerizing or dimerizing domains disclosed herein (e.g., heterodimerizing domains). The Cas protein can be coupled (e.g., by fusion, covalent coupling, or non-covalent coupling) to heteropeptides that provide increased or decreased stability. The Cas protein can be coupled (e.g., by fusion, covalent coupling, or non-covalent coupling) to sequences that can promote the degradation of the Cas protein or complexes containing the Cas protein, such sequences as degradation determinants, like inducible degradation determinants (e.g., auxin-inducible).
[0259] The Cas protein can be conjugated with any suitable number of chaperones (e.g., fusion, covalent coupling, or non-covalent coupling), such as at least 1, at least 2, at least 3, at least 4, or at least 5, at least 6, at least 7, or at least 8 chaperones. In some embodiments, the Cas protein of this disclosure is conjugated (e.g., fusion, covalent coupling, or non-covalent coupling) to at most 2, at most 3, at most 4, at most 5, at most 6, at most 7, at most 8, or at most 10 chaperones. In some embodiments, the Cas protein of this disclosure is conjugated with 1-5, 1-4, 1-3, 1-2, 2-5, 2-4, 2-3, 3-5, 3-4, or 4-5 chaperones (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is conjugated with 1 chaperone (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is coupled with two chaperones (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is coupled with three chaperones (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is coupled with four chaperones (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is coupled with five chaperones (e.g., fusion, covalent coupling, or non-covalent coupling). In some embodiments, the Cas protein of this disclosure is coupled with six chaperones (e.g., fusion, covalent coupling, or non-covalent coupling).
[0260] Cas proteins can be fusion proteins. The fusion domain or heterologous polypeptide can be located at the N-terminus, C-terminus, or inside the Cas protein.
[0261] Cas proteins can be provided in any form. For example, Cas proteins can be provided as proteins, such as Cas proteins alone, or Cas proteins complexed with guide nucleic acids to form ribonucleoproteins. Cas proteins can be provided as complexes, such as complexed with guide nucleic acids and / or one or more heterologous gene effectors of this disclosure. Cas proteins can be provided as nucleic acids encoding Cas proteins, such as RNA (e.g., messenger RNA (mRNA)) or DNA. The nucleic acids encoding Cas proteins can be codon-optimized for efficient translation into proteins in specific cells or organisms.
[0262] Nucleic acids encoding Cas proteins, fragments thereof, or derivatives can be stably integrated into the cellular genome. Nucleic acids encoding Cas proteins can be operatively linked to promoters, such as constitutively or inducibly active promoters in the cell. Nucleic acids encoding Cas proteins can be operatively linked to promoters in expression constructs. Expression constructs can include any nucleic acid construct capable of guiding the expression of a gene of interest or other nucleic acid sequences of interest (e.g., the Cas gene) and can transfer such nucleic acid sequences of interest to target cells.
[0263] In some implementations, the Cas protein, its variants, or derivatives are nuclease-dead Cas (dCas) proteins. Dead Cas proteins can be proteins lacking nucleic acid cleavage activity.
[0264] Cas proteins can comprise modified forms of wild-type Cas proteins. Modified forms of wild-type Cas proteins can comprise amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the Cas protein. For example, modified forms of Cas proteins can have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of wild-type Cas proteins (e.g., Cas9 from Streptococcus pyogenes). Modified forms of Cas proteins may lack substantial nucleic acid cleavage activity. When a Cas protein is a modified form lacking substantial nucleic acid cleavage activity, it may be referred to as enzymatically inactive, “inactivated,” and / or “dead” (abbreviated as “d”). Dead Cas proteins (e.g., dCas, dCas9) can bind to target polynucleotides but may not cleave or cleave target polynucleotides minimally. In some respects, a dead Cas protein is a dead Cas9 protein.
[0265] The dCas9 polypeptide can bind to a single guide RNA (sgRNA) to activate or repress the transcription of a target gene (e.g., an endogenous target gene), for example, by binding to a heterologous gene effector disclosed herein. The sgRNA can be introduced into cells expressing the Cas or guide fraction components disclosed herein. In some cases, such cells may contain one or more different sgRNAs that target the same target gene (e.g., an endogenous target gene) or a target gene regulatory sequence. In other cases, the sgRNA targets different nucleic acids in the cell (e.g., different target genes, different target gene regulatory sequences, or different sequences within the same target gene or a target gene regulatory sequence).
[0266] Enzymatic inactivity can refer to a nuclease's ability to bind to the nucleic acid sequence in a sequence-specific manner but not to cleave the target polynucleotide, or to cleave it at a significantly reduced frequency. The guide portion of enzymatic inactivity may include an enzymatically inactive domain (e.g., a nuclease domain). Enzymatic inactivity can mean no activity. Enzymatic inactivity can mean essentially no activity. Enzymatic inactivity can mean intrinsically inactive. Enzymatic inactivity can mean activity not exceeding 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the equivalent wild-type activity (e.g., nucleic acid cleavage activity, wild-type Cas9 activity).
[0267] In some implementations, the guide portion does not contain a nucleic acid-guided targeting system. For example, the guide portion may include proteins that bind to target genes (e.g., target endogenous genes) or target gene regulatory sequences based on protein structural features, such as certain nucleases disclosed herein.
[0268] In some embodiments, the guide portion comprises a zinc finger nuclease (ZFN) or a variant, fragment, or derivative thereof. The ZFN may be a fusion of a cleavage domain (such as the cleavage domain of Fok1) with at least one zinc finger motif (e.g., at least two, three, four, or five zinc finger motifs), said fusion being capable of binding polynucleotides such as DNA and RNA. In some embodiments, the targeting portion of this disclosure uses a ZFN to bind polynucleotides (e.g., a target gene or a target gene regulatory sequence), but the ZFN does not cleave or substantially does not cleave polynucleotides, for example, a nuclease-killing ZFN. ZFNs or variants, fragments, or derivatives thereof may be fused or bound to one or more heterologous gene effectors to form the complexes of this disclosure.
[0269] Two separate ZFNs heterodimerize at specific locations on a polynucleotide at specific orientations and with specific spacing, which can induce cleavage of that polynucleotide in a nuclease-active ZFN. For example, a ZFN bound to DNA can induce double-strand breaks in DNA. To allow the two cleavage domains to dimerize and cleave DNA, two separate ZFNs can bind to opposite strands of DNA with their C-termini spaced apart. In some cases, the linker sequence between the zinc finger domain and the cleavage domain may require the 5' edges of each binding site to be separated by 5, 6, or 7 base pairs. In some cases, the cleavage domain is fused to the C-terminus of each zinc finger domain.
[0270] In some embodiments, the cleavage domain comprising the ZFN guide portion comprises a modified form of the wild-type cleavage domain. The modification of the cleavage domain may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the cleavage domain. For example, the modified cleavage domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the corresponding wild-type cleavage domain. The modified cleavage domain may not have substantial nucleic acid cleavage activity. In some embodiments, the cleavage domain is enzymatically inactive.
[0271] In some embodiments, the guide portion comprises “TALEN” or “TAL effector nuclease” or variants, fragments, or derivatives thereof. TALEN refers to an engineered transcription activator-like effector nuclease, typically comprising a central domain and a cleavage domain of a DNA-binding tandem repeat sequence. TALEN can be generated by fusing the DNA-binding domain of a TAL effector with the DNA-cleavage domain. In some cases, the DNA-binding tandem repeat sequence is 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13, which recognize at least one specific DNA base pair. Transcription activator-like effector (TALE) proteins can be fused with nucleases such as wild-type or mutant Fok1 endonucleases, or the catalytic domain of Fok1. In some embodiments, TALEN is used in the targeting portion of this disclosure to bind polynucleotides (e.g., target genes or target gene regulatory sequences), but the TALEN does not cleave or substantially does not cleave polynucleotides, e.g., nuclease-death TALEN. TALEN or variants, fragments, or derivatives thereof can be fused or bound to one or more heterologous gene effectors to form the complexes of this disclosure.
[0272] In some embodiments, TALEN is engineered to reduce nuclease activity. In some embodiments, the nuclease domain of TALEN comprises a modified form of the wild-type nuclease domain. The modification of the nuclease domain may comprise amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is enzymatically inactive. TALEN or variants, fragments, or derivatives thereof may be fused or bound to one or more heterologous gene effectors to form the complexes disclosed herein.
[0273] Several mutations have been performed on Fok1 for use in TALENs, for example, to improve cleavage specificity or activity. Such TALENs can be engineered to bind to any desired DNA sequence. TALENs can be used to generate genetic modifications (e.g., nucleic acid sequence editing) by inducing double-strand breaks in a target DNA sequence, which then undergoes NHEJ or HDR.
[0274] TALE or its variants, fragments, or derivatives can be fused or bound to one or more heterologous gene effectors to form the complexes disclosed herein. In some embodiments, a transcription activator-like effector (TALE) protein is fused to a heterologous gene effector and does not contain a nuclease. In some embodiments, TALEN does not cleave or substantially does not cleave polynucleotides; for example, a nuclease kills TALE. TALE or its variants, fragments, or derivatives can be fused or bound to one or more heterologous gene effectors to form the complexes disclosed herein.
[0275] In some embodiments, the complex of a transcription activator-like effector (TALE) protein with a heterologous gene effector is designed to function as a transcription activator. In other embodiments, the complex of a transcription activator-like effector (TALE) protein with a heterologous gene effector is designed to function as a transcription repressor. For example, the DNA-binding domain of the TALE protein can be fused (e.g., linked) with one or more heterologous gene effectors containing a transcription activation domain, or fused (e.g., linked) with one or more heterologous gene effectors containing a transcription repression domain.
[0276] In some embodiments, the guide portion comprises a broad-spectrum nuclease. A broad-spectrum nuclease generally refers to a rare cleaving endonuclease or homing endonuclease that can be highly sequence-specific. A broad-spectrum nuclease can recognize DNA target sites of at least 12 base pairs in length, for example, 12 to 40 base pairs, 12 to 50 base pairs, or 12 to 60 base pairs in length. A broad-spectrum nuclease can be a modular DNA-binding nuclease, such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA-binding domain of a specified nucleic acid target sequence or protein. The DNA-binding domain may contain at least one motif recognizing single-stranded or double-stranded DNA. A broad-spectrum nuclease with nuclease activity can produce double-strand breaks. In some embodiments, a broad-spectrum nuclease is used in the targeting portion of this disclosure to bind polynucleotides (e.g., target genes or target gene regulatory sequences), but the broad-spectrum nuclease does not cleave or substantially does not cleave polynucleotides; for example, the nuclease-killing broad-spectrum nuclease. A wide range of nucleases or their variants, fragments or derivatives can be fused or bound to one or more heterologous gene effectors to form complexes disclosed herein.
[0277] The macronuclease can be a monomer or a dimer. In some embodiments, the macronuclease is naturally occurring (found in nature) or wild-type, while in others it is non-natural, artificial, engineered, synthetic, rationally designed, or man-made. In some embodiments, the macronucleases of this disclosure include I-CreI macronuclease, I-CeuI macronuclease, I-Msol macronuclease, or I-SceI macronuclease, variants thereof, derivatives thereof, or functional fragments thereof.
[0278] In some embodiments, the nuclease domain of the macronuclease comprises a modified form of the wild-type nuclease domain. The modification of the nuclease domain may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce or eliminate the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is enzymatically inactive. In some embodiments, the macronuclease can bind DNA but cannot cleave DNA. In some embodiments, a nuclease-inactivated macronuclease is fused or bound to one or more heterologous gene effectors to produce the complex of this disclosure.
[0279] In some embodiments, the guide portion can regulate the expression and / or activity of a target gene (e.g., a target endogenous gene). In some embodiments, the guide portion can edit the sequence of nucleic acids (e.g., genes and / or gene products). Cas proteins with nuclease activity can edit nucleic acid sequences by creating double-strand breaks or single-strand breaks in target polynucleotides.
[0280] In some embodiments, the guide portion containing a nuclease can induce double-strand breaks in target polynucleotides, such as DNA. Double-strand breaks in DNA can trigger DNA break repair, which allows for the introduction of gene modifications (e.g., nucleic acid editing). In some embodiments, the nuclease induces site-specific single-strand DNA breaks or gaps, thereby causing HDR.
[0281] Double-strand breaks in DNA can trigger DNA break repair, which allows for the introduction of gene modifications (e.g., nucleic acid editing). DNA break repair can occur through non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor DNA repair template or template polynucleotide containing flanking homologous arms of the target DNA site can be provided.
[0282] In some implementations, the guide portion or complex containing the nuclease does not produce double-strand breaks in the target polynucleotide (such as DNA).
[0283] In some respects, this document discloses one or more complexes or systems comprising heterologous polypeptides and heterologous polynucleotides. In some cases, the complex may comprise a heterologous gene effector and a guide motif, such as a guide nucleic acid and / or a nuclease (e.g., a nuclease lacking or substantially lacking cleavage activity).
[0284] The complex disclosed herein can be used, for example, to bring one or more heterologous gene effectors into close proximity to a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence, thereby promoting the regulation of the expression, epigenetic modification, or activity level of the target gene.
[0285] In some embodiments, the complexes of this disclosure bind to DNA (e.g., genomic DNA). In some embodiments, the complexes of this disclosure bind to RNA (e.g., mRNA, microRNA, siRNA, or non-coding RNA). In some embodiments, the complexes of this disclosure bind to both DNA and RNA.
[0286] In some implementations, the complex can regulate (e.g., increase or decrease) the expression and / or activity of a target gene (e.g., target endogenous gene) by physically blocking polynucleotide sequences (e.g., promoters, enhancers, repressors, operons, or silencers, insulators, cis-regulatory elements, trans-regulatory elements, epigenetic modification (e.g., DNA methylation) sites, coding sequences).
[0287] In some implementations, the complex can regulate (e.g., increase or decrease) the expression and / or activity of the target gene (e.g., the target endogenous gene) by recruiting additional factors that effectively inhibit or enhance the expression of the target gene.
[0288] In some embodiments, the complexes of this disclosure are used to introduce epigenetic modifications into a target gene (e.g., a target endogenous gene) or a target gene regulatory sequence (e.g., a promoter, enhancer, silencer, insulator, cis-regulatory element, trans-regulatory element, or epigenetic modification (e.g., DNA methylation) site). In some embodiments, the complexes of this disclosure are used to generate a three-dimensional structure, topological association domain, or genome boundary containing a target gene or a target gene regulatory sequence (e.g., a gene distal or proximal to the target gene).
[0289] In some embodiments, the complex or system comprises a heterologous gene effector and a guide portion. In some embodiments, the complex or system comprises one heterologous gene effector and one guide portion. In some embodiments, the complex or system comprises two heterologous gene effectors and one guide portion. In some embodiments, the complex or system comprises three or more heterologous gene effectors and one guide portion.
[0290] In some embodiments, the complex or system comprises a heterologous gene effector and a guide nucleic acid. In some embodiments, the complex or system comprises one heterologous gene effector and one guide nucleic acid. In some embodiments, the complex or system comprises two heterologous gene effectors and one guide nucleic acid. In some embodiments, the complex or system comprises three or more heterologous gene effectors and one guide nucleic acid.
[0291] The two components present in the complex or system may be covalently linked (e.g., in a fusion protein), cross-linked (e.g., treated with a cross-linking agent), or linked via peptide or non-peptide linkers disclosed herein.
[0292] In some implementations, the two components present in the complex or system are part of the same fusion protein. The components may optionally be linked by a linker, such as a peptide linker or a non-peptide linker.
[0293] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is linked to a heterologous gene effector via a adapter. In some embodiments, the guide portion or a portion thereof is further linked to a second heterologous gene effector via the same or a different second adapter. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is fused to the heterologous gene effector without an adapter.
[0294] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is linked to an oligomerization domain or a dimerization (e.g., heterodimerization) domain via a linker. In some embodiments, the guide portion or a portion thereof is further linked to a second oligomerization domain or a dimerization (e.g., heterodimerization) domain via the same or different second linkers. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is fused to the second oligomerization domain or a dimerization (e.g., heterodimerization) domain without a linker.
[0295] In some embodiments, the heterogene effector is linked to a second heterogene effector via a adapter. In some embodiments, the heterogene effector is further linked to a third heterogene effector via the same or different second adapters. In some embodiments, the heterogene effector is fused to the second heterogene effector without an adapter.
[0296] In some embodiments, the heterogeneous effector is linked to an oligomerization domain or a dimerization (e.g., heterodimerization) domain via a linker. In some embodiments, the heterogeneous effector is further linked to a second oligomerization domain or a dimerization (e.g., heterodimerization) domain via the same or different second linkers. In some embodiments, the heterogeneous effector is fused to the second oligomerization domain or a dimerization (e.g., heterodimerization) domain without a linker.
[0297] Any suitable linker can be used. Flexible linkers can have sequences containing segments of glycine and serine residues. The small size of glycine and serine residues provides flexibility and allows for the mobility of the linked functional domains. The incorporation of serine or threonine can maintain the stability of the linker in aqueous solution by forming hydrogen bonds with water molecules, thereby reducing unfavorable interactions between the linker and protein moieties. Flexible linkers can also contain additional amino acids (such as threonine and alanine) to maintain flexibility, and polar amino acids (such as lysine and glutamine) to improve solubility. Rigid linkers can have, for example, an α-helical structure. α-helical rigid linkers can serve as spacers between protein domains.
[0298] The length of the linker sequence can be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 amino acid residues.
[0299] In some embodiments, the linker has at least 1, at least 2, at least 3, at least 5, at least 7, at least 9, at least 11, at least 13, at least 15, or at least 20 amino acids. In some embodiments, the linker has at most 5, at most 7, at most 9, at most 11, at most 13, at most 15, at most 20, at most 25, at most 30, at most 40, or at most 50 amino acids.
[0300] In some embodiments, non-peptide linkers are used. Non-peptide linkers can be, for example, chemical linkers. Two parts of the complex or system disclosed herein can be linked by chemical linkers. The various chemical linkers of this disclosure can be alkylene, alkenylene, ynylene, heteroalkylene, cycloalkylene, heterocycloalkylene, arylene, or heteroarylene, any of which may optionally be substituted. In some embodiments, the chemical linkers of this disclosure can be esters, ethers, amides, thioethers, or polyethylene glycol (PEG). In some embodiments, the linker can reverse the order of the amino acid sequences in the compound, for example, making the amino acid sequences linked by the linker head-to-head rather than head-to-tail. Non-limiting examples of such linkers include dicarboxylic acid diesters, such as oxaloyl diester, malonyl diester, succinyl diester, glutaryl diester, adipyl diester, heptayl diester, fumarate, maleic acid diester, phthaloyl diester, isophthaloyl diester, and terephthaloyl diester. Non-limiting examples of such connectors include dicarboxylic acid diamides, such as oxaloyldiamide, malonyldiamide, succinyldiamide, glutaryldiamide, adipyldiamide, heptayldiamide, fumaryldiamide, maleyldiamide, phthaloyldiamide, isophthaloyldiamide, and terephthaloyldiamide. Non-limiting examples of such connectors include diamides with a diamino connector, such as ethylenediamine, 1,2-di(methylamino)ethane, 1,3-diaminopropane, 1,3-di(methylamino)propane, 1,4-di(methylamino)butane, 1,5-di(methylamino)pentane, 1,6-di(methylamino)hexane, or piperazine. Non-limiting examples of optional substituents include hydroxyl, mercapto, halogen, amino, nitro, nitroso, cyano, azide, sulfoxide, sulfone, sulfonamide, carboxyl, aldehyde, imino, alkyl, haloalkyl, alkenyl, haloalkenyl, alkynyl, haloalkynyl, alkoxy, aryl, aryloxy, arylalkyl, arylalkoxy, heterocyclic, acyl, acyloxy, carbamate, amide, urea, epoxy, or ester.
[0301] The two components present in a complex or system can be non-covalently coupled, for example through ionic bonds, hydrogen bonds, or interactions mediated by oligomerization or dimerization domains disclosed herein.
[0302] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is linked to a heterogene effector via non-covalent coupling. In some embodiments, the guide portion or a portion thereof is further linked to a second heterogene effector via non-covalent coupling. In some embodiments, the guide portion or a portion thereof is covalently linked to a first heterogene effector (e.g., as a fusion protein, optionally having a linker), and the guide portion or a portion thereof is further linked to a second heterogene effector via non-covalent coupling.
[0303] In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is connected to an oligomerization domain or a dimerization (e.g., heterodimerization) domain via non-covalent coupling. In some embodiments, the guide portion or a portion thereof is further connected to a second oligomerization domain or a dimerization (e.g., heterodimerization) domain via non-covalent coupling. In some embodiments, the guide portion or a portion thereof (e.g., a nuclease such as dCas9) is fused to a first oligomerization domain or a dimerization (e.g., heterodimerization) domain via covalent coupling (e.g., fusion, optionally via a linker) and connected to a second oligomerization domain or a dimerization (e.g., heterodimerization) domain via non-covalent coupling.
[0304] In some embodiments, the first component of the guide portion (e.g., guide nucleic acid) is non-covalently linked to the second component of the guide portion (e.g., nuclease). In some embodiments, the first component of the guide portion (e.g., guide nucleic acid) is covalently linked to the second component of the guide portion (e.g., nuclease).
[0305] Any combination of covalent and non-covalent coupling can be used in the complexes or systems disclosed herein. For example, one or more heterologous gene effectors can be non-covalently fused to the guide portion, and one or more oligomerizing domains can be covalently bound to components of the complexes or systems (e.g., nucleases).
[0306] In some embodiments, a peptide that provides increased or decreased stability is fused to or otherwise bound to a component of the disclosed complex or system (e.g., a guide moiety or a heterologous gene effector). The fused peptide may be located at the N-terminus, C-terminus, or interior of the fusion protein.
[0307] In some embodiments, one or more components of the disclosed complex or system are fused with a domain that guides desired subcellular localization, such as a nuclear localization signal, or a protein that targets the inner nuclear membrane, outer nuclear membrane, Cajal body, nuclear speckle, nuclear pore complex, PML body, nucleolus, P granule, GW body, stress granule, corpus cavernosum, endoplasmic reticulum, mitochondria, etc.
[0308] In some embodiments, the complex or system of this disclosure comprises a first protein linked to a first oligomerization (e.g., dimerization) domain and a second protein linked to a second oligomerization (e.g., dimerization) domain. In some embodiments, the oligomerization or dimerization domain may comprise a peptide-interacting domain, such as using systems utilizing sgRNA2.0, SAM, SunTag, RAB, FLAG-Biotin, or the inducible oligomerization (e.g., dimerization) systems disclosed herein.
[0309] deliver
[0310] One or more genes encoding any of the heterologous polypeptides disclosed herein (e.g., heterologous gene effectors) and any additional molecules operatively coupled thereto (e.g., heterologous polynucleotides, such as one or more guide nucleic acid molecules) may be integrated into the cell's genome, whereby aberrant expression of the target gene will be modified. Alternatively, the one or more genes may not be integrated and need not be integrated into the cell's genome. The one or more genes may be a single gene (e.g., a single expression cassette) or multiple genes (e.g., multiple expression cassettes). The one or more genes may be heterologous to the cell.
[0311] Any heterologous polypeptide (e.g., heterologous gene effector) and any additional molecule operatively coupled thereto (e.g., heterologous polynucleotide, such as one or more guide nucleic acid molecules) can be introduced (e.g., delivered, expressed, etc.) into cells through a variety of methods (e.g., viral and non-viral delivery methods). Viral vector delivery systems can include DNA viruses and RNA viruses, which may have a free or integrated genome after delivery to cells. Non-viral vector delivery systems can include DNA plasmids, RNA (e.g., transcripts of the vectors described herein), naked nucleic acids, and nucleic acids complexed with a delivery medium (e.g., liposomes).
[0312] One or more genes may further include one or more promoters to control system expression. The promoters disclosed herein may be active in eukaryotic cells, mammalian cells, non-human mammalian cells, or human cells. Promoters may be inducible or constitutively active promoters. Alternatively or supplementally, promoters may be tissue- or cell-specific. Non-limiting examples of suitable eukaryotic promoters (i.e., promoters that function in eukaryotic cells) may include: cytomegalovirus (CMV) immediate early promoter, herpes simplex virus (HSV) thymidine kinase promoter, SV40 early and late promoters, long terminal repeat (LTR) promoters of retroviruses, human elongation factor-1 promoter (EF1), hybrid constructs containing a CMV enhancer fused to a chicken β-active promoter (CAG), mouse stem cell virus promoters (MSCV), phosphoglycerate kinase-1 locus promoters (PGK), or mouse metallothionein-I promoters. Promoters may be fungal promoters. Promoters may be plant promoters. Plant promoter databases (e.g., PlantProm) are available. Expression vectors may also contain ribosome binding sites for translation initiation and transcription terminators. Expression vectors may also contain suitable sequences for amplifying expression. In some cases, the promoters disclosed herein may be any tissue-specific promoter provided herein, or any cell type-specific promoter provided herein.
[0313] A single gene provided herein (e.g., encoding the system disclosed herein) may have the following sizes: at least or at most about 2.5 kilobases, at least or at most about 2.6 kilobases, at least or at most about 2.7 kilobases, at least or at most about 2.8 kilobases, at least or at most about 2.9 kilobases, at least or at most about 3.0 kilobases, at least or at most about 3.1 kilobases, at least or at most about 3.2 kilobases, at least or at most about 3.3 kilobases, at least or at most about 3.4 kilobases, at least or at most about 3.5 kilobases, at least or at most about 3.6 kilobases, at least or at most about 3.7 kilobases, at least or at most about 3.8 kilobases, at least or at most about 3.9 kilobases, at least or at most about 4.0 kilobases. The number of bases is at least or at most about 4.1 kilobases, at least or at most about 4.2 kilobases, at least or at most about 4.3 kilobases, at least or at most about 4.4 kilobases, at least or at most about 4.5 kilobases, at least or at most about 4.6 kilobases, at least or at most about 4.7 kilobases, at least or at most about 4.8 kilobases, at least or at most about 4.9 kilobases, at least or at most about 5.0 kilobases, at least or at most about 5.5 kilobases, at least or at most about 6.0 kilobases, at least or at most about 6.5 kilobases, at least or at most about 7.0 kilobases, at least or at most about 7.5 kilobases, at least or at most about 8.0 kilobases, at least or at most about 9.0 kilobases, or at least or at most about 10 kilobases. In some cases, a single gene can have the following sizes: between about 3 kilobases and about 5 kilobases; between about 3 kilobases and about 4.8 kilobases; between about 3 kilobases and about 4.6 kilobases; between about 3 kilobases and about 4.4 kilobases; between about 3 kilobases and about 4.2 kilobases; between about 3 kilobases and about 4.0 kilobases; between about 3 kilobases and about 3.5 kilobases; between about 3.5 kilobases and about 5 kilobases; between about 3.5 kilobases and about 4.8 kilobases; between about 3.5 kilobases and about 4.6 kilobases; between about 3.5 kilobases and about 4.4 kilobases. Between, between about 3.5 kilobases and about 4.2 kilobases, between about 3.5 kilobases and about 4 kilobases, between about 4 kilobases and about 5 kilobases, between about 4 kilobases and about 4.9 kilobases, between about 4 kilobases and about 4.8 kilobases, between about 4 kilobases and about 4.7 kilobases, between about 4 kilobases and about 4.6 kilobases, between about 4 kilobases and about 4.5 kilobases, between about 4 kilobases and about 4.4 kilobases, between about 4 kilobases and about 4.3 kilobases, between about 4 kilobases and about 4.2 kilobases, or between about 4 kilobases and about 4.1 kilobases.
[0314] RNA or DNA virus-based systems can be used to target specific cells and deliver viral payloads to the cell nucleus. Viral vectors can be used to contact cells in vitro, and modified cells can be optionally (ex vivo) administered to a subject (e.g., a human). Alternatively, viral vectors can be administered directly (in vivo) to a subject. Virus-based systems may contain retroviral, lentiviral, adenovirus, adeno-associated virus, or herpes simplex virus vectors for gene transfer. Integration into the host genome can occur using retroviral, lentiviral, or adeno-associated virus gene transfer methods, which can lead to long-term expression of the inserted transgene.
[0315] In some embodiments, the compositions and systems provided herein are delivered to a subject using a viral vector. In some cases, the viral vector is an adeno-associated virus (AAV) vector. The term “AAV” is an abbreviation for adeno-associated virus and can be used to refer to the virus itself or its derivatives. Unless otherwise required, the term covers all serotypes, subtypes, and naturally occurring and recombinant forms. The abbreviation “rAAV” refers to recombinant adeno-associated virus, also known as a recombinant AAV vector (or “rAAV vector”). The term “AAV” includes AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, rh10 and their hybrids, avian AAV, bovine AAV, canine AAV, equine AAV, primate AAV, non-primate AAV, or sheep AAV. The genomic sequences of different serotypes of AAV, as well as the sequences of the natural terminal repeat (TR), Rep protein, and capsid subunit, are known in the art. These sequences can be found in the literature or in public databases such as GenBank. The term "rAAV vector" as used in this article refers to an AAV vector containing a polynucleotide sequence of non-AAV origin (i.e., a polynucleotide heterologous to AAV), typically the sequence of interest for cellular genetic transformation. This heterologous polynucleotide is usually flanked by at least one, and usually two, AAV inverted terminal repeat (ITR) sequences. The term rAAV vector encompasses both rAAV vector particles and rAAV vector plasmids. rAAV vectors can be single-stranded (ssAAV) or self-complementary (scAAV). "AAV virus," "AAV virus particle," or "rAAV vector particle" refers to a viral particle composed of at least one AAV capsid protein and an encapsulated polynucleotide rAAV vector. If the particle contains a heterologous polynucleotide (i.e., a polynucleotide other than the wild-type AAV genome, such as a transgene to be delivered to mammalian cells), it is generally referred to as an "rAAV vector particle" or simply "rAAV vector." Therefore, the production of rAAV particles must include the production of an rAAV vector, as this vector is contained within the rAAV particle. In some cases, AAV vectors are selected based on the tropism of viral vectors. In some embodiments, AAV vectors that are tropism-dependent on the tissue of interest (e.g., AAV2 for muscle tissue) can be used to deliver polynucleotides encoding the compositions and systems provided herein to that tissue.
[0316] RNA or DNA virus-based systems can be used to target specific cells in vivo and deliver viral payloads to the cell nucleus. Viral vectors can be administered directly (in vivo) to subjects, or they can be used to contact cells in vitro, and modified cells can optionally (ex vivo) be administered to subjects (e.g., humans). Virus-based systems can contain retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, or herpes simplex virus vectors for gene transfer. Integration into the host genome can occur using retroviral, lentiviral, or adeno-associated virus gene transfer methods, which can lead to long-term expression of the inserted transgene. High transduction efficiency has been observed in many different cell types and target tissues.
[0317] Retroviral tropism can be altered by incorporating exogenous envelope proteins, thereby expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and produce high viral titers. The choice of retroviral gene transfer system can be dependent on the target tissue. Retroviral vectors can contain cis-acting long terminal repeats (LTRs) with the ability to package exogenous sequences up to 6–10 kb. Minimal cis-acting LTRs are sufficient to replicate and package the vector, which can be used to integrate therapeutic genes into target cells to provide permanent transgenic expression. Retroviral vectors can include vectors based on murine leukemia virus (MuLV), gibberish leukemia virus (GaLV), simultaneous immunodeficiency virus (SIV), or human immunodeficiency virus (HIV), and combinations thereof.
[0318] Adenovirus-based systems can be used. Adenovirus-based systems can induce transient expression of transgenes. Adenovirus-based vectors can exhibit high transduction efficiency in cells and can be used without cell division. High titers and high levels of expression can be obtained using adenovirus-based vectors. Adeno-associated virus (“AAV”) vectors can be used for transducing cells with target nucleic acids (e.g., in the in vitro production of nucleic acids and peptides), as well as for in vivo and in vitro gene therapy procedures.
[0319] Packaging cells can be used to form viral particles capable of infecting host cells. These cells may include 293 cells (e.g., for packaging adenoviruses), or Psi2 cells, PA317 cells (e.g., for packaging retroviruses). Viral vectors can be generated by producing cell lines that package nucleic acid vectors into viral particles. The vector may contain the minimum viral sequence required for packaging and subsequent integration into the host. The vector may contain additional viral sequences replaced by expression cassettes of the polynucleotides to be expressed. Missing viral functions can be provided trans-present by the packaging cell line. For example, an AAV vector may contain an ITR sequence from the AAV genome, which is essential for packaging and integration into the host genome. Viral DNA can be packaged in cell lines that may contain helper plasmids encoding other AAV genes (i.e., rep and cap) but lacking the ITR sequence. This cell line can also be infected using adenovirus as a helper. The helper virus can promote the replication of the AAV vector and the expression of AAV genes from the helper plasmid. Adenovirus contamination can be reduced, for example, by heat treatment (adenoviruses are more sensitive to heat treatment than AAVs).
[0320] Host cells can be transiently or non-transiently transfected using one or more vectors described herein. Cells can be transfected while they are naturally present in the subject. Cells can be taken from or derived from the subject and transfected. Cells can be derived from cells taken from the subject (e.g., cell lines). In some embodiments, cells transfected with one or more vectors described herein are used to establish new cell lines containing one or more vector-derived sequences. In some embodiments, transient transfection with the compositions disclosed herein (e.g., transient transfection via one or more vectors, or transfection with RNA) is performed... Device Cells with partial (e.g., CRISPR complex) activity modifications are used to establish new cell lines containing cells that contain the modifications but lack any other exogenous sequences.
[0321] Any suitable vector compatible with the host cell can be used with the methods disclosed herein. Non-limiting examples of vectors for use with eukaryotic host cells include pXT1, pSG5 (Stratagene), etc. TM ), pSVK3, pBPV, pMSG or pSVLSV40 (Pharmacia) TM ).
[0322] In some embodiments, additional components of the compositions disclosed herein may include excipients. Non-limiting examples of excipients may include solvents, dispersion media, diluents or other liquid solvents, dispersing or suspending agents, surfactants, isotonic agents, thickeners or emulsifiers, preservatives, lipids, liposomes, lipid nanoparticles, polymers, lipoplexes, core-shell nanoparticles, peptides, proteins, hyaluronidase, nanoparticle mimics, inactive diluents, buffers, lubricants or oils, and combinations thereof. In some instances, the compositions disclosed herein may contain one or more excipients, each in an amount that together increase (i) the heterologous polypeptide or a heterologous gene encoding the heterologous polypeptide and / or (ii) the stability of cells or modified cells.
[0323] In some aspects, this disclosure provides kits and instructions for use, the kits comprising such compositions, the instructions for use instructing (i) to contact cells with the composition (e.g., in vitro, ex vivo, or in vivo), or (ii) to administer cells comprising any of the compositions disclosed herein to a subject. The subject may have or may be suspected of having a condition, such as a hereditary disease.
[0324] In some embodiments, any of the compositions disclosed herein may be administered to a subject orally, intraperitoneally, intravenously, intraarterially, percutaneously, intramuscularly, in liposome form, locally via catheter or stent, subcutaneously, intrafacially, or intrathecally. In a particular aspect, the compositions and systems provided herein, including polynucleotides encoding said compositions and systems, such as those contained in an AAV carrier, may be administered to a subject via intravenous administration.
[0325] Non-limiting examples of viral vectors that can be used to deliver heterologous peptides and / or heterologous polynucleotides (or genes encoding one or more of them) may include, but are not limited to, retroviral vectors, lentiviral vectors, adenovirus vectors, poxvirus vectors, herpesvirus vectors, or adeno-associated virus (AAV) vectors. Non-limiting examples of AAV vectors may include AAV1, AAV10, AAV106.1 / hu.37, AAV11, AAV114.3 / hu.40, AAV12, AAV127.2 / hu.41, AAV127.5 / hu.42, AAV128.1 / hu.43, AAV128.3 / hu.44, AAV130.4 / hu.48, AAV145.1 / hu.53, AAV145.5 / hu.54, AAV145.6 / hu.55, AAV16.12 / hu.11, AAV16.3, and AAV16.8 / hu. 10. AAV161.10 / hu.60, AAV161.6 / hu.61, AAV1-7 / rh.48, AAV1-8 / rh.49, AAV2, AAV2.5T, AAV2-15 / rh.62, AAV223.1, AAV223.2, AAV223. 4. AAV223.5, AAV223.6, AAV223.7, AAV2-3 / rh.61, AAV24.1, AAV2-4 / rh.50, AAV2-5 / rh.51, AAV27.3, AAV29.3 / bb.1, AAV29.5 / bb.2, AA V2G9, AAV-2-pre-miRNA-101, AAV3, AAV3.1 / hu.6, AAV3.1 / hu.9, AAV3-11 / rh.53, AAV3-3, AAV33.12 / hu.17, AAV33.4 / hu.15, AAV33.8 / hu.16, AAV3-9 / rh.52, AAV3a, AAV3b, AAV4, AAV4-19 / rh.55, AAV42.12, AAV42-10, AAV42-11, AAV42-12, AAV42-13, AAV42-15, AAV42-1 b. AAV42-2, AAV42-3a, AAV42-3b, AAV42-4, AAV42-5a, AAV42-5b, AAV42-6b, AAV42-8, AAV42-aa, AAV43-1, AAV43-12, AAV43-20, AAV43- 21. AAV43-23, AAV43-25, AAV43-5, AAV4-4, AAV44.1, AAV44.2, AAV44.5, AAV46.2 / hu.28, AAV46.6 / hu.29, AAV4-8 / r11.64, AAV4-8 / rh.64、AAV4-9 / rh.54、AAV5、AAV52.1 / hu.20、AAV52 / hu.19、AAV5-22 / rh.58、AAV5-3 / rh.57、AAV54.1 / hu.21、AAV54.2 / hu.22、AAV54.4R / hu.27、AAV54.5 / hu.23、AAV54.7 / hu.24、AAV58.2 / hu.25、AAV6、AAV6.1、AAV6.1.2、AAV6.2、AAV7、AAV7.2、AAV7.3 / hu.7、AAV8、AAV-8b、AAV-8h、AAV9、AAV9.11、AAV9. 13、AAV9.16、AAV9.24、AAV9.45、AAV9.47、AAV9.61、AAV9.68、AAV9.84、AAV9.9、AAVA3.3、AAVA3.4、AAVA3.5、AAVA3.7、AAV-b、AAVC1、AAVC2、AAVC5 VCh.5、AAVCh.5R1、AAVcy.2、AAVcy.3、AAVcy.4、AAVcy.5、AAVCy.5R1、AAVCy.5R2、AAVCy.5R3、AAVCy.5R4、AAVcy.6、AAV-DJ、AAV-DJ8、AAVF3、AAVF5、AAVcy. V-h、AAVH-1 / hu.1、AAVH2、AAVH-5 / hu.3、AAVH6、AAVhE1.1、AAVhER1.14、AAVhEr1.16、AAVhEr1.18、AAVhER1.23、AAVhEr1.35、AAVhEr1.36、AAVhEr1.5 、AAVhEr1.7、AAVhEr1.8、AAVhEr2.16、AAVhEr2.29、AAVhEr2.30、AAVhEr2.31、AAVhEr2.36、AAVhEr2.4、AAVhEr3.1、AAVhu.1、AAVhu.10、AAVhu.11、AAVhu. hu.11、AAVhu.12、AAVhu.13、AAVhu.14 / 9、AAVhu.15、AAVhu.16、AAVhu.17、AAVhu.18、AAVhu.19、AAVhu.2、AAVhu.20、AAVhu.21、AAVhu.22、AAVhu.23. 2、AAVhu.24、AAVhu.25、AAVhu.27、AAVhu.28、AAVhu.29、AAVhu.29R、AAVhu.3、AAVhu.31、AAVhu.32、AAVhu.34、AAVhu.35、AAVhu.37、AAVhu.39、AAVhu.4、AAVhu.40、AAVhu.41、AAVhu.42、AAVhu.43、AAVhu.44、AAVhu.44R1、AAVhu.44R2、AAVhu.44R3、AAVhu.45、AAVhu.46、AAVhu.47、AAVhu.48、AAVhu.48 R1、AAVhu.48R2、AAVhu.48R3、AAVhu.49、AAVhu.5、AAVhu.51、AAVhu.52、AAVhu.53、AAVhu.54、AAVhu.55、AAVhu.56、AAVhu.57、AAVhu.58、AAVhu.6、AAVhu. hu.60、AAVhu.61、AAVhu.63、AAVhu.64、AAVhu.66、AAVhu.67、AAVhu.7、AAVhu.8、AAVhu.9、AAVhu.t19、AAVLG-10 / rh.40、AAVLG-4 / rh.38、AAVLG-9 / hu. 39、AAVLG-9 / hu.39、AAV-LKO1、AAV-LK02、AAVLK03、AAV-LK03、AAV-LK04、AAV-LK05、AAV-LKO6、AAV-LK07、AAV-LK08、AAV-LK09、AAV-LK10、AAV-LK11 AV-LK12、AAV-LK13、AAV-LK14、AAV-LK15、AAV-LK17、AAV-LK18、AAV-LK19、AAVN721-8 / rh.43、AAV-PAEC、AAV-PAEC11、AAV-PAEC12、AAV-PAEC2、AAV-PAEC EC4、AAV-PAEC6、AAV-PAEC7、AAV-PAEC8、AAVpi.1、AAVpi.2、AAVpi.3、AAVrh.10、AAVrh.12、AAVrh.13、AAVrh.13R、AAVrh.14、AAVrh.17、AAVrh.18、AAV rh.19、AAVrh.2、AAVrh.20、AAVrh.21、AAVrh.22、AAVrh.23、AAVrh.24、AAVrh.25、AAVrh.2R、AAVrh.31、AAVrh.32、AAVrh.33、AAVrh.34、AAVrh.35、AAVrh. rh.36、AAVrh.37、AAVrh.37R2、AAVrh.38、AAVrh.39、AAVrh.40、AAVrh.43、AAVrh.44、AAVrh.45、AAVrh.46、AAVrh.47、AAVrh.48、AAVrh.48、AAVrh.48.1. AAVrh.48.1.2, AAVrh.48.2, AAVrh.49, AAVrh.50, AAVrh.51, AAVrh.52, AAVrh.53, AAVrh.54, AAVrh.55, AAVrh.56, AAVrh.57, AAVrh.58, AAVrh.59, AAVrh.60, AAVrh.61, AAVrh.62, AAVrh.64, AAVrh.64R1, AAVrh.64R2, AAVrh.65, AAVrh.67, AAVrh.68, AAVrh .69, AAVrh.70, AAVrh.72, AAVrh.73, AAVrh.74, AAVrh.8, AAVrh.8R, AAVrh8R, AAVrh8R A586R mutant, AAVrh8R R533A mutant, BAAV, BNP61 AAV, BNP62 AAV, BNP63 AAV, bovine AAV, goat AAV, Japanese AAV10, true AAV (ttAAV), UPENN AAV10, AAV-LK16, AAAV, AAV Shuffle 100-1, AAV Shuffle 100-2, AAV Shuffle 100-3, AAV Shuffle 100-7, AAV Shuffle 10-2, AAV Shuffle 10-6, AAV Shuffle 10-8, AAV SM 100-10, AAV SM 100-3, AAV SM 10-1, AAV SM 10-2 or AAV SM 10-8. For example, AAVrh.74 can be used as a viral vector to deliver polynucleotide sequences encoding heterologous polypeptides and heterologous polynucleotides (such as Cas protein-gene effector fusions and one or more guide nucleic acid molecules).
[0326] Non-viral delivery methods for nucleic acids may include lipid transfection, nuclear transfection, microinjection, gene gun, virions, liposomes, immunoliposomes, polycationic or lipid-nucleic acid conjugates, lipid nanoparticles (LNPs), naked DNA, artificial viral particles, or reagent-enhanced DNA uptake. Highly efficient receptors suitable for polynucleotides can be used to recognize both cationic and neutral lipids transfected with lipids.
[0327] Any of the compositions disclosed herein (or one or more genes encoding any part of the composition), such as heterologous gene effectors and / or guide nucleic acid molecules, can be administered via any suitable route of administration, including but not limited to parenteral (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intraventricular, intra-articular, intraperitoneal, or intracranial), intranasal, buccal, sublingual, oral, or rectal routes. In some cases, the pharmaceutical compositions are formulated for parenteral (e.g., intravenous, intratumoral, subcutaneous, intramuscular, intracerebral, intraventricular, intra-articular, intraperitoneal, or intracranial) administration.
[0328] The compositions disclosed herein (e.g., pharmaceutical compositions) are suitable for administration to humans. Furthermore, these compositions are suitable for administration to any other animal, such as non-human animals (e.g., non-human mammals). It is well understood that pharmaceutical compositions suitable for administration to humans may be modified to make them suitable for administration to a variety of animals, and such modifications can generally be designed and / or performed by a skilled veterinary pharmacologist using only routine experiments (if any). Subjects intended to administer the pharmaceutical compositions include, but are not limited to, humans and / or other primates; mammals (including commercially relevant mammals such as cattle, pigs, horses, sheep, cats, dogs, mice, and / or rats); and / or birds (including commercially relevant birds such as poultry, chickens, ducks, geese, and / or turkeys).
[0329] target genes
[0330] This disclosure provides compositions, methods, and systems for regulating the expression of target genes (e.g., target endogenous genes). For example, this document discloses complexes or systems comprising a guide portion and one or more heterologous gene effectors that can increase or decrease the activity or expression level of a target gene.
[0331] In some implementations, the target gene or its regulatory sequence is endogenous to the subject, for example, present in the subject's genome. In some implementations, the target gene or its regulatory sequence is not part of an engineered reporter subsystem.
[0332] In some embodiments, the target gene is exogenous to the host subject, such as a pathogen target gene or an exogenous gene expressed as a result of a therapeutic intervention (such as gene therapy and / or cell therapy). In some embodiments, the target gene is an exogenous reporter gene. In some embodiments, the target gene is an exogenous synthetic gene.
[0333] In some embodiments, the target gene (e.g., a target endogenous gene) is a gene that is overexpressed or underexpressed in a disease or condition. In some embodiments, the target gene is a gene that is overexpressed or underexpressed in a hereditary genetic disease.
[0334] In some embodiments, the target gene (e.g., a target endogenous gene) is a gene that is overexpressed or underexpressed in autoimmune diseases. In some embodiments, the target gene is a gene that is overexpressed or underexpressed in the following: acute disseminated encephalomyelitis, acute motor axonal neuropathy, Addison's disease, painful obesity, adult-onset Still's disease, alopecia areata, ankylosing spondylitis, antiglomerular basement membrane nephritis, antineutrophil cytoplasmic antibody-associated vasculitis, anti-N-methyl-D-aspartate receptor encephalitis, antiphospholipid syndrome, antisyntheticase syndrome, aplastic anemia, autoimmune angioedema, autoimmune encephalitis, autoimmune enteropathy, autoimmune hemolytic anemia, autoimmune hepatitis, autoimmune inner otopathy, autoimmune lymphoproliferative syndrome, autoimmune neutropenia, and other autoimmune diseases. Autoimmune oophoritis, autoimmune orchitis, autoimmune pancreatitis, autoimmune polyendocrine syndrome, type 2 autoimmune polyendocrine syndrome, type 3 autoimmune polyendocrine syndrome, autoimmune progesterone dermatitis, autoimmune retinopathy, autoimmune thrombocytopenic purpura, autoimmune thyroiditis, autoimmune urticaria, autoimmune uveitis, Balo concentric sclerosis, Behçet's disease, Bickerstaff encephalitis, bullous pemphigoid, celiac disease, chronic fatigue syndrome, chronic inflammatory demyelinating polyneuropathy, Churg-Strauss syndrome, cicatricial pemphigoid, Cogan syndrome, cryotherapy Graves' disease, complicated regional pain syndrome, CREST syndrome, Crohn's disease, herpetic dermatitis, dermatomyositis, type 1 diabetes, discoid lupus erythematosus, endometriosis, enthesitis, enthesitis-associated arthritis, eosinophilic esophagitis, eosinophilic fasciitis, acquired epidermolysis bullosa, erythema nodosum, primary mixed cryoglobulinemia, Evans syndrome, Felty syndrome, fibromyalgia, gastritis, pemphigoid of pregnancy, giant cell arteritis, Goodpasture syndrome, Graves' disease, Graves' ophthalmopathy, Guillain-Barré syndrome, Hashimoto's encephalopathy, Hashimoto's thyroiditis, allergic purpura, purulent sweating Adenitis, idiopathic dilated cardiomyopathy, idiopathic inflammatory demyelinating disease, IgA nephropathy, IgG4-related systemic diseases, inclusion body myositis, inflammatory bowel disease (IBD), intermediate uveitis, interstitial cystitis, juvenile arthritis, Kawasaki disease, Lambert-Eaton myasthenic syndrome, leukocytoclastic vasculitis, lichen planus, lichen sclerosus, woody conjunctivitis, linear IgA disease, lupus nephritis, lupus vasculitis, Lyme disease, Ménière's disease, microscopic colitis, microscopic polyangiitis, mixed connective tissue disease, Mooren's ulcer, morphine scleroderma, Mucha-Habermann disease, multiple sclerosis, myasthenia gravis, myocarditis.Myositis, neuromyelitis optica, neurotic myotonia, strabismus oculoclonus myoclonus syndrome, optic neuritis, Ord's thyroiditis, relapsing rheumatism, paraneoplastic cerebellar degeneration, Parry-Romberg syndrome, Parsonage-Turner syndrome, streptococcal-associated childhood autoimmune neuropsychiatric disorders, pemphigus vulgaris, pernicious anemia, acute pityriasis lichenoides, POEMS syndrome, polyarteritis nodosa, polymyalgia rheumatica, polymyositis, post-myocardial infarction syndrome, post-pericardiotomy syndrome, primary biliary cirrhosis, primary immunodeficiency disease, primary sclerosing cholangitis, progressive inflammatory neuropathy, psoriasis, psoriatic arthritis, pure red cell aplasia, pyoderma gangrenosa, Raynaud's phenomenon, reaction Rheumatoid arthritis, relapsing polychondritis, restless legs syndrome, retroperitoneal fibrosis, rheumatic fever, rheumatoid arthritis, rheumatoid vasculitis, sarcoidosis, Schnitzler syndrome, scleroderma, Sjögren's syndrome, stiff-person syndrome, subacute bacterial endocarditis, Susac's syndrome, Sydenham's chorea, sympathetic ophthalmia, systemic lupus erythematosus, systemic scleroderma, thrombocytopenia, Tolosa-Hunt syndrome, transverse myelitis, ulcerative colitis, undifferentiated connective tissue disease, urticaria, urticarial vasculitis, vasculitis, or vitiligo.
[0335] In some implementations, the target gene (e.g., the target endogenous gene) is a gene that is overexpressed or underexpressed in cancers, such as acute leukemia, astrocytoma, cholangiocarcinoma, bone cancer, breast cancer, brainstem glioma, bronchioloalveolar cell lung cancer, adrenal cancer, anal region cancer, bladder cancer, endocrine system cancer, esophageal cancer, head and neck cancer, kidney cancer, parathyroid cancer, penile cancer, pleural / peritoneal cancer, salivary gland cancer, small bowel cancer, thyroid cancer, ureteral cancer, cervical cancer, endometrial cancer, fallopian tube cancer, renal pelvis cancer, vaginal cancer, vulvar cancer, cervical cancer, chronic leukemia, colon cancer, colorectal cancer, melanoma, ependymoma, epidermoid tumor, Ewing's sarcoma, gastric cancer, glioblastoma, etc. Glioblastoma multiforme, glioma, hematologic malignancies, hepatocellular carcinoma, Hodgkin's disease, intraocular melanoma, Kaposi's sarcoma, lung cancer, lymphoma, medulloblastoma, melanoma, meningioma, mesothelioma, multiple myeloma, muscle cancer, central nervous system (CNS) vegetations, neuronal carcinoma, small cell lung cancer, non-small cell lung cancer, osteosarcoma, ovarian cancer, pancreatic cancer, pediatric malignancies, Schwann cell carcinoma, prostate cancer, rectal cancer, renal cell carcinoma, soft tissue sarcoma, schwannoma, skin cancer, spinal axis tumors, squamous cell carcinoma, gastric cancer, synovial sarcoma, testicular cancer, uterine cancer, or tumors and their metastases, including refractory forms of any of the above cancers, or combinations thereof.
[0336] In some implementations, the target genes (e.g., target endogenous genes) are differentiation-related genes, such as SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30, CD50, AHR, Aiolos / IKZF3, CDX4, CREB, DNMT3A, DNMT3B, EGR1, FoxO3, GATA-1, GATA-2, GATA-3, H elios, HES-1, HHEX, HIF-1α / HIF1A, HMGB1 / HMG-1, HMGB3, Ikaros, c-Jun, LMO2, LMO4, c-Maf, MafB, MEF2C, MYB, c-Myc, NFATC2, NFIL3 / E4BP4, Nrf2, p53, PITX2, PRDM16 / MEL1, Prox1, PU.1 / Spi-1, RUNX1 / CBFA2, SALL4, SCL / Tal1, Smad2, Smad2 / 3, Smad4, Smad7, Spi-B, STAT activator, STAT inhibitor, STAT3, STAT4 STAT5a, STAT6, TSC22, DUX4, DUX4 / DUX4c, DUX4c, EBF-1, EBF-2, EBF-3, ETV5, FoxC2, FoxF1, GATA-4, GATA-6, HMGA2, c-Jun, MYF-5, Cardiomyin, MyoD, Myopoietin, NFATC2, p53, Pax3, PDX-1 / IPF1, PLZF, PRDM16 / MEL1, RUNX2 / CBFA1, Smad1, Smad3, Smad4, Smad5, Smad8, Smad9, Snail, SOX2, SOX9, SOX11, STAT activator, STA T-inhibitor, STAT1, STAT3, TBX18, Twist-1, Twist-2, Brachyury, EOMES, FoxC2, FoxD3, FoxF1, FoxH1, FoxO1 / FKHR, GATA-2, GATA-3, GBX2, Goosecoid, HES-1, HNF-3α / FoxA1, c-Jun, KLF2, KLF4, KLF5, c-Maf, Max, MEF2C, MIXL1, MTF2, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NFkB2, Oct-3 / 4, Otx2p53, Pax2, Pax6, PRDM14, Rex-1 / ZFP42, SALL1, SALL4, Smad1, Smad2, Smad2 / 3, Smad3, Smad4, Smad5, Smad8, Snail, SOX2, SOX7, SOX15, SOX17, STAT activator, STAT inhibitor, STAT3, SUZ12, TBX6, TCF-3 / E2A, THAP11, UTF1, WDR5, WT1, ZNF206, ZNF281, KLF2, KLF4, c-Maf, c-Myc, Nanog, Oct-3 / 4, p53, SOX1, SOX2, SOX3, SOX15, SOX18, TBX18, ASCL2 / Mash2, CDX2, DNMT1, ELF3, Ets-1, FoxM1, FoxN1, GATA-6, Hairless, HNF-4α / NR2A1, IRF6, c-Maf, MITF, Miz-1 / ZBTB17, MSX1, MSX2, MYB, c-Myc, Neurogenin-3, NFATC1, NKX3.1, Nrf2, p53, p63 / TP73L, Pax2, Pax3, RUNX1 / CBFA2 RUNX2 / CBFA1, RUNX3 / CBFA3, Smad1, Smad2, Smad2 / 3, Smad4, Smad5, Smad7, Smad8, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, SUZ12, TCF-3 / E2A, TCF7 / TCF1, androgen R / NR3C4, AP-2γ, β-catenin, β-catenin inhibitor, Brachyury, CREB, ERα / NR3A1, ERβ / NR3A2, FoxM1, FoxO3, FRA-1, GLI-1, G LI-2, GLI-3, HIF-1α / HIF1A, HIF-2α / EPAS1, HMGA1B, c-Jun, JunB, KLF4, c-Maf, MCM2, MCM7, MITF, c-Myc, Nanog, NFkB / IkB activator, NFkB / IkB inhibitor, NFkB1, NKX3.1, Oct-3 / 4, p53, PRDM14, Snail, SOX2, SOX9, STAT activator, STAT inhibitor, STAT3, TAZ / WWTR1, TBX3, Twist-1, Twist-2, WT1 or ZEB1.
[0337] In some implementations, regulation of the expression level and / or epigenetic level (e.g., methylation level) of the target gene in the target cell (e.g., muscle cell) can achieve modification (e.g., upregulation or downregulation) of downstream genes (e.g., one or more downstream genes) of the target gene. In some cases, the target gene may be encoded by a D4Z4 repeat array (e.g., the target gene is DUX4), and the downstream genes whose expression is subsequently modified (e.g., downregulated) may include, but are not limited to, ZSCAN4, LEUTX, MBD3L2, TRIM48, TRIM43, DEFB103, ZFN217, RNASEL, EIF2AK2, BMP2, SP1, P21, MYC, MURF1, ATROOGIN1, CRYM, PRAMEF1, RFPL2, KHDC1, SPRYD5, TPRX1, HSPA2, FGFR3, SLC2A14, ID2, PVRL3, SFRS2B, THOC4, ZNHIT6, DBR1, TFIP11, FBXO33, USP29, TRIM23, SLC34A2, CSAG3, and / or PNMA6B.
[0338] In some implementations, regulation of the expression level and / or epigenetic level (e.g., methylation level) of a target gene in target cells can induce apoptosis in target cells (e.g., muscle cells). In some cases, such regulation of the target gene can reduce stress in the target cells. For example, regulation of a target gene (e.g., DUX4) can induce downregulation of one or more stress-related markers in the target cells. Non-limiting examples of one or more stress-related markers may include ACTH, glucocorticoid receptor, CRHR-1 / 2, POMC, prolactin, arginine vasopressin receptor V1a, superoxide dismutase 1, superoxide dismutase 2, thioredoxin peroxidase-3, CCR5, iNOS, eNOS, heme oxygenase 2, cyclooxygenase 2, HSP27, HSP40, HSP60, HSP70, HSP70i, HSP90, HSP110, GRP78 / BIP, AIF, annexin II, annexin IV, caspase 1, caspase 2, caspase 3, caspase 6, cytokeratin, E-cadherin and / or annexin V, caspase 5, caspase 7, caspase 8, caspase 9, caspase 10, BAD, BAX, BAK, BCL2, BID, PARP-1, NOXA, PUMA, RIPK3, RIPK1, FADD, APAF1, DFF40, DFF45, or ROCK. One or more stress-related biomarkers disclosed herein may be apoptosis biomarkers. Without being limited by theory, higher expression levels of one or more stress-related biomarkers provided herein in cells may indicate higher levels of apoptosis in cells.
[0339] In some embodiments, the heterologous gene effector is derived from a gene product, said gene product being a hematopoietic stem cell transcription factor. In some embodiments, the target gene is a mesenchymal stem cell transcription factor. In some embodiments, the target gene is an embryonic stem cell transcription factor. In some embodiments, the target gene is an induced pluripotent stem cell (iPSC) transcription factor. In some embodiments, the target gene is an epithelial stem cell transcription factor. In some embodiments, the target gene is a cancer stem cell transcription factor.
[0340] In some implementations, the target gene is an age-related gene. In some implementations, the target gene is a aging-related protein. In some implementations, the target gene is a drug target.
[0341] In some implementations, the target gene (e.g., a target endogenous gene) is a cancer-related gene. Non-limiting examples of cancer-related genes include A1CF, ABI1, ABL1, ABL2, ACKR3, ACSL3, ACSL6, ACVR1, ACVR2A, AFDN, AFF1, AFF3, AFF4, AKAP9, AKT1, AKT2, AKT3, ALDH2, ALK, AMER1, ANK1, APC, APOBEC3B, AR, ARAF, ARHGAP26, ARHGAP5, ARHGEF10, ARHGEF10L, ARHGEF12, ARID1A, ARID1B, ARID2, ARNT, ASPSCR1, ASXL1, ASXL2, AT F1, ATIC, ATM, ATP1A1, ATP2B3, ATR, ATRX, AXIN1, AXIN2, B2M, BAP1, BARD1, BAX, BAZ1A, BCL10, BCL11A, BCL11B, BCL2, BCL2L12, BCL3, BCL6, BCL7A, BCL9, BCL9L, BCLAF1, BCOR, BCORL1, BCR, BIRC3, BIRC6, BLM, BMP5, BMPR1A, BRAF, BRCA1, BRCA2, BRD3, BRD4, BRIP1, BTG1, BTK, BUB1B, C15orf65, CA CNA1D, CALR, CAMTA1, CANT1, CARD11, CARS, CASP3, CASP8, CASP9, CBFA2T3, CBFB, CBL, CBLB, CBLC, CCDC6, CCNB1IP1, CCNC, CCND1, CCND2, CCND3, C CNE1, CCR4, CCR7, CD209, CD274, CD28, CD74, CD79A, CD79B, CDC73, CDH1, CDH10, CDH11, CDH17, CDK12, CDK4, CDK6, CDKN1A, CDKN1B, CDKN2A, CDKN2C , CDX2, CEBPA, CEP89, CHCHD7, CHD2, CHD4, CHEK2, CHIC2, CHST11, CIC, CIITA, CLIP1, CLP1, CLTC, CLTCL1, CNBD1, CNBP, CNOT3, CNTNAP2, CNTRL, COL 1A1, COL2A1, COL3A1, COX6C, CPEB3, CREB1, CREB3L1, CREB3L2, CREBBP, CRLF2, CRNKL1, CRTC1, CRTC3, CSF1R, CSF3R, CSMD3, CTCF, CTNNA2, CTNNB1,CTNND1, CTNND2, CUL3, CUX1, CXCR4, CYLD, CYP2C8, CYSLTR2, DAXX, DCAF12L2, DCC, DCTN1, DDB2, DDIT3, DDR2, DDX10, DDX3X, DDX5, DDX6, DEK, DGCR8, DICER1, DNAJB1, DNM2, DNMT3A, DROSHA, DUX4L1 (or Dux4), EBF1, ECT2L, EED, EGFR, EIF1AX, EIF3E, EIF4A2, ELF3, ELF4, ELK4, ELL, ELN, EML4, EP300, EPAS1, EPHA3, EPHA7, EPS15, ERBB2, ERBB3, ERBB4, ERC1, ERCC2, ERCC3, ERCC4, ERCC5, ERG, ESR1, ETNK1, ETV1, ETV4, ETV5, ETV6, EWSR1, EXT1, EXT2, EZH2, EZR, FAM131B, FAM135B, FAM47C, FANCA, FANCC, FANCD2, FANCE, FANCF, FANCG, FAS, FAT1, FAT3, FAT4, FBLN2, FBXO11, FBXW7, FCGR2B, FCRL4, FEN1, FES, FEV, FGFR1, FGFR1OP, FGFR2, FGFR3, FGFR4, FH, FHIT, FIP1L1, FKBP9, FLCN, FLI1, FLNA, FLT3, FLT4, FNBP1, FOXA1, FOXL2, FOXO1, FOXO3, FOXO4, FOXP1, FOXR1, FSTL3, FUBP1, FUS, GAS7, GATA1, GATA2, GATA3, GLI1, GMPS, GNA11, GNAQ, GNAS, GOLGA5, GOPC, GPC3, GPC5, GPHN, GRIN2A, GRM3, H3F3A, H3F3B, HERPUD1, HEY1, HIF1A, HIP1, HIST1H3B, HIST1H4I, HLA-A, HLF, HMGA1, HMGA2, HMGN2P46, HNF1A, HNRNPA2B1, HOOK3, HOXA11, HOXA13, HOXA9, HOXC11, HOXC13, HOXD11, HOXD13, HRAS, HSP90AA1, HSP90AB1, ID3, IDH1, IDH2, IGF2BP2, IGH, IGK, IGL, IKBKB, IKZF1, IL2, IL21R, IL6ST, IL7R, IRF4, IRS4, ISX, ITGAV, ITKJAK1, JAK2, JAK3, JAZF1, JUN, KAT6A, KAT6B, KAT7, KCNJ5, KDM5A, KDM5C, KD M6A, KDR, KDSR, KEAP1, KIAA1549, KIF5B, KIT, KLF4, KLF6, KLK2, KMT2A, KMT2 C, KMT2D, KNL1, KNSTRN, KRAS, KTN1, LARP4B, LASP1, LATS1, LATS2, LCK, LCP 1 LEF1, LEPROTL1, LHFPL6, LIFR, LMNA, LMO1, LMO2, LPP, LRIG3, LRP1B, LSM1 4A, LYL1, LZTR1, MACC1, MAF, MAFB, MALAT1, MALT1, MAML2, MAP2K1, MAP2K2 MAP2K4, MAP3K1, MAP3K13, MAPK1, MAX, MB21D2, MDM2, MDM4, MDS2, MECOM, MED 12. MEN1, MET, MGMT, MITF, MLF1, MLH1, MLLT1, MLLT10, MLLT11, MLLT3, MLLT 6, MN1, MNX1, MPL, MRTFA, MSH2, MSH6, MSI2, MSN, MTCP1, MTOR, MUC1, MUC16, M UC4, MUTYH, MYB, MYC, MYCL, MYCN, MYD88, MYH11, MYH9, MYO5A, MYOD1, N4BP2 、NAB2、NACA、NBEA、NBN、NCKIPSD、NCOA1、NCOA2、NCOA4、NCOR1、NCOR2、NDRG1 NF1, NF2, NFATC2, NFE2L2, NFIB, NFKB2, NFKBIE, NIN, NKX2-1, NONO, NOTCH 1, NOTCH2, NPM1, NR4A3, NRAS, NRG1, NSD1, NSD2, NSD3, NT5C2, NTHL1, NTRK1 NTRK3, NUMA1, NUP214, NUP98, NUTM1, NUTM2B, NUTM2D, OLIG2, OMD, P2RY8, P ABPC1, PAFAH1B2, PALB2, PATZ1, PAX3, PAX5, PAX7, PAX8, PBRM1, PBX1, PCBP1 PCM1, PDCD1LG2, PDE4DIP, PDGFB, PDGFRA, PDGFRB, PER1, PHF6, PHOX2B, PI CALM, PIK3CA, PIK3CB, PIK3R1, PIM1, PLAG1, PLCG1, PML, PMS1, PMS2, POLD1.POLE、POLG、POLQ、POT1、POU2AF1、POU5F1、PPARG、PPFIBP1、PPM1D、PPP2R1A 、PPP6C、PRCC、PRDM1、PRDM16、PRDM2、PREX2、PRF1、PRKACA、PRKAR1A、PRKCB 、PRPF40B、PRRX1、PSIP1、PTCH1、PTEN、PTK6、PTPN11、PTPN13、PTPN6、PTPRB 、PTPRC、PTPRD、PTPRK、PTPRT、PWWP2A、QKI、RABEP1、RAC1、RAD17、RAD21、RAD 51B、RAF1、RALGDS、RANBP2、RAP1GDS1、RARA、RB1、RBM10、RBM15、RECQL4、RE L、RET、RFWD3、RGPD3、RGS7、RHOA、RHOH、RMI2、RNF213、RNF43、ROBO2、ROS1、R PL10, RPL22, RPL5, RPN1, RSP2, RSP3, RUNX1, RUNX1T1, S100A7, SALL4, SBDS, SDC4, SDHA, SDHAF2, SDHB, SDHC, SDHD, 44444, 44445, 44448, SET, SETBP1 、SETD1B、SETD2、SETDB1、SF3B1、SFPQ、SFRP4、SGK1、SH2B3、SH3GL1、SHTN1、 SIRPA、SIX1、SIX2、SKI、SLC34A2、SLC45A3、SMAD2、SMAD3、SMAD4、SMARCA4、 SMARCB1、SMARCD1、SMARCE1、SMC1A、SMO、SND1、SNX29、SOCS1、SOX2、SOX21、 SPECC1、SPEN、SPOP、SRC、SRGAP3、SRSF2、SRSF3、SS18、SS18L1、SSX1、SSX2、S SX4、STAG1、STAG2、STAT3、STAT5B、STAT6、STIL、STK11、STRN、SUFU、SUZ12、 SYK、TAF15、TAL1、TAL2、TBL1XR1、TBX3、TCEA1、TCF12、TCF3、TCF7L2、TCL1A、 TEC、TENT5C、TERT、Tet1、Tet2、TFE3、TFEB、TFG、TFPT、TFRC、TGFBR2、THRAP 3、TLX1、TLX3、TMEM127、TMPRSS2、TNC、TNFAIP3、TNFRSF14、TNFRSF17、TOP1、TP53, TP63, TPM3, TPM4, TPR, TRA, TRAF7, TRB, TRD, TRIM24, TRIM27, TRIM33, TRIP11, TRRAP, TSC1, TSC2, TSHR, U2AF1, UBR5, USP44, USP6, USP8, VAV1, VHL, VTI1A, W AS, WDCP, WIF1, WNK2, WRN, WT1, WWTR1, XPA, ,
[0342] In some embodiments, the target gene provided herein may be Dux4. Target polynucleotide sequences that can be targeted (e.g., bound) by the systems disclosed herein (e.g., Cas or dCas proteins and / or guide nucleic acid molecules) may be located at or adjacent to a D4Z4 repeat array. For example, guide nucleic acid molecules (e.g., guide RNA) may comprise: (i) a scaffold sequence configured to complex with a Cas protein, and (ii) a spacer sequence exhibiting specific binding to the target polynucleotide sequence. Any suitable scaffold sequence may be used to complex with a Cas protein. Non-limiting examples of suitable scaffold sequences are disclosed, for example, in International Publication No. WO2023 / 168242, which is incorporated herein by reference in its entirety. In another example, nucleic acids that are cellularly heterologous and function in the absence of Cas / dCas proteins (e.g., antisense oligonucleotides, small interfering RNAs, ribozymes, etc.) may exhibit specific binding to the target polynucleotide sequence. In various other examples, non-Cas / dCas proteins (e.g., ZFN, Talen, etc.) may exhibit specific binding to the target polynucleotide sequence.
[0343] In some cases, the spacer sequence of the target polynucleotide sequence or the guide nucleic acid molecule may contain a polynucleotide sequence (or the spacer sequence is encoded by it) that shows at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with the polynucleotide sequence listed in Table 1 (e.g., one or more polynucleotide sequences of SEQ ID NO: 1-14) or their complementary sequences. In some cases, the spacer sequence of the target polynucleotide sequence or the guide nucleic acid molecule may contain a polynucleotide sequence (or the spacer sequence is encoded by it) that shows at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with one or more polynucleotide sequences in Table 2 (e.g., one or more polynucleotide sequences in SEQ ID NO: 45-726) or their complementary sequences. In some cases, the spacer sequence of the target polynucleotide sequence or the guide nucleic acid molecule may contain a polynucleotide sequence (or the spacer sequence is encoded by it) that shows at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with one or more polynucleotide sequences in Table 4 (e.g., one or more polynucleotide sequences in SEQ ID NO: 800-881) or their complementary sequences.
[0344] In some cases, target polynucleotide sequences or spacer sequences of guide nucleic acid molecules can be selected to achieve a threshold level reduction of Dux4 expression in muscle cells by at least or at least about 50% compared to control cells. In some cases, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may comprise a polynucleotide sequence (or the spacer sequence is encoded by it) that shows at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or 100% sequence identity with polynucleotide sequences selected from one or more members of the group consisting of the following and their complementary sequences: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 854, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 854, SEQ ID NO: 855, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 854, SEQ ID NO: 855, SEQ ID NO: 856 ... NO:824, SEQ ID NO:829, SEQ ID NO:840, SEQ ID NO:848, SEQ ID NO:860, SEQ ID NO:825, SEQ ID NO:828, SEQ ID NO:826, SEQ ID NO:864, SEQ ID NO:830, SEQ ID NO:833 and SEQ ID NO:865.
[0345] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to achieve a threshold level (e.g., a predetermined threshold level) reduction in the expression of a target gene (e.g., Dux4) in muscle cells (e.g., FSHD muscle cells) compared to control cells (e.g., a system lacking a repressed target gene). This threshold level can be at least or at least about 40%, at least or at least about 50%, at least or at least about 55%, at least or at least about 60%, at least or at least about 65%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 95%, at least or at least about 99%, or substantially about 100%. For example, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to achieve a threshold level reduction in Dux4 expression in muscle cells of at least or at least about 50% compared to control cells. In some instances, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may comprise a polynucleotide sequence (or the spacer sequence is encoded therein) that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with polynucleotide sequences selected from one or more members of the group consisting of the following and their complementary sequences: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 854, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 855, SEQ ID NO: 856, SEQ ID NO: 857, SEQ ID NO: 858, SEQ ID NO: 859, SEQ ID NO: 854, SEQ ID NO: 855, SEQ ID NO: 853, SEQ ID NO: 812, SEQ ID NO: 874, SEQ ID NO: 859 ... NO: 841.
[0346] In some cases, target polynucleotide sequences or spacer sequences of guide nucleic acid molecules can be selected to achieve reduced expression of the target gene (e.g., Dux4) compared to control cells (e.g., systems lacking the repressed target gene), while exhibiting off-target effects below a threshold level (e.g., a predetermined threshold level) in muscle cells (e.g., FSHD muscle cells). Off-target effects can be determined in vitro, ex vivo, in vivo, or through computer simulation. For example, in computer simulation off-target prediction analyses, off-target sites can be identified when similar genomic sequences with a smaller than predetermined edit distance or "edit distance" (e.g., the sum of identified mismatches, deletions, and insertions) (e.g., less than 3 edit distances) are found if: (i) they are found in target cells of interest (e.g., muscle cells, such as adult skeletal muscle cells); (ii) they are found in inactive or "resting" chromatin regions of the target cells; and / or (iii) they are not found in promoter or exon regions of the target cell chromosomes (e.g., non-intergenic, non-intronic, etc.). Based on this identification of off-target sites, the level of off-target activity (e.g., the predicted level of off-target activity) can be determined (e.g., based on one or more methods described in Muhammad Naeem et al., Cells, 9(7), 16008 (2020), which is incorporated herein by reference in its entirety). In some instances, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to have no more than about 5, no more than about 4, no more than about 3, no more than about 2, no more than about 1, or substantially no identified off-target sites. In some instances, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule can be selected to show an off-target activity level of no more than about 30%, no more than about 25%, no more than about 20%, no more than about 15%, no more than about 14%, no more than about 13%, no more than about 12%, no more than about 11%, no more than about 10%, no more than about 9%, no more than about 8%, no more than about 7%, no more than about 6%, no more than about 5%, no more than about 4%, no more than about 3%, no more than about 2%, no more than about 1%, or substantially about 0%.In some instances, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may comprise a polynucleotide sequence (or the spacer sequence is encoded therein) that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with polynucleotide sequences selected from one or more members of the group consisting of the following and their complementary sequences: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 824, SEQ ID NO: 829 ... NO:840, SEQ ID NO:825, SEQ ID NO:828.
[0347] In some instances, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule may be selected to have at most about one off-target site or substantially no off-target sites, with an edit distance of 2. In some instances, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule may be selected to have at most about one off-target site or substantially no off-target sites, with an edit distance of 1. In some instances, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule may be selected to have at most about one off-target site or substantially no off-target sites, with an edit distance of 0.
[0348] In some cases, the target polynucleotide sequence or the spacer sequence of the guide nucleic acid molecule may be selected to (i) achieve a reduction in the expression of the target gene (e.g., Dux4) at the threshold level described above (e.g., a reduction in the expression level of the target gene by at least or at least about 95%, or at least or at least about 98%), and (ii) exhibit off-target effects below the threshold level described above (e.g., off-target activity of no more than about 10% or substantially 0%). In some instances, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may comprise a polynucleotide sequence (or the spacer sequence is encoded thereby) that exhibits at least or at least about 50%, at least or at least about 60%, at least or at least about 70%, at least or at least about 75%, at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with polynucleotide sequences selected from one or more members of the group consisting of the following and their complementary sequences: SEQ ID NO: 800, SEQ ID NO: 867, SEQ ID NO: 836, SEQ ID NO: 851, SEQ ID NO: 874, SEQ ID NO: 841, SEQ ID NO: 830, SEQ ID NO: 833 and SEQ ID NO: 841. NO: 865. In some instances, the spacer sequence of the target polynucleotide sequence or guide nucleic acid molecule may comprise a polynucleotide sequence (or the spacer sequence is encoded by it) that shows at least or at least about 80%, at least or at least about 85%, at least or at least about 90%, at least or at least about 91%, at least or at least about 92%, at least or at least about 93%, at least or at least about 94%, at least or at least about 95%, at least or at least about 96%, at least or at least about 97%, at least or at least about 98%, at least or at least about 99%, or substantially about 100% sequence identity with a polynucleotide sequence selected from one or more members of the group consisting of SEQ ID NO: 800, SEQ ID NO: 836, SEQ ID NO: 851 and their complementary sequences.
[0349] Table 4. Examples of target polynucleotide sequences or guide nucleic acid spacer sequences
[0350] SEQ ID NO: Guide nucleic acid spacer sequence (without scaffold sequence) SEQ ID NO: Guide nucleic acid spacer sequence (without scaffold sequence) 800 TTTATTTTTTCACCCAGAACAGTAACT 841 TTTGCCTTCAACTTCACCTTAACAACA 801 TTTAAAGAGATCTGGGGATCTATACAG 842 TTTGTATGTTGTTAAGGTGAAGTTGAA 802 TTTAACTTGGAAACACAGCGAAGTCCA 843 TTTAATTTTTCAACATAATTAACTCTC 803 TTTGCACTGGAGCAGAGATGACCACAG 844 TTTAAAAATAACGAATGAGCAAAATAT 804 TTTGCCTGTGAGTTCGAATGCACTTTA 845 TTTAGATTCTATTGTCTATTTTCTTCC 805 TTTGGGAATGTGTTTGTGAAGCACCTA 846 TTTGTGAAGCACCTAGAATCTATAGCC 806 TTTGGCTTTTTTGATAAATTGTCTAATG 847 TTTGATAAATTGTCTAATGACTAGATT 807 TTTATCAAAAAGCCAAACATTTCAACA 848 TTTGCTCATTCGTTATTTTTAAATTTC 808 TTTGATGAAGTCTGGCTTACGCCTGT 849 TTTAAGTTCTCCATCAGATATGCAAAA 809 TTTACTTCCGTCACTTTCTTAACATTA 850 TTTGCCTAGACAGCGTCGGAAGGTGGG 810 TTTGTAATGTTAAGAAAGTGACGGAAG 851 TTTATAAATTCACTACAGAGACACAAC 811 TTTAAGATTCTGGGAGGGAGAGAAAAA 852 TTTGCCCTGGGGACCTTAGCAATGGGC 812 TTTATATATGATTTGTATTTTCACAGA 853 TTTGTATTTTCACAGAGATTTAAGAAT 813 TTTAATGCACCATTATAGTAGAAAATT 854 TTTAAAAACCCAACAGAATCATAGAG 814 TTTGAAATCTGGAAAGTTCTTAGCATC 855 TTTAAAAAAAAAAATCACAAGGCACA 815 TTTAAAAGAATAGAGGGAGAAAATGG 856 TTTGCCCGCTTCCTGGCTAGACCTGCG 816 TTTAAAGAATGGGAAAATTACGGGTGA 857 TTTATGATGCTGTCCAGCATCATTTAA 817 TTTAAAATATTAGTTTCCAGGACTCAA 858 TTTACATTGAACAGAGAGCTTTATATT 818 TTTATCTCTTTGTTGATATTTTGCTCA 859 TTTGGCATTGCTTTTGGGGATCTGGGA 819 TTTGTTGATATTTTGCTCATTCGTTAT 860 TTTATGTTCTCACAAGATTCTGGGAGG 820 TTTAAATTTCACTCAGTTGTCTCTTTC 861 TTTGGGGATCTGGGAAAATCTGTGCAC 821 TTTATGTTTTTCTTCCAATGGGGAATA 862 TTTGGTTTCCGCGTGGCTTTGCCCTCC 822 TTTAAAGACTGGCTCAGTAAAGGGGGA 863 TTTGTCCCGGAGGAAACCGCCCACTCC 823 TTTGCTACAGCACTAGTGAAACTGCAA 864 TTTACAAGGGCGGCTGGCTGGCTGGCT 824 TTTAATTCTCTCCTGAAGGAGATACTG 865 TTTGCCCTCCGCAAGGCGGCCTGTTGC 825 TTTGAATATACTGTGGTCATCTCTGCT 866 TTTGCTCCCGGAGCTCTGCGGGCACCC 826 TTTATAAATAATGGCATGACAAGGGTC 867 TTTGGAACCTGGCAAGGAGAGCGAAGG 827 TTTAGCATTTTTTTTCCTAGGGTTATT 868 TTTGAGAAGGATCGCTTTCCAGGCATC 828 TTTGCTCACTGAGAATGCATAAGATGA 869 TTTGAGCGGAACCCGTACCCGGGCATC 829 TTTGATGAGTGCTGTATAGATCCCCAG 870 TTTGGTTTCAGAATGAGAGGTCACGCC 830 TTTGTGTCTGCTGAGAAGAAAGATGAG 871 TTTGGACCCCGAGCCAAAGCGAGGCCC 831 TTTAAGAATTTAATGCACCATTATAGT 872 TTTGGCTCGGGGTCCAAACGAGTCTCC 832 TTTGAAATACAGTATTTCCCAGATCAA 873 TTTAGGACGCGGGGTTGGGACGGGGTC 833 TTTATGCCATTTTCTCCCTCTATTCTT 874 TTTAGGACGCGGGGTTGGGACGGGGTC 834 TTTACTGAGCCAGTCTTTAAATGCTAG 875 TTTATATTTTCATGTGGTTTTATGATG 835 TTTAAATGCTAGATTTGATGAGTGCTG 876 TTTACAAGAGAAAAACAAAAAACCCTA 836 TTTAAAAGACTCTATCTCTGAATGTAT 877 TTTAATAGGGTTTTTTGTTTTTCTCTT 837 TTTGCATATCTGATGGAGAACTTAAAA 878 TTTCACCCAGAACAGTAACT 838 TTTATTTGTTAAAATTCAGTTTCTGAA 879 CCCAGAACAGTAACT 839 TTTATAAATCTATTGTGCCTCAAGTCA 880 AACAGTAACT 840 TTTGACCGCCAGGCGCTCCGTGCTGGC 881 TAACT
[0351] cell
[0352] The compositions, methods, and systems disclosed herein can be applied to a wide variety of cell types and populations thereof. For example, the complexes or systems disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in specific cell types or populations thereof. The methods disclosed herein can be used to identify complexes capable of inducing changes in the expression or activity of target genes (e.g., target endogenous genes) in specific cell types or populations thereof.
[0353] In some embodiments, the complexes, systems, or heterologous gene effectors identified by the methods of this disclosure induce desired changes in the expression of target genes (e.g., target endogenous genes) specific to a particular cell type. In some embodiments, the complexes, systems, or heterologous gene effectors identified by the methods of this disclosure induce desired changes in the expression of target genes (e.g., target endogenous genes) applicable to two or more cell types. In some embodiments, the complexes, systems, or heterologous gene effectors identified by the methods of this disclosure induce desired changes in the expression of target genes (e.g., target endogenous genes) applicable to three or more cell types. In some embodiments, the complexes, systems, or heterologous gene effectors identified by the methods of this disclosure induce desired changes in the expression of target genes (e.g., target endogenous genes) applicable to a class of cell types (e.g., cell types with overlapping functions present in similar tissues or derived from the same or similar differentiation lineages, such as stem cells, immune cells, T cells, T effector cells, etc.). In some embodiments, the complex or system or heterologous gene effector identified by the methods of this disclosure causes desired changes in the expression of target genes (e.g., target endogenous genes) that are broadly applicable to a variety of cell types; for example, when introduced into cells using appropriate methods, it causes the expression levels of target genes in a variety of target cell types to be above or below a specific threshold.
[0354] In some embodiments, the compositions, complexes, systems, or methods of this disclosure are used to induce changes in the expression, epigenetic modification, or activity levels of target genes in primary cells. In some embodiments, the compositions, complexes, systems, or methods of this disclosure are used to induce changes in the expression, epigenetic modification, or activity levels of target genes in cell lines. In some embodiments, the compositions, complexes, systems, or methods of this disclosure are used to induce changes in the expression, epigenetic modification, or activity levels of target genes in immortalized cells.
[0355] In some embodiments, the compositions, complexes, systems, or methods of this disclosure are used to induce changes in the expression, epigenetic modification, or activity level of target genes (e.g., target endogenous genes) in mammalian cells (e.g., human cells, non-human primate cells, non-rodent mammalian cells, non-human mammalian cells, pig cells, lagomorph cells, canine cells, etc.). In some embodiments, the compositions, complexes, systems, or methods of this disclosure are used to induce changes in the expression, epigenetic modification, or activity level of target genes in plant cells, avian cells, reptile cells, bacterial cells, or archaea cells.
[0356] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in human cells.
[0357] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in stem cells.
[0358] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in differentiated cells.
[0359] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in disease-related cells.
[0360] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in cancer cells.
[0361] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in non-cancerous cells.
[0362] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in the following cells: lymphocytes, such as B cells; T cells (cytotoxic T cells, natural killer T cells, regulatory T cells, helper T cells); natural killer cells; cytokine-induced killer (CIK) cells (see, for example, US20080241194); myeloid cells, such as granulocytes (basophils, eosinophils, neutrophils / multisegmented neutrophils). Cells from the endocrine system include: neutrophils, monocytes / macrophages, erythrocytes, reticulocytes, mast cells, platelets / megakaryocytes, and dendritic cells; thyroid cells (thyroid epithelial cells, parafollicular cells), parathyroid cells (parathyroid chief cells, eosinophils), adrenal cells (chromaffin cells), and pineal cells (pineal gland cells); and cells from the nervous system, including glial cells (astrocytes, microglia) and large neurosecretory cells. Cells of the respiratory system, including lung cells (type I and type II lung cells), Clara cells, goblet cells, and dust cells; cells of the circulatory system, including cardiomyocytes and pericytes; cells of the digestive system, including gastric cells (chief cells and parietal cells), goblet cells, Paneth cells, G cells, D cells, ECL cells, I cells, K cells, and S cells; and enteroendocrine cells, including enterochromaffin cells and APUD cells. Cells include: liver cells (e.g., hepatocytes or Kupffer cells), cartilage / bone / muscle; bone cells, including osteoblasts, osteocytes, osteoclasts, and tooth cells (cementoblasts, ameloblasts); chondrocytes, including chondroblasts and chondrocytes; skin cells, including hair cells, keratinocytes, and melanocytes (nevus cells); muscle cells, including muscle cells; urinary system cells, including podocytes, juxtaglomerular cells, intraglomerular / extraglomerular mesangial cells, proximal tubule brush border cells, and maculae densa cells; reproductive system cells, including sperm, Sertoli cells, interstitial cells, and ovum.And other cells, including adipocytes, fibroblasts, tendon cells, epidermal keratinocytes, epidermal basal cells, keratinocytes of the nail and toenail, nail bed basal cells, medullary hair stem cells, cortical hair stem cells, epidermal hair stem cells, epidermal root sheath cells, Huxley's layer root sheath cells, Henle's layer root sheath cells, outer root sheath cells, hair matrix cells, wet stratified barrier epithelial cells, surface epithelial cells of the stratified squamous epithelium of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra, and vagina, epithelial basal cells of the cornea, tongue, oral cavity, esophagus, anal canal, distal urethra, and vagina, urothelial cells, exocrine epithelial cells, salivary gland mucin cells, salivary gland serous cells, and von Willebrand cells of the tongue. Ebner's gland cells, mammary gland cells, lacrimal gland cells, ceruminous gland cells of the ear, dark cells of eccrine sweat glands, clear cells of eccrine sweat glands, apocrine sweat gland cells, Moll's gland cells of the eyelids, sebaceous gland cells, Bowman's gland cells of the nose, Brunner's gland cells of the duodenum, seminal vesicle cells, prostate cells, bulbourethral gland cells, Bartholin's gland cells, Littre gland cells, endometrial cells, isolated goblet cells of the respiratory and digestive tracts, gastric mucus cells, gastric zymogen cells, gastric acid-secreting cells, pancreatic acinar cells, Panthen cells of the small intestine, type II lung cells, Clara cells of the lungs, hormone-secreting cells, anterior pituitary cells, somatic cells, prolactin cells, thyroid-stimulating hormone cells, gonadotropin cells, corticotropin-stimulating hormone cells. Pituitary cells, pituitary intermediate cells, large neurosecretory cells, intestinal and respiratory cells, thyroid cells, thyroid epithelial cells, parafollicular cells, parathyroid cells, parathyroid chief cells, eosinophils, adrenal cells, chromaffin cells, testicular interstitial cells, theca interna cells of follicles, luteal cells of ruptured follicles, granulosa luteal cells, membranous luteal cells, juxtaglomerular cells, renal macula densa cells, metabolic and storage cells, barrier function cells (e.g., lungs, intestines, exocrine glands and urogenital tract), kidneys, type I lung cells, pancreatic duct cells (alopecis cells), non-striated ductal cells (of sweat glands, salivary glands, mammary glands, etc.), ductal cells (of seminal vesicles, prostate gland, etc.), epithelial cells lining closed internal body cavities, propulsive ciliated cells, extracellular matrix secretory cells, contractile cells;Skeletal muscle cells, stem cells, cardiomyocytes, blood and immune system cells, erythrocytes, megakaryocytes, monocytes, connective tissue macrophages (various types), epidermal Langerhans cells, osteoclasts, dendritic cells, microglia, neutrophils, eosinophils, basophils, mast cells, helper T cells, repressive T cells, cytotoxic T cells, natural killer T cells, B cells, natural killer cells, reticulocytes, stem cells and directed progenitor cells of the blood and immune system (various types), pluripotent stem cells, totipotent stem cells, induced pluripotent stem cells. Cells, adult stem cells, sensory conduction cells, neurons, autonomic neurons, supporting cells of sensory organs and peripheral neurons, neurons and glial cells of the central nervous system, lens cells, pigment cells, melanocytes, retinal pigment epithelial cells, germ cells, oogonia / oocytes, spermatids, spermatocytes, spermatogonia, sperm, trophoblasts, follicular cells, supporting cells, thymic epithelial cells, interstitial cells, renal interstitial cells, common myeloid progenitor cells, common lymphoid progenitor cells, or stem cells differentiated into or to be differentiated into any cell type disclosed herein.
[0363] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in stem cells (e.g., isolated stem cells (e.g., ESCs) or induced stem cells (e.g., iPSCs)).
[0364] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) of hematopoietic stem cells (e.g., hematopoietic stem cells from a subject, such as those from bone marrow or peripheral blood (e.g., mobilized peripheral blood apheresis products, such as those mobilized by administration of GCSF, GM-CSF, mozobil, or combinations thereof)).
[0365] In some cases, the pluripotency of stem cells (e.g., ESCs or iPSCs) can be determined in part by assessing the pluripotency characteristics of the cells. Pluripotency characteristics may include, but are not limited to: pluripotent stem cell morphology; unlimited self-renewal potential; expression of pluripotent stem cell markers, including but not limited to SSEA1, SSEA3 / 4, SSEA5, TRA1-60 / 81, TRA1-85, TRA2-54, GCTM-2, TG343, TG30, CD9, CD29, CD133 / prominin, CD140a, CD56, CD73, CD90, CD105, OCT4, NANOG, SOX2, CD30, and / or CD50; the ability to differentiate into all three somatic cell lineages (ectoderm, mesoderm, and endoderm); the ability to form teratomas containing the three somatic cell lineages; and / or (vi) the formation of embryoids containing cells from the three somatic cell lineages.
[0366] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of target genes (e.g., target endogenous genes) in the following cells: immune cells, such as lymphocytes, T cells, CD4+ T cells, CD8+ T cells, α-β T cells, γ-δ T cells, regulatory T cells (Tregs), cytotoxic T lymphocytes, Th1 cells, Th2 cells, Th17 cells, Th9 cells, naive T cells, memory T cells, effector T cells, effector memory T cells (TEM), central memory T cells (TCMs), resident memory T cells (TRMs), follicular helper T cells (TFHs), natural killer T cells (NKTs), tumor-infiltrating lymphocytes (TILs), natural killer cells (NKs), intrinsic lymphocytes (ILCs), ILC1 cells, ILC2 cells, and ILC3 cells. Cells, lymphoid tissue-induced (LTi) cells, B cells, B1 cells, B1a cells, B1b cells, B2 cells, plasma cells, regulatory B cells, memory B cells, marginal zone B cells, follicular B cells, germinal center B cells, antigen-presenting cells (APCs), monocytes, macrophages, M1 macrophages, M2 macrophages, tissue-associated macrophages, dendritic cells, plasmacytoid dendritic cells, neutrophils, mast cells, basophils, eosinophils, common myeloid progenitor cells, common lymphoid progenitor cells, or any combination thereof.
[0367] The compositions, complexes, systems, or methods disclosed herein can be used to induce changes in the expression, epigenetic modification, or activity levels of engineered cells used to manufacture biological products (e.g., antibodies or other protein-based therapeutic agents).
[0368] Additional implementation schemes
[0369] Non-limiting embodiments of this disclosure are also provided by the following numbered options.
[0370] 1. A system for regulating the aberrant expression of target genes in muscle cells, the system comprising:
[0371] Heteropeptides, wherein the heteropeptides comprise nucleases; and
[0372] A guide nucleic acid molecule, configured to form a complex with a heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or near the D4Z4 repeat array in muscle cells.
[0373] After the complex is formed, it can bind to the target polynucleotide sequence, thereby modifying the expression level and / or methylation level of the target gene in muscle cells. The target gene is located within a D4Z4 repeat array.
[0374] in:
[0375] (1) The guide nucleic acid molecule contains a spacer sequence with specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of one or more members in Table 4; optionally, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or
[0376] (2) Configure the system to reduce the expression level of target genes in muscle cells by at least approximately 80% compared to control cells; and / or
[0377] (3) The system is configured to modify the expression level and / or methylation level of target genes in muscle cells, while exhibiting off-target effects below a predetermined threshold in muscle cells; and / or
[0378] (4) Configure the system to continuously modulate the expression level and / or methylation level of downstream genes of the target gene for at least approximately 5 days; and / or
[0379] (5) Configure the system to normalize the apoptosis level of muscle cells to a level comparable to that of healthy muscle cells; and / or
[0380] (6) Configure the system to improve the survival rate of muscle cells in vitro or in vivo; and / or
[0381] (7) The system is configured to have minimal impact on the expression profile of at least one muscle cell-specific gene in muscle cells, wherein the muscle cell-specific gene is different from the target gene; and / or
[0382] (8) Configure the system so that it has minimal impact on at least one health indicator of the subject after treatment with the system.
[0383] 2. The system as described in Option 1, wherein, after the formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cells is maintained for at least about 2 days.
[0384] 3. The system as described in Option 2, wherein the modified expression level and / or methylation level of the target gene is maintained for at least about 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks or 2 months.
[0385] 4. The system as described in Option 2, wherein the modified expression level and / or methylation level of the target gene is maintained for at least about 17 days.
[0386] 5. The system as described in Option 2, wherein the modified expression level and / or methylation level of the target gene is maintained for at least about 18 days.
[0387] 6. The system as described in Option 1, wherein the muscle cells are in a subject who has or is suspected of having facioscapulohumeral muscular dystrophy (FSHD).
[0388] 7. The system as described in option 1, wherein the target gene is Dux4.
[0389] 8. The system as described in option 1, wherein the nuclease has a length of less than or equal to about 800 amino acids.
[0390] 9. The system as described in option 1, wherein the nuclease has a length of less than or equal to about 750 amino acids.
[0391] 10. The system as described in option 1, wherein the nuclease is Un1Cas12f1 or a modified variant thereof.
[0392] 11. The system of option 1, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:43.
[0393] 12. The system of option 1, wherein the nuclease comprises an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO: 44.
[0394] 13. The system as described in option 1, wherein the heterologous polypeptide further comprises a transcriptional regulatory factor.
[0395] 14. The system of option 13, wherein the transcriptional regulatory factor comprises at least one methyltransferase.
[0396] 15. The system as described in option 14, wherein the transcriptional regulatory factor comprises at least one DNA methyltransferase (DNMT).
[0397] 16. The system as described in option 15, wherein the transcriptional regulatory factor comprises DNMT-A or DNMT-L.
[0398] 17. The system as described in option 14, wherein the transcriptional regulatory factor comprises: (i) DNMT-A or DNMT-L, and (ii) KRAB or a variant of KRAB.
[0399] 18. The system as described in option 14, wherein the transcriptional regulator comprises (i) DNMT-L and (ii) KRAB or a variant of KRAB.
[0400] 19. The system as described in option 14, wherein the transcriptional regulator comprises DNMT-A, DNMT-L, and KRAB or a variant of KRAB.
[0401] 20. The system as described in option 13, wherein the transcriptional regulator comprises DNMT-L, or KRAB or a variant of KRAB.
[0402] 21. The system as described in option 13, wherein the transcriptional regulator comprises KRAB or a variant of KRAB.
[0403] 22. The system as described in option 13, wherein the transcriptional regulatory factors comprise a variety of different transcriptional regulatory factors.
[0404] 23. The system of option 1, wherein modification of the expression level and / or methylation level of the target gene causes downregulation of downstream genes of the target gene, wherein the downstream genes comprise one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48 and TRIM43.
[0405] 24. The system as described in option 1, wherein modification of the target gene expression level and / or methylation level causes downregulation of apoptosis markers in muscle cells.
[0406] 25. The system as described in option 24, wherein the apoptosis marker includes caspase 3.
[0407] 26. The system as described in option 1, wherein the complex causes modification of the expression level of the target gene in the muscle gene.
[0408] 27. The system as described in option 26, wherein modification of expression levels causes downregulation of the target gene.
[0409] 28. The sys...
Claims
1. A system for regulating the aberrant expression of a target gene in muscle cells, the system comprising: Heterogeneous polypeptide, wherein the heteropeptide contains a nuclease; as well as A guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule exhibits specific binding to a target polynucleotide sequence at or near the D4Z4 repeat array in the muscle cells. Wherein, after the formation of the complex, the complex is able to bind to the target polynucleotide sequence, thereby modifying the expression level and / or methylation level of the target gene in the muscle cells, wherein the target gene is located within the D4Z4 repeat array, and in: (1) The guide nucleic acid molecule comprises a spacer sequence having the specific binding, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence, the polynucleotide sequence showing at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of one or more members in Table 4; optionally, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence, the polynucleotide sequence showing at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of SEQ ID NO: 800, SEQ ID NO: 836, or SEQ ID NO: 851; and / or (2) The system is configured to reduce the expression level of the target gene in the muscle cells by at least approximately 80% compared to control cells; and / or (3) The system is configured to modify the expression level and / or methylation level of the target gene in the muscle cells, while exhibiting off-target effects below a predetermined threshold level in the muscle cells; and / or (4) Configure the system to continuously regulate the expression level and / or methylation level of downstream genes of the target gene for at least approximately 5 days; and / or (5) Configure the system to normalize the apoptosis level of the muscle cells to a level comparable to that of healthy muscle cells; and / or (6) Configuring the system to improve the survival rate of the muscle cells in vitro or in vivo; and / or (7) The system is configured to have minimal impact on the expression profile of at least one muscle cell-specific gene in the muscle cells, wherein the muscle cell-specific gene is different from the target gene; and / or (8) Configure the system so that after treating the subject with the system, it has minimal impact on at least one of the subject's health indicators.
2. The system as claimed in claim 1, wherein, After the formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cells are maintained for at least about 2 days.
3. The system as described in claim 2, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months.
4. The system as described in claim 2 or 3, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 17 days.
5. The system as described in claim 2 or 3, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 18 days.
6. The system as described in any one of claims 1-5, wherein, The muscle cells were in subjects who had or were suspected of having facioscapulohumeral muscular dystrophy (FSHD).
7. The system as described in any one of claims 1-6, wherein, The target gene is Dux4.
8. The system as claimed in any one of claims 1-7, wherein, The nuclease has a length of less than or equal to about 800 amino acids.
9. The system as described in any one of claims 1-8, wherein, The nuclease has a length of less than or equal to about 750 amino acids.
10. The system as claimed in any one of claims 1-9, wherein, The nuclease is Un1Cas12f1 or a modified variant thereof.
11. The system as claimed in any one of claims 1-10, wherein, The nuclease contains an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:
43.
12. The system as claimed in any one of claims 1-10, wherein, The nuclease contains an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:
44.
13. The system as claimed in any one of claims 1-12, wherein, The heterologous polypeptide further includes transcriptional regulatory factors.
14. The system of claim 13, wherein, The transcriptional regulatory factor contains at least one methyltransferase.
15. The system of claim 14, wherein, The transcriptional regulatory factor contains at least one DNA methyltransferase (DNMT).
16. The system of claim 15, wherein, The transcriptional regulatory factors include DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).
17. The system of claim 14, wherein, The transcriptional regulatory factors include: (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L), and (ii) KRAB or a variant of KRAB.
18. The system of claim 14, wherein, The transcriptional regulatory factors include (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.
19. The system of claim 14, wherein, The transcriptional regulatory factors include DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or variants of KRAB.
20. The system of claim 13, wherein, The transcriptional regulatory factors include DNMT-L (or DNMT3L), or KRAB or variants of KRAB.
21. The system of claim 13, wherein, The transcriptional regulatory factors include KRAB or variants of KRAB.
22. The system of claim 13, wherein, The transcriptional regulatory factors include a variety of different transcriptional regulatory factors.
23. The system as described in any one or more of claims 1-22, wherein, The modification of the expression level and / or methylation level of the target gene causes downregulation of downstream genes of the target gene, wherein the downstream genes comprise one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48 and TRIM43.
24. The system as claimed in any one of claims 1-23, wherein, The modification of the expression level and / or methylation level of the target gene caused downregulation of apoptosis markers in the muscle cells.
25. The system of claim 24, wherein, The apoptosis markers include caspase 3.
26. The system of claim 1, wherein, The complex causes the modification of the expression level of the target gene in the muscle gene.
27. The system of claim 26, wherein, The modification of the expression level causes downregulation of the target gene.
28. The system as claimed in any one of claims 1-27, wherein, The complex causes the modification of the methylation level of the target gene in the muscle gene.
29. The system of claim 28, wherein, The modification of the methylation level causes downregulation of the target gene.
30. The system as claimed in any one of claims 1-29, wherein, The nuclease in question is an inactivated nuclease.
31. A composition comprising the system of any one of the preceding claims.
32. A viral vector comprising the system of any one of the preceding claims.
33. The viral vector as described in claim 32, wherein, The viral vectors include adeno-associated virus (AAV), retrovirus, lentivirus, poxvirus, or adenovirus.
34. The viral vector as described in claim 33, wherein, The AAV includes AAV serotype RH74 AAV.
35. A method for regulating the abnormal expression of a target gene in muscle cells, the method comprising: (a) Contact the muscle cells with a complex or system comprising (i) a heterologous polypeptide comprising a nuclease, and (ii) a guide nucleic acid molecule that exhibits specific binding to a target polynucleotide sequence at or near a D4Z4 repeat array in the muscle cells. as well as (b) Following the contact, the target gene is bound to the complex or system to modify the expression level and / or methylation level of the target gene in the muscle cells, wherein the target gene is located within the D4Z4 repeat array. in: (1) The guide nucleic acid molecule comprises a spacer sequence having the specific binding capability, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of one or more members in Table 4; optionally, wherein the spacer sequence comprises or is encoded by a polynucleotide sequence that shows at least about 80%, at least about 90%, or at least about 95% sequence identity with the polynucleotide sequences of SEQ ID NO:800, SEQ ID NO:836, or SEQ ID NO:851; and / or (2) The complex or system is configured to reduce the expression level of the target gene in the muscle cells by at least about 80% compared with that in control cells; and / or (3) The complex or system is configured to modify the expression level and / or methylation level of the target gene in the muscle cells, while exhibiting off-target effects below a predetermined threshold level in the muscle cells; and / or (4) The complex or system is configured to continuously modulate the expression level and / or methylation level of downstream genes of the target gene for at least about 5 days; and / or (5) Configuring the complex or system to normalize the apoptosis level of the muscle cells to a level comparable to that of healthy muscle cells; and / or (6) Configuring the complex or system to improve the survival rate of the muscle cells in vitro or in vivo; and / or (7) The complex or system is configured to have minimal impact on the expression profile of at least one muscle cell-specific gene in the muscle cells, wherein the muscle cell-specific gene is different from the target gene; and / or (8) The complex or system is configured to have minimal impact on at least one health indicator of the subject after treatment with the complex or system.
36. The method of claim 35, wherein, After the formation of the complex, the modified expression level and / or methylation level of the target gene in the muscle cells are maintained for at least about 2 days.
37. The method of claim 35 or 36, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 2 weeks, 4 weeks, or 2 months.
38. The method of claim 36, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 17 days.
39. The method of claim 36, wherein, The modified expression level and / or methylation level of the target gene are maintained for at least approximately 18 days.
40. The method according to any one of claims 35-39, wherein, The contact includes injecting a composition containing the complex into a subject in need, wherein the subject has or is suspected of having facioscapulohumeral muscular dystrophy (FSHD).
41. The method according to any one of claims 35-40, wherein, The target gene is Dux4.
42. The method according to any one of claims 35-41, wherein, The nuclease has a length of less than or equal to about 800 amino acids.
43. The method of claim 42, wherein, The nuclease has a length of less than or equal to about 750 amino acids.
44. The method according to any one of claims 35-43, wherein, The nuclease is Un1Cas12f1 or a modified variant thereof.
45. The method according to any one of claims 35-44, wherein, The nuclease contains an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:
43.
46. The method according to any one of claims 35-45, wherein, The nuclease contains an amino acid sequence that is at least about 80%, at least about 90%, at least about 95%, or at least about 99% identical to the polypeptide sequence of SEQ ID NO:
44.
47. The method according to any one of claims 35-46, wherein, The heterologous polypeptide further includes transcriptional regulatory factors.
48. The method of claim 47, wherein, The transcriptional regulatory factor contains at least one methyltransferase.
49. The method of claim 48, wherein, The transcriptional regulatory factor contains at least one DNA methyltransferase (DNMT).
50. The method of claim 49, wherein, The transcriptional regulatory factors include DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L).
51. The method of claim 49, wherein, The transcriptional regulatory factors include: (i) DNMT-A (or DNMT3A) or DNMT-L (or DNMT3L), and (ii) KRAB or a variant of KRAB.
52. The method of claim 49, wherein, The transcriptional regulatory factors include (i) DNMT-L (or DNMT3L) and (ii) KRAB or a variant of KRAB.
53. The method of claim 49, wherein, The transcriptional regulatory factors include DNMT-A (or DNMT3A), DNMT-L (or DNMT3L), and KRAB or variants of KRAB.
54. The method of claim 47, wherein, The transcriptional regulatory factors include DNMT-L (or DNMT3L), or KRAB or variants of KRAB.
55. The method of claim 47, wherein, The transcriptional regulatory factors include KRAB or variants of KRAB.
56. The method of claim 47, wherein, The transcriptional regulatory factors include a variety of different transcriptional regulatory factors.
57. The method according to any one of claims 35-56, wherein, The modification of the expression level and / or methylation level of the target gene causes downregulation of downstream genes of the target gene, wherein the downstream genes comprise one or more members selected from the group consisting of ZSCAN4, LEUTX, MBD3L2, TRIM48 and TRIM43.
58. The method according to any one of claims 35-57, wherein, The modification of the expression level and / or methylation level of the target gene caused downregulation of apoptosis markers in the muscle cells.
59. The method of claim 58, wherein, The apoptosis markers include caspase 3.
60. The method according to any one of claims 35-59, wherein, The complex causes the modification of the expression level of the target gene in the muscle gene.
61. The method according to any one of claims 35-60, wherein, The modification of the expression level causes downregulation of the target gene.
62. The method according to any one of claims 35-61, wherein, The complex causes the modification of the methylation level of the target gene in the muscle gene.
63. The method of claim 62, wherein, The modification of the methylation level causes downregulation of the target gene.
64. The method according to any one of claims 35-63, wherein, The nuclease in question is an inactivated nuclease.
65. The system as described in any one of claims 1-30, the composition as described in claim 31, or the carrier as described in any one of claims 32-34, used as a medicament, for example for inhibiting, improving, or treating cancer, autoimmune diseases, or facioscapulohumeral muscular dystrophy (FSHD).
66. A system as described in any one of claims 1-30, a composition as described in claim 31, or a vector as described in any one of claims 32-34 for editing genes in cells in vitro or in vivo.
67. A method of delivering a system to cells, the system regulating the aberrant expression of a target gene, the method comprising introducing the system of any one of claims 1-30, the composition of claim 31, or the vector of any one of claims 32-34 into cells, preferably into the cells of a subject, for example, into the cells of a person suffering from cancer, an autoimmune disease, or facioscapulohumeral muscular dystrophy (FSHD).
68. The system, composition, or viral vector as claimed in any of the preceding claims, wherein, The system includes: A heterologous polypeptide comprising a nuclease, the nuclease comprising an amino acid sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728; and A guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:
865.
69. The system, composition, or viral vector as claimed in any of the preceding claims, wherein, The system includes: A heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, or having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, or having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727; and A guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:
865.
70. The method as described in any of the preceding claims, wherein, The complex or system includes: A heterologous polypeptide comprising a nuclease, the nuclease comprising an amino acid sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:43, SEQ ID NO:44, and SEQ ID NO:728; and A guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:
865.
71. The method as described in any of the preceding claims, wherein, The complex or system includes: A heterologous polypeptide comprising a nuclease operatively coupled to a heterologous gene effector, wherein the nuclease comprises an amino acid sequence having, or having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:728, and the heterologous gene effector comprises an amino acid sequence having, or having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with SEQ ID NO:727; and A guide nucleic acid molecule configured to form a complex with the heterologous polypeptide, wherein the guide nucleic acid molecule comprises a spacer sequence encoded by a polynucleotide sequence having, having about or having at least 80%, 85%, 90%, 95%, 97%, 98%, 99%, or about 100% sequence identity with any one of SEQ ID NO:800, SEQ ID NO:867, SEQ ID NO:836, SEQ ID NO:851, SEQ ID NO:874, SEQ ID NO:841, SEQ ID NO:830, SEQ ID NO:833, and SEQ ID NO:865.
Citation Information
Patent Citations
SE44444C1
SE44445C1
Compositions and methods for protecting organs, tissue and cells from immune system-mediated damage
US20080241194A1
Engineered nucleases, compositions, and methods of use thereof
WO2023168242A1