Engineered CASX repressor system
Patent Information
- Application Number
- JP2024517501
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-03-18
- Filing Date
- 2022-09-21
- Publication Date
- 2025-10-01
AI Technical Summary
を及ぼすことが可能である、ある量の薬物又は生物学的物質を単独で、又は組成物の一部として指す。そのような効果は、有益であるために絶対的である必要はない。
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application Nos. 63 / 246,543, filed September 21, 2021, and 63 / 321,517, filed March 18, 2022, the contents of each of which are incorporated herein by reference in their entirety.
[0002] Incorporation by reference of sequence listing The contents of the electronic sequence listing (SCRB_034_02WO_SeqList_ST26.xml, size: 63,394,386 bytes, and creation date: September 16, 2022) are incorporated herein by reference in their entirety. [Background technology]
[0003] There are a variety of methods for regulating target gene expression in cells. In mammalian systems, cells use a system of chromatin regulators (CRs) and associated histone and DNA modifications to regulate gene expression and establish long-term epigenetic memory. This system is important in development, aging, and disease and may provide essential capabilities for incorporating regulation into synthetic biology. In experimental systems, methods such as RNA interference (RNAi) are useful for targeted gene knockdown and are widely used in large-scale library screens. However, RNAi has several limitations. In particular, RNAi-based knockdown suffers from off-target effects, as well as incomplete knockdown of the target (Jackson AL, et al., Expression profiling reveals off-target gene regulation by RNAi. Nat Biotechnol. 21:635 (2003)); Sigoillot FD, et al., A bioinformatics method identifies prominent off-target transcripts in RNAi screens. Nat Methods. 19:9(4):363 (2012)). Tailored DNA-binding proteins, such as zinc finger proteins or transcription activator-like effectors (TALEs) linked to transcriptional repressor domains, can mediate selective gene repression but are limited by the fact that each desired target gene requires the generation of new protein.
[0004] The emergence of CRISPR / Cas systems and their programmable nature have facilitated their use as versatile technologies for genome manipulation and engineering. Certain CRISPR proteins are particularly suited for such manipulation. For example, certain Class 2 CRISPR / Cas systems have a compact size, providing ease of delivery, and the nucleotide sequences encoding the proteins are relatively short, an advantage of their incorporation into viral vectors for cellular delivery. However, for certain disease indications, gene silencing or suppression is preferred over gene editing. The ability to render CasX catalytically inactive (dCasX) has been demonstrated ( WO 2020 / 247882 A1 ), making it an attractive platform for generating fusion proteins capable of gene silencing. Thus, there is a need in the art for additional gene repressor systems (e.g., dCas protein + repressor domain) that are optimized and / or offer improvements over earlier-generation gene repressor systems, such as those based on Cas9, for use in a variety of therapeutic, diagnostic, and research applications. Summary of the Invention
[0005] Aspects of the present disclosure are directed to compositions and methods for modulating the expression of a target nucleic acid in a cell.
[0006] The present disclosure provides gene repressor systems comprising a catalytically inactive Class 2 CRISPR protein, e.g., a Class 2 V-type CRISPR protein, linked to one or more transcriptional repressor domains as a fusion protein and a targeting sequence complementary to a target nucleic acid of a gene in a cell (dXR:gRNA system), nucleic acids encoding the fusion proteins, vectors encoding or comprising components of the dXR:gRNA system, and lipid nanoparticles encoding or encapsulating components of the dXR:gRNA system, as well as methods of making and using the dXR:gRNA system. The dXR:gRNA systems of the present disclosure have utility in methods of gene silencing or gene suppression in diseases where suppression of a gene product is useful for reversing the underlying cause of the disease or ameliorating signs or symptoms of the disease, and methods are also provided.
[0007] Further features and advantages of certain embodiments of the present disclosure will become more fully apparent in the following description of the embodiments and drawings thereof, and from the claims.
[0008] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief explanation of the drawings]
[0009] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0010] [Figure 1]1 shows the results of an assay evaluating targeted and non-targeted dXR molecules using non-targeted and targeted spacers (left and right bars, respectively), along with the percent loss of target mRNA measured by qPCR in HEK293T cells, as described in Example 1. Data represent ΔΔCt values from biological replicates. [Figure 2] Figure 1 shows the dose-response results of diphtheria toxin titration of cells transduced with either catalytically active CasX editors, i.e., CasX-34.19 and CasX-34.21, with gRNA targeting the gene encoding heparin-binding EGF-like growth factor (HBEGF), a catalytically inactive CasX (dCasX) protein linked to a repressor domain as a fusion protein targeted to HBEGF (a dXR fusion protein, i.e., dXR1-34.28), or a non-targeting dXR molecule (CasX-NT or dXR-NT), as described in Example 2. Data represent the mean and standard deviation of two biological replicates. [Figure 3] Figure 1 shows cell counts from an array study of dXR and three spacers targeting the 5'UTR sequence of the C9orf72 locus in the TK-GFP cell line after ganciclovir treatment, as described in Example 3. Data represent single-point cell counts after treatment with ganciclovir. NT: non-targeting spacer. [Figure 4] A schematic diagram showing the plasmids utilized to generate the XDP construct is shown, in which dXR is encoded on a separate plasmid, the plasmid encoding the Gag components also encodes the MS2 coat protein, and the protease cleavage sequence site is indicated by an arrow. [Figure 5] A schematic diagram showing the plasmids utilized in generating the XDP constructs is shown, in which dXR is encoded on the plasmid encoding the Gag component and the protease cleavage sequence site is indicated by an arrow. [Figure 6]1 is a bar graph showing Western blot quantification of PTBP1 protein levels in mouse astrocytes harvested 11 days after transduction with lentiviral particles containing dXR with the indicated PTBP1-targeting spacers, as described in Example 5. Cells treated with XDP containing CasX ribonucleoprotein (RNP) using lentiviral particles with spacer 28.10 or NT spacers served as experimental controls. The ratio of PTBP1 protein to total protein was normalized to that determined for the NT control in the graph. [Figure 7] Figure 1 shows a schematic diagram of various constructs of epigenetic long-range CasX repressor (ELXR) molecules tested for gene repression activity. D3A and D3L represent DNA methyltransferase 3 alpha (DNMT3A) and DNMT3A-like protein (DNMT3L), respectively, as described in Example 6. CD = catalytic domain, ID = interaction domain. L1-L4 are linkers. NLS is a nuclear localization signal. See Tables 24 and 25 for ELXR sequences. [Figure 8A] 1 shows the results of a time course experiment comparing the beta-2-microglobulin (B2M) suppressive activity (expressed as percentage of HLA-negative cells) of ELXR protein numbers 1-3, as described in Example 6. Data are expressed as mean with standard deviation, N=3. [Figure 8B] 8A shows the results of the same time course experiment as shown in FIG. 8A, but showing the B2M inhibitory activity of ELXR proteins 1-3 containing the ZIM3-KRAB domain, benchmarked against the same experimental control, as described in Example 6. Data are presented as mean with standard deviation, N=3. [Figure 9A] 1 shows the results of a time course experiment comparing the B2M silencing activity (expressed as percentage of HLA-negative cells) of ELXR proteins #1, #4, and #5, as described in Example 6. Data are presented as mean with standard deviation, N=3. [Figure 9B]9A shows the results of the same time course experiment as shown in FIG. 9A, but shows the B2M silencing activity of ELXR proteins #1, #4, and #5 containing the ZIM3-KRAB domain, benchmarked against the same experimental control, as described in Example 6. Data are presented as mean with standard deviation, N=3. [Figure 10] 1 is a violin plot of the percent CpG methylation of CpG sites surrounding the transcription start site of the B2M locus for each experimental condition indicated, as described in Example 6. [Figure 11] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 21) versus specificity (percent off-target CpG methylation at the B2M locus quantified at day 5) of ELXR Proteins 1-3 benchmarked against catalytically active CasX491 and dCas9-ZNF10-DNMT3A / L, as described in Example 6. [Figure 12] 1 is a violin plot of percent CpG methylation of CpG sites downstream of the transcription start site of the VEGFA locus for each experimental condition indicated, as described in Example 6. [Figure 13A] 1 is a violin plot of the percent CpG methylation of CpG sites surrounding the transcription start site of the VEGFA locus for each indicated experimental condition evaluating ELXRs #1, #4, and #5 with a B2M targeting spacer, as described in Example 6. [Figure 13B] 1 is a violin plot of the percent CpG methylation of CpG sites surrounding the transcription start site of the VEGFA locus for each indicated experimental condition evaluating ELXRs #1, #4, and #5 with non-targeting spacers, as described in Example 6. [Figure 14]Scatter plot showing the relative activity (mean percentage of HLA-negative cells at day 21) versus specificity (median percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR proteins 1-5 carrying either the ZNF10-KRAB domain or the ZIM-KRAB domain. ELXR proteins were benchmarked against catalytically active CasX491 and dCas9-ZNF10-DNMT3A / L as described in Example 6. [Figure 15] Quantification of percent editing, measured as indel frequency detected by NGS at the B2M locus, for each of the indicated catalytically inactive CasXs with either a B2M-targeted or non-targeted spacer, as described in Example 8. Catalytically active CasX491, catalytically inactive CasX9 (dCas9), and mock transfection served as experimental controls. [Figure 16] 1 provides a violin plot showing the log2 (fold change) of sequences before and after selection for their ability to support dXR repression of the HBEGF locus, as described in Example 4. The plot shows results for the entire KRAB domain library, a negative control sequence set, a positive control set of known KRAB repressors, the top 1597 KRAB domains tested with a log2 (fold change) greater than 2 and a p-value less than 0.01, and the top 95 KRAB domains tested. [Figure 17] 1 shows the B2M silencing activity (expressed as percentage of HLA-negative cells) of dXR proteins with various KRAB domains as described in Example 4. Data are expressed as mean with standard deviation, N=3. [Figure 18] 1 shows the B2M silencing activity (expressed as percentage of HLA-negative cells) of dXR proteins with various KRAB domains as described in Example 4. Data are expressed as mean with standard deviation, N=3. [Figure 19A] A logo of KRAB domain motif 1, as described in Example 4, is provided. [Figure 19B]A logo of KRAB domain motif 2, as described in Example 4, is provided. [Figure 19C] A logo for KRAB domain motif 3 (SEQ ID NO: 59345), as described in Example 4, is provided. [Figure 19D] A logo for KRAB domain motif 4 (SEQ ID NO: 59346), as described in Example 4, is provided. [Figure 19E] A logo for KRAB domain motif 5 (SEQ ID NO: 59347), as described in Example 4, is provided. [Figure 19F] A logo for KRAB domain motif 6 (SEQ ID NO: 59348), as described in Example 4, is provided. [Figure 19G] A logo for KRAB domain motif 7 (SEQ ID NO: 59349), as described in Example 4, is provided. [Figure 19H] A logo of KRAB domain motif 8, as described in Example 4, is provided. [Figure 19I] A logo of KRAB domain motif 9, as described in Example 4, is provided. [Figure 20A] Schematic diagram showing the relative positions of CD151 sequences targeted by the spacer of the assayed ELXR molecules and the dCas9-ZNF10-DNMT3A / L control, as described in Example 9. The positions targeted by the gRNA are indicated by light gray (paired with ELXR) and dark gray (paired with dCas9-ZNF10-DNMT3A / L) bars. [Figure 20B] 1 is a bar graph showing the results of a time course experiment comparing the CD151 suppressive activity (expressed as a percentage of total cells with CD151 knockdown) of ELXR proteins #1, #4, and #5 containing the ZIM3-KRAB domain with the indicated targeting spacers, as described in Example 9. Data for each time point (day 6, day 15, and day 22) are overlaid and presented as the mean with standard deviation, N=3. [Figure 21A]21B is a schematic diagram showing the locations of various B2M-targeting gRNAs tiled over a 1 KB window in the promoter region of the B2M gene, as described in Example 10. Numbers correspond to the specific B2M-targeting spacers shown in FIG. 21B. Targeting gRNAs are indicated by gray bars. [Figure 21B] 1 is a bar graph showing quantification of B2M suppression, expressed as the mean percentage of HLA-negative cells, mediated by either dXR1 or ELXR#1 with the indicated B2M-targeting spacer, as described in Example 10. Data are presented as mean with standard deviation, N=3. NT=non-targeting spacer. [Figure 22]
[0033] Figure 1 shows the results of a time course experiment comparing the B2M inhibitory activity (expressed as percentage of HLA-negative cells) of the indicated ELXR5-ZIM3 and its variants with B2M-targeting gRNA using spacer 7.37, as described in Example 11. Data are presented as mean with standard deviation, N=3. CD = catalytic domain of DNMT3A. [Figure 23] 22 shows the results of the same time course experiment as shown in Figure 22, but showing the B2M inhibitory activity of the indicated ELXR5-ZIM3 variants with a B2M-targeting gRNA using spacer 7.160, as described in Example 11. Data are presented as mean with standard deviation, N=3. [Figure 24] 22 shows the results of the same time course experiment as shown in Figure 22, but showing the B2M inhibitory activity of the indicated ELXR5-ZIM3 variants with a B2M-targeting gRNA using spacer 7.165, as described in Example 11. Data are presented as mean with standard deviation, N=3. [Figure 25] 22 shows the results of the same time course experiment as shown in Figure 22, but showing the B2M inhibitory activity of the indicated ELXR5-ZIM3 variants with a non-targeting gRNA, as described in Example 11. Data are presented as mean with standard deviation, N=3. [Figure 26]10 is a violin plot of the percent CpG methylation of CpG sites downstream of the transcription start site of the VEGFA locus for each of the indicated ELXR5-ZIM3 variants for three B2M-targeting gRNAs and a non-targeting gRNA, as described in Example 11. [Figure 27] FIG. 12 is a scatter plot showing relative activity (mean percentage of HLA-negative cells at day 21 for spacer 7.160) versus specificity (percent off-target CpG methylation at the VEGFA locus quantified at day 7 for spacer 7.160) for the indicated ELXR5-ZIM3 variants, as described in Example 11. [Figure 28] 1 is a bar graph showing the percentage of mouse Hepa1-6 cells treated with either dXR1 or ELXR1-ZIM3 mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 on day 6, as described in Example 14. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control. [Figure 29] 1 is a time course plot showing the percentage of mouse Hepa1-6 cells treated with dXR1 mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 at days 6, 13, and 25 after delivery, as described in Example 14. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control, and treatment with water served as a negative control. [Figure 30] 1 is a time course plot showing the percentage of mouse Hepa1-6 cells treated with ELXR1-ZIM3 mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 at days 6, 13, and 25 after delivery, as described in Example 14. Spacer 6.7, which targets the human PCSK9 locus, served as a non-targeting control, and treatment with water served as a negative control. [Figure 31]
[0033] Figure 1 is a bar graph showing the percentage of mouse Hepa1-6 cells treated with ELXR1-ZIM3, ELXR5-ZIM3, catalytically active CasX491, or dCas9-ZNF10-DNMT3A / 3L mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 at day 7, as described in Example 14. Spacer 6.7, targeting the human PCSK9 locus, served as a non-targeting control. In-house IVT or third-party mRNA production is indicated in parentheses. [Figure 32] FIG. 10 is a time course plot showing the percentage of mouse Hepa1-6 cells treated with IVT-produced ELXR1-ZIM3 versus ELXR5-ZIM5 mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 at 7 and 14 days post-delivery, as described in Example 14. [Figure 33] FIG. 10 is a time course plot showing the percentage of mouse Hepa1-6 cells treated with third-party-produced ELXR1-ZIM3 versus dCas9-ZNF10-DNMT3A / 3L mRNA paired with the indicated PCSK9-targeting gRNA that stained negative for intracellular PCSK9 at 7 and 14 days post-delivery, as described in Example 14. [Figure 34] 10 is a plot showing the percentage of HEK293T cells transfected with plasmids encoding the indicated CasX or ELXR:gRNA constructs that expressed B2M 6 days after treatment with the DNMT1 inhibitor 5-azadC at various concentrations, as described in Example 12. [Figure 35] 10 is a plot juxtaposing quantification of B2M repression in HEK293T cells transfected with plasmids encoding the indicated CasX or ELXR:gRNA constructs and cultured for 58 days with quantification of B2M reactivation upon treatment of cells transfected with 5-azadC, as described in Example 12. [Figure 36]1 shows a schematic diagram of various ELXR No. 5 architectures incorporating additional DNMT3A domains, as described in Example 11. The additional DNMT3A domains were the DNMT3A ADD domain ("D3A ADD") and the DNMT3A PWWP domain ("D3A PWWP"). "D3A endo" encodes the endogenous sequence occurring between the DNMT3A PWWP domain and the ADD domain. "D3A CD" and "D3L ID" indicate the DNMT3A catalytic domain and the DNMT3L interaction domain, respectively. "L1-L3" are linkers. "NLS" is a nuclear localization signal. See Table 33 for ELXR sequences. [Figure 37] 1 shows a schematic diagram of the general architecture of ELXR molecules with ADD domains of ELXR constructs #1, #4, and #5 tested in Example 13. "D3A ADD," "D3A CD," and "D3L ID" represent the ADD domain of DNMT3A, the catalytic domain of DNMT3A, and the interaction domain of DNMT3L, respectively, as described in Example 13. "L1-L4" are linkers. "NLS" is a nuclear localization signal. See Table 35 for ELXR sequences. [Figure 38] Figure 2 shows a schematic of the general dXR construct as described in Example 1. NLS is the nuclear localization signal and L3 is linker 3 (see Table 24 for AA sequences). [Figure 39A]
[0033] Figure 1 shows the results of a time course experiment comparing the B2M-inhibitory activity (expressed as percentage of HLA-negative cells) of ELXRs with ZIM3-KRAB domains having constructs #1, #4, or #5, with or without a DNMT3A ADD domain, when paired with a B2M-targeting gRNA with spacer 7.160, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 39B]39B is a plot showing the results of the same time course experiment as shown in FIG. 39A, but showing the B2M inhibitory activity of ELXR#5, which has the ZNF10-KRAB domain or the ZIM3-KRAB domain with or without the DNMT3A ADD domain, paired with a B2M-targeting gRNA with spacer 7.160, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 39C] 39B is a plot showing the results of the same time course experiment as shown in FIG. 39A, but showing the B2M inhibitory activity for ELXR5-ZIM3 with or without the DNMT3A ADD domain paired with B2M-targeting gRNAs with the indicated spacers, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 40A] 1 is a plot showing the results of B2M suppression activity at 27 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 1, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 40B] 1 is a plot showing the results of B2M suppression activity at 27 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 4, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 40C]1 is a plot showing the results of B2M suppression activity at 27 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 5, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as mean with standard deviation, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 41A] 10 is a plot showing the results of bisulfite sequencing used to determine off-target methylation at the VEGFA locus 5 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 1, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as the average percentage of CpG methylation at CpG sites near the VEGFA locus, with the standard error of the mean also shown, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 41B] 10 is a plot showing the results of bisulfite sequencing used to determine off-target methylation at the VEGFA locus 5 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 4, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as the average percentage of CpG methylation at CpG sites near the VEGFA locus, with the standard error of the mean also shown, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 41C]10 is a plot showing the results of bisulfite sequencing used to determine off-target methylation at the VEGFA locus 5 days post-transfection for ELXR carrying either the ZNF10-KRAB domain or the ZIM3-KRAB domain with construct number 5, with or without the DNMT3A ADD domain of the indicated gRNA, as described in Example 13. Data are presented as the average percentage of CpG methylation at CpG sites near the VEGFA locus, with the standard error of the mean also shown, N=3. "NT" is a gRNA with a non-targeting spacer. [Figure 42A] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZIM3-KRAB domains with constructs #1, #4, and #5 for B2M-targeting gRNAs with spacer 7.160, as described in Example 13. [Figure 42B] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZNF10-KRAB domains with constructs #1, #4, and #5 for B2M-targeting gRNA with spacer 7.160, as described in Example 13. [Figure 43A] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZIM3-KRAB domains with constructs #1, #4, and #5 for B2M-targeting gRNA with spacer 7.37, as described in Example 13. [Figure 43B]1 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZNF10-KRAB domains with constructs no. 1, no. 4, and no. 5 for a B2M-targeting gRNA with spacer 7.37, as described in Example 13. [Figure 44A] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZIM3-KRAB domains with constructs #1, #4, and #5 for B2M-targeting gRNAs with spacer 7.165, as described in Example 13. [Figure 44B] 10 is a dot plot showing the relative activity (mean percentage of HLA-negative cells at day 27) versus specificity (percentage of off-target CpG methylation at the VEGFA locus quantified at day 5) of ELXR molecules having ZNF10-KRAB domains with constructs no. 1, no. 4, and no. 5 for B2M-targeting gRNAs with spacer 7.165, as described in Example 13. [Figure 45] Schematic diagrams of various configurations of ELXR molecules with the incorporation of DNMT3A ADD are shown. "D3A ADD," "D3A CD," and "D3L ID" represent the DNMT3A ADD domain, the DNMT3A catalytic domain, and the DNMT3L interaction domain, respectively. L1-L3 are linkers. NLS is a nuclear localization signal. DETAILED DESCRIPTION OF THE INVENTION
[0011] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. The following claims define the scope of the invention, and methods and structures within the scope of these claims and their equivalents are intended to be covered.
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present embodiments, suitable methods and materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the invention herein.
[0013] definition The terms "polynucleotide" and "nucleic acid," used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. Thus, the terms "polynucleotide" and "nucleic acid" encompass single-stranded DNA, double-stranded DNA, multi-stranded DNA, single-stranded RNA, double-stranded RNA, multi-stranded RNA, genomic DNA, cDNA, DNA-RNA hybrids, and polymers that contain purine and pyrimidine bases, or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases.
[0014] "Hybridizable" and "complementary" are used interchangeably to mean that a nucleic acid (e.g., RNA, DNA) comprises a sequence of nucleotides that, under appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength, allows it to noncovalently bind to another nucleic acid in a sequence-specific antiparallel manner (i.e., the nucleic acid specifically binds to the complementary nucleic acid), i.e., form Watson-Crick and / or G / U base pairs. It is understood that a polynucleotide sequence need not be 100% complementary to its target nucleic acid sequence to be specifically hybridizable; it can have at least about 70%, at least about 80%, at least about 90%, or at least about 95% sequence identity and still be able to hybridize to the target nucleic acid sequence. Furthermore, a polynucleotide can hybridize across one or more segments (e.g., a loop or hairpin structure, a "bulge," a "bubble," etc.) such that intervening or adjacent segments are not involved in the hybridization event.
[0015] For purposes of this disclosure, a "gene" includes a DNA region that encodes a gene product (e.g., a protein or RNA) as well as all DNA regions that regulate the production of a gene product, regardless of whether such regulatory element sequences flank the coding and / or transcribed sequence. Thus, a gene can include regulatory sequences, including, but not limited to, promoter sequences, terminators, translational control sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, origins of replication, matrix attachment sites, and locus control regions. A coding sequence encodes a gene product upon transcription or transcription and translation, and coding sequences of the present disclosure can include fragments and need not contain a full-length open reading frame. A gene can include both the transcribed strand, e.g., the strand containing the coding sequence, and the complementary strand.
[0016] The term "downstream" refers to a nucleotide sequence located 3' of a reference nucleotide sequence. In certain embodiments, the downstream nucleotide sequence relates to the sequence following the start of transcription. For example, the translation start codon of a gene is located downstream of the transcription start site.
[0017] The term "upstream" refers to a nucleotide sequence located 5' of a reference nucleotide sequence. In certain embodiments, an upstream nucleotide sequence relates to a coding region or a sequence located 5' of the transcription start site. For example, most promoters are located upstream of the transcription start site.
[0018] The term "regulatory element" is used interchangeably herein with the term "regulatory sequence" and is intended to include promoters, enhancers, and other expression control elements (e.g., transcription termination signals such as polyadenylation signals and polyU sequences). Exemplary regulatory elements include transcription promoters, such as, but not limited to, CMV, CMV + intron A, SV40, RSV, HIV-Ltr, elongation factor 1 alpha (EF1α), MMLV-ltr, internal ribosome entry sites (IRES) or P2A peptides to allow translation of multiple genes from a single transcript, metallothionein, transcription enhancer elements, transcription termination signals, polyadenylation sequences, sequences for optimizing translation initiation, and translation termination sequences. It will be understood that the selection of appropriate regulatory elements will depend on the encoded component (e.g., protein or RNA) to be expressed or whether the nucleic acid includes multiple components that require different polymerases or are not intended to be expressed as a fusion protein.
[0019] The term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, a transcription initiation site, a TATA box, and / or a B recognition element and that assists or facilitates the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). A promoter can be produced synthetically or derived from a known or naturally occurring promoter sequence, or from another promoter sequence. A promoter can be proximal or distal to the gene to be transcribed. A promoter can also include chimeric promoters, which contain a combination of two or more heterologous sequences to confer certain properties. Promoters of the present disclosure can include variants of promoter sequences that are similar in composition to, but not identical to, other promoters known or provided herein. Promoters can be classified according to criteria related to the expression pattern of the associated coding or transcribable sequence or gene operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc.
[0020] The term "enhancer" refers to a regulatory element DNA sequence that, when bound by specific proteins called transcription factors, regulates the expression of an associated gene. Enhancers may be located in the introns of a gene or 5' or 3' of the coding sequence of a gene. Enhancers may be proximal to the gene (i.e., within tens or hundreds of base pairs (bp) of the promoter) or distal to the gene (i.e., thousands, hundreds of thousands, or even millions of bp away from the promoter). A single gene may be regulated by more than one enhancer, all of which are contemplated within the scope of this disclosure.
[0021] "Operably linked," in reference to the juxtaposition of two or more components (such as sequence elements), means that the components are positioned so that both components can function normally, with at least one of the components mediating a function exerted on the other components, e.g., at least one of a promoter and a coding sequence.
[0022] A "repressor domain" refers to a polypeptide factor that acts as a regulatory element on DNA that inhibits, suppresses, or blocks transcription of the DNA, resulting in the suppression of gene expression. A repressor domain can be a subunit of a repressor, and individual domains can have different functional properties. In the context of the present disclosure, the linkage of a repressor domain to a catalytically inactive CRISPR protein paired as a ribonucleoprotein complex (RNP) with a guide RNA that has binding affinity for a specific region of a target nucleic acid can prevent transcription from a promoter or otherwise inhibit expression of a gene when bound to the target nucleic acid. Without wishing to be bound by theory, it is believed that transcriptional repressors can function by a variety of mechanisms, including physically blocking the passage of RNA polymerase by steric hindrance, altering the post-translational modification state of the polymerase, modifying the epigenetic state of the nascent RNA, altering the epigenetic state of DNA through methylation, altering the epigenetic state of DNA through histone deacetylation or modulating nucleosome remodeling, or preventing enhancer-promoter interactions, thereby resulting in gene silencing or reduced gene expression levels.
[0023] As used herein, "CRISPR protein without catalytic activity" refers to CRISPR protein that lacks endonuclease activity.Those skilled in the art will understand that CRISPR protein may not have catalytic activity, but can still perform additional protein functions, such as DNA binding.Similarly, "CasX without catalytic activity" refers to CasX protein that lacks endonuclease activity, but can still perform additional protein functions, such as DNA binding.
[0024] As used herein, "recombinant" means that a particular nucleic acid (DNA or RNA) is the product of various combinations of cloning, restriction, and / or ligation steps, resulting in a construct having structural coding or non-coding sequences distinguishable from the endogenous nucleic acid found in natural systems. Generally, DNA sequences encoding structural coding sequences can be assembled from cDNA fragments and short oligonucleotide linkers, or from a series of synthetic oligonucleotides, to provide synthetic nucleic acids capable of expression from recombinant transcription units contained in cells or cell-free transcription and translation systems. Such sequences can be provided in the form of an open reading frame uninterrupted by internal untranslated sequences, or introns, typically present in eukaryotic genes. Genomic DNA containing the relevant sequences can also be used to form recombinant genes or transcription units. Sequences of untranslated DNA can be present 5' or 3' from the open reading frame; such sequences do not interfere with the manipulation or expression of the coding region and may, in fact, act to regulate the production of a desired product by various mechanisms (see "enhancer" and "promoter" above).
[0025] The term "recombinant polynucleotide" or "recombinant nucleic acid" refers to something that does not exist in nature, e.g., something that is created by the artificial combination of two otherwise separated segments of sequence through human intervention. This artificial combination is often accomplished either by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques. Such is usually done to replace codons with redundant codons that encode the same or a conservative amino acid, typically while introducing or removing sequence recognition sites. Alternatively, this is done by joining nucleic acid segments of desired functions together to produce a desired combination of functions. This artificial combination is often accomplished either by chemical synthesis means or by the artificial manipulation of isolated segments of nucleic acid, e.g., by genetic engineering techniques.
[0026] Similarly, the term "recombinant polypeptide" or "recombinant protein" refers to one that does not exist in nature, e.g., a polypeptide or protein that is made by the artificial combination of two otherwise separated segments of amino acid sequence through human intervention. Thus, for example, a protein that includes a heterologous amino acid sequence is recombinant.
[0027] As used herein, the term "contacting" refers to establishing a physical connection between two or more entities. For example, contacting a target nucleic acid sequence with a guide nucleic acid means that the target nucleic acid sequence and the guide nucleic acid are made to share a physical connection, e.g., can hybridize if the sequences share sequence similarity.
[0028] "Dissociation constant" or "K d " is used interchangeably and refers to the affinity between a ligand "L" and a protein "P", i.e., how tightly the ligand binds to a particular protein. This is expressed in the formula K d = [L][P] / [LP], where [P], [L], and [LP] represent the molar concentrations of the protein, ligand, and complex, respectively.
[0029] As used herein, "homology-directed repair" (HDR) refers to a form of DNA repair that occurs during the repair of double-strand breaks in cells. This process requires nucleotide sequence homology and uses a donor template to repair or knock out target DNA, leading to the transfer of genetic information from the donor (e.g., donor template) to the target. If the donor template differs from the target DNA sequence and some or all of the sequence of the donor template is incorporated into the target DNA, homology-directed repair can result in a change in the sequence of the target nucleic acid sequence due to insertion, deletion, or mutation.
[0030] As used herein, "non-homologous end joining" (NHEJ) refers to the repair of double-stranded breaks in DNA by direct ligation of the broken ends to each other without the need for a homologous template (as opposed to homology-directed repair, which requires a homologous sequence to guide the repair). NHEJ often results in the loss (deletion) of nucleotide sequences near the site of the double-stranded break.
[0031] As used herein, "microhomology-mediated end joining" (MMEJ) refers to a mutagenic DSB repair mechanism that does not require a homologous template (as opposed to homology-directed repair, which requires a homologous sequence to guide the repair) and is always associated with a deletion adjacent to the break site. MMEJ often results in the loss (deletion) of nucleotide sequences near the site of the double-strand break.
[0032] A polynucleotide or polypeptide (or protein) has a certain percentage of "sequence similarity" or "sequence identity" with another polynucleotide or polypeptide, meaning that, when aligned, a percentage of bases or amino acids are the same and are in the same relative positions when comparing the two sequences. Sequence similarity (sometimes referred to as percentage similarity, percentage identity, or homology) can be determined in several different ways. To determine sequence similarity, sequences can be aligned using methods and computer programs known in the art, including BLAST, available on the World Wide Web at ncbi.nlm.nih.gov / BLAST. Any convenient method can be used to determine the percentage of complementarity between specific stretches of nucleic acid sequences within a nucleic acid. Exemplary methods include using the BLAST program (Basic Local Alignment Search Tool) and PowerBLAST program (Altschul et al., J. Mol. Biol., 1990, 215, 403-410; Zhang and Madden, Genome Res., 1997, 7, 649-656), or the Gap program (Wisconsin Sequence Analysis Package, Version 8 for Unix, Genetics Computer Group, University Research Park, Madison Wis.), e.g., by using default settings using the algorithm of Smith and Waterman (Adv. Appl. Math., 1981, 2, 482-489).
[0033] The terms "polypeptide" and "protein" are used interchangeably herein to refer to polymeric forms of amino acids of any length, which can include coded and non-coded amino acids, chemically or biochemically modified or derivatized amino acids, and polypeptides with modified peptide backbones. The term includes fusion proteins, including, but not limited to, fusion proteins with heterologous amino acid sequences.
[0034] A "vector" or "expression vector" is a replicon, such as a plasmid, phage, virus, or cosmid, to which another DNA segment, i.e., an "insert," may be attached so as to bring about the replication or expression of the attached segment in a cell.
[0035] The terms "naturally-occurring" or "unmodified" or "wild-type" as used herein as applied to a nucleic acid, polypeptide, cell, or organism refers to a nucleic acid, polypeptide, cell, or organism that is found in nature.
[0036] As used herein, a "mutation" refers to an insertion, deletion, substitution, duplication, or inversion of one or more amino acids or nucleotides compared to a wild-type or reference amino acid sequence or a wild-type or reference nucleotide sequence.
[0037] As used herein, the term "isolated" is intended to describe a polynucleotide, polypeptide, or cell that is in an environment that is different from the environment that the polynucleotide, polypeptide, or cell naturally occurs in. An isolated genetically modified host cell may be present in a mixed population of genetically modified host cells.
[0038] "Host cell," as used herein, means a eukaryotic cell, a prokaryotic cell, or a cell derived from a multicellular organism cultured as a unicellular entity (e.g., in a cell line), which eukaryotic or prokaryotic cell has been used as a recipient of a nucleic acid (e.g., an expression vector), including the progeny of the original cell that has been genetically modified with the nucleic acid. The progeny of a single cell may not necessarily be completely identical in morphology or in genomic or total DNA complement to the original parent due to natural, accidental, or deliberate mutation. A "recombinant host cell" (also called a genetically modified host cell) is a host.
[0039] As used herein, the term "tropism" refers to the preferential entry of a virus-like particle (VLP or XDP) into a particular cell or tissue type and / or preferential interaction with the cell surface that facilitates entry into a particular cell or tissue type, and optionally and preferably subsequent expression (e.g., transcription and optionally translation) of sequences carried into the cell by the VLP or XDP.
[0040] As used herein, the term "pseudotype" or "pseudotyping" refers to a viral envelope protein being replaced with that of another virus having favorable characteristics. For example, HIV can be pseudotyped with the vesicular stomatitis virus G protein (VSV-G) envelope protein (particularly as described herein below), which allows HIV to infect a wider range of cells because the HIV envelope protein primarily targets the virus to CD4+ presenting cells.
[0041] As used herein, the term "targeting agent" refers to a moiety incorporated into the surface of an XDP or VLP that confers targeting to a particular cell or tissue type. Non-limiting examples of targeting agents include glycoproteins, antibody fragments (e.g., scFvs, nanobodies, linear antibodies, etc.), receptors, and ligands for target cell markers.
[0042] "Target cell marker" refers to a molecule expressed by a target cell, including, but not limited to, a cell surface receptor, cytokine receptor, antigen, tumor-associated antigen, glycoprotein, oligonucleotide, enzyme substrate, antigenic determinant, or binding site that may be present on the surface of a target tissue or cell that may serve as a ligand for a tropism factor.
[0043] The term "conservative amino acid substitution" refers to the interchangeability of amino acid residues in proteins with similar side chains. For example, the group of amino acids with aliphatic side chains consists of glycine, alanine, valine, leucine, and isoleucine; the group of amino acids with aliphatic-hydroxyl side chains consists of serine and threonine; the group of amino acids with amide-containing side chains consists of asparagine and glutamine; the group of amino acids with aromatic side chains consists of phenylalanine, tyrosine, and tryptophan; the group of amino acids with basic side chains consists of lysine, arginine, and histidine; and the group of amino acids with sulfur-containing side chains consists of cysteine and methionine. Exemplary conservative amino acid substitution groups are valine-leucine-isoleucine, phenylalanine-tyrosine, lysine-arginine, alanine-valine, and asparagine-glutamine.
[0044] As used herein, "treatment" or "treating" are used interchangeably herein and refer to an approach for obtaining a beneficial or desired result, including, but not limited to, a therapeutic benefit and / or a prophylactic benefit. Therapeutic benefit refers to the eradication or amelioration of the underlying disease or disorder being treated. Therapeutic benefit can also be achieved by the eradication or amelioration of one or more symptoms associated with the underlying disorder, or the improvement of one or more clinical parameters, such that an improvement is observed in a subject, even though the subject may still be afflicted with the underlying disorder.
[0045] The terms "therapeutically effective amount" and "therapeutically effective dose," as used herein, refer to a quantity of a drug or biological substance, alone or as part of a composition, that, when administered in single or repeated doses to a subject, such as a human or experimental animal, is capable of exerting any detectable beneficial effect on any symptom, aspect, measured parameter, or characteristic of a disease state or condition. Such an effect need not be absolute to be beneficial.
[0046] As used herein, "administering" refers to the method of giving a dosage of a compound (e.g., a composition of the present disclosure) or composition (e.g., a pharmaceutical composition) to a subject.
[0047] A "subject" is a mammal, including, but not limited to, domestic animals, non-human primates, humans, rabbits, mice, rats, and other rodents.
[0048] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.
[0049] I. General Methods The practice of the present invention will employ, unless otherwise indicated, conventional techniques of immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and recombinant DNA, which may be practiced in accordance with the teachings of such publications as may be found in Molecular Cloning: A Laboratory Manual, 3rd Ed. (Sambrook et al., Harbor Laboratory Press 2001), Short Protocols in Molecular Biology, 4th Ed. (Ausubel et al. eds., John Wiley & Sons 1999), Protein Methods (Bollag et al., John Wiley & Sons 1996), Nonviral Vectors for Gene Therapy (Wagner et al. eds., Academic Press 1999), Viral Vectors (Kaplift & Loewy eds., Academic Press 1995), Immunology Methods Manual (I. Lefkovits ed., Academic Press 1997), and Cell and Tissue Culture: Laboratory These can be found in standard textbooks such as Procedures in Biotechnology (Doyle & Griffiths, John Wiley & Sons 1998), the disclosures of which are incorporated herein by reference.
[0050] Where a range of values is provided, the endpoints are included, and unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, is included between the upper and lower limit of that range, and between any other stated or intervening value in that stated range. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also included subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0051] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited.
[0052] It must be noted that as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise.
[0053] It will be understood that certain features of the present disclosure that are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. In other cases, various features of the present disclosure that are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable subcombination. All combinations of the embodiments of the present disclosure are specifically embraced by the present disclosure and are intended to be disclosed herein to the same extent as if each and every combination were individually and explicitly disclosed. In addition, all subcombinations of the various embodiments and elements thereof are also specifically embraced by the present disclosure and are intended to be disclosed herein to the same extent as if each and every such subcombination were individually and explicitly disclosed herein.
[0054] II. Repressors and the Epigenetic Long-Term X-Repressor (ELXR) System In a first aspect, the present disclosure provides a gene repressor system comprising a catalytically inactive CRISPR protein linked to one or more repressor domains and one or more guide ribonucleic acids (gRNAs) comprising targeting sequences complementary to a target nucleic acid sequence of a gene targeted for repression, silencing, or downregulation, wherein the system is capable of binding to the target nucleic acid of the gene and repressing transcription of the gene.
[0055] In the context of the present disclosure, and with respect to genes, the terms "suppression," "suppressing," "inhibition of gene expression," "downregulation," and "silencing" are used interchangeably herein to refer to the inhibition or blocking of transcription of a gene or a portion thereof. Gene products that can be suppressed by the system of the present disclosure include mRNA, rRNA, tRNA, structural RNA, or proteins encoded by mRNA. Thus, gene suppression can result in a decrease in the production of a gene product. Examples of gene suppression processes that reduce transcription include, but are not limited to, those that inhibit the formation of a transcription initiation complex, those that reduce the rate of transcription initiation, those that reduce the rate of transcription elongation, those that reduce transcription processivity, and those that antagonize transcription activation (e.g., by blocking the binding of transcription activators). Gene suppression can constitute, for example, the prevention of activation as well as the inhibition of expression below existing levels. Transcriptional suppression includes both reversible and irreversible inactivation of gene transcription. In some embodiments, suppression by the disclosed system includes any detectable decrease in production of the gene product in cells, preferably at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 99%, or any integer decrease therebetween, in gene product production compared to untreated cells or cells treated with an equivalent system containing a non-targeting spacer. Most preferably, gene suppression results in complete inhibition of gene expression such that no gene product is detectable. In some embodiments, suppression of transcription by the disclosed system is maintained for at least about 8 hours, at least about 1 day, at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 3 months, or at least about 6 months, as assessed in in vitro assays, including cell-based assays. In some embodiments, repression of transcription by the gene repressor system of the embodiments is maintained for at least about 1 day, at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 3 months, or at least about 6 months, when assessed in a subject administered a therapeutically effective dose of the system of the embodiments described herein.In some embodiments, gene silencing by the system results in no or minimal detectable off-target methylation or off-target activity when assessed in in vitro assays, hi other embodiments, gene silencing by the system results in no or minimal detectable off-target methylation or off-target activity when assessed in a subject administered a therapeutically effective dose of a system of embodiments described herein.
[0056] In some embodiments, the present disclosure provides a system of a catalytically inactive CRISPR protein linked to one or more repressor domains as a fusion protein and one or more guide ribonucleic acids (gRNAs) for use in suppressing target nucleic acids, including coding and non-coding regions.
[0057] In some embodiments, the present disclosure provides a system of a catalytically inactive CasX (dCasX) protein linked to one or more repressor domains (dXR) as a fusion protein and one or more guide ribonucleic acids (gRNAs), collectively referred to as a dXR:gRNA system, for use in suppressing target nucleic acids, including coding and non-coding regions. The gRNA variant and targeting sequence, as well as the dCasX variant protein and linked repressor domain of any of the embodiments, can form a complex and associate through non-covalent interactions, referred to herein as a ribonucleoprotein (RNP) complex. In some embodiments, the use of a pre-complexed dXR:gRNA RNP provides advantages in delivering system components to cells or target nucleic acids for suppression of the target nucleic acid. In the RNP, the gRNA can provide target specificity to the RNP complex by including a targeting sequence (also referred to as a "spacer") having a nucleotide sequence complementary to that of the target nucleic acid. In RNPs, the dCasX protein and linked repressor domain of a pre-complexed dXR:gRNA provide site-specific activity and are directed to (and further stabilized at) a target site within the target nucleic acid sequence, which is modified by its association with the gRNA. The dCasX protein and linked repressor domain of the RNP complex provide the site-specific activity of the complex, such as binding of the target sequence by the dCasX protein, and the linked repressor domain provides repression activity, either directly or by recruiting other cellular factors.
[0058] Provided herein are dXR:gRNA gene suppression pairs, including dXR variant proteins and linked repressor domains (dXR), gRNA variants, and any combination of dXR and gRNA; compositions comprising or encoding nucleic acids encoding dXR and gRNA; and delivery modalities comprising dXR:gRNA or the encoding nucleic acid. Also provided herein are methods for producing dCasX proteins and linked repressor domains and gRNAs, as well as methods for using CasX and gRNAs, including gene suppression and therapeutic methods. The dCasX protein and linked repressor domain and gRNA components of the dXR:gRNA system, and their characteristics, as well as delivery modalities and methods for using the compositions for gene suppression, downregulation, or silencing, are described more fully below.
[0059] III. dXR:gRNA-based repressor domain fusion proteins In one aspect, the present disclosure relates to a fusion protein comprising one or more repressor domains operably linked to a catalytically inactive CRISPR protein, for example, a catalytically inactive class 2 CRISPR protein. In some embodiments, the catalytically inactive class 2 CRISPR protein is a catalytically inactive class 2, type V CRISPR protein. In some embodiments, the catalytically inactive CRISPR protein comprises a class 2, type II CRISPR / Cas nuclease, for example, Cas9. In other cases, CRISPR proteins without catalytic activity include class 2, type V CRISPR / Cas nucleases such as Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas12l, Cas14, and / or CasΦ. In some embodiments, the catalytically inactive Class 2, Type V CRISPR protein is a catalytically inactive CasX protein (dCasX) selected from the group of SEQ ID NOs: 17-36 and 59353-59358 in Table 4, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, linked to one or more repressor domains to result in a dXR fusion protein. In some embodiments, the catalytically inactive Class 2, Type V CRISPR protein is a catalytically inactive CasX protein (dCasX) selected from the group of SEQ ID NOs: 17-36 and 59353-59358 in Table 4, linked to one or more repressor domains, resulting in a dXR fusion protein.
[0060] In some embodiments, the present disclosure provides a fusion protein comprising a first repressor domain as a fusion protein, wherein the first repressor domain is a Krüppel-associated box (KRAB) domain that can be fused to a catalytically inactive CRISPR protein by a linker peptide disclosed herein. In some embodiments, the present disclosure provides a dXR fusion protein comprising a first repressor domain as a fusion protein, wherein the first repressor domain is a Krüppel-associated box (KRAB) domain that can be fused to dCasX by a linker peptide disclosed herein, resulting in a dXR fusion protein.
[0061] Among repressor domains capable of repressing or silencing genes, the Krüppel-associated box (KRAB) repressor domain is one of the most powerful in the human genome system (Alerasool, N., et al. An efficient KRAB domain for CRISPRi applications. Nat. Methods 17:1093 (2020)). Upon binding of dXR to target nucleic acids, the KRAB domain can recruit additional repressor domains, such as, but not limited to, Trim28 (also known as Kap1 or Tif1-beta), present in approximately 400 human zinc finger protein-based transcription factors, which then assemble protein complexes with chromatin regulators such as CBX5 / HP1α and SETDB1 to induce repression of gene transcription. SETDB1 is a histone methyltransferase that deposits the H3K9me3 mark on histones, a mark of heterochromatin (the complex that acetylates histones and deposits the active H3K9ac mark is displaced). In some cases, DNA methyltransferases (DNMT domains DNMT3A and DNMT3L) are subsequently recruited to deposit methylation marks on DNA, resulting in persistent gene silencing after the complex is no longer bound to the target nucleic acid. Methylation of CpG dinucleotides (CpG) in mammalian cells is catalyzed by the DNA methyltransferases DNMT3a and 3b, which establish DNA methylation patterns, and DNMTL, which maintains the methylation patterns after DNA replication (Zhang, Y., et al. Chromatin methylation activity of Dnmt3a and Dnmt3a / 3L is guided by interaction of the ADD domain with the histone H3 tail. Nucleic Acids Research 38:4246 (2010)).Thus, SETDB1 and DNMT3, recruited by the KRAB domain, act as corepressors of the dXR fusion protein (Tatsumi, D., et al. DNMTs and SETDB1 function as co-repressors in MAX-mediated repression of germ cell-related genes in mouse embryonic stem cells. PLoS ONE 13(11):e0205969(2018)).
[0062] Other repressor domains suitable for inclusion in the dXR of the present disclosure include DNA methyltransferase 3 alpha (DNMT3A or a subdomain thereof), DNMT3A-like protein (DNMT3L or a subdomain thereof), DNA methyltransferase 3 beta (DNMT3B), DNA methyltransferase 1 (DNMT1), Friend of GATA-1 (FOG), Mad mSIN3 interacting domain (SID), enhanced SID (SID4X), nuclear receptor corepressor (NcoR), nuclear effector protein (NuE), KOX1 repression domain, ERF repressor domain (ERD), SRDX repression domain, histone lysine methyltransferases, such as PR / SET domain-containing protein (Pr-SET) 7 / 8, lysine methyltransferase 5B (SUV4-20H1), PR / SET domain-containing protein (Pr-SET) 7 / 8, and lysine methyltransferase 5B (SUV4-20H1). SET domain 2 (RIZ1), histone lysine demethylases, e.g., lysine demethylase 4A (JMJD2A / JHDM3A), lysine demethylase 4B (JMJD2B), lysine demethylase 4C (JMJD2C / GASC1), lysine demethylase 4D (JMJD2D), lysine demethylase 5A (JARID1A / RBP2), lysine demethylase 5B (JARID1B / PLU-1), lysine demethylase 5C (JARID 1C / SMCX), lysine demethylase 5D (JARID1D / SMCY), sirtuin 1 (SIRT1), SIRT2, DNA methylases, e.g., HhaI DNA mc-methyltransferase (M.HhaI), methyltransferase 1 (MET1), histone H3 lysine 9 methyltransferase G9a (G9a), S-adenosyl-L-methionine-dependent methyltransferase superfamily protein (DRM3), DNA cytosine methyltransferase MET2a (ZMET2), methyl-CpG (mCpG) binding domain 2 (meCP2), switch-independent 3 transcriptional regulator family member A (SIN3A), histone deacetylase HDT1 (HDT1), n-terminal truncation of methyl-CpG binding domain protein 2 (MBD2B), nuclear inhibitor of protein phosphatase-1 (NIPP1), GLP, chromomethylase 1 (CMT1), chromomethylase 2 (CMT2), heterochromatin protein 1 (HP1A), mixed lineage leukemia protein-5 (MLL5), histone-lysine N-methyltransferase SETDB1 (SETB1 ), suppressor of variegation 3-9 homolog 1 (SUV39H1), SUV39H2, euchromatin histone lysine methyltransferase 1 (EHMT1), histone-lysine N-methyltransferase EZH1 (EZH1), EZH2, nuclear receptor-associated SET domain protein 1 (NSD1), NSD2, NSD3, ASH1-like histone lysine methyltransferase (ASH1L), tripartite motif-containing 28 (TRIM28), methyltransferase-like 3 (METTL3), METTL4, family 208 member A with sequence similarity (FAM208A), M-phase phosphoprotein 8 (MPHOSPH8), SET domain-containing 2 (SETD2), histone deacetylase 1 (HDAC1), HDAC2, HDAC3, HDAC8, HDAC4, HDAC5, HDAC7, HDAC9, periphilin 1 (PPHLN1), and their subdomains. .
[0063] Human genes encoding KRAB zinc finger proteins include KOX1 / ZNF10, KOX8 / ZNF708, ZNF43, ZNF184, ZNF91, HPF4, HTF10, HTF34, and the sequences of SEQ ID NOs: 355 to 888. In some embodiments, the dXR:gRNA system includes a KRAB transcriptional repressor domain, ZNF343, ZNF10, ZNF337, ZNF334, ZNF215, ZNF519, ZNF485, ZNF214, ZNF33B, ZNF287, ZNF705A, ZNF37A, KRBOX4, ZKSCAN3, ZKSCAN4, ZNF57, ZNF557, ZNF705B, ZNF662, ZNF77, ZNF500, ZNF558, ZNF620, ZNF713, ZNF823, ZNF440, ZNF540, ZNF541, ZNF542, ZNF543, ZNF544, ZNF545, ZNF546, ZNF547, ZNF548, ZNF549, ZNF662, ZNF77, ZNF500, ZNF558, ZNF620, ZNF713, ZNF823, ZNF440, ZNF545, ZNF546, ZNF547, ZNF548, ZNF549 ...549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, ZNF549, , ZNF441, ZNF136, small nuclear ribonucleoprotein polypeptide B and B1 (SNRPB), ZNF735, ZKSCAN2, ZNF619, ZNF627, ZNF333, ATP-binding cassette subfamily A member 11 (ABCA11P), PLD5 pseudogene 1 (PLD5P1), ZNF25, ZNF727, ZNF595, ZNF14, ZNF33A, ZNF101, ZNF253, ZNF56, ZNF720, ZNF85, ZNF66, ZNF722P, ZNF486, ZNF 682, ZNF626, ZNF100, ZNF93, ZKSCAN1, ZNF257, ZNF729, ZNF208, ZNF90, ZNF430, ZNF676, ZNF91, ZNF429, ZNF675, ZNF681, ZNF99, ZNF431, ZNF98, ZNF708, ZNF732, SSX family member 2 (SSX2), ZNF721, ZNF726, ZNF730, ZNF506, ZNF728, ZNF141, ZNF723, ZNF302, ZNF484, SSX2B , ZNF718, ZNF74, ZNF157, ZNF790, ZNF565, ZNF705G, vomeronasal 1 receptor 107 pseudogene (VN1R107P), solute carrier family 27 member 5 (SLC27A5), ZNF737, SSX4, ZNF850, ZNF717, ZNF155, ZNF283, ZNF404, ZNF114, ZNF716, ZNF230, ZNF45, ZNF222, ZNF286A, ZNF624, ZNF223, ZNF284, ZNF790-AS1, ZNF382,ZNF749, ZNF615, ZFP90, ZNF225, ZNF234, ZNF568, ZNF614, ZNF584, ZNF432, ZNF461, ZNF182, ZN F630, ZNF630-AS1, ZNF132, ZNF420, ZNF324B, ZNF616, ZNF471, ZNF227, ZNF324, ZNF860, ZFP28 zinc finger protein (ZFP28), ZNF470, ZNF586, ZNF235, ZNF274, ZNF446, ZFP1, ZIM3, ZNF212, ZNF766, ZNF264, ZNF480, ZNF667, ZNF805, ZNF610, ZNF783, ZNF621, ZNF8-DT, ZNF880, ZNF213-AS1, ZNF213, ZNF263, zinc finger and SCAN domain-containing 32 (ZSCAN32), ZIM2, ZNF597 , ZNF786, KRAB-A domain containing 1 (KRBA1), ZNF460, ZNF8, ZNF875, ZNF543, ZNF133, ZNF229, ZNF528, SSX1, ZNF81, ZNF578, ZNF862, ZNF777, ZNF425, ZNF548, ZNF746, ZNF282, ZNF398, ZNF599, ZNF251, ZNF195, ZNF181, RBAK-RBAKDN readthrough (RBAK-RBAKDN), ZFP37 , RNA, 7SL, cytoplasm526, pseudogene (RN7SL526P), ZNF879, ZNF26, ZSCAN21, ZNF3, ZNF354C, ZNF10, ZNF75D, ZNF426, ZNF561, ZNF562, ZNF 846, ZNF782, ZNF552, ZNF587B, ZNF814, ZNF587, ZNF92, ZNF417, ZNF256, ZNF473, ZFP14, ZFP82, ZNF529, ZNF605, ZFP57, ZNF7 24, ZNF43, ZNF354A, ZNF547, SSX4B, ZNF585A, ZNF585B, ZNF792, ZNF789, ZNF394, ZNF655, ZFP92, ZNF41, ZNF674, ZNF546, ZNF 780B, ZNF699, ZNF177, ZNF560, ZNF583, ZNF707, ZNF808, ZKSCAN5, ZNF137P, ZNF611, ZNF600, ZNF28, ZNF773, ZNF549, ZNF550,ZNF416, ZIK1, ZNF211, ZNF527, ZNF569, ZNF793, ZNF571-AS1, ZNF540, ZNF571, ZNF607, ZNF75A, ZNF205, ZNF175, ZNF268, ZNF354B, ZNF135, ZNF221, ZNF 285, ZNF419, ZNF30, ZNF304, ZNF254, ZNF701, ZNF418, ZNF71, ZNF570, ZNF705E, KRBOX1, ZNF510, ZNF778, PR / SET domain 9 (PRDM9), ZNF248, ZNF845, ZNF52 5, ZNF765, ZNF813, ZNF747, ZNF764, ZNF785, ZNF689, ZNF311, ZNF169, ZNF483, ZNF493, ZNF189, ZNF658, ZNF564, ZNF490, ZNF791, ZNF678, ZNF454, ZNF3 4, ZNF7, ZNF250, ZNF705D, ZNF641, ZNF2, ZNF554, ZNF555, ZNF556, ZNF596, ZNF517, ZNF331, ZNF18, ZNF829, ZNF772, ZNF17, ZNF112, ZNF514, ZNF688, PR DM7, ZNF695, ZNF670-ZNF695, ZNF138, ZNF670, ZNF19, ZNF316, ZNF12, ZNF202, RBAK, ZNF83, ZNF468, ZNF479, ZNF679, ZNF736, ZNF680, ZNF273, ZNF107, ZNF267, ZKSCAN8, ZNF84, ZNF573, ZNF23, ZNF559, ZNF44, ZNF563, ZNF442, ZNF799, ZNF443, ZNF709, ZNF566, ZNF69, ZNF700, ZNF763, ZNF433-AS1, ZNF43 3, ZNF878, ZNF844, ZNF788P, ZNF20, ZNF625-ZNF20, ZNF625, ZNF606, ZNF530, ZNF577, ZNF649, ZNF613, ZNF350, ZNF317, ZNF300, ZNF180, ZNF415, vomeronasal 1 receptor 1 (VN1R1), ZNF266, ZNF738, ZNF445, ZNF852, ZKSCAN7, ZNF660, myosin phosphatase Rho-interacting protein pseudogene 1 (MPRIPP1), ZNF197, ZNF567, ZNF582, ZNF439, ZFP30,ZNF559-ZNF177, ZNF226, ZNF841, ZNF544, ZNF233, ZNF534, ZNF836, ZNF320, KRBA2, ZNF761, ZNF383, ZNF224, ZNF551, ZNF154, ZNF671, ZNF776, ZNF780A, ZNF888, ZNF816-ZNF321P, ZNF3 21P, ZNF816, ZNF347, ZNF665, ZNF677, ZNF160, ZNF184, ZNF140, ZNF589, ZNF891, ZFP69B, ZNF436, pogo transposable element derived from KRAB domain (POGK), ZNF669, ZFP69, ZNF684, ZNF124, and ZNF496, or and sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence variants (in all cases, ZNF = zinc finger protein, KRBOX = KRAB box domain-containing, ZKSCAN = zinc finger with KRAB and SCAN domains, SSX = SSX family member, KRBA = KRAB-A domain-containing, ZFP = zinc finger protein).
[0064] In some embodiments, the gene repressor system comprises a single KRAB domain operably linked to a catalytically inactive CRISPR protein as a fusion protein, wherein the KRAB domain is selected from the group of sequences consisting of SEQ ID NOs: 889-2100 and 2332-33239, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the system comprises a single KRAB domain operably linked to a catalytically inactive CRISPR protein, wherein the KRAB domain is selected from the group of sequences consisting of SEQ ID NOs: 889-2100 and 2332-33239. In some embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a catalytically inactive CRISPR protein, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-59342, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a catalytically inactive CRISPR protein, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57840, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In some embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a catalytically inactive CRISPR, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In certain embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a catalytically inactive Cas9 protein, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.
[0065] In some embodiments, the fusion protein of the system comprises a single KRAB domain operably linked by a peptide linker to a catalytically inactive CRISPR protein, the KRAB domain being one of: a) PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, X6 is L or M, and X7 is G, K, Q, or R); b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is G, K, Q, or R); X is A, G, L, T, or V, X is A, F, or S, X is L or V, X is C, F, H, I, L, or Y, X is A, C, P, Q, or S, X is A, F, G, I, S, or V, X is A, P, S, or T, and X is K or R; c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X is K or R, X is A, D, E, G, N, S, or T, X is D, E, or S, and X is L or R); d) X1X2X3FX4DVX5X6X7FX8X9X 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10 (X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X1 is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-59342. In other embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a CRISPR protein without catalytic activity, wherein the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, and X9 is A, G, I, L, T, or V; X 10 X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-59342. In other embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a CRISPR protein without catalytic activity, wherein the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, and X9 is A, G, I, L, T, or V; X10 X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-57840. In yet other embodiments, the fusion protein of the system comprises a single KRAB domain operably linked to a CRISPR protein without catalytic activity, wherein the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, and X 8は , L or V, X9 is A, G, I, L, T or V, and X 10 X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-57755.
[0066] In some embodiments, the dXR:gRNA system comprises a single KRAB domain operably linked to a catalytically inactive Class 2, Type V CRISPR protein as a fusion protein, wherein the catalytically inactive Class 2, Type V CRISPR protein has a sequence similar to SEQ ID NOs: 17-36 and 59353-59358 in Table 4, or a sequence similar to SEQ ID NOs: 17-36, 59353-59358 ... In some embodiments, the system comprises a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOS: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group of sequences set forth in SEQ ID NOS: 889-2100 and 2332-33239, or sequences having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOs: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-59342, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In some embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOs: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57840, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOS: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOS: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In certain embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX of SEQ ID NO: 18 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In another specific embodiment, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX of SEQ ID NO: 25 listed in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another specific embodiment, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX of SEQ ID NO: 59357 listed in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another specific embodiment, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX of SEQ ID NO: 59358 listed in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.
[0067] In some embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked by a peptide linker to a dCasX selected from the group consisting of SEQ ID NOS: 17-36 and 59353-59358 set forth in Table 4, wherein the KRAB domain is selected from the group consisting of: a) PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, X6 is L or M, and X7 is G, K, Q, or R); b) X1X2X3X4GX5X6X7X8X9 (X1 is , L, or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, X7 is A, F, G, I, S, or V, X8 is A, P, S, or T, and X9 is K or R); c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R); d) X1X2X3FX4DVX5X6X7FX8X9X 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X1 is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-59342. In other embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX, and the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, and X9 is A, G, I, L, T, or V; X 10 X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-59342. In other embodiments, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX, and the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, and X9 is A, G, I, L, T, or V; X 10X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-57840. In yet another embodiment, the dXR fusion protein of the system comprises a single KRAB domain operably linked to dCasX, and the KRAB domain comprises the sequence LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, and X9 is A, G, I, L, T, or V; X 10X1 is A, F, or S), the second sequence motif comprises the sequence FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), and the KRAB domain comprises a sequence selected from the group consisting of SEQ ID NOs: 57746-57755. In certain embodiments, the dXR fusion protein comprises a sequence selected from the group consisting of SEQ ID NOs: 59508-59567 and 59673-60012. In the foregoing embodiments of this paragraph, the dXR fusion protein, when assayed in an in vitro cell assay together with a gRNA targeting the reporter gene, can suppress reporter gene expression to a greater extent than an equivalent fusion protein comprising the ZNF10 KRAB domain (SEQ ID NO: 59626). In some embodiments, the reporter gene is the B2M locus in a eukaryotic cell, such as, but not limited to, HEK293 cells. In some embodiments, reporter gene expression is suppressed by at least about 75%, at least about 80%, at least about 85%, or at least about 90% in an in vitro assay at day 7 of the assay. Exemplary Methods of Measuring Reporter Gene Suppression are provided in the examples, e.g., Example 4.
[0068] In some embodiments, the dXR fusion protein can form a ribonucleoprotein complex (RNP) with the gRNA, and upon binding to a target nucleic acid in a cell in a cellular assay, the dXR:gRNA system can suppress transcription of a gene encoded by the target nucleic acid by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99%. In some embodiments, the dXR fusion protein can form a ribonucleoprotein complex (RNP) with the gRNA, and upon binding to a target nucleic acid in a cell in a cellular assay, the system can suppress transcription of a gene encoded by the target nucleic acid, and suppression of gene transcription is maintained for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 2 months.
[0069] In some embodiments, the present disclosure provides a system comprising first and second repressor domains linked as a fusion protein to a catalytically inactive CRISPR protein and one or more gRNAs comprising a targeting sequence complementary to a target nucleic acid sequence of a gene targeted for silencing, wherein the system is capable of binding to the target nucleic acid in a manner that results in long-term epigenetic modification of the gene, such that repression persists even after the system is no longer present on the target nucleic acid. In some embodiments, the first and second repressor domains are operably linked as a fusion protein, such as to dCasX in embodiments described herein. As used herein, "epigenetic modification" refers to a modification to either DNA or histones associated with DNA, whether direct modification by components of the system or indirectly through the recruitment of one or more additional cellular components, but where the DNA target nucleic acid sequence itself is not edited. For example, while DNMT3A (or its catalytic domain) directly modifies DNA by methylating it, KRAB recruits the KAP-1 / TIF1β corepressor complex, which can act as a potent transcriptional repressor and further recruit factors associated with DNA methylation and the formation of repressive chromatin, such as heterochromatin protein 1 (HP1), histone deacetylases, and histone methyltransferases (Ying, Y., et al. The Kruppel-associated box repressor domain induces reversible and irreversible regulation of endogenous mouse genes by mediating different chromatin states. Nucleic Acids Res. 43(3):1549 (2015)). Together, the first and second repressor components of the system act in concert to produce additive or synergistic effects on the transcriptional silencing of target genes.In some embodiments, the disclosure provides a system comprising first and second repressor domains operably linked to dCasX, wherein the first repressor is the KRAB domain of any of the preceding embodiments, and the second repressor is selected from the group consisting of DNMT3A, DNMT3L, DNMT3B, DNMT1, FOG, SID4X, SID, NcoR, NuE, histone H3 lysine 9 methyltransferase G9a (G9a), methyl-CpG (mCpG) binding domain 2 (meCP2), switch Independence 3 transcription regulator family member A (SIN3A), histone deacetylase HDT1 (HDT1), n-terminal truncation of methyl-CpG-binding domain-containing protein 2 (MBD2B), nuclear inhibitor of protein phosphatase-1 (NIPP1), GLP, heterochromatin protein 1 (HP1A), mixed lineage leukemia protein-5 (MLL5), histone-lysine N-methyltransferase SETDB1 (SETB1), suppressor of variegation 3-9 homolog 1 (SUV39H1) , SUV39H2, euchromatic histone lysine methyltransferase 1 (EHMT1), histone-lysine N-methyltransferase EZH1 (EZH1), EZH2, nuclear receptor-binding SET domain protein 1 (NSD1), NSD2, NSD3, ASH1-like histone lysine methyltransferase (ASH1L), tripartite motif-containing 28 (TRIM28), methyltransferase-like 3 (METTL3), METTL4, and family 208 members with sequence similarity. A (FAM208A), M-phase phosphoprotein 8 (MPHOSPH8), SET domain-containing 2 (SETD2), histone deacetylase 1 (HDAC1), HDAC2, HDAC3, periphilin 1 (PPHLN1), and subdomains thereof (e.g., the DNMT3A catalytic domain and the ATRX-DNMT3-DNMT3L (ADD) domain are subdomains of DNMT3A, and the DNMT3L interaction domain is a subdomain of DNMT3L).
[0070] In some embodiments, the disclosure provides a dXR:gRNA system comprising first and second repressor domains operably linked to dCasX. In some embodiments, the disclosure provides a dXR fusion protein comprising a dCasX selected from the group consisting of SEQ ID NOS:17-36 and 59353-59358 set forth in Table 4, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; wherein the first repressor is selected from the group consisting of SEQ ID NOS:889-2100 and 2332-33239, or sequence variants having at least about 70%, at least about 80%, at least about 99% identity thereto. the second repressor is a DNMT3A domain lacking the regulatory subdomain and maintaining only a catalytic domain selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, wherein the transcriptional repressor domain is linked by a linker peptide sequence to a catalytically inactive CasX protein or another repressor domain. In some embodiments, the dXR containing the DNMT3A catalytic domain affects methylation only at CpG sequences.In certain embodiments, the present disclosure provides a system comprising first and second repressor domains operably linked to dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the first repressor is a KRAB domain selected from the group of sequences consisting of SEQ ID NOs: 57746-59342, or selected from the group consisting of SEQ ID NOs: 57746-57840, or SEQ ID NOs: 57746-57755, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91% identity thereto. the second repressor domain is a DNMT3A catalytic domain selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the transcriptional repressor domain is linked to the catalytically inactive CasX protein or another repressor domain by a linker peptide sequence.In certain embodiments, the disclosure provides a system comprising first and second repressor domains operably linked to dCasX comprising the sequence of SEQ ID NO: 18, wherein the first repressor is a KRAB domain selected from the group consisting of SEQ ID NOs: 57746-59342 or selected from the group consisting of SEQ ID NOs: 57746-57840, and the second repressor domain is a DNMT3A catalytic domain selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, and the transcriptional repressor domain is linked to a catalytically inactive CasX protein or another repressor domain by a linker peptide sequence. In the foregoing embodiments, the fusion protein comprises KRAB, and the second transcriptional repressor domain comprises a DNMT3A catalytic domain. Upon binding of the fusion protein and the gRNA RNP to the target nucleic acid, the system can recruit one or more additional cellular repressor domains, including those listed herein, to affect repression of transcription of the gene encoded by the target nucleic acid, such that upon binding of the fusion protein and the gRNA RNP to the target nucleic acid, transcription of the gene in the cell is suppressed by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99%, or any percentage therebetween, as assayed in an in vitro assay, including a cell-based assay. Most preferably, the epigenetic modification results in complete silencing of gene expression, such that the gene product is undetectable. In some embodiments, the repression of transcription by the system of the embodiments is maintained for at least about 1 day, at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 3 months, or at least about 6 months, as assessed in an in vitro assay.In some embodiments, transcriptional repression by the system of embodiments is maintained for at least about 1 day, at least about 1 week, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 3 months, at least about 6 months, or at least about 1 year, as assessed in a subject administered a therapeutically effective dose of the system of embodiments described herein. In some embodiments, use of the system results in no or minimal detectable off-target methylation or off-target activity, as assessed in an in vitro assay. In some embodiments, use of the system results in less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% off-target methylation or off-target activity in cells. In other embodiments, use of the system results in no or minimal detectable off-target methylation or off-target activity, as assessed in a subject administered a therapeutically effective dose of the system of embodiments described herein.
[0071] In other embodiments, the present disclosure provides a gene repressor system in which the fusion protein comprises first, second, and third transcriptional repressor domains, the third transcriptional repressor domain being different from the first and second transcriptional repressor domains. In some embodiments, the present disclosure provides a dXR:gRNA system in which dXR comprises a KRAB domain of any of the embodiments described herein as the first repressor domain, a DNMT3A catalytic domain as the second repressor domain, and a DNMT3L domain as the third repressor domain. It has been discovered that such dXR fusion proteins, when used in a dXR:gRNA system, result in epigenetic long-term repression of transcription of a target nucleic acid (such fusion proteins are alternatively referred to herein as "ELXR"). In the foregoing, DNMT3L helps maintain methylation patterns after DNA replication. In the foregoing exemplary embodiment, the Class 2 protein without catalytic activity is a Class 2 and a dCasX selected from the group consisting of a V-type CRISPR protein, for example, a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, including sequences of SEQ ID NOs: 17-36 and 59353-59358 set forth in Table 4, and a first repressor domain selected from the group consisting of a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto .... and the second repressor domain is a KRAB repressor domain selected from the group of sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to a sequence of SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%,the DNMT3A catalytic domain of DNMT3A, or a sequence variant thereof, comprising a sequence variant having at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor domain is a DNMT3L interaction domain and is set forth in SEQ ID NO: 59625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein, and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid. In certain embodiments, the disclosure provides a system comprising first, second, and third repressor domains operably linked to dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the first domain is a) PX1X2X3X4X5X6EX7 (wherein X1 is A, D, E, or is N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, X6 is L or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, X7 is A, F, G, I, S, or V, X8 is A, P, S, or T, and X9 isX1 is K or R), c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X; 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10 (X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), wherein the KRAB domain is set forth in SEQ ID NOs: 57746 to 59342, or at least one sequence thereto. and the second repressor domain comprises a sequence selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, or sequences having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of a sequence variant having 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein; and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid.In certain embodiments, the disclosure provides first, second, and third replicators operably linked to dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. The present invention provides a system comprising a PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, X6 is L or M, and X7 is G, K, Q, or R), or b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, and X3 is A, F, or S), X4 is L or V, X5 is C and F, H, I, L, or Y, X6 is A, C, P, Q, or S, X7 is A, F, G, I, S, or V, X8 is A, P, S, or T, and X9 is K or R); c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R); d) X1X2X3FX4DVX5X6X7FX8X9X 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), wherein the KRAB domain is set forth in SEQ ID NOs: 57746 to 57840, or a variant thereof. and the second repressor domain comprises a sequence selected from the group consisting of sequences having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of sequence variants having at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein, and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid.In certain embodiments, the disclosure provides a system comprising first, second, and third repressor domains operably linked to a dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the KRAB domain is a) PX1X2X3X4X5X6EX7, where X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, and X6 is , L, or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, and X7 is A, F, G , I, S, or V, and X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (sequence number 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X. 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S g) FX1DVX2X3X4FX5X6X7EWX8 (SEQ ID NO: 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R); h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain is set forth in SEQ ID NOs: 57746-57755, or at least about 70 amino acids thereto. 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto; the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of sequences having 0%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein; and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid.In another embodiment, the disclosure provides a system comprising first, second, and third repressor domains operably linked to a dCasX comprising the sequence of SEQ ID NO:25, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the KRAB domain is selected from the group consisting of: a) PX1X2X3X4X5X6EX7, where X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, and X6 is X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, X7 is A, F, G, X1 is I, S, or V, X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (sequence number 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X. 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain is set forth in SEQ ID NOs: 57746-57755, or at least about 70 amino acids thereto. 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto; the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of sequences having 0%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein; and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid. In another embodiment, the present disclosure provides a sequence of SEQ ID NO: 59357, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about The system includes first, second, and third repressor domains operably linked to dCasX, the first, second, and third repressor domains comprising sequences having about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the KRAB domain, and the KRAB domain is selected from the group consisting of: a) PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, X5 is H, K, L, Q, R, or W, X6 is L or M, and X7 is G, K, Q, or R); b) X1X2X3X4GX5X6X7X8X9 (X1 is L or or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, X7 is A, F, G, I, S, or V, X8 is A, P, S, or T, and X9 is K or R); c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R); d) X1X2X3FX4DVX5X6X7FX8X9X 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain is set forth in SEQ ID NOs: 57746-57755, or at least about 70 amino acids thereto. 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto; the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of sequences having 0%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein; and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid.In another embodiment, the disclosure provides a system comprising first, second, and third repressor domains operably linked to a dCasX comprising the sequence of SEQ ID NO: 59358, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the KRAB domain is selected from the group consisting of: a) PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N; X2 is L or V; X3 is I or V; X4 is S, T, or F; X5 is H, K, L, Q, R, or W; and X6 is H, K, L, Q, R, or W). is L or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, and X7 is A, F, G , I, S, or V, and X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (sequence number 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X. 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, and X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain is set forth in SEQ ID NOs: 57746-57755, or at least about 70 amino acids thereto. 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identical thereto; the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of sequences having 0%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor is a DNMT3L interacting domain comprising a sequence of SEQ ID NO: 59625 or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein; and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid. In some embodiments of the system, the fusion protein components of the system are configured according to the configuration shown generally in Figure 7. In some embodiments, the repressor domain and dCasX are each operably linked, optionally via a linker, as described herein.In some embodiments, the dXR fusion protein is configured as follows: configuration 1 (NLS-linker4-DNMT3A-linker2-DNMT3L-linker1-linker3-dCasX-linker3-KRAB-NLS), configuration 2 (NLS-linker3-dCasX-linker3-KRAB-NLS-linker1-DNMT3A-linker2-DNMT3L), configuration 3 (NLS-linker3-dCasX-linker1-DNMT3 The DNMT3A-linker2-DNMT3L-linker3-KRAB-NLS constructs have the following configurations from N-terminus to C-terminus (see components in Table 45): configuration 4 (NLS-KRAB-linker3-DNMT3A-linker2-DNMT3L-linker1-dCasX-linker3-NLS), configuration 5 (NLS-DNMT3A-linker2-DNMT3L-linker3-KRAB-linker1-dCasX-linker3-NLS). In some embodiments, the dXR fusion protein comprises a sequence selected from the group consisting of SEQ ID NOs: 59508-59517, 59528-59537, 59548-59557, and 59673-59842, or a sequence having or having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the fusion protein is in configuration 1, 4, or 5.
[0072] In some embodiments, the dXR fusion protein comprises an ADD domain as a fourth domain, the C-terminus of the ADD domain operable relative to the N-terminus of the DNMT3A catalytic domain, a representative configuration of which is shown schematically in Figure 45. In some embodiments, the dXR comprises dCasX and a first, second, third, and fourth repressor, wherein the dXR comprises a sequence selected from the group consisting of SEQ ID NOs: 59508-59567 and 59673-60012, or a sequence having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the fusion protein comprises one or more linker peptides described herein, and wherein the fusion protein is capable of forming an RNP with a gRNA of the system that binds to a target nucleic acid. In some embodiments of the system comprising a dCasX variant and first, second, and third repressor domains, including constructs 1-5, upon binding of the fusion protein and gRNA RNP to the target nucleic acid, the gene is epigenetically modified and transcription of the gene is repressed. In some embodiments, transcription of the gene is repressed by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99% when assayed in an in vitro assay, including a cell-based assay. In some embodiments, repression of gene transcription by the system compositions comprising constructs 1-5 is maintained for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 2 months when assayed in an in vitro assay, including a cell-based assay.In certain embodiments, dXR configurations 4 and 5, when used in a dXR:gRNA system, result in less off-target methylation or off-target activity in in vitro assays compared to configuration 1 (as shown in Figures 7 and 45). In some embodiments, the use of dXR configurations 4 and 5 (as shown in Figures 7 and 45), when used in a dXR:gRNA system, results in less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% off-target methylation or off-target activity in cells.
[0073] In yet other embodiments, the disclosure provides a dXR:gRNA system, wherein the dXR comprises dCasX and first, second, third, and fourth repressor domains. In some embodiments, the dXR comprises a dCasX selected from the group of SEQ ID NOS:17-36 and 59353-59358 set forth in Table 4, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the first repressor domain is a dCasX selected from the group consisting of SEQ ID NOS:17-36 and 59353-59358 set forth in Table 4, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. The KRAB repressor domain is a KRAB repressor domain selected from the group consisting of SEQ ID NOS: 889-2100 and 2332-33239, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the second repressor domain is the first repressor domain is a DNMT3A catalytic domain selected from the group of SEQ ID NOs: 33625-57543 and 59450, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor domain is a DNMT3A catalytic domain selected from the group of SEQ ID NOs: 33625-57543 and 59450, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; 9625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and a fourth domain having a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto;In certain embodiments, the dXR comprises a DNMT3A ADD domain of a sequence variant having at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the first repressor domain is a sequence variant of SEQ ID NO: 18 as set forth in Table 4, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the second repressor domain is a sequence variant of SEQ ID NO: a KRAB repressor domain selected from the group consisting of sequences 889-2100 and 2332-33239, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the third repressor domain is a DNMT3A catalytic domain selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor domain is a DNMT3L interacting domain having a sequence of SEQ ID NO: 59625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto;The fourth domain is the DNMT3A ADD domain of SEQ ID NO: 59452, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the dXR comprises the sequence of SEQ ID NO:25 set forth in Table 4, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the first repressor domain is selected from the group of sequences of SEQ ID NOs:889-2100 and 2332-33239, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the second repressor domain is a DNMT3A catalytic domain selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor domain is a DNMT3A catalytic domain selected from the group consisting of SEQ ID NO: 59625, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%,or a sequence variant having at least about 99% identity thereto, and the fourth domain is the DNMT3A ADD domain of SEQ ID NO: 59452, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the dXR comprises the sequence of SEQ ID NO:59357 set forth in Table 4, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the first repressor domain comprises the sequence of SEQ ID NOs:889-2100 and 2332-33239, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the second repressor domain is a DNMT3A catalytic domain selected from the group of SEQ ID NOs: 33625-57543 and 59450, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor domain is a DNMT3A catalytic domain selected from the group of SEQ ID NO: 59625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%,the fourth domain is a DNMT3A ADD domain of SEQ ID NO: 59452, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In another embodiment, the dXR comprises the sequence of SEQ ID NO: 59358 set forth in Table 4, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, and the first repressor domain comprises the sequence of SEQ ID NOs: 889-2100 and 2332-33239, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97% or at least about 98% identity thereto. the first repressor domain is a KRAB repressor domain selected from the group of sequence variants having at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor domain is a DNMT3A catalytic domain selected from the group of sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto;DNMT3L interactors having sequence variants with at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity; The fourth domain is a DNMT3A ADD domain of SEQ ID NO: 59452, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. The ADD domain is known to have two major functions: 1) allosterically regulating the catalytic activity of DNMT3A by functioning as a methyltransferase autoinhibitory domain, and 2) recognizing unmethylated H3K4 (H3K4me0). Without wishing to be bound by theory, it is believed that the interaction of the ADD domain with the H3K4me0 mark reveals the catalytic site of DNMT3A, thereby recruiting active DNMT3A to chromatin for de novo methylation at these sites. In a surprising result, it has been discovered that the addition of the DNMT3A ADD domain to a dXR construct containing the DNMT3A catalytic domain and the DNMT3L interaction domain significantly enhances repression of the target nucleic acid compared to a dXR construct lacking the ADD domain. Exemplary data for improved repression are provided in the Examples.
[0074] In certain embodiments, the disclosure provides a system comprising first, second, third, and fourth repressor domains operably linked to a dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the KRAB domain is selected from the group consisting of: a) PX1X2X3X4X5X6EX7 (X1 is A, D, E, or N; X2 is L or V; X3 is I or V; X4 is S, T, or F; and X5 is H, K, L, Q, R, or C). or W, X6 is L or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, and X7 is , A, F, G, I, S, or V, and X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10(X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), and the KRAB domain is SEQ ID NOs: 57746 to 59342, or at least about 70%, at least about 80%, at least about 85%, at least about 90% identical thereto; and the second repressor domain comprises a sequence selected from the group consisting of sequences having at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95% or at least about 96% identity thereto. and the third repressor is a DNMT3A catalytic domain sequence selected from the group consisting of a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor is a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the fourth repressor is an ADD domain comprising the sequence of SEQ ID NO: 59452 or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the fusion protein isThe fusion protein, comprising one or more linker peptides described herein, is capable of forming an RNP with the gRNA of the system that binds to a target nucleic acid. The present disclosure provides a system comprising first, second, third, and fourth repressor domains operably linked to dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the first repressor domain is a) PX1X2X3X4X5X6EX7, where X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, and X5 is H, K, L, Q, R, or is W, X6 is L or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, and X7 is A , F, G, I, S, or V, and X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X, 10 X 11 (SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10 (X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, and X6 is I, L, P, or X7 is D, E, K, or V; X8 is E, G, K, P, or R; X9 is A, D, R, G, K, Q, or V; and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12 (X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), wherein the KRAB domain comprises a KRAB domain comprising one or more motifs selected from the group consisting of SEQ ID NOs: 57746-57840, or at least about 70%, at least about 80%, at least about 100%, at least about 15 ... and the second repressor domain comprises a sequence selected from the group consisting of sequences having at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to SEQ ID NOs: 33625-57543 and 59450, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. a DNMT3A catalytic domain comprising a sequence selected from the group consisting of an amino acid sequence of SEQ ID NO: 59625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and a third repressor comprising an amino acid sequence of SEQ ID NO: 59625, or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94% identity thereto. , or a sequence variant having at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the fourth repressor is a DNMT3L interaction domain comprising a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto;or an ADD domain comprising a sequence variant having at least about 99% identity, wherein the fusion protein comprises one or more linker peptides described herein, and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to a target nucleic acid. The present disclosure provides a system comprising first, second, third, and fourth repressor domains operably linked to dCasX comprising the sequence of SEQ ID NO: 18, or a sequence having at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the first repressor domain is a) PX1X2X3X4X5X6EX7, where X1 is A, D, E, or N, X2 is L or V, X3 is I or V, X4 is S, T, or F, and X5 is H, K, L, Q, R, or is W, X6 is L or M, and X7 is G, K, Q, or R), b) X1X2X3X4GX5X6X7X8X9 (X1 is L or V, X2 is A, G, L, T, or V, X3 is A, F, or S, X4 is L or V, X5 is C, F, H, I, L, or Y, X6 is A, C, P, Q, or S, and X7 is A , F, G, I, S, or V, and X8 is A, P, S, or T, and X9 is K or R), c) QX1X2LYRX3VMX4 (SEQ ID NO: 59345) (X1 is K or R, X2 is A, D, E, G, N, S, or T, X3 is D, E, or S, and X4 is L or R), d) X1X2X3FX4DVX5X6X7FX8X9X, 10 X 11(SEQ ID NO: 59346) (X1 is A, L, P, or S, X2 is L or V, X3 is S or T, X4 is A, E, G, K, or R, X5 is A or T, X6 is I or V, X7 is D, E, N, or Y, X8 is S or T, X9 is E, P, Q, R, or W, and X 10 is E or N, and X 11 is E or Q), e)X1X2X3PX4X5X6X7X8X9X 10 (X1 is E, G, or R, X2 is E or K, X3 is A, D, or E, X4 is C or W, X5 is I, K, L, M, T, or V, X6 is I, L, P, or V, X7 is D, E, K, or V, X8 is E, G, K, P, or R, X9 is A, D, R, G, K, Q, or V, and X 10 is D, E, G, I, L, R, S, or V), f)LYX1X2VMX3EX4X5X6X7X8X9X 10 (SEQ ID NO: 59348) (X1 is K or R, X2 is D or E, X3 is L, Q, or R, X4 is N or T, X5 is F or Y, X6 is A, E, G, Q, R, or S, X7 is H, L, or N, X8 is L or V, X9 is A, G, I, L, T, or V, and X 10 is A, F, or S), g) FX1DVX2X3X4FX5X6X7EWX8 (sequence number 59349) (X1 is A, E, G, K, or R, X2 is A, S, or T, X3 is I or V, X4 is D, E, N, or Y, X5 is S or T, X6 is E, L, P, Q, R, or W, X7 is D or E, and X8 is A, E, G, Q, or R), h) X1PX2X3X4X5X6LEX7X8X9X 10 X 11 X 12(X1 is K or R, X2 is A, D, E, or N, X3 is I, L, M, or V, X4 is I or V, X5 is F, S, or T, X6 is H, K, L, Q, R, or W, X7 is K, Q, or R, X8 is E, G, or R, X9 is D, E, or K, and X 10 is A, D, or E, and X 11 is L or P, and X 12 X is C or W), or i) X1LX2X3X4QX5X6 (X1 is C, H, L, Q, or W, X2 is D, G, N, R, or S, X3 is L, P, S, or T, X4 is A, S, or T, X5 is K or R, and X6 is A, D, E, K, N, S, or T), wherein the KRAB domain comprises a KRAB domain comprising one or more motifs selected from the group consisting of SEQ ID NOs: 57746-57755, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 100%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%, at least about 135%, at least about 136%, at least about 137%, at least about 138%, at least about 139%, at least about 140%, the second repressor domain is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of SEQ ID NOs: 33625-57543 and 59450, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; and the third repressor is a DNMT3A catalytic domain comprising a sequence selected from the group consisting of SEQ ID NOs: 59625, or sequence variants having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. the fourth repressor is an ADD domain comprising the sequence of SEQ ID NO: 59452 or a sequence variant having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto; the fusion protein comprises one or more linker peptides described herein, and the fusion protein is capable of forming an RNP with a gRNA of the system that binds to the target nucleic acid. In some embodiments, the dXR fusion protein comprises an ADD domain and a DNMT3A catalytic domain, wherein the C-terminus of the ADD domain is operably linked to the N-terminus of the DNMT3A catalytic domain. In some embodiments, the repressor domain and dCasX are each operably linked, optionally via a linker, as described herein.In some embodiments, the dXR fusion protein is selected from the group consisting of configuration 1 (NLS-ADD-DNMT3A-linker2-DNMT3A-linker1-linker3-dCasX-linker3-KRAB-NLS), configuration 2 (NLS-linker3-dCasX-linker3-KRAB-NLS-linker1-ADD-DNMT3A-linker2-DNMT3L), configuration 3 (NLS-linker3-dCasX-linker1-ADD- In some embodiments of the system, the fusion protein components of the system have the N- to C-terminal configurations of: DNMT3A-linker2-DNMT3L-linker3-KRAB-NLS), configuration 4 (NLS-KRAB-linker3-ADD-DNMT3A-linker2-DNMT3L-linker1-dCasX-linker3-NLS), or configuration 5 (NLS-ADD-DNMT3A-linker2-DNMT3L-linker3-KRAB-linker1-dCasX-linker3-NLS). In some embodiments of the system, the fusion protein components of the system are configured as shown generally in Figure 45. In some embodiments, the dXR fusion protein comprises a sequence selected from the group consisting of SEQ ID NOs: 59518-59526, 59538-59547, 59558-59567, and 59843-60012, or a sequence having or having at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto, wherein the fusion protein is in configuration 1, 4, or 5.
[0075] In some embodiments of the system comprising a dCasX variant and first, second, third, and fourth repressor domains, upon binding of the fusion protein and gRNA RNP to the target nucleic acid, the gene encoded by the target nucleic acid is epigenetically modified and transcription of the gene is repressed. In some embodiments, transcription of the gene is repressed by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99% when assayed in an in vitro assay, including a cell-based assay. In some embodiments, repression of gene transcription by the system compositions is maintained for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 2 weeks, at least about 3 weeks, at least about 1 month, or at least about 2 months when assayed in an in vitro assay, including a cell-based assay. In certain embodiments, dXR configurations 4 and 5, when used in a dXR:gRNA system, result in less off-target methylation or off-target activity in in vitro assays compared to configuration 1. In some embodiments, the use of dXR configurations 4 and 5, when used in a dXR:gRNA system, results in less than about 10%, less than about 9%, less than about 8%, less than about 7%, less than about 6%, less than about 5%, less than about 4%, less than about 3%, less than about 2%, or less than about 1% off-target methylation or off-target activity in cells.
[0076] In some embodiments, the transcriptional repressor domains are linked to each other or to the catalytically inactive CRISPR protein or catalytically inactive Class 2, type V CRISPR protein (e.g., dCasX) within the fusion protein by a linker peptide sequence. In some cases, one or more transcriptional repressor domains are linked at or near the C-terminus of the catalytically inactive Class 2, type V CRISPR protein (e.g., dCasX) by a linker peptide sequence. In other cases, one or more transcriptional repressor domains are linked at or near the N-terminus of the catalytically inactive Class 2, type V CRISPR protein (e.g., dCasX) by a linker peptide sequence. In still other cases, a first transcriptional repressor domain is linked by a linker peptide sequence at or near the C-terminus of a catalytically inactive Class 2, Type V CRISPR protein (e.g., dCasX), and a second, third, and optionally, fourth transcriptional repressor domain are linked at or near the N-terminus of the catalytically inactive Class 2, Type V CRISPR protein. Representative, but non-limiting, configurations are shown schematically in Figures 7, 38, and 45. In the foregoing, the linker peptide may be RS, (G)n (SEQ ID NO: 33240), (GS)n (SEQ ID NO: 33241), (GGS)n (SEQ ID NO: 33242), (GSGGS)n (SEQ ID NO: 33243), (GGSGGS)n (SEQ ID NO: 33244), (GGGS)n (SEQ ID NO: 33245), GGSG (SEQ ID NO: 33246), GGSGG (SEQ ID NO: 33247), GSGSG (SEQ ID NO: 33248), GSGGG (SEQ ID NO: 33249), GGGS G (SEQ ID NO: 33250), GSSSG (SEQ ID NO: 33251), (GP)n (SEQ ID NO: 33252), GPGP (SEQ ID NO: 33253), GGSGGGS (SEQ ID NO: 33254), GSGSGGG (SEQ ID NO: 57628), GCGGTTCCGGCGGAGGAAGC (SEQ ID NO: 57624), GCGGTTCCGGCGGAGGTTCC (SEQ ID NO: 57625), GGATCAGGCTCTGGAGGTGGA (SEQ ID NO: 57627),GGAGGGCCGAGCTCTGGCGCACCCCCACCAAGTGGAGGGTCCTGCCGGGTCCCCAACATCTACTGAAGAAGGCACCAGCGAATCCGCAACGCCCGAGTCAGGCCCTGGTACCTCCACAGAACCATCTGAAGGTAGTGCGCCTGGTTCCCCAGCTGGAAGCCCTACTTCCACCGAAG AAGGCACGTCAACCGAACCAAGTGAAGGATCTGCCCCTGGGACCAGCACTGAACCATCTGAG (SEQ ID NO: 57620), SSGNSNANSRGPSFSSGLVPLSLRGSH (SEQ ID NO: 57623), GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTE PSEGSAPGTSTEPSE (SEQ ID NO: 57621), TCTAGCGGCAATAGTAACGCTAACAGCCGCGGGCCGAGCTTCAGCAGCGGCCTGGTGCCGTTAAGCTTGCGCGGCAGCCAT (SEQ ID NO: 57622), GGP, PPP, PPAPPA (SEQ ID NO: 33255), PPPGPPP (SEQ ID NO: 33256), PPPG (SEQ ID NO: 33257), PPP(GGGS)n (SEQ ID NO: 33258), (GGGS)nPPP (SEQ ID NO: 33259), AEAAAKEAAAKEAAAKA (SEQ ID NO: 33260), AEAAAKEAAAKA (SEQ ID NO: 33261), SGSETPGTSESATPES (SEQ ID NO: 33262), and TPPKTKRKVEFE (SEQ ID NO: 33263), wherein n is an integer of 1 to 5.
[0077] IV. Guide ribonucleic acid (gRNA) of the system In another aspect, the present disclosure provides a guide ribonucleic acid (gRNA) utilized in the gene repressor system of the present disclosure, which is useful, together with other components of the gene repressor system, in suppressing the transcription of a gene targeted by the design of the gRNA. As a component of the gene repression system, the present disclosure provides a specifically designed gRNA having a targeting sequence (or "spacer") that is complementary to (and therefore capable of hybridizing with) the target nucleic acid, and the gRNA can form a ribonucleoprotein (RNP) complex with a CRISPR protein (e.g., dCasX) without the catalytic activity of the fusion protein. In the case of a dCasX variant with a linked repressor domain employed in the system of the present disclosure, the dCasX variant has specificity for a protospacer adjacent motif (PAM) sequence containing a TC motif in the complementary non-target strand, with the PAM sequence located one nucleotide 5' from the sequence of the non-target strand that is complementary to the target nucleic acid sequence in the target strand of the target nucleic acid. The use of precomplexed RNPs provides advantages in delivering system components to cells or target nucleic acid sequences for the repression of transcription of the target nucleic acid sequence. The dCasX variant protein component of the RNP provides site-specific activity that is directed to (e.g., stabilized at) a target site within the target nucleic acid sequence by its association with a guide RNA that includes a targeting sequence that is complementary to a desired specific location of the target nucleic acid and proximal to a PAM sequence.
[0078] In some embodiments, it is contemplated that multiple gRNAs (e.g., multiple gRNAs) are delivered by the system for repression at different regions of a gene, increasing the efficiency and / or duration of repression, as described more fully below.
[0079] a. Reference gRNA and gRNA variants In designing gRNAs for incorporation into the gene repressor system of the present disclosure, a comprehensive approach called deep mutational evolution (DME), deep mutation scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping was utilized to systematically introduce mutations and variations into the nucleic acid sequence of an initial naturally occurring gRNA ("reference gRNA"), resulting in gRNA variants with improved properties, and then reapplied the approach to the gRNA variants to further evolve and improve the resulting gRNA variants. gRNA variants also include variants containing one or more chemical modifications. The activity of the reference gRNA can be used as a benchmark against which the activity of a gRNA variant can be compared, thereby measuring improvements in function or other characteristics of the gRNA variant. In other embodiments, the reference gRNA or gRNA variant can undergo one or more deliberate, targeted mutations to produce a gRNA variant, e.g., a rationally designed variant.
[0080] The gRNA of the present disclosure comprises two segments: a targeting sequence and a protein-binding segment. The targeting segment of a gRNA comprises a nucleotide sequence (interchangeably referred to as a guide sequence, spacer, targeting factor, or targeting sequence) that is complementary to (and therefore hybridizes with) a specific sequence (target site) within a target nucleic acid sequence (e.g., a target ssRNA, a target ssDNA, a strand of a double-stranded target nucleic acid, etc.), as described more fully below. The targeting sequence of a gRNA is capable of binding to a target nucleic acid sequence, including coding sequences, complements of coding sequences, non-coding sequences, and regulatory elements. The protein-binding segment (or "activator" or "protein-binding sequence") interacts with (e.g., binds to) the dCasX protein as a complex to form an RNP (described more fully below). The protein-binding segment is alternatively referred to herein as a "scaffold," which is composed of several regions, as described more fully below.
[0081] In the case of a dual guide RNA (dgRNA), the targeter portion and the activator portion each have a duplex-forming segment, and the targeter duplex-forming segment and the activator duplex-forming segment are complementary to each other and hybridize to form a double-stranded duplex (the gRNA dsRNA duplex). Herein, the term "targeter" or "targeter RNA" is used to refer to the crRNA-like molecule (crRNA: "CRISPR RNA") of a CasX dual guide RNA (and thus of a CasX single guide RNA when the "activator" and "targeter" are linked together, e.g., by intervening nucleotides). The crRNA has a tracrRNA followed by a 5' region that anneals to the nucleotides of the targeting sequence. Thus, for example, a guide RNA (dgRNA or sgRNA) comprises a guide sequence and a duplex-forming segment of the crRNA, which may also be referred to as a crRNA repeat. The corresponding tracrRNA-like molecule (activator) also contains a duplex-forming stretch of nucleotides that forms the other half of the dsRNA duplex of the protein-binding segment of the guide RNA. Thus, the targeting agent and activator hybridize as a matched pair to form a dual guide RNA, referred to herein as a "dual molecule gRNA" or "dgRNA." Site-specific binding of a target nucleic acid sequence (e.g., genomic DNA) by the dCasX protein and linked repressor domain can occur at one or more positions (e.g., the sequence of the target nucleic acid) determined by base-pairing complementarity between the targeting sequence of the gRNA and the target nucleic acid sequence. Thus, for example, a gRNA of the present disclosure can have complementarity to, and thus hybridize with, a target nucleic acid adjacent to a TC PAM motif or sequence complementary to a PAM sequence, e.g., ATC, CTC, GTC, or TTC. Because the targeting sequence of the guide sequence hybridizes with the sequence of the target nucleic acid sequence, the targeting sequence can be modified by the user to hybridize with a specific target nucleic acid sequence, as long as the position of the PAM sequence is taken into consideration.In other embodiments, the activator and targeting element of the gRNA are covalently linked to each other (rather than hybridized to each other) and comprise a single molecule, referred to herein as a "single-molecule gRNA," "single-molecule guide RNA," "single guide RNA," "single-molecule guide RNA," "sgRNA," or "single-molecule guide RNA."
[0082] Collectively, the assembled gRNA of the present disclosure comprises four distinct regions or domains: an RNA triplex, a scaffold stem, an extension stem, and a targeting sequence that, in embodiments of the present disclosure, is specific to the target nucleic acid and is located at the 3' end of the gRNA. The RNA triplex, scaffold stem, and extension stem are collectively referred to as the "scaffold" of the gRNA. The foregoing components of the gRNA are described in WO2020 / 247882A1 and WO2022 / 120095, which are incorporated herein by reference.
[0083] b. Targeting sequence In some embodiments of gRNAs of the present disclosure, the extended stem-loop is followed by a region that forms part of a triplex, followed by a targeting sequence (or "spacer") at the 3' end of the gRNA, with the scaffold being that region 5' to the targeting sequence. The targeting sequence targets the CasX ribonucleoprotein holocomplex to a specific region of the target nucleic acid sequence of the gene to be repressed, 3' to the binding of the RNP. Thus, for example, the gRNA targeting sequence of the present disclosure has sequence complementarity to, and can therefore hybridize to, a portion of a gene in a nucleic acid of a eukaryotic cell (e.g., a eukaryotic chromosome, chromosomal sequence, eukaryotic RNA, etc.) as a component of an RNP when the TC PAM motif or any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand sequence complementary to the target sequence. The targeting sequence of a gRNA can be modified to enable the gRNA to target a desired sequence of any desired target nucleic acid sequence, as long as the position of the PAM sequence is taken into consideration. In some embodiments, the PAM motif sequence recognized by the RNP nuclease is TC. In other embodiments, the PAM sequence recognized by the RNP nuclease is NTC. In other embodiments, the PAM sequence recognized by the RNP nuclease is TTC. In other embodiments, the PAM sequence recognized by the RNP nuclease is ATC. In other embodiments, the PAM sequence recognized by the RNP nuclease is CTC. In other embodiments, the PAM sequence recognized by the RNP nuclease is GTC.
[0084] The gene repressor system of the present disclosure can be designed to target any region of or proximal to a gene or region of a gene for which transcription repression is desired. When the entire gene is to be repressed, it is contemplated by the present disclosure to design guides with targeting sequences that encompass the transcription start site (TSS) or are complementary to sequences proximal to it. TSS selection occurs at different positions within the promoter region, depending on the promoter sequence and initiating substrate concentration. Core promoters serve as binding platforms for the transcription machinery, including Pol II and its associated general transcription factors (GTFs) (Haberle, V. et al. Eukaryotic core promoters and the functional basis of transcription initiation (Nat Rev Mol Cell Biol. 19(10):621(2018)). Variability in TSS selection has been proposed to involve DNA "scrunching" and "antiscrunching," characterized by (i) forward and reverse movement of the leading edge of RNA polymerase relative to DNA, rather than the trailing edge, and (ii) expansion and contraction of the transcription bubble. In some embodiments, the target nucleic acid sequence bound by the RNP of the dXR:gRNA system is within 1 kb of the transcription start site (TSS) in a gene. In some embodiments, the target nucleic acid sequence bound by the RNP of the system is within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, or 300 bp upstream of the TSS of a gene. In some embodiments, the target nucleic acid sequence bound by an RNP of the system is within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, or 1 kb downstream of the TSS of the gene. In some embodiments, the target nucleic acid sequence bound by an RNP of the system is within 500 bp upstream to 500 bp downstream, or within 300 bp upstream to 300 bp downstream of the TSS of the gene. In some embodiments, the target nucleic acid sequence bound by an RNP of the system is within 20 bp, 50 bp, 100 bp, 150 bp, 200 bp, 250 bp, 500 bp, or 1 kb of an enhancer of the gene.In some embodiments, the target nucleic acid sequence bound by the RNP of the disclosed system is within 1 kb 3' to the 5' untranslated region of the gene. In other embodiments, the target nucleic acid sequence bound by the RNP of the system is within the open reading frame of the gene, including introns (if present). In some embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be specific to an exon of the gene of the target nucleic acid. In certain embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be specific to exon 1 of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be specific to an intron of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be specific to an intron-exon junction of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be specific to a regulatory element of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gRNA of the disclosed system is designed to be complementary to a sequence in an intergenic region of the gene of the target nucleic acid. In other embodiments, the targeting sequence of the gRNA of the disclosed system is specific to the junction of an exon, intron, and / or regulatory element of a gene. When the targeting sequence is specific to a regulatory element, such regulatory elements include, but are not limited to, promoter regions, enhancer regions, intergenic regions, 5' untranslated regions (5'UTRs), 3' untranslated regions (3'UTRs), conserved elements, and regions containing cis-regulatory elements. The promoter region is intended to encompass nucleotides within 5 kb from the start of the coding sequence, or in the case of a gene enhancer element or conserved element, it may be thousands, hundreds of thousands, or even millions of bp away from the coding sequence of the gene of the target nucleic acid. In the foregoing, the target is intended to suppress the coding gene of the target so that the gene product is not expressed or is expressed at a lower level in the cell. In some embodiments, when the RNP of the disclosed system binds to the binding site of the target nucleic acid, the system can suppress transcription of the gene 5' to the binding site of the RNP.In other embodiments, when the RNP of the system binds to the binding site of the target nucleic acid, the system is capable of repressing transcription of a gene 3' to the binding site of the RNP. In some embodiments, when the RNP of the system binds to the binding site of the target nucleic acid, the system is capable of repressing transcription of the gene by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least 99% greater than the untreated gene, as assessed in an in vitro assay. In some embodiments, when the RNP of the system binds to the binding site of the target nucleic acid, the system is capable of repressing transcription of the gene for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 2 weeks, at least about 3 weeks, at least about 1 month, at least about 2 months, or at least about 6 months, or at least about 1 year.
[0085] In some embodiments, the targeting sequence of the gRNA of the system has 14-20 contiguous nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, or 20 contiguous nucleotides. In some embodiments, the targeting sequence of the gRNA of the system consists of 20 contiguous nucleotides. In some embodiments, the targeting sequence consists of 19 contiguous nucleotides. In some embodiments, the targeting sequence consists of 18 contiguous nucleotides. In some embodiments, the targeting sequence consists of 17 contiguous nucleotides. In some embodiments, the targeting sequence consists of 16 contiguous nucleotides. In some embodiments, the targeting sequence consists of 15 contiguous nucleotides. In some embodiments, the targeting sequence has 14, 15, 16, 17, 18, 19, or 20 contiguous nucleotides, and the targeting sequence contains 0 to 5, 0 to 4, 0 to 3, or 0 to 2 mismatches to the target nucleic acid sequence, and can retain sufficient binding specificity such that an RNP containing a gRNA that includes the targeting sequence can form a complementary bond to the target nucleic acid.
[0086] In some embodiments, the repressor system dXR:gRNA of the present disclosure comprises a first gRNA and further comprises a second (and optionally a third, fourth, fifth, or more) gRNA, where the second or additional gRNA has a targeting sequence complementary to a different or overlapping portion of the target nucleic acid sequence compared to the targeting sequence of the first gRNA, such that multiple points in the target nucleic acid are targeted, thereby increasing the ability of the system to effectively repress transcription. In such cases, it will be understood that the second or additional gRNA is complexed with an additional copy of dXR. Depending on the selection of the gRNA targeting sequence, defined regions of the target nucleic acid sequence can be repressed using the systems described herein.
[0087] c. gRNA scaffold Except for the targeting sequence region, the remaining region of the gRNA is referred to herein as the scaffold. In some embodiments, a gRNA scaffold is a variant of a reference gRNA in which mutations, insertions, deletions, or domain substitutions are introduced to confer desired properties to the gRNA.
[0088] In some embodiments, the reference gRNA comprises a sequence isolated or derived from Deltaproteobacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Deltaproteobacteria may include ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 6) and ACAUCUGGCGCGUUUAUUCCAUUACUUUGGAGCCAGUCCCAGCGACUAUGUCGUAUGGACGAAGCGCUUAUUUAUCGG (SEQ ID NO: 7). An exemplary crRNA sequence isolated or derived from Deltaproteobacteria may comprise the sequence CCGAUAAGUAAAACGCAUCAAAG (SEQ ID NO: 33271).
[0089] In some embodiments, the reference guide RNA comprises a sequence isolated or derived from Planctomycetes. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary reference tracrRNA sequences isolated or derived from Planctomycetes include UACUGGCGCUUUUAUCUCAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAUGGGUAAAGCGCUUAUUUAUCGGAGA (SEQ ID NO: 8) and
[0090] An exemplary crRNA sequence isolated or derived from Planctomycetes may include the sequence UCUCCGAUAAAUAAGAAGCAUCAAAG (SEQ ID NO: 33272).
[0091] In some embodiments, the reference gRNA comprises a sequence isolated or derived from Candidatus Sungbacteria. In some embodiments, the sequence is a CasX tracrRNA sequence. Exemplary CasX reference tracrRNA sequences isolated or derived from Candidatus Sungbacteria may include the sequences GUUUACACACUCCCUCUCAUAGGGU (SEQ ID NO: 10), GUUUACACACUCCCUCUCAUGAGGU (SEQ ID NO: 11), UUUUACAUACCCCCUCUCAUGGGAU (SEQ ID NO: 12), and GUUUACACACUCCCUCUCAUGGGGG (SEQ ID NO: 13). Table 1 provides sequences of reference gRNA tracr, cr, and scaffold sequences that, in some embodiments, are modified to generate the system gRNA. In some embodiments, the disclosure provides gRNA variant sequences, where the gRNA has a scaffold comprising a sequence with one or more nucleotide modifications relative to a reference gRNA sequence having any one of SEQ ID NOs: 4-16 in Table 1. In embodiments in which the vector comprises a DNA coding sequence for a gRNA, it will be understood that thymine (T) bases can be substituted for uracil (U) bases in any of the gRNA sequence embodiments described herein. [Table 1]
[0092] d. gRNA variants In another aspect, the present disclosure relates to guide ribonucleic acid variants (herein referred to as "gRNA variants") that contain one or more modifications relative to a reference gRNA scaffold. As used herein, "scaffold" refers to all portions of a gRNA that are necessary for gRNA function, excluding spacer sequences.
[0093] In some embodiments, gRNA variants contain one or more nucleotide substitutions, insertions, deletions, or exchanged or replaced regions relative to a reference gRNA sequence of the present disclosure. In some embodiments, mutations can occur in any region of the reference gRNA scaffold to produce a gRNA variant. In some embodiments, the scaffold of the gRNA variant sequence has at least about 50%, at least about 60%, or at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity to the sequence of SEQ ID NO:4 or SEQ ID NO:5.
[0094] In some embodiments, the gRNA variant comprises one or more nucleotide changes in one or more regions of the reference gRNA scaffold that improve the characteristics of the reference gRNA. Exemplary regions include an RNA triplex, a pseudoknot, a scaffold stem-loop, and an extended stem-loop. In some cases, the variant scaffold stem further comprises a bubble. In other cases, the variant scaffold further comprises a triplex loop region. In still other cases, the variant scaffold further comprises a 5' unstructured region. In some embodiments, the gRNA variant scaffold comprises a scaffold stem-loop with at least 60% sequence identity, at least 70% sequence identity, at least 80% sequence identity, at least 90% sequence identity, at least 95% sequence identity, or at least 99% sequence identity to SEQ ID NO: 14. In some embodiments, the gRNA variant scaffold comprises a scaffold stem-loop with at least 60% sequence identity to SEQ ID NO: 14. In other embodiments, the gRNA variant comprises a scaffold stem-loop with the sequence CCAGCGACUAUGUCGUAGUGG (SEQ ID NO: 33273). In other embodiments, the present disclosure comprises a modified extended stem-loop relative to SEQ ID NO: 5, with a C18G substitution, a G55 insertion, a U1 deletion, and the original 6-nt loop and 13 most-loop-proximal base pairs (32 nucleotides total) replaced by a Uvsx hairpin (4-nt loop and 5 loop-proximal base pairs, 14 nucleotides total), and the loop-distal bases of the extended stem converted to a fully base-paired stem contiguous with the new Uvsx hairpin by deletion of A99 and substitution of G65U. In the foregoing embodiment, the gRNA scaffold comprises the sequence ACUGGCGCUUUUAUCUGAUUACUUUGAGAGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAAAG (SEQ ID NO: 33274).
[0095] All gRNA variants that have one or more improved characteristics or add one or more new functions are contemplated within the scope of the present disclosure, provided that the variant gRNA is compared to the reference gRNA described herein. A representative example of such a gRNA variant suitable for a gene repressor system is gRNA variant 174 (SEQ ID NO: 2238). Another representative example of such a gRNA variant suitable for a gene repressor system is gRNA variant 235 (SEQ ID NO: 2292). In some embodiments, the gRNA variant adds a new function to the RNP containing the gRNA variant. In some embodiments, the gRNA variant has an improved characteristic selected from improved stability, improved solubility, improved gRNA transcription, improved resistance to nuclease activity, increased gRNA folding rate, reduced by-product formation during folding, increased productive folding, improved binding affinity for a dXR fusion protein and linked repressor domain, improved binding affinity for a target nucleic acid when complexed with a dXR fusion protein and linked repressor domain, and improved ability to utilize a larger spectrum of one or more PAM sequences, including ATC, CTC, GTC, or TCC, in binding a target nucleic acid when complexed with a dXR fusion protein, and any combination thereof. In some cases, one or more of the improved characteristics of the gRNA variant is improved by at least about 1.1 to about 100,000-fold relative to the reference gRNA of SEQ ID NO:4 or SEQ ID NO:5. In other cases, one or more improved characteristics of the gRNA variant are improved by at least about 1.1, at least about 10, at least about 100, at least about 1000, at least about 10,000, at least about 100,000, or more times relative to the reference gRNA of SEQ ID NO:4 or SEQ ID NO:5.In other cases, one or more of the improved features of the gRNA variant may be about 1.1-100,00 fold, about 1.1-10,00 fold, about 1.1-1,000 fold, about 1.1-500 fold, about 1.1-100 fold, about 1.1-50 fold, about 1.1-20 fold, about 10-100,00 fold, about 10-10,00 fold, about 10-1,000 fold, about 10-500 fold, about 10-100 fold, about 10-50 fold, about 10-20 fold, about 2-70 fold, about 2-50 fold, about 2-30 fold, about 2-20 fold, about 2-10 fold, or about 5-50 fold relative to a reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. , about 5 to 30 times, about 5 to 10 times, about 100 to 100,00 times, about 100 to 10,000 times, about 100 to 1,000 times, about 100 to 500 times, about 500 to 100,00 times, about 500 to 10,000 times, about 500 to 1,000 times, about 500 to 750 times, about 1,000 to 100,00 times, about 10,000 to 100,00 times, about 20 to 500 times, about 20 to 250 times, about 20 to 200 times, about 20 to 100 times, about 20 to 50 times, about 50 to 10,000 times, about 50 to 1,000 times, about 50 to 500 times, about 50 to 200 times, or about 50 to 100 times improvement. In other cases, one or more improved characteristics of the gRNA variant may be about 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2-fold, 3-fold, 4-fold, 5-fold, 6-fold, 7-fold, 8-fold, 9-fold, 10-fold, 11-fold, 12-fold, 13-fold, 14-fold, 15-fold, 16-fold, 17-fold, 18-fold, 19-fold, 20-fold, 25-fold, 30-fold, 40-fold, 45-fold, 50-fold, 55-fold, 60-fold, or greater than the reference gRNA of SEQ ID NO: 4 or SEQ ID NO: 5. , 70x, 80x, 90x, 100x, 110x, 120x, 130x, 140x, 150x, 160x, 170x, 180x, 190x, 200x, 210x, 220x, 230x, 240x, 250x, 260x, 270x, 280x, 290x, 300x, 310x, 320x, 330x, 340x, 350x, 360x, 370x, 380x, 390x, 400x, 425x, 450x, 475x, or 500x improvement.
[0096] In some embodiments, gRNA variants can be generated by subjecting them to one or more mutagenesis methods, such as those described herein below, which may include deep mutational evolution (DME), deep mutation scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping, to generate gRNA variants of the present disclosure. The activity of a reference gRNA can be used as a benchmark against which the activity of a gRNA variant can be compared, thereby measuring the improvement in function of the gRNA variant. In other embodiments, the reference gRNA can be subjected to one or more deliberate, targeted mutations, substitutions, or domain swapping to produce a gRNA variant, e.g., a rationally designed variant. Exemplary gRNA variants produced by such methods are shown in Table 2.
[0097] In some embodiments, the gRNA variant comprises one or more modifications compared to a reference guide ribonucleic acid scaffold sequence, the one or more modifications being selected from at least one nucleotide substitution in a region of the reference gRNA, at least one nucleotide deletion in a region of the reference gRNA, at least one nucleotide insertion in a region of the reference gRNA, a substitution of all or a portion of a region of the reference gRNA, a deletion of all or a portion of a region of the reference gRNA, or any combination of the foregoing. In some cases, the modification is a substitution of 1 to 15 contiguous or non-contiguous nucleotides in the reference gRNA in one or more regions. In other cases, the modification is a deletion of 1 to 10 contiguous or non-contiguous nucleotides in the reference gRNA in one or more regions. In other cases, the modification is an insertion of 1 to 10 contiguous or non-contiguous nucleotides in the reference gRNA in one or more regions. In other cases, the modification is a replacement of a scaffold stem-loop or extended stem-loop with an RNA stem-loop sequence from a heterologous RNA source having proximal 5' and 3' ends. In some cases, the gRNA variants of the present disclosure comprise two or more modifications in one region relative to a reference gRNA. In other cases, the gRNA variants of the present disclosure comprise modifications in two or more regions. In other cases, the gRNA variants comprise any combination of the modifications described in this paragraph.
[0098] In some embodiments, a 5' G is added to a gRNA variant sequence for in vivo expression relative to a reference gRNA because transcription from the U6 promoter is more efficient and more consistent with respect to the start site when the +1 nucleotide is a G. In other embodiments, two 5' Gs are added to generate a gRNA variant sequence for in vitro transcription to increase production efficiency because T7 polymerase strongly prefers a G at the +1 position and a purine at the +2 position. In some cases, a 5' G base is added to a reference scaffold in Table 1. In other cases, a 5' G base is added to a variant scaffold in Table 2.
[0099] Table 2 provides exemplary gRNA variant scaffold sequences. In some embodiments, the gRNA variant scaffold comprises any one of the sequences listed in Table 2, or a sequence having at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In embodiments in which the vector comprises a DNA coding sequence for a gRNA, it will be understood that thymine (T) bases can be substituted for uracil (U) bases in any of the gRNA sequence embodiments described herein. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9]
[0100] In some embodiments, the gRNA variant of the gene repressor system comprises the sequence of any one of SEQ ID NOs: 2238-2331, 57544-57589, and 59352 listed in Table 2. In some embodiments, the gRNA variant comprises the sequence of any one of SEQ ID NOs: 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, the gRNA variant comprises the sequence of any one of SEQ ID NOs: 2238, 2241, 2244, 2248, 2249, or 2259-2280. In some embodiments, the gRNA variant comprises the sequence of any one of SEQ ID NOs: 2281-2331. In some embodiments, the gRNA variant comprises the sequence of any one of SEQ ID NOs: 57544-57589 and 59352. In some embodiments, the gRNA variant comprises one or more chemical modifications to the sequence.
[0101] Additional representative gRNA variant scaffold sequences for use with the gene repressor system of the present disclosure are included as SEQ ID NOs: 2101-2237.
[0102] e.gRNA316 Guide scaffolds can be produced by several methods, including recombinant or solid-phase RNA synthesis. However, scaffold length can affect manufacturability when using solid-phase RNA synthesis; longer lengths result in increased production costs, decreased purity and yield, and a higher rate of synthesis failure. For use in lipid nanoparticle (LNP) formulations, solid-phase RNA synthesis of the scaffold is preferred to produce the quantities necessary for commercial development. Previous experiments identified gRNA scaffold 235 (SEQ ID NO: 2292) as having enhanced properties relative to gRNA scaffold 174 (SEQ ID NO: 2238), but its increased length made its use for LNP formulation problematic. Therefore, alternative sequences were sought. In some embodiments, the present disclosure provides gRNAs in which the gRNA and linked targeting sequence have sequences of less than about 120 nucleotides, less than about 110 nucleotides, or less than about 100 nucleotides.
[0103] In one embodiment, a scaffold was designed in which the scaffold 235 sequence was modified by a domain swap in which the extended stem loop of scaffold 174 replaced the extended stem loop of scaffold 235, resulting in chimeric RNA scaffold 316 having the sequence ACUGGCGCUUCUAUCUGAUUACUCUGAGCGCCAUCACCAGCGACUAUGUCGUAGUGGGUAAAGCUCCCUCUUCGGAGGGAGCAUCAGAG (SEQ ID NO: 59352), which has 89 nucleotides compared to 99 nucleotides in gRNA scaffold 235. In addition to improved manufacturability, the 316 scaffold was determined to perform comparably or better than gRNA scaffold 174 in editing assays, as described in the Examples. The resulting 316 scaffold had the additional advantage in that the extended stem loop did not contain a CpG motif, a property enhancement described more fully below.
[0104] f. Chemically modified scaffolds In another aspect, the present disclosure relates to a gRNA having chemical modifications.In some embodiments, the chemical modification is the addition of 2'O-methyl groups to one or more nucleotides of the sequence.In some embodiments, the chemical modification is the replacement of phosphorothioate bonds between two or more nucleotides of the sequence.
[0105] g. Stem-loop modification In some embodiments, a gRNA variant of a gene repressor system comprises an exogenous extended stem-loop; such differences from the reference gRNA are described below. In some embodiments, the exogenous extended stem-loop has little or no identity with a reference stem-loop region disclosed herein (e.g., SEQ ID NO: 15). In some embodiments, the exogenous stem-loop is at least 10 bp, at least 20 bp, at least 30 bp, at least 40 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1,000 bp. In some embodiments, the gRNA variant comprises an extended stem-loop region comprising at least 10, at least 100, at least 500, or at least 1,000 nucleotides. In some embodiments, the heterologous stem loop increases the stability of the gRNA. In some embodiments, the heterologous RNA stem loop is capable of binding to a protein, an RNA structure, a DNA sequence, or a small molecule.In some embodiments, the exogenous stem-loop region is one or more RNA stem-loops or hairpins, e.g., thermostable RNAs, such as MS2 binding (or tagging) sequences (ACAUGAGGAUCACCCAUGU (SEQ ID NO: 33276), Qβ hairpin (AUGCAUGUCUAAGACAGCAU (SEQ ID NO: 33277)), U1 hairpin II (GGAAUCCAUUGCACUCCGGAUUUCACUAG (SEQ ID NO: 33278)), Uvsx (CCUCUUCGGAGG (SEQ ID NO: 33279)), PP7 (AAGGAGUUUAUAUAU GGAAACCCUU (SEQ ID NO: 33280)), phage replication loop (AGGUGGGACGACCUCUCGGUCGUCCUAUCU (SEQ ID NO: 33281)), kissing loop a (UGCUCGCUCCGUUCGAGCA (SEQ ID NO: 33282)), kissing loop b1 (UGCUCGACGCGUCCUCGAGCA (SEQ ID NO: 33283)), kissing loop b2 (UGCUCGUUUGCGGCUACGAGCA (SEQ ID NO: 33284)), G-quadriplex (M3q (AGGGAGGGAGGGAG AGG (SEQ ID NO: 33285), G-quadriplex telomere basket (GGUUAGGGUUAGGGUUAGG (SEQ ID NO: 33286)), sarcin-ricin loop (CUGCUCAGUACGAGAGGAACCGCAG (SEQ ID NO: 33287)), pseudoknot (UACACUGGGAUCGCUGAAUUAGAGAUCGGCGUCCUUUCAUUCUAUAUACUUUGGAGUUUUAAAAUGUCUCUAAGUACA (SEQ ID NO: 33288)), transactivation response element (TAR) (GGCUCGU GUAGCUCAUUAGCUCCGAGCC (SEQ ID NO: 57741)), iron response element (IRE) CCGUGUGCAUCCGCAGUGUCGGAUCCACGG (SEQ ID NO: 57742)), transactivation response element (TAR) GGCUCGUGUAGCUCAUUAGCUCCGAGCC (SEQ ID NO: 57743)), phage GA hairpin (AAAACAUAAGGAAAACCUAUGUU (SEQ ID NO: 57744)), phage ΛN hairpin (GCCCUGAAGAAGGGC (SEQ ID NO: 57745)), or sequence variants thereof.In some embodiments, one of the aforementioned hairpin sequences is incorporated into the stem-loop to aid in directing incorporation of the gRNA (and associated CasX in the RNP complex) into the budding XDP (described more fully below).
[0106] In some embodiments, sgRNA variants of the gene repressor system of the present disclosure contain one or more additional changes to a previously generated variant, which itself serves as a reference sequence. In some embodiments, the sgRNA variants contain one or more additional changes to the sequence of SEQ ID NO:2238, SEQ ID NO:2239, SEQ ID NO:2240, SEQ ID NO:2241, SEQ ID NO:2241, SEQ ID NO:2274, SEQ ID NO:2275, SEQ ID NO:2279, or SEQ ID NO:59352.
[0107] In some exemplary embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO: 2238 (variant scaffold 174, see Table 2).
[0108] In some exemplary embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO: 2239 (variant scaffold 175, see Table 2).
[0109] In some exemplary embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO: 2275 (variant scaffold 215, see Table 2).
[0110] In some exemplary embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO: 2292 (variant scaffold 235, see Table 2).
[0111] In some exemplary embodiments, the sgRNA variant comprises one or more additional changes to the sequence of SEQ ID NO: 59352 (variant scaffold 316, see Table 2).
[0112] h. Complex formation with dCasX protein In some embodiments, a gRNA variant of the present disclosure has improved affinity for dCasX and the linked repressor domain compared to a reference gRNA, thereby improving its ability to form a ribonucleoprotein (RNP) complex with the dCasX protein and the linked repressor domain. In some embodiments, improved ribonucleoprotein complex formation can improve the efficiency with which functional RNPs are assembled. In some embodiments, greater than 90%, greater than 93%, greater than 95%, greater than 96%, greater than 97%, greater than 98%, or greater than 99% of RNPs comprising the gRNA variant and spacer are capable of binding to a target nucleic acid.
[0113] An exemplary nucleotide change that can improve the ability of a gRNA variant to form a complex with dXR may, in some embodiments, include replacing the scaffold stem with a thermostable stem-loop. Without wishing to be bound by any theory, replacing the scaffold stem with a thermostable stem-loop may increase the overall binding stability of the gRNA variant with dXR. Alternatively or additionally, removing a large section of the stem-loop may alter the folding kinetics of the gRNA variant, for example, by reducing the extent to which the gRNA variant may "entangle" itself, making it easier and more rapid to structurally assemble a functionally folded gRNA. In some embodiments, the selection of the scaffold stem-loop sequence may vary with different spacers utilized in the gRNA. In some embodiments, the scaffold sequence may be tailored to the spacer, and therefore the target sequence. Biochemical assays, including those in the Examples, can be used to assess the binding affinity of dXR to gRNA variants to form RNPs. For example, one skilled in the art can measure the change in the amount of fluorescently labeled gRNA bound to immobilized dXR in response to increasing concentrations of additional unlabeled "cold competitor" gRNA. Alternatively or additionally, the fluorescent signal can be monitored or seen to change as different amounts of fluorescently labeled gRNA are flowed over the immobilized dXR. Alternatively, the ability to form RNPs can be assessed using an in vitro assay for a defined target nucleic acid sequence.
[0114] i. Addition or alteration of gRNA function In some embodiments, the gRNA variants in the system can include larger structural changes that alter the topology of the gRNA variant relative to the reference gRNA, thereby enabling different gRNA functionality. For example, in some embodiments, the gRNA variant swaps the endogenous stem-loop of the reference gRNA scaffold with a stem-loop that can interact with a previously identified stable RNA structure, or a protein or RNA-binding partner, to recruit additional moieties to the dCasX variant or to a specific location, such as the inside of the XDP capsid, that has a binding partner for the RNA structure. The RNA-binding domain can be a retroviral Psi packaging element inserted into the gRNA, or a stem-loop or hairpin (e.g., MS2 hairpin, Qβ hairpin, U1 hairpin II, Uvsx, or PP7 hairpin) with affinity for a protein selected from the group consisting of MS2 coat protein, PP7 coat protein, Qβ coat protein, U1A protein, or phage R-loop, which can promote gRNA binding to the dCasX variant. Similar RNA components with affinity for protein structures incorporated into dCasX variants include kissing loop a, kissing loop_b1, kissing loop_b2, G-quadriplex M3q, G-quadriplex telomeric basket, sarcin-ricin loop, and pseudoknot. In some embodiments, gRNA variants of the present disclosure contain multiple of the foregoing components or multiple copies of the same component.
[0115] V. CRISPR Proteins in Gene Repressor Systems Provided herein are gene repressor systems comprising fusion proteins containing catalytically inactive CRISPR proteins. In some embodiments, the catalytically inactive CRISPR proteins are catalytically inactive class 2 CRISPR proteins. Class 2 systems have a single multidomain effector protein and are further divided into type II, type V, or type VI systems, as described in Makarova, et al. Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nature Rev. Microbiol. 18:67 (2020), incorporated herein by reference. In some embodiments, the catalytically inactive CRISPR protein is a class 2, type II CRISPR / Cas nuclease, such as Cas9. In other cases, the CRISPR without catalytic activity is a class 2, type V CRISPR / Cas nuclease, such as Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas12l, Cas14, and / or CasΦ.
[0116] Unlike type II effectors (e.g., Cas9), which contain two core domains, each responsible for cleaving one strand of target DNA, type V nucleases incorporate an HNH nuclease inserted within a Ruv-C-like nuclease domain sequence. Type V nucleases possess a single RNA-guided RuvC domain-containing effector rather than an HNH domain, and they recognize a T-rich protospacer adjacent motif (PAM) located 5' upstream of the target region on the non-target strand, unlike the Cas9 system, which relies on a G-rich PAM 3' to the target sequence. Unlike Cas9, which generates blunt ends at the proximal site close to the PAM, type V nucleases generate staggered double-strand breaks distal to the PAM sequence. Additionally, when activated by target dsDNA or ssDNA binding in cis, type V nucleases degrade ssDNA in trans. In some embodiments, Type V nucleases utilized in XDP embodiments recognize a 5'TC PAM motif and produce staggered ends cleaved by the RuvC domain. Type V systems (e.g., Cas12) contain only a RuvC-like nuclease domain that cleaves both strands. Type VI (Cas13) is independent of the effectors of Type II and Type V systems, contains two HEPN domains, and targets RNA.
[0117] As used herein, the term "CasX protein" refers to a family of proteins and encompasses all naturally occurring CasX proteins ("reference CasXs") as well as CasX variants that have one or more improved characteristics relative to the naturally occurring reference CasX proteins. In the context of the present disclosure, catalytically inactive CasX variants are prepared from reference CasX and CasX variant proteins; exemplary dCasX variant sequences are set forth in SEQ ID NOS: 17-36 and 59353-59358 in Table 4. CasX and dCasX proteins of the present disclosure comprise at least one of a non-target strand binding (NTSB) domain, a target strand loading (TSL) domain, a helical I domain, a helical II domain, an oligonucleotide binding domain (OBD), and a RuvC domain, the last of which may be modified or deleted to generate catalytically inactive CasX variants, as described more fully below.
[0118] a. Reference CasX protein The present disclosure provides a reference CasX protein that exists in nature and was the starting material for the aforementioned protocols for introducing sequence modifications for the generation of dCasX variants. For example, the reference CasX protein can be isolated from a naturally occurring prokaryote, such as a Deltaproteobacteria, Planctomycetes, or Candidatus Sungbacteria species. The reference CasX protein (sometimes referred to herein as a reference CasX polypeptide) is a type II CRISPR / Cas endonuclease that belongs to the CasX (sometimes referred to as Cas12e) family of proteins that can interact with a guide RNA to form a ribonucleoprotein (RNP) complex.
[0119] In some cases, the reference CasX protein is isolated from or derived from a Deltaproteobacteria having the following sequence: [ka]
[0120] In some cases, the reference CasX protein is isolated or derived from a Planctomycetes having the following sequence: [ka]
[0121] In some cases, the reference CasX protein is isolated from or derived from Candidatus Sungbacteria having the following sequence: [ka]
[0122] b. CasX variant protein without catalytic activity (dCasX variant) In gene repressor systems, the CasX protein is catalytically deficient (dCasX) but retains the ability to bind to target nucleic acids. The present disclosure provides catalytically deficient variants (interchangeably referred to herein as "dCasX variants" or "dCasX variant proteins") that contain at least one modification in at least one domain relative to the catalytically deficient versions of the sequences of SEQ ID NOS: 1-3 (above). Exemplary catalytically deficient CasX proteins contain one or more mutations in the active site of the RuvC domain of the CasX protein. In some embodiments, the catalytically deficient reference CasX protein contains substitutions at residues 672, 769, and / or 935 with reference to SEQ ID NO: 1. In one embodiment, the catalytically deficient reference CasX protein contains substitutions D672A, E769A, and / or D935A with reference to SEQ ID NO: 1. In other embodiments, the reference CasX protein without catalytic activity comprises substitutions at amino acids 659, 756, and / or 922 with reference to SEQ ID NO:2. In some embodiments, the reference CasX protein without catalytic activity comprises substitutions D659A, E756A, and / or D922A with reference to SEQ ID NO:2. Exemplary RuvC domains of dCasX of the present disclosure comprise amino acids 661-824 and 935-986 of SEQ ID NO:1, or amino acids 648-812 and 922-978 of SEQ ID NO:2, with one or more amino acid modifications relative to the RuvC cleavage domain sequence, wherein the dCasX variant exhibits one or more improved characteristics compared to the reference dCasX. In further embodiments, the CasX variant protein without catalytic activity comprises a deletion of all or part of the RuvC domain of the reference CasX protein. It will be understood that the same aforementioned substitutions or deletions can be similarly introduced into any of the CasX variants of SEQ ID NOs: 33352-33624 or 57647-57735 of the present disclosure at the corresponding positions of the starting variant (allowing for any insertions or deletions) to result in a dCasX variant (see, e.g., Table 4 for exemplary sequences).
[0123] In some embodiments, a dCasX variant with a linked repressor domain exhibits at least one improved characteristic compared to a reference dCasX protein with a linked repressor domain configured in a comparable manner, e.g., a catalytically inactive version of any one of the CasX variants set forth in SEQ ID NOs: 33352-33624 or 57647-57735. All variants that improve one or more functions or characteristics of the dCasX variant protein when it has a linked repressor domain compared to the reference dCasX protein with a linked repressor domain described herein are contemplated within the scope of this disclosure. In some embodiments, the modification is a mutation of one or more amino acids of the reference dCasX. In some embodiments, the modification is a mutation of one or more amino acids of a dCasX variant that has been subjected to additional sequence mutations or alterations. In other embodiments, the modification is a substitution of one or more domains of the reference dCasX with one or more domains from a different CasX. In some embodiments, the insertion includes insertion of part or all of a domain from a different CasX protein. Mutations can occur in any one or more domains of a reference dCasX protein or dCasX variant, including, for example, deletion of part or all of one or more domains, or one or more amino acid substitutions, deletions, or insertions in any domain. Domains of a CasX protein include the non-target strand binding (NTSB) domain, the target strand loading (TSL) domain, the helical I domain, the helical II domain, the oligonucleotide binding domain (OBD), and the RuvC DNA cleavage domain, which may further include the subdomains described below. Any change in the amino acid sequence of a reference dCasX protein that results in improved protein characteristics is considered a dCasX variant protein of the present disclosure. For example, a dCasX variant can include one or more amino acid substitutions, insertions, deletions, or swapped domains, or any combination thereof, relative to the reference dCasX protein sequence.
[0124] Suitable mutagenesis methods for generating dCasX variant proteins of the present disclosure may include, for example, deep mutational evolution (DME), deep mutation scanning (DMS), error-prone PCR, cassette mutagenesis, random mutagenesis, staggered extension PCR, gene shuffling, or domain swapping. In some embodiments, dCasX variants are designed, for example, by selecting one or more desired mutations in a reference dCasX. In certain embodiments, the activity of the reference dCasX protein is used as a benchmark against which the activity of one or more dCasX variants is compared, thereby measuring the improvement in the function of the dCasX variants.
[0125] In some embodiments of the dCasX variants described herein, at least one modification comprises (a) a substitution of 1 to 100 consecutive or non-consecutive amino acids in the dCasX variant, (b) a deletion of 1 to 100 consecutive or non-consecutive amino acids in the dCasX variant, (c) an insertion of 1 to 100 consecutive or non-consecutive amino acids in dCasX, or (d) any combination of (a)-(c). In some embodiments, at least one modification comprises (a) a substitution of 5 to 10 consecutive or non-consecutive amino acids in the dCasX variant, (b) a deletion of 1 to 5 consecutive or non-consecutive amino acids in the dCasX variant, (c) an insertion of 1 to 5 consecutive or non-consecutive amino acids in dCasX, or (d) any combination of (a)-(c).
[0126] Any amino acid can be substituted for any other amino acid in the substitutions described herein. Substitutions can be conservative (e.g., a basic amino acid is substituted for another basic amino acid). Substitutions can be non-conservative (e.g., a basic amino acid is substituted for an acidic amino acid, or vice versa). For example, a proline in a reference dCasX protein can be substituted for any of arginine, histidine, lysine, aspartic acid, glutamic acid, serine, threonine, asparagine, glutamine, cysteine, glycine, alanine, isoleucine, leucine, methionine, phenylalanine, tryptophan, tyrosine, or valine to generate a dCasX variant protein of the present disclosure.
[0127] Any permutation of the substitution, insertion, and deletion embodiments described herein can be combined to generate a dCasX variant protein of the present disclosure. For example, a dCasX variant protein can include at least one substitution and at least one deletion relative to a reference dCasX protein sequence, at least one substitution and at least one insertion relative to a reference dCasX protein sequence, at least one insertion and at least one deletion relative to a reference dCasX protein sequence, or at least one substitution, one insertion, and one deletion relative to a reference dCasX protein sequence.
[0128] In some embodiments, the dCasX variant protein comprises between 700 and 1200 amino acids, between 800 and 1100 amino acids, or between 900 and 1000 amino acids.
[0129] The disclosed dCasX and linked repressor domains exhibit enhanced ability to efficiently bind to target nucleic acids when complexed with a gRNA as an RNP, relative to the RNP of a reference dCasX protein and a reference gRNA, by utilizing a PAM TC motif containing a PAM sequence selected from TTC, ATC, GTC, or CTC, where the PAM sequence is located at least one nucleotide 5' to the non-target strand of a protospacer that has identity to the targeting sequence of a gRNA in an assay system, relative to the binding of an RNP containing a reference dCasX protein and a reference gRNA in an equivalent assay system.
[0130] In some embodiments, RNPs comprising a dCasX variant protein and a gRNA having a linked repressor domain of the present disclosure can bind to double-stranded DNA targets with at least 70%, at least 80%, at least 85%, at least 90%, or at least 95% efficiency at a concentration of 20 pM or less. In one embodiment, RNPs comprising a dCasX variant and a gRNA having a linked repressor domain exhibit greater binding of a target sequence in a target nucleic acid compared to RNPs comprising a reference dCasX protein and a reference gRNA having a linked repressor domain in an equivalent assay system, and the PAM sequence of the target nucleic acid is TTC. In another embodiment, RNPs comprising a dCasX variant and a gRNA having a linked repressor domain exhibit greater binding affinity of a target sequence in a target nucleic acid compared to RNPs comprising a reference dCasX protein and a reference gRNA having a linked repressor domain in an equivalent assay system, and the PAM sequence of the target nucleic acid is ATC. In another embodiment, an RNP of a dCasX variant and a gRNA variant having a linked repressor domain exhibits greater binding affinity for a target sequence in a target nucleic acid compared to an RNP comprising a reference dCasX protein having a linked repressor domain and a reference gRNA in an equivalent assay system, wherein the PAM sequence of the target nucleic acid is CTC. In another embodiment, an RNP of a dCasX variant and a gRNA variant having a linked repressor domain exhibits greater binding affinity for a target sequence in a target nucleic acid compared to an RNP comprising a reference dCasX protein having a linked repressor domain and a reference gRNA in an equivalent assay system, wherein the PAM sequence of the target nucleic acid is GTC. In the foregoing embodiment, the increase in binding affinity for one or more PAM sequences is at least 1.5-fold or greater for the PAM sequence compared to the binding affinity of an RNP of any one of the reference dCasX proteins having a linked repressor domain and a gRNA in Table 1 (modified from SEQ ID NOS: 1-3).
[0131] c. dCasX variant proteins with domains derived from multiple source proteins In certain embodiments, the present disclosure provides chimeric dCasX variant proteins for use in dXR systems, which comprise protein domains from two or more naturally occurring CasX proteins or two or more different CasX proteins, such as two or more CasX variant protein sequences described herein. As used herein, "chimeric dCasX protein" refers to a catalytically inactive CasX that comprises at least two domains isolated from or derived from different sources, such as two naturally occurring proteins, which in some embodiments may be isolated from different species. For example, in some embodiments, the chimeric dCasX variant protein comprises a first domain from a first CasX protein and a second domain from a second, different CasX protein. In some embodiments, the first domain may be selected from the group consisting of NTSB, TSL, helix II, helix I-II, helix II, OBD-I, OBD-II, RuvC-I, and RuvC-II domains. In some embodiments, the second domain is selected from the group consisting of NTSB, TSL, helix II, helix I-II, helix II, OBD-I, OBD-II, RuvC-I, and RuvC-II domains, and the second domain is different from the first domain. A chimeric dCasX variant protein may comprise the NTSB, TSL, helix II, helix I-II, helix II, OBD-I, and OBD-II domains from the CasX protein of SEQ ID NO:2 and the RuvC-I and / or RuvC-II domains from the CasX protein of SEQ ID NO:1, or vice versa, with mutations or other sequence modifications introduced to generate catalytically inactive variants with improved variant properties relative to the reference dCasX protein. As an example of the foregoing, a chimeric RuvC domain comprises amino acids 661-824 of SEQ ID NO:1 and amino acids 922-978 of SEQ ID NO:2. As an alternative to the foregoing, the chimeric RuvC domain comprises amino acids 648 to 812 of SEQ ID NO:2 and amino acids 935 to 986 of SEQ ID NO:1.In certain embodiments, dCasX for use in dXR comprises the NTSB domain and helical I-II domain from SEQ ID NO: 1 and the helical II domain from SEQ ID NO: 2, the latter being a chimeric domain. The coordinates of the CasX domains in the reference CasX proteins of SEQ ID NO: 1 and SEQ ID NO: 2 are provided in Table 3 below. [Table 3]
[0132] In some embodiments, the improved characteristics of the dCasX variant are at least about 1.1 to about 100,000-fold improved relative to a reference dCasX protein. In some embodiments, the improved characteristics of the CasX variant are at least about 1.1 to about 10,000-fold improved, at least about 1.1 to about 1,000-fold improved, at least about 1.1 to about 500-fold improved, at least about 1.1 to about 400-fold improved, at least about 1.1 to about 300-fold improved, at least about 1.1 to about 200-fold improved, at least about 1.1 to about 100-fold improved, at least about 1.1 to about 50-fold improved, at least about 1.1 to about 40-fold improved, at least about 1.1 to about 30-fold improved, at least about 1.1 to about 20-fold improved, at least about 1.1 to about 10-fold improved, at least about 1.1 to about 9-fold improved, or ... or at least about 1. The improved characteristics of the dCasX variant are at least about 1.1 to about 8-fold improved, at least about 1.1 to about 7-fold improved, at least about 1.1 to about 6-fold improved, at least about 1.1 to about 5-fold improved, at least about 1.1 to about 4-fold improved, at least about 1.1 to about 3-fold improved, at least about 1.1 to about 2-fold improved, at least about 1.1 to about 1.5-fold improved, at least about 1.5 to about 3-fold improved, at least about 1.5 to about 4-fold improved, at least about 1.5 to about 5-fold improved, at least about 1.5 to about 10-fold improved, at least about 5 to about 10-fold improved, at least about 10 to about 20-fold improved, at least 10 to about 30-fold improved, at least 10 to about 50-fold improved, or at least 10 to about 100-fold improved. In some embodiments, the improved characteristics of the dCasX variant are at least about 10 to about 1000-fold improved relative to the reference dCasX protein.
[0133] In some embodiments, a dCasX variant protein utilized in a gene repressor system of the disclosure comprises one or more of the sequences set forth in SEQ ID NOs: 33352-33624 or 57647-57735, including one or more insertions, substitutions, or deletions described above that inactivate the catalytic domain of the CasX variant to produce the dCasX variant. In some embodiments, a dCasX variant protein utilized in a gene repressor system of the disclosure comprises the sequence set forth in SEQ ID NOs: 17-36 and 59353-59358 listed in Table 4. In some embodiments, a dCasX variant protein consists of the sequence set forth in SEQ ID NOs: 17-36 and 59353-59358 listed in Table 4. In other embodiments, the dCasX variant protein comprises a sequence that is at least 70% identical, at least 75% identical, at least 80% identical, at least 81% identical, at least 82% identical, at least 83% identical, at least 84% identical, at least 85% identical, at least 86% identical, at least 86% identical, at least 87% identical, at least 88% identical, at least 89% identical, at least 89% identical, at least 90% identical, at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, at least 99% identical, or at least 99.5% identical to the sequence of SEQ ID NOs: 17-36 and 59353-59358 set forth in Table 4. [Table 4-1] [Table 4-2] [Table 4-3] [Table 4-4] [Table 4-5] [Table 4-6] [Table 4-7] [Table 4-8] [Table 4-9] [Table 4-10] [Table 4-11]
[0134] d. Affinity for gRNA In some embodiments, dCasX with an attached repressor domain has improved affinity for the gRNA relative to the reference dCasX protein, resulting in the formation of a ribonucleoprotein complex. Increased affinity of dXR for the gRNA can result in, for example, a lower K for the generation of RNP complexes. d This may result in a more stable ribonucleoprotein complex formation in some cases. In some embodiments, the K d is increased by at least about 1.1, at least about 1.2, at least about 1.3, at least about 1.4, at least about 1.5, at least about 1.6, at least about 1.7, at least about 1.8, at least about 1.9, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, or at least about 100-fold relative to a reference dCasX protein. In some embodiments, the dCasX variant has an increased binding affinity for a gRNA by about 1.1 to about 10-fold compared to a catalytically inactive variant of the reference CasX protein of SEQ ID NO:2.
[0135] In some embodiments, increasing the affinity of dCasX with a linked repressor domain for gRNA results in increased stability of the ribonucleoprotein complex when delivered to mammalian cells, including in vivo delivery to a subject. This increased stability can affect the function and availability of the complex in the subject's cells and can result in improved pharmacokinetic properties in the blood when delivered to a subject. In some embodiments, increasing the affinity of dXR and the resulting increased stability of the ribonucleoprotein complex allows for lower doses of dXR to be delivered to a subject or cell while still maintaining the desired activity, such as gene repression in vivo or in vitro. The increased ability to form RNPs and maintain them in a stable form can be assessed using in vitro assays known in the art.
[0136] In some embodiments, the higher affinity (tighter binding) of the dCasX variant protein and linked repressor domain for the gRNA allows for a greater number of repression events when both the dCasX variant protein and the gRNA remain in the RNP complex. This increase in repression events can be assessed using the repression assays described herein.
[0137] Methods for measuring dXR fusion protein binding affinity for gRNA include in vitro methods using purified dXR fusion protein and gRNA. If the gRNA or dXR fusion protein is tagged with a fluorophore, binding affinity for the reference dXR can be measured by fluorescence polarization. Alternatively or additionally, binding affinity can be measured by biolayer interferometry, electrophoretic mobility shift assay (EMSA), or filter binding. Additional standard techniques for quantifying the absolute affinity of an RNA-binding protein, such as the reference dCasX and variant proteins of the present disclosure, for a particular gRNA, such as a reference gRNA and its variants, include, but are not limited to, isothermal calorimetry (ITC) and surface plasmon resonance (SPR), as well as methods in the Examples.
[0138] e. Improved specificity to the target site In some embodiments, a dCasX variant protein having an attached repressor domain has improved specificity for a target nucleic acid sequence relative to a reference dCasX protein having an attached repressor domain. As used herein, "specificity," sometimes referred to as "target specificity," refers to the degree to which a CRISPR / Cas ribonucleoprotein complex binds to off-target sequences that are similar, but not identical, to the target nucleic acid sequence; for example, a dXR RNP with higher specificity exhibits reduced off-target methylation of sequences relative to a reference dXR protein. The specificity of a CRISPR / Cas protein and the reduction of potentially harmful off-target effects can be crucial for achieving an acceptable therapeutic index for use in mammalian subjects.
[0139] In some embodiments, dCasX variant proteins with linked repressor domains have improved specificity for target sites within a target sequence complementary to the targeting sequence of a gRNA. Without wishing to be bound by theory, amino acid changes in the helical I and II domains that increase the specificity of dXR for a target nucleic acid strand may increase the specificity of dXR for the entire target nucleic acid. In some embodiments, amino acid changes that increase the specificity of dXR for a target nucleic acid may also result in a decrease in the affinity of dXR for DNA.
[0140] f. Protospacer and PAM sequences As used herein, a protospacer is defined as a DNA sequence complementary to the targeting sequence of a guide RNA and a DNA sequence complementary to that sequence, referred to as the target strand and non-target strand, respectively. As used herein, a PAM is a nucleotide sequence proximal to the protospacer that, together with the targeting sequence of a gRNA, helps to orient and position CasX on the DNA strand.
[0141] PAM sequences can be degenerate, and particular RNP constructs can have different preferred and tolerated PAM sequences that support different efficiencies of binding and, in the case of catalytically active nucleases, cleavage. By convention, unless otherwise noted, the present disclosure refers to both the PAM sequence and the protospacer sequence, and their orientation relative to the orientation of the non-target strand. This does not imply that the PAM sequence of the non-target strand, rather than the target strand, determines cleavage or is mechanistically involved in target recognition. For example, if reference is made to a TTC PAM, it may actually be the complementary GAA sequence required for target binding, or some combination of nucleotides from both strands. In the case of the CasX proteins disclosed herein, the PAM is located 5' of the protospacer, with a single nucleotide separating the PAM from the first nucleotide of the protospacer. Thus, for reference CasX, TTC PAM should be understood to mean a sequence according to the formula 5'-...NNTTCN(protospacer)NNNNNN...3', where "N" is any DNA nucleotide and "(protospacer)" is a DNA sequence having identity to the targeting sequence of a guide RNA. For CasX variants with extended PAM recognition, TTC, CTC, GTC, or ATC PAM should be understood to mean a sequence according to the following formula: 5'-...NNTTCN(protospacer)NNNNNN...3', 5'-...NNCTCN(protospacer)NNNNNN...3', 5'-...NNGTCN(protospacer)NNNNNN...3', or 5'-...NNATCN(protospacer)NNNNNN...3'. Alternatively, TC PAM should be understood to mean a sequence according to the formula 5'-...NNNTCN(protospacer)NNNNNN...3'.
[0142] In some embodiments, a dCasX variant exhibits greater silencing efficiency and / or binding of a target sequence in a target nucleic acid when any one of the PAM sequences TTC, ATC, GTC, or CTC is located one nucleotide 5' to the non-target strand of a protospacer having identity to the targeting sequence of a gRNA in a cellular assay system compared to the silencing efficiency and / or binding of an RNP containing a reference dCasX protein in an equivalent assay system. In some embodiments, the PAM sequence is TTC. In some embodiments, the PAM sequence is ATC. In some embodiments, the PAM sequence is CTC. In some embodiments, the PAM sequence is GTC.
[0143] g.dCasX fusion protein In some embodiments, the present disclosure provides a dXR fusion protein comprising a heterologous protein.
[0144] In some cases, a heterologous polypeptide (fusion partner) for use with dXR provides for subcellular localization, i.e., the heterologous polypeptide comprises a subcellular localization sequence (e.g., a nuclear localization signal (NLS) for targeting to the nucleus, a sequence for keeping the fusion protein outside the nucleus, e.g., a nuclear export sequence (NES), a sequence for keeping the fusion protein retained in the cytoplasm, a mitochondrial localization signal for targeting to mitochondria, a chloroplast localization signal for targeting to chloroplasts, an ER retention signal, etc.).
[0145] In some cases, the dXR fusion protein comprises (is fused to) a nuclear localization signal (NLS). In some cases, the dXR fusion protein is fused to two or more, three or more, four or more, five or more, six or more, seven or more, eight or more NLSs. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus and / or C-terminus of the dXR fusion protein. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the N-terminus of the dXR fusion protein. In some cases, one or more NLSs (two or more, three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) the C-terminus of the dXR fusion protein. In some cases, one or more NLSs (three or more, four or more, or five or more NLSs) are located at or near (e.g., within 50 amino acids of) both the N-terminus and C-terminus of the dXR fusion protein. In some cases, one NLS is located at the N-terminus and one NLS is located at the C-terminus of the dXR fusion protein. Representative configurations of dXRs with NLSs are shown in Figures 7, 38, and 45.
[0146] In some cases, non-limiting examples of NLSs suitable for use with dXR include the NLS of the SV40 virus large T-antigen having the amino acid sequence PKKKRKV (SEQ ID NO: 33289), an NLS from nucleoplasmin (e.g., the nucleoplasmin bisecting NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 33290)), a c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 33291) or RQRRNELKRSP (SEQ ID NO: 33292), and a hRNPAl M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 33293). NLS, the sequence of the IBB domain from importin-alpha RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 33294), the sequences of the sarcoma T protein VSRKRPRP (SEQ ID NO: 33295) and PPKKARED (SEQ ID NO: 33296), the sequence of human p53 PQPKKKPL (SEQ ID NO: 33297), the sequence of mouse c-abl IV SALIKKKKKMAP (SEQ ID NO: 33298), the sequences of influenza virus NS1 DRLRR (SEQ ID NO: 33299) and PKQKKRK (SEQ ID NO: 33300), the sequence of hepatitis virus delta antigen RKLKKKIKKL (SEQ ID NO: 33301), the sequence of mouse Mxl protein REKKKFLKRR (SEQ ID NO: 33302), the sequence of human poly(ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK ( SEQ ID NO: 33303), steroid hormone receptor (human) glucocorticoid sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 33304), Borna disease virus P protein (BDV-P1) sequence PRPRKIPR (SEQ ID NO: 33305), hepatitis C virus nonstructural protein (HCV-NS5A) sequence PPRKKRTVV (SEQ ID NO: 33306), LEF1 sequence NLSKKKKRKREK (SEQ ID NO: 33307), ORF57 simirae sequence RRPSRPFRKP (SEQ ID NO: 33308), EBV LANA sequence KRPRSPSS (SEQ ID NO: 33309), influenza A protein sequence KRGINDRNFWRGENERKTR (SEQ ID NO: 33310),The sequence PRPPKMARYDN (SEQ ID NO: 33311) of human RNA helicase A (RHA), the sequence KRSFSKAF (SEQ ID NO: 33312) of nucleolar RNA helicase II, the sequence KLKIKRPVK (SEQ ID NO: 33313) of TUS-protein, the sequence PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 33314) related to importin-alpha, the sequence PKTRRRPRRSQRKRPPT (SEQ ID NO: 33315) derived from the Rex protein in HTLV-1, Caenorhabditis The sequence SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 33316) from the EGL-13 protein of C. elegans, as well as the sequences KTRRRPRRSQRKRPPT (SEQ ID NO: 33317), RRKKRRPRRKKRR (SEQ ID NO: 33318), PKKKSRKPKKKSRK (SEQ ID NO: 33319), HKKKHPDASVNFSEFSK (SEQ ID NO: 33320), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 33321), LSPSLSPLLSPSLSPL (SEQ ID NO: 33322), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 33323), PKRGRGRPKRGRGR (SEQ ID NO: 33324), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 33325), PKKKRKVPPPPKKKRKV (SEQ ID NO: 33326), PAKRARRGYKC (SEQ ID NO: 33327) , KLGPRKATGRW (SEQ ID NO: 33328), PRRKREE (SEQ ID NO: 33329), PYRGRKE (SEQ ID NO: 33330), PLRKRPRR (SEQ ID NO: 33331), PLRKRPRRGSPLRKRPRR (SEQ ID NO: 33332), PAAKRVKLDGGKRTADGSEFESPKKKRKV (SEQ ID NO: 33333), PAAKRVKLDGGKRTADGSEFESPKKKRKVGIHGVPAA (SEQ ID NO: 33334), PAAKRVKLDGGKRTADGSEFESPKKKRKVAEAAAKEAAAKEAAAKA (SEQ ID NO: 33335), PAAKRVKLDGGKRTADGSEFESPKKKRKVPG (SEQ ID NO: 33336), KRKGSPERGERKRHW (SEQ ID NO: 33337), KRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 33338),and PKKKRKVGGSKRTADSQHSTPPKTKRKVEFEPKKKRKV (SEQ ID NO: 33339). In some embodiments, one or more NLSs are linked to the dXR or to an adjacent NLS with a linker peptide, and the linker peptide is selected from the group consisting of RS, (G)n (SEQ ID NO: 33240), (GS)n (SEQ ID NO: 33241), (GGS)n (SEQ ID NO: 33242), (GSGGS)n (SEQ ID NO: 33243), (GGSGGS)n (SEQ ID NO: 33244), (GGGS)n (SEQ ID NO: 33245), GGSG (SEQ ID NO: 33246), GGSGG (SEQ ID NO: 33247), (GGGS)n (SEQ ID NO: 33248), (GGGS)n (SEQ ID NO: 33249), (GGSGGS)n (SEQ ID NO: 33250), (GGGS)n (SEQ ID NO: 33251), GGSG (SEQ ID NO: 33252), GGSGG (SEQ ID NO: 33253), (GGGS)n (SEQ ID NO: 33254), (GGGS)n (SEQ ID NO: 33255), GGSG (SEQ ID NO: 33256), GGSGG (SEQ ID NO: 33257), (GGGS)n (SEQ ID NO: 33258), (GGGS)n (SEQ ID NO: 33259), (GGGS)n (SEQ ID NO: 33260), (GGGS)n (SEQ ID NO: 33261), (GGGS)n (SEQ ID NO: 33262), (GGGS)n (SEQ ID NO: 33263), (GGGS)n (SEQ ID 7), GSGSG (SEQ ID NO: 33248), GSGGG (SEQ ID NO: 33249), GGGSG (SEQ ID NO: 33250), GSSSG (SEQ ID NO: 33251), (GP)n (SEQ ID NO: 33252), GPGP (SEQ ID NO: 33253), GGSGGGS (SEQ ID NO: 33254), GSGSGGG (SEQ ID NO: 57628), GCGGTTCCGGCGGAGGAAGC (SEQ ID NO: 57624), GCGGTTCCGGCGGAGGTTCC (SEQ ID NO: 57625), GGATCAGGCTCTGGAGGTGGA (SEQ ID NO: 57627), GGAGGGCCGAGCTCTGGCGCACCCCCACCAAGTGGAGGGTCTCCTGCCGGGTCCCCAACATCTACTGAAGAAGGCACCAGCGAATCCGCAACGCCCGAGTCAGGCCCTGGTACCTCCACAGAACCATCTGAAGGTAGTGCGCCTGGTTCCCCAGCTGGAAGCCCTACTTCCACCGAAGAAGGCACGTCAACCGAACCAAGTGAAGGATCTGCCCCTGGGACCAGCACTGAACCATCTGAG (SEQ ID NO: 57620), SSGNSNANSRGPSFSSGLVPLSLRGSH (SEQ ID NO: 57623), GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSE (SEQ ID NO: 57621),TCTAGCGGCAATAGTAACGCTAACAGCCGCGGGCCGAGCTTCAGCAGCGGCCTGGTGCCGTTAAGCTTGCGCGGCAGCCAT (SEQ ID NO: 57622), GGP, PPP, PPAPPA (SEQ ID NO: 33255), PPPGPPP (SEQ ID NO: 33256), PPPG (SEQ ID NO: 33257), PPP(GGGS)n (SEQ ID NO: 33258), (GGGS)nPPP (SEQ ID NO: 33259), AEAAAKEAAAKEAAAKA (SEQ ID NO: 33260), AEAAAKEAAAKA (SEQ ID NO: 33261), SGSETPGTSESATPES (SEQ ID NO: 33262), and TPPKTKRKVEFE (SEQ ID NO: 33263), wherein n is an integer of 1 to 5.
[0147] Generally, the NLS (or NLSs) have sufficient strength to promote the accumulation of the reference or dCasX variant fusion protein in the nucleus of a eukaryotic cell. Detection of nuclear accumulation can be performed by any suitable technique. For example, a detectable marker can be fused to the reference or dCasX variant fusion protein so that its location within the cell can be visualized. Cell nuclei can also be isolated from cells, and their contents can then be analyzed by any suitable process for detecting proteins, such as immunohistochemistry, Western blot, or enzyme activity assay. Nuclear accumulation can also be determined indirectly.
[0148] In some embodiments, the dXR comprising an N-terminal NLS comprises any one of SEQ ID NOs: 37 to 112 in Tables 5 and 6, and SEQ ID NOs: 59359 to 59432 in Table 7. [Table 5-1] [Table 5-2] [Table 6-1] [Table 6-2] Table 7-1 Table 7-2 Table 7-3
[0149] In some cases, the dXR fusion protein comprises a "protein transduction domain" or PTD (also known as a CPP - cell penetrating peptide), which refers to a protein, polynucleotide, carbohydrate, or organic or inorganic compound that facilitates crossing a lipid bilayer, micelle, cell membrane, organelle membrane, or vesicle membrane. The PTD, attached to another molecule, which can range from small polar molecules to large macromolecules and / or nanoparticles, facilitates the molecule's passage across a membrane, e.g., from the extracellular space to the intracellular space, or from the cytosol into an organelle. In some embodiments, the PTD is covalently attached to the amino terminus of the dXR fusion protein. In some embodiments, the PTD is covalently attached to the carboxyl terminus of the dXR fusion protein. Examples of PTDs include the peptide transduction domain of HIV TAT, including YGRKKRRQRRR (SEQ ID NO: 33340), RKKRRQRR (SEQ ID NO: 33341), YARAAARQARA (SEQ ID NO: 33342), THRLPRRRRRR (SEQ ID NO: 33343), and GGRRARRRRRR (SEQ ID NO: 33344); a polyarginine sequence containing a sufficient number of arginines for direct cell entry (e.g., 3, 4, 5, 6, 7, 8, 9, 10, or 10-50 arginines (SEQ ID NO: 33345)); a VP22 domain (Zender et al. (2002) Cancer Gene Ther. 9(6):489-96); a Drosophila Antennapedia protein transduction domain (Noguchi et al. (2003) Diabetes 52(7):1732-1737); a truncated human calcitonin peptide (Trehin et al. (2003) Cancer Gene Ther. 9(6):489-96); al. (2004) Pharm. Research 21:1248-1256), polylysine (Wender et al. (2000) Proc. Natl. Acad. Sci. USA 97:13003-13008), RRQRRTSKLMKR (SEQ ID NO: 33346), transportan GWTLNSAGYLLGKINLKALAALAKKIL (SEQ ID NO: 33347), KALAWEAKLAKALAKALAKHLAKALAKALKCEA (SEQ ID NO: 33348), and RQIKIWFQNRRMKWKK (SEQ ID NO: 33349).
[0150] In some embodiments, the individual components of dXR can be linked via a linker polypeptide (e.g., one or more linker polypeptides). Linker polypeptides can have any of a variety of amino acid sequences. Proteins can be linked by a generally flexible spacer peptide, although other chemical bonds are not excluded. Suitable linkers include polypeptides between 4 and 40 amino acids in length, or between 4 and 25 amino acids in length. These linkers are generally produced by linking proteins using oligonucleotides encoding synthetic linkers. Peptide linkers with some degree of flexibility can be used. The linking peptide can have virtually any amino acid sequence, keeping in mind that preferred linkers generally have sequences that result in flexible peptides. The use of small amino acids such as glycine, serine, proline, and alanine is useful for generating flexible peptides. The generation of such sequences is routine for those skilled in the art. A variety of different linkers are commercially available and are considered suitable for use. Exemplary linker polypeptides include RS, (G)n (SEQ ID NO: 33240), (GS)n (SEQ ID NO: 33241), (GGS)n (SEQ ID NO: 33242), (GSGGS)n (SEQ ID NO: 33243), (GGSGGS)n (SEQ ID NO: 33244), (GGGS)n (SEQ ID NO: 33245), GGSG (SEQ ID NO: 33246), GGSGG (SEQ ID NO: 33247), GSGSG (SEQ ID NO: 33248), GSGGG (SEQ ID NO: 33249), GGGSG (SEQ ID NO: 33250), GSSSG (SEQ ID NO: 33251), (GP)n (SEQ ID NO: 33252), ), GPGP (SEQ ID NO: 33253), GGSGGGS (SEQ ID NO: 33254), GGP, PPP, PPAPPA (SEQ ID NO: 33255), PPPGPPP (SEQ ID NO: 33256), PPPG (SEQ ID NO: 33257), PPP(GGGS)n (SEQ ID NO: 33258), (GGGS)nPPP (SEQ ID NO: 33259), AEAAAKEAAAKEAAAKA (SEQ ID NO: 33260), AEAAAKEAAAKA (SEQ ID NO: 33261), SGSETPGTSESATPES (SEQ ID NO: 33262), TPPKTKRKVEFE (SEQ ID NO: 33263),GSGSGGG (SEQ ID NO: 57628), GGGCGGTTCCGGCGGAGGAAGC (SEQ ID NO: 57624), GGCGGTTCCGGCGGAGGTTCC (SEQ ID NO: 57625), GGATCAGGCTCTGGAGGTGGA (SEQ ID NO: 57627), GGAGGCCGAGCTCTGGCGCACCCCCACCAAGTGGAGGGT CTCCTGCCGGGTCCCCAACATCTACTGAAGAAGGCACCAGCGAATCCGCAACGCCCGAGTCAGGCCCTGGTACCTCCACAGAACCATCTGAAGGTAGTGCGCCTGGTTCCCAGCTGGAAGCCCTACTTCCACCGAAGAAGGCACGTCAACCGAACCA AGTGAAGGATCTGCCCCTGGGACCAGCACTGAACCATCTGAG (SEQ ID NO: 57620), SSGNSNANSRGPSFSSGLVPLSLRGSH (SEQ ID NO: 57623), GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSE (SEQ ID NO: 57621), and TCTAGCGGCAATAGTAACGCTAACAGCCGCGGGCCGAGCTTCAGCAGCGGCCTGGTGCCGTTAAGCTTGCGCGGCAGCCAT (SEQ ID NO: 57622), where n is an integer from 1 to 5. Those skilled in the art will recognize that the design of peptides conjugated to any of the above elements can include a linker that is fully or partially flexible, such that the linker can include one or more moieties that impart a flexible linker as well as a less flexible structure.
[0151] VI. gRNA and dCRISPR protein-repressor domain gene suppression pair In another aspect, provided herein are compositions comprising a gene suppression pair, wherein the gene suppression pair comprises a catalytically inactive CRISPR protein having one or more linked repressor domains and a guide RNA. In some embodiments, the gene suppression pair comprises a catalytically inactive Class 2 CRISPR-Cas having one or more linked repressor domains. In some embodiments, the gene suppression pair comprises a catalytically inactive Class 2, Type II, Type V, or Type VI CRISPR protein. In some embodiments, the gene suppression pair comprises a catalytically inactive Class 2, Type II CRISPR / Cas protein, such as Cas9. In other cases, the gene suppression pair includes a class 2, type V CRISPR / Cas nuclease such as Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas12l, Cas14, and / or CasΦ protein without catalytic activity.
[0152] In certain embodiments, the gene suppression pair comprises a dCasX variant protein described herein (e.g., any one of the sequences described in Table 4) linked to one or more repressor domains (e.g., any one of the sequences of SEQ ID NOs: 889-2100, 2332-33239, 33625-57543, and 59450), while the guide RNA is a gRNA variant described herein (e.g., SEQ ID NOs: 2238-2331, 57544-57589, and 59352, or a sequence described in Table 2), or a sequence variant having at least 60%, or at least 70%, or at least about 80%, or at least about 90%, or at least about 95%, or at least about 96%, or at least about 97%, or at least about 98%, or at least about 99% sequence identity thereto, and the gRNA comprises a targeting sequence complementary to the target nucleic acid. In some embodiments, the gene suppression pair comprises a dCasX selected from any one of SEQ ID NOs: 17-36 and 59353-59358 listed in Table 4, one or more repressor domains linked to a dCasX selected from any one of SEQ ID NOs: 889-2100, 2332-33239, 33625-57543, and 59450, and a gRNA selected from any one of SEQ ID NOs: 2238, 2239, and 2292. In some embodiments, the gene suppression pair comprises a dCasX selected from any one of SEQ ID NOs: 17-36 and 59353-59358 listed in Table 4, one or more repressor domains linked to a dCasX selected from any one of SEQ ID NOs: 355-888, 33625-57543, and 59450, and a gRNA selected from any one of SEQ ID NOs: 2238, 2239, and 2292, wherein the gRNA comprises a targeting sequence complementary to a target nucleic acid.In some embodiments, the gene suppression pair comprises a dCasX selected from any one of SEQ ID NOs: 17-36 and 59353-59358, one or more repressor domains linked to a dCasX selected from any one of SEQ ID NOs: 355-888, 33625-57543, and 59450, and a gRNA selected from any one of SEQ ID NOs: 2238-2331, 57544-57589, and 59352, wherein the gRNA comprises a targeting sequence complementary to a target nucleic acid.
[0153] In some embodiments, the gene suppression pair includes a dCasX of SEQ ID NO: 18, a KRAB domain sequence of SEQ ID NOs: 57746-57755, a DNMT3A catalytic domain of SEQ ID NOs: 33625-57543 and 59450, a DNMT3L interaction domain of SEQ ID NO: 59625, and a dXR comprising an ADD domain of SEQ ID NO: 59452 (the dXR has configuration 1, 4, or 5 of Figure 45), and a gRNA of SEQ ID NO: 2292 or 59352 (the gRNA includes a targeting sequence complementary to the target nucleic acid).
[0154] In other embodiments, the gene suppression pair includes a dCasX protein selected from any one of SEQ ID NOS: 17-36 and 59353-59358 in Table 4, and one or more repressor domains linked to dCasX; a first gRNA having a targeting sequence (a gRNA variant described herein (e.g., SEQ ID NOS: 2238-2331, 57544-57589, and 59352, or a sequence described in Table 2)); and a second gRNA variant and dXR, wherein the second gRNA variant has a targeting sequence that is complementary to a different or overlapping portion of the target nucleic acid compared to the targeting sequence of the first gRNA.
[0155] In some embodiments, when a gene suppression pair includes both a dCasX variant protein and a linked repressor domain and a gRNA variant described herein, one or more characteristics of the gene suppression pair are improved beyond those that could be achieved by altering the dCasX protein or the gRNA alone. In some embodiments, the dCasX variant protein and the gRNA variant act additively to improve one or more characteristics of the gene suppression pair. In some embodiments, the dCasX variant protein and the gRNA variant act synergistically to improve one or more characteristics of the gene suppression pair. In the foregoing embodiments, the improvement is at least about 2-fold, at least about 5-fold, at least about 10-fold, at least about 50-fold, at least about 100-fold, at least about 500-fold, at least about 1000-fold, at least about 5000-fold, at least about 10,000-fold, or at least about 100,000-fold compared to the characteristics of a reference dCasX protein and reference gRNA pair.
[0156] VII. Vectors In some embodiments, vectors are provided herein that include polynucleotides encoding the catalytically unactivated CRISPR proteins described herein, and linked repressor domains and gRNA variants. In some cases, the vectors are used to express and restore the catalytically unactivated CRISPR proteins (e.g., dXR) and gRNA components of gene suppression pairs or RNPs. In other cases, the vectors are used to deliver the encoding polynucleotides to target cells for suppression of target nucleic acids, as described more fully below.
[0157] In some embodiments, provided herein are polynucleotides encoding the gRNA variants described herein. In some embodiments, the polynucleotide is DNA. In other embodiments, the polynucleotide is RNA. In other embodiments, the polynucleotide is mRNA. In some embodiments, provided herein are vectors comprising polynucleotide sequences encoding the gRNA variants described herein. In some embodiments, vectors comprising the polynucleotide include bacterial plasmids, viral vectors, etc. In some embodiments, the dXR and gRNA variants are encoded on the same vector. In some embodiments, the dXR and gRNA variants are encoded on different vectors.
[0158] In some embodiments, the present disclosure provides vectors comprising nucleotide sequences encoding components of a dXR:gRNA system. For example, in some embodiments, provided herein are recombinant expression vectors comprising: a) a nucleotide sequence encoding a dXR fusion protein; and b) a nucleotide sequence encoding a gRNA variant described herein. In some cases, the nucleotide sequence encoding the dXR fusion protein and / or the nucleotide sequence encoding the gRNA variant is operably linked to a promoter that is operable in a cell type of choice (e.g., prokaryotic cells, eukaryotic cells, plant cells, animal cells, mammalian cells, primate cells, rodent cells, human cells). Suitable promoters for inclusion in vectors are described herein below.
[0159] In some embodiments, the nucleotide sequence encoding the dXR fusion protein is codon-optimized. This type of optimization can involve altering the dCasX-encoding nucleotide sequence to mimic the codon preferences of the intended host organism or cell while still encoding the same protein. Thus, codons can be changed, but the encoded protein remains unchanged. For example, if the intended target cell is a human cell, a human-codon-optimized dCasX variant-encoding nucleotide sequence can be used. As another non-limiting example, if the intended host cell is a mouse cell, a mouse-codon-optimized dCasX variant-encoding nucleotide sequence can be generated. As another non-limiting example, if the intended host cell is a bacterial cell, a bacterial-codon-optimized dXR fusion protein-encoding nucleotide sequence can be generated.
[0160] In some embodiments, the nucleotide sequence encoding the dXR fusion protein is an mRNA designed for incorporation into an LNP. In some embodiments, the mRNA encoding a dXR fusion protein of the disclosure is chemically modified, where the chemical modification is a substitution of one or more uridine nucleotides of the sequence with N1-methyl-pseudouridine. In some embodiments, the mRNA encoding a dXR fusion protein of the disclosure is codon-optimized. In some embodiments, the mRNA encoding a dXR fusion protein of the disclosure comprises one or more sequences selected from the group consisting of SEQ ID NOs: 59584, 59585, 59610, 59611, 59622, and 59623. In some embodiments, an mRNA encoding a dXR fusion protein of the disclosure comprises one or more sequences encoded by a sequence selected from the group consisting of 59444-59449, 59455-59456, 59488-59497, 59568-59583, 59595-59609, and 59612-59621.
[0161] In some embodiments, provided herein are one or more recombinant expression vectors including (i) a nucleotide sequence encoding a gRNA described herein (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell), and (ii) a nucleotide sequence encoding a dXR fusion protein (e.g., operably linked to a promoter operable in a target cell, such as a eukaryotic cell). In some embodiments, the sequences encoding the gRNA and the dXR fusion protein are in different recombinant expression vectors, while in other embodiments, the gRNA and the dXR fusion protein are in the same recombinant expression vector. In some embodiments, either the gRNA in the recombinant expression vector, the dXR fusion protein encoded by the recombinant expression vector, or both, are variants of the reference dCasX protein or gRNA described herein. In the case of a nucleotide sequence encoding a gRNA, the recombinant expression vector can be in vitro transcribed using, for example, T7 promoter regulatory sequences and T7 polymerase to produce the gRNA, which can then be recovered by conventional methods, for example, purification via gel electrophoresis. Once synthesized, the gRNA can be utilized in a gene suppression pair to directly contact the target nucleic acid, or can be introduced into a cell by any of the well-known techniques for introducing nucleic acids into cells (e.g., microinjection, electroporation, transfection, etc.).
[0162] Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc., may be used in the expression vector.
[0163] In some embodiments, the nucleotide sequence encoding the dXR and / or gRNA is operably linked to a regulatory element, e.g., a transcriptional regulatory element, such as a promoter. In some embodiments, the nucleotide sequence encoding the dXR fusion protein is operably linked to a regulatory element, e.g., a transcriptional regulatory element, such as a promoter. In some cases, the promoter is a constitutively active promoter. In some cases, the promoter is a regulatable promoter. In some cases, the promoter is an inducible promoter. In some cases, the promoter is a tissue-specific promoter. In some cases, the promoter is a cell-type-specific promoter. In some cases, the transcriptional regulatory element (e.g., promoter) is functional in a targeted cell type or targeted cell population. For example, in some cases, the transcriptional regulatory element may be functional in eukaryotic cells, e.g., hematopoietic stem cells (e.g., mobilized peripheral blood (mPB) CD34(+) cells, bone marrow (BM) CD34(+) cells, etc.). By transcriptional activation it is intended that transcription be increased by about 10-fold, about 100-fold, more usually about 1000-fold above basal levels in the target cell.
[0164] Non-limiting examples of Pol II promoters include EF-1 alpha, EF-1 alpha core promoter, Jens Tornoe (JeT), promoter from cytomegalovirus (CMV), CMV immediate early (CMVIE), CMV enhancer, herpes simplex virus (HSV) thymidine kinase, early and late simian virus 40 (SV40), SV40 enhancer, long terminal repeats (LTR) from retroviruses, mouse metallothionein-I, adenovirus major late promoter (Ad MLP), CMV promoter full length promoter, minimal CMV promoter, chicken β-actin promoter (CBA), CBA hybrid (CBh), chicken β-actin promoter with cytomegalovirus enhancer (CB7), chicken beta-actin promoter and rabbit beta-globin splice acceptor site fusion (CAG), Rous sarcoma virus (RSV) promoter, HIV-Ltr promoter, hPGK promoter, HSVTK promoter, 7SK promoter, Mini-TK promoter, human synapsin I (SYN) promoter conferring neuron-specific expression, beta-actin promoter, supercore promoter 1 (SCP1), Mecp2 promoter for selective expression in neurons, minimal IL-2 promoter, Rous sarcoma virus enhancer / promoter (single), spleen focus-forming virus long terminal repeat (LTR) promoter, TBG promoter, promoter from human thyroxine-binding globulin gene (liver-specific), PGK promoter, human ubiquitin C promoter (UBC), UCOE promoter (promoter of HNRPA2B1-CBX3), synthetic CAG promoter, histone H2 promoter, histone H3 promoter, U1a1 micronuclear R promoter Examples of promoters include, but are not limited to, the NA promoter (226 nt), U1a1 small nuclear RNA promoter (226 nt), U1b2 small nuclear RNA promoter (246 nt), GUSB promoter, CBh promoter, rhodopsin (Rho) promoter, silencing-prone spleen focus-forming virus (SFFV) promoter, human H1 promoter (H1), POL1 promoter, TTR minimal enhancer / promoter, b-kinesin promoter, mouse mammary tumor virus long terminal repeat (LTR) promoter, human eukaryotic initiation factor 4A (EIF4A1) promoter, ROSA26 promoter, glyceraldehyde 3-phosphate dehydrogenase (GAPDH) promoter, tRNA promoter, and truncated versions and sequence variants of the foregoing. In certain embodiments, the Pol II promoter is EF-1 alpha, and the promoter enhances transfection efficiency, transgene transcription or expression of CRISPR nuclease, the percentage of expression-positive clones, and the copy number of the episomal vector in long-term culture. Pol II promoters are also useful.Non-limiting examples of Pol III promoters include, but are not limited to, U6, mini-U6, U6 truncated promoter, BiH1 (bidirectional H1 promoter), BiU6, Bi7SK, BiH1 (bidirectional U6, 7SK, and H1 promoter), gorilla U6, rhesus U6, human 7SK, human H1 promoter, and truncated versions and sequence variants thereof. In the foregoing embodiments, the Pol III promoter enhances transcription of the gRNA.
[0165] Selection of an appropriate vector and promoter is well within the level of ordinary skill in the art. The expression vector may also contain a ribosome binding site for translation initiation and a transcription terminator. The expression vector may also contain appropriate sequences for amplifying expression. The expression vector may also contain a nucleotide sequence encoding a protein tag (e.g., a 6xHis tag, a hemagglutinin tag, a fluorescent protein, etc.) that can be fused to the dXR fusion protein, thus resulting in a chimeric CasX variant polypeptide.
[0166] The recombinant expression vectors of the present disclosure can also include elements that promote robust expression of the dXR and / or variant gRNAs of the present disclosure. For example, the recombinant expression vectors can include one or more of a polyadenylation signal (poly(A), an intron sequence, or a post-transcriptional regulatory element such as a woodchuck hepatitis post-transcriptional regulatory element (WPRE). Exemplary poly(A) sequences include the hGH poly(A) signal (short), the HSV TK poly(A) signal, a synthetic polyadenylation signal, the SV40 poly(A) signal, the β-globin poly(A) signal, and the like. Additionally, vectors used to provide cells with nucleic acids encoding the gRNA and / or dXR protein can include a nucleic acid sequence encoding a selectable marker in target cells to identify cells that have incorporated the gRNA and / or dXR protein. One of skill in the art would be able to select appropriate elements for inclusion in the recombinant expression vectors described herein.
[0167] Recombinant expression vector sequences can be packaged into viruses or virus-like particles (also referred to herein as "particles" or "virions") for subsequent infection and transformation of cells, ex vivo, in vitro, or in vivo. Such particles or virions typically contain proteins that encapsidate or package the vector genome. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant adeno-associated virus (AAV) vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant lentivirus vector. In some embodiments, a recombinant expression vector of the present disclosure is a recombinant retrovirus vector.
[0168] a. Recombinant AAV for delivery of dXR:rRNA Adeno-associated viruses (AAV) are small (20 nm), non-pathogenic viruses that are useful for the treatment of human diseases in situations where viral vectors are used for delivery to cells, such as eukaryotic cells, either in vivo or ex vivo, where the cells are prepared for administration to a subject. Constructs, such as those encoding the fusion proteins and gRNA embodiments described herein, are generated and are flanked by AAV inverted terminal repeat (ITR) sequences, thereby allowing packaging of the AAV vector into an AAV viral particle, using the AAV cap coding region sequence described below.
[0169] An "AAV" vector can refer to a naturally occurring wild-type virus itself or its derivatives. The term encompasses all subtypes, serotypes, and pseudotypes, as well as both naturally occurring and recombinant forms, unless otherwise specified. As used herein, the term "serotype" refers to an AAV that is identified and distinguished from other AAVs based on capsid protein reactivity with a defined antiserum, such as the many known serotypes of primate AAV. In some embodiments, the AAV vector is selected from AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV9.45, AAV9.61, AAV-Rh74 (rhesus macaque-derived AAV), and AAVRh10, and modified capsids of these serotypes. For example, serotype AAV-2 is used to refer to an AAV containing a genome containing capsid proteins encoded from the AAV-2 cap gene and 5' and 3' ITR sequences from the same AAV-2 serotype. Pseudotyped AAV refers to an AAV containing a viral genome containing capsid proteins from one serotype and 5'-3' ITRs from a second serotype. Pseudotyped rAAVs are expected to have the cell surface binding characteristics of the capsid serotype and genetic characteristics consistent with the ITR serotype. Pseudotyped recombinant AAVs (rAAVs) are produced using standard techniques described in the art. As used herein, for example, rAAV1 can be used to refer to an AAV having both capsid proteins and 5'-3' ITRs from the same serotype, or it can refer to an AAV having capsid proteins from serotype 1 and 5'-3' ITRs from a different AAV serotype, such as AAV serotype 2. For each example exemplified herein, the description of vector design and production describes the serotype of the capsid and 5'-3' ITR sequences.
[0170] "AAV virus" or "AAV viral particle" refers to a viral particle composed of at least one AAV capsid protein (preferably all of the capsid proteins of wild-type AAV) and an encapsidated polynucleotide. When the particle further comprises a heterologous polynucleotide (i.e., a polynucleotide other than the wild-type AAV genome that is delivered to a mammalian cell, called a "transgene"), it is typically referred to as an "rAAV." An exemplary heterologous polynucleotide is a polynucleotide comprising a dXR protein and / or sgRNA of any of the embodiments described herein. Because AAV is naturally replication-deficient and can transduce nearly every cell type in the human body, it represents a suitable vector for therapeutic applications in gene therapy or vaccine delivery. Typically, when producing a recombinant AAV vector, the sequence between the two ITRs is replaced with one or more sequences of interest (e.g., a transgene), and the Rep and Cap sequences are provided in trans, making the ITRs the only viral DNA remaining in the vector. The resulting recombinant AAV vector genome construct contains two cis-acting 130-145 nucleotide ITRs flanking an expression cassette encoding the transgene sequence of interest, providing at least 4.7 kb for packaging of foreign DNA, which can include a transgene, one or more promoters, and associated elements, so that the total vector size is below 5-5.2 kb, compatible with packaging within the AAV capsid (it is understood that vector packaging efficiency decreases when the construct size exceeds this threshold). In the context of the present disclosure, the transgene can be used to suppress transcription of a defective gene in a cell of interest. However, in the context of CRISPR-mediated gene suppression, the size limitation of the expression cassette is a challenge for most CRISPR systems (e.g., Cas9) given the large size of the nuclease. However, it has been discovered that the small size of dCasX and gRNA allows for the generation of "integrated" constructs capable of delivering dXR:gRNA capable of gene suppression in cells.
[0171] "Adeno-associated virus inverted terminal repeats" or "AAV ITRs" refer to art-recognized regions found at each end of the AAV genome that function together in cis as an origin of DNA replication and as a viral packaging signal. The AAV ITRs, together with the AAV rep coding region, provide for efficient removal and rescue of the nucleotide sequence intervening between the two adjacent ITRs and integration into the mammalian cell genome. The nucleotide sequences of the AAV ITR regions are known. See, e.g., Kotin, RM (1994) Human Gene Therapy 5:793-801; Berns, KI "Parvoviridae and their Replication" in Fundamental Virology, 2 ndEdition, (BN Fields and DMKnipe, eds.). As used herein, AAV ITRs need not have the wild-type nucleotide sequence shown, but may be modified, for example, by the insertion, deletion, or substitution of nucleotides. Furthermore, AAV ITRs can be derived from any of several AAV serotypes, including, but not limited to, AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV-Rh74, and AAVRh10, as well as modified capsids of these serotypes. Furthermore, the 5' and 3' ITRs flanking a selected nucleotide sequence in an AAV vector need not necessarily be identical or derived from the same AAV serotype or isolate, so long as they function as intended, i.e., to allow removal and rescue of the sequence of interest from the host cell genome or vector, and to allow integration of the heterologous sequence into the recipient cell genome if the AAV Rep gene product is present in the cell. The use of AAV serotypes for the integration of heterologous sequences into host cells is known in the art (see, e.g., WO2018 / 195555A1 and US2018 / 0258424A1, which are incorporated by reference herein). In one particular embodiment, the ITRs are derived from serotype AAV1. In another specific embodiment of an AAV of the present disclosure, the ITRs are derived from serotype AAV2, a 5' ITR having the sequence CCTGCAGGCAGCTGCGCGCTCGCTCGCTCACTGAGGCCGCCCGGGCGTCGGGCGACCTTTGGTCGCCCGGCCTCAGTGAGCGAGCGAGCGCGCAGAGAGGGAGTGGCCAACTCCATCACTAGGGGTTCCT (SEQ ID NO: 33350), and a 3' ITR having the sequence AGGAACCCCTAGTGATGGAGTTGGCCACTCCCTCTCTGCGCGCTCGCTCGCTCACTGAGGCCGGGCGACCAAAGGTCGCCCGACGCCCGGGCTTTGCCCGGGCGGCCTCAGTGAGCGAGCGAGCGCGCAGCTGCCTGCAGG (SEQ ID NO: 33351).
[0172] "AAV rep coding region" refers to the region of the AAV genome that encodes the replication proteins Rep78, Rep68, Rep52, and Rep40. These Rep expression products have been shown to have many functions, including recognition, binding, and nicking of the AAV origin of DNA replication, DNA helicase activity, and regulation of transcription from AAV (or other heterologous) promoters. The Rep expression products are collectively required to replicate the AAV genome.
[0173] "AAV cap coding region" refers to the region of the AAV genome that encodes the capsid proteins VP1, VP2, and VP3, or functional homologs thereof. These Cap expression products provide the packaging functions collectively required for packaging the viral genome.
[0174] In some embodiments, the AAV capsid utilized to deliver a transgene comprising the coding sequence for the dXR and gRNA of the present disclosure into a host cell can be derived from any of several AAV serotypes, including but not limited to AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV44.9, AAV-Rh74 (AAV derived from rhesus monkeys), and AAVRh10, and the AAV ITRs are derived from AAV serotype 1 or serotype 2.
[0175] To produce rAAV viral particles, an AAV expression vector is introduced into a suitable host cell using known techniques, such as by transfection. Packaging cells are typically used to form the viral particles, and such cells include HEK293 cells (and other cells known in the art) that package adenovirus. Several transfection techniques are generally known in the art. See, for example, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York. Particularly suitable transfection methods include calcium phosphate coprecipitation, direct microinjection into cultured cells, electroporation, liposome-mediated gene transfer, lipid-mediated transduction, and nucleic acid delivery using high-velocity microprojectiles.
[0176] In some embodiments, host cells transfected with the above-described AAV expression vectors are enabled to provide AAV helper functions to replicate and encapsidate nucleotide sequences flanked by AAV ITRs to produce rAAV viral particles. AAV helper functions are generally AAV-derived coding sequences that can be expressed to provide AAV gene products, which in turn function in trans for productive AAV replication. AAV helper functions are used herein to complement necessary AAV functions missing from an AAV expression vector. Thus, AAV helper functions include one or both of the major AAV ORFs (open reading frames) encoding the rep and cap coding regions, or functional homologs thereof. The accessory functions can be introduced into and then expressed in host cells using methods known to those skilled in the art. Typically, accessory functions are provided by infection of the host cells with an unrelated helper virus. In some embodiments, accessory functions are provided using accessory function vectors. Depending on the host / vector system utilized, any of a number of suitable transcriptional and translational control elements, including constitutive and inducible promoters, transcriptional enhancer elements, transcriptional terminators, etc., may be used in the expression vector.
[0177] The present disclosure provides an AAV comprising a transgene encoding dXR and a gRNA, where dXR comprises a dCasX and a KRAB domain as a single repressor, given the size limitations of the transgene. In some embodiments, the transgene encodes a dXR fusion protein of the system comprising a single KRAB domain operably linked to dCasX selected from the group of SEQ ID NOS: 17-36 and 59353-59358 listed in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOS: 57746-59342, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In some embodiments, the transgene encodes a dXR fusion protein of the system comprising a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOS: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOS: 57746-57840, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto.In some embodiments, the transgene encodes a dXR fusion protein of the system comprising a single KRAB domain operably linked to a dCasX selected from the group of sequences set forth in SEQ ID NOS: 17-36 and 59353-59358 in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOS: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. In certain embodiments, the transgene encodes a dXR fusion protein of the system comprising a single KRAB domain operably linked to dCasX of SEQ ID NO: 18 listed in Table 4, wherein the KRAB domain is selected from the group consisting of SEQ ID NOs: 57746-57755, or a sequence having at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% identity thereto. The transgene of the foregoing embodiments further encodes a gRNA having a scaffold comprising the sequence of SEQ ID NO: 2292 or 59352, or a sequence having at least about 70%, at least about 80%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto, wherein the gRNA comprises a targeting sequence complementary to a target nucleic acid sequence of a gene targeted for inhibition. In the foregoing embodiments, the dXR and gRNA are each operably linked to a promoter, embodiments of which are described herein.
[0178] b. VLP and XDP for delivery of dXR:gRNA In other embodiments, retroviruses, e.g., lentiviruses, may be suitable for use as vectors for delivery of nucleic acids encoding the gene repressor system of the present disclosure. Commonly used retroviral vectors are "defective," e.g., unable to produce viral proteins necessary for productive infection, and may be referred to as virus-like particles (VLPs) or delivery particles (XDPs), depending on the components utilized. Rather, vector replication requires growth in a packaging cell line. To generate viral particles containing a nucleic acid of interest, the retroviral nucleic acid containing the nucleic acid is packaged into a VLP or XDP capsid by the packaging cell line. Different packaging cell lines offer different envelope proteins (ecotropic, amphotropic, or xenotropic) that are incorporated into the capsid, and this envelope protein determines the specificity of the viral particle for cells (ecotropic for mice and rats, amphotropic for most mammalian cell types, including human, dog, and mouse, and xenotropic for most mammalian cell types except mouse cells). An appropriate packaging cell line can be used to ensure that cells are targeted by the packaged viral particles. Methods for introducing an expression vector of interest into a packaging cell line and harvesting the viral particles produced by the packaging cell line are well known in the art.
[0179] In some embodiments, the present disclosure provides a vector encoding or comprising a gene repressor system comprising a dXR fusion protein, wherein the dXR fusion protein comprises a first transcriptional repressor domain, wherein the dXR comprises a catalytically inactive CasX of any of the embodiments described herein linked to a KRAB domain of any of the embodiments described herein as the first repressor domain.
[0180] In other embodiments, the disclosure provides a vector encoding or comprising a gene repressor system comprising a fusion protein, the fusion protein comprising a catalytically inactive CasX of any of the embodiments described herein linked to a first, second, and third transcriptional repressor domain, wherein the first transcriptional repressor domain is a KRAB domain of any of the embodiments described herein, the second domain is a DNMT3A catalytic domain of any of the embodiments described herein, and the third transcriptional repressor domain is a DNMT3L-interacting domain, and the fusion protein comprises one or more NLS peptides and a linker peptide. In some embodiments, the fusion protein is composed, from N- to C-terminus, of: NLS-linker4-DNMT3A CD-linker2-DNMT3L ID-linker1-linker3-dCasX-linker3-KRAB-NLS; NLS-linker3-dCasX-linker3-KRAB-NLS-linker1-DNMT3A CD-linker2-DNMT3L ID; NLS-linker3-dCasX-linker1-DNMT3A CD-linker2-DNMT3L ID-linker3-KRAB-NLS; NLS-KRAB-linker3-DNMT3A CD-linker2-DNMT3L ID-linker1-dCasX-linker3-NLS; or NLS-DNMT3A CD-linker2-DNMT3L ID-linker3-KRAB-linker1-dCasX-linker3-NLS.
[0181] In other embodiments, the disclosure provides a vector encoding or comprising a gene repressor system comprising a fusion protein, the fusion protein comprising a catalytically inactive CasX of any of the embodiments described herein linked to first, second, third, and fourth transcriptional repressor domains, wherein the first transcriptional repressor domain is a KRAB domain of any of the embodiments described herein, the second domain is a DNMT3A catalytic domain of any of the embodiments described herein, the third transcriptional repressor domain is a DNMT3L-interacting domain, and the fourth transcriptional repressor domain is an ATRX-DNMT3-DNMT3L (ADD) domain linked at its N-terminus to the DNMT3A catalytic domain, and the fusion protein comprises one or more NLS peptides and a linker peptide. In some embodiments, the fusion protein comprises, from N-terminus to C-terminus, NLS-linker4-ADD-DNMT3A CD-linker2-DNMT3L ID-linker1-linker3-dCasX-linker3-KRAB-NLS, NLS-linker3-dCasX-linker3-KRAB-NLS-linker1-ADD-DNMT3A CD-linker2-DNMT3L ID, NLS-linker3-dCasX-linker1-DNMT3A CD-linker2-DNMT3L ID-linker3-KRAB-NLS, NLS-KRAB-linker3-ADD-DNMT3A CD-linker2-DNMT3L ID-linker1-dCasX-linker3-NLS, or NLS-ADD-DNMT3A CD-linker2-DNMT3L It is composed of ID-linker3-KRAB-linker1-dCasX-linker3-NLS.
[0182] In some embodiments, the present disclosure provides an XDP comprising components selected from all or a portion of a retroviral gag polyprotein, a gag-poly polyprotein, a dXR:gRNA RNP, an RNA transport component, and one or more targeting factors having binding affinity for a cell surface marker of a target cell to facilitate entry of the XDP into a target cell.
[0183] In some embodiments, the retroviral component of the XDP system is derived from an Orthretrovirinae virus or a Spumaretrovirinae virus, wherein the Orthretrovirinae virus is selected from the group consisting of an alpharetrovirus, a betaretrovirus, a deltaretrovirus, an epsilonretrovirus, a gamaretrovirus, and a lentivirus, and the Spumaretrovirinae virus is selected from the group consisting of a bovispumavirus, an equispumavirus, a felispumavirus, a prosimispumavirus, a simispumavirus, and a spumavirus.
[0184] XDPs for use with dXR:gRNA systems can be constructed in different configurations based on the components utilized. In some embodiments, the XDP comprises one or more retroviral components selected from the Gag polyprotein, Gag-transframe region-pol protease polyprotein (Gag-TFR-PR), matrix protein (MA), nucleocapsid protein (NC), capsid protein (CA), p1 peptide, p6 peptide, p2A peptide, p2B peptide, p10 peptide, p12 peptide, p21 / 24 peptide, p12 / p3 / p8 peptide, p20 peptide, protease cleavage site, and a protease capable of cleaving the protease cleavage site, which can be encoded on one or more nucleic acids for production of the XDP in a packaging cell. The remaining components, such as the encapsidated payload of dXR and gRNA (complexed as an RNP), the RNA transport component (described below) used to increase incorporation of the RNP into the XDP, and the targeting factor, can be incorporated into the nucleic acid encoding the retroviral components or can be encoded on separate nucleic acids. In some embodiments, the components of the XDP system are encoded on a single nucleic acid, two nucleic acids, three nucleic acids, four nucleic acids, or five nucleic acids, and then incorporated into a plasmid used in transfection to generate the XDP in packaging cells. Representative, non-limiting configurations of plasmids used to generate XDPs in packaging cells are shown in Figures 4 and 5. In a specific embodiment of the configuration of Figure 4, the Gag polyprotein of Plasmid 1 and the Gag-TFR-PR polyprotein of Plasmid 2 are derived from a lentivirus (harboring HIV-1 protease), the encoded MS2 of Plasmid 1 comprises the sequence of SEQ ID NO: 33276, the encoded dXR fusion protein of Plasmid 3 comprises any of the dXR embodiments described herein, the VSV-G plasmid encodes the VSV-G sequence of SEQ ID NO: 113, and the gRNA plasmid encodes the scaffold of SEQ ID NO: 2292 or 59352. In some embodiments, the components of the XDP system can self-assemble into an XDP with an integrated dXR:gRNA RNP when one or more nucleic acids are introduced into and expressed in a eukaryotic host cell.In the foregoing embodiments, the dXR:gRNA RNP is encapsidated within the XDP during self-assembly of the XDP. In certain embodiments, the targeting factor is incorporated onto the surface of the XDP during self-assembly of the XDP. XDP compositions and methods for making XDPs are described in WO2021 / 113772A1 and PCT / US22 / 32579, which are incorporated herein by reference.
[0185] The polynucleotides encoding Gag, dXR, and gRNA of any of the embodiments described herein can further comprise paired components designed to assist in the transport of the components from the host cell nucleus and promote the recruitment of the complexed CasX:gRNA to the budding XDP. Non-limiting examples of such non-covalent transport components include hairpin RNAs or loops, such as the MS2 hairpin, PP7 hairpin, Qβ hairpin, box B, transactivation response element (TAR), Rev response element, phage GA hairpin, and U1 hairpin II, that have binding affinity for MS2 coat protein, PP7 coat protein, Qβ coat protein, protein N, protein Tat, Rev, phage GA coat protein, and U1A signal recognition particle, respectively, fused to the Gag polyprotein. It has been discovered that incorporating a binding partner and packaging recruiter inserted into a guide RNA into a nucleic acid containing a Gag polypeptide promotes packaging of XDP particles, in part due to the affinity of CasX for the gRNA that results in RNPs. As a result, both the gRNA and CasX associate with Gag during the XDP encapsidation process, increasing the proportion of XDPs containing RNPs compared to constructs lacking the binding partner and packaging recruiter. In other embodiments, the gRNA may contain a Rev response element (RRE) or a portion thereof that has binding affinity for Rev and can be linked to the Gag polyprotein. In other embodiments, the gRNA may contain one or more RRE sequences and one or more MS2 hairpin sequences. The RRE may be selected from the group consisting of stem IIB of the Rev response element (RRE), stems II-V of the RRE, stem II of the RRE, the Rev binding element (RBE) of stem IIB, and a full-length RRE.In the foregoing embodiment, the components are UGGGCGCAGCGUCAAUGACGCUGACGGUACA (Stem IIB, SEQ ID NO: 57736), GCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGC (Stem II, SEQ ID NO: 57737), CAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCUGUUGCAACUCACAGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAAUCC UG (Stem II-V, SEQ ID NO: 57738), GCUGACGGUACAGGC (RBE, SEQ ID NO: 57739), and the sequence AGGAGCUUUGUUCCUUGGGUUCUUGGGAGCAGCAGGAAGCACUAUGGGCGCAGCGUCAAUGACGCUGACGGUACAGGCCAGACAAUUAUUGUCUGGUAUAGUGCAGCAGCAGAACAAUUUGCUGAGGGCUAUUGAGGCGCAACAGCAUCUGUUGCAACUCACAGUCUGGGGCAUCAAGCAGCUCCAGGCAAGAAUCCUGGCUGUGGAAAGAUACCUAAAGGAUCAACAGCUCCU (Full-length RRE, SEQ ID NO: 57740). In other embodiments, the gRNA may comprise one or more RRE sequences and one or more MS2 hairpin sequences. In certain embodiments, the gRNA comprises an MS2 hairpin variant that is optimized to increase binding affinity to the MS2 coat protein, thereby enhancing incorporation of the gRNA and associated CasX into the budding XDP.
[0186] In some embodiments, the tropism factor incorporated onto the XDP surface is selected from the group consisting of glycoproteins, antibody fragments, receptors, and ligands for target cell markers. In one such embodiment, the tropism factor is a glycoprotein having a sequence selected from the group consisting of a sequence set forth in Table 8, or a sequence having at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto. In a specific embodiment, the glycoprotein is VSV-G. [Table 8-1] [Table 8-2] [Table 8-3] [Table 8-4] [Table 8-5] [Table 8-6] [Table 8-7]
[0187] In some embodiments, the protease encoded by a nucleic acid utilized in the XDP system is selected from the group consisting of HIV-1 protease, tobacco etch virus protease (TEV), potyvirus HC protease, potyvirus P1 protease, PreScission (HRV3C protease), b virus NIa protease, B virus RNA-2 encoded protease, aphthovirus L protease, enterovirus 2A protease, rhinovirus 2A protease, picorna 3C protease, comovirus 24K protease, nepovirus 24K protease, RTSV (rice tungro spherical virus) 3C-like protease, parsnip yellow mottle virus protease, 3C-like protease, heparin, cathepsin, thrombin, factor Xa, metalloproteinase, and enterokinase.
[0188] In some embodiments, the disclosure provides a eukaryotic cell transfected with a plasmid encoding the XDP system of any one of the preceding embodiments, wherein the cell is a packaging cell capable of promoting expression of the encoded dXR:gRNA and XDP components and assembly of XDP particles encapsidating the dXR and gRNA RNPs. In some embodiments, the eukaryotic cell is selected from the group consisting of HEK293 cells, HEK293T cells, Lenti-X293T cells, BHK cells, HepG2, Saos-2, HuH7, NS0 cells, SP2 / 0 cells, YO myeloma cells, A549 cells, P3X63 mouse myeloma cells, PER cells, PER.C6 cells, hybridoma cells, VERO, NIH3T3 cells, COS, WI38, MRC5, A549, HeLa cells, CHO cells, and HT1080 cells. In some embodiments, packaging host cells are modified to reduce or eliminate cell surface markers or receptors that would otherwise be incorporated into the XDP, thereby reducing an immune response to the cell surface marker or receptor by a subject receiving the XDP. Such markers may include receptors or proteins that can be bound by MHC receptors or that would otherwise elicit an immune response in a subject. In some embodiments, packaging host cells are modified to reduce or eliminate expression of cell surface markers selected from the group consisting of B2M, CIITA, PD1, and HLA-E KI, such that marker incorporation is reduced on the surface of the XDP. In some embodiments, packaging host cells are modified to express one or more cell surface markers selected from the group consisting of CD46, CD47, CD55, CD59, CD24, CD58, SLAMF4, and SLAMF3 (which function as "don't eat me" signals), such that the cell surface markers are incorporated on the surface of the XDP, such that the incorporation renders the XDP incapable of engulfment and phagocytosis by host surveillance cells, such as macrophages and monocytes.
[0189] For non-viral delivery, vectors can also be delivered, and vectors encoding and / or containing dXR and gRNA can be formulated in nanoparticles; contemplated nanoparticles include, but are not limited to, nanospheres, liposomes, quantum dots, polyethylene glycol particles, hydrogels, and micelles. As described more fully below, lipid nanoparticles are generally composed of an ionizable cationic lipid and three or more additional components, such as cholesterol, DOPE, polylactic-co-glycolic acid, and polyethylene glycol (PEG)-containing lipids. In some embodiments, mRNA encoding a dXR variant of the embodiments disclosed herein is formulated in a lipid nanoparticle. In some embodiments, the nanoparticle comprises a gRNA of the embodiments disclosed herein. In some embodiments, the nanoparticle comprises an mRNA and gRNA encoding dXR. In some embodiments, the components of the dXR:gRNA system are formulated in separate nanoparticles for delivery to cells or for administration to a subject in need thereof.
[0190] c. Lipid nanoparticles (LNPs) In another aspect, the present disclosure provides lipid nanoparticles (LNPs) for delivery of a gRNA and an mRNA encoding a fusion protein of any of the system embodiments disclosed herein. In certain embodiments, the compositions described herein comprise LNPs that encapsulate a gene repressor system of the present disclosure (i.e., an mRNA encoding a fusion protein (e.g., dXR) and a gRNA having a targeting sequence for a target nucleic acid) that represses transcription of a target gene.
[0191] In some embodiments, the LNPs of the present disclosure are tissue- or organ-specific, have excellent biocompatibility, and can deliver a system comprising an mRNA encoding dXR and a gRNA having a targeting sequence to a target nucleic acid with high efficiency, and therefore can be effectively used to suppress or silence a target nucleic acid of a gene in cells of a subject with a disease or disorder.
[0192] In their native form, nucleic acid polymers are unstable in biological fluids and cannot penetrate the membrane of target cells to be delivered to the cytoplasm, thus requiring a delivery system capable of entering cells. Lipid nanoparticles (LNPs) have proven useful for both protecting nucleic acids and delivering them to tissues and cells. Furthermore, the use of mRNA in LNPs to encode CRISPR nucleases eliminates the possibility of undesired genomic integration compared to DNA vectors. Furthermore, because mRNA exerts its function in the cytoplasmic compartment and does not require nuclear entry, it efficiently transfects both mitotic and non-mitotic cells. LNPs as a delivery platform offer the additional advantage that both the mRNA encoding the CRISPR nuclease and the gRNA can be co-formulated into a single LNP particle.
[0193] Thus, in various embodiments, the present disclosure encompasses LNPs and compositions that can be used for a variety of purposes, including delivery of an encapsulated dXR:gRNA system to cells, both in vitro and in vivo. In some embodiments, the gRNA for use in the LNPs is SEQ ID NO: 59352. In some embodiments, the gRNA for use in the LNPs comprises one or more chemical modifications to its sequence. In some embodiments, the mRNA for incorporation into the LNPs of the present disclosure encodes any of the dXR embodiments described herein. In some embodiments, the mRNA for incorporation into the LNPs of the present disclosure is codon-optimized. In some embodiments, the mRNA encoding the dXR fusion protein of the present disclosure is chemically modified, the chemical modification being a substitution of one or more uridine nucleotides of the sequence for N1-methyl-pseudouridine. In some embodiments, the mRNA for incorporation into the LNPs of the present disclosure comprises one or more sequences selected from the group consisting of SEQ ID NOs: 59584-59585, 59610, 59611, 59622, and 59623. In some embodiments, mRNA for incorporation into a LNP of the present disclosure comprises one or more sequences encoded by a sequence selected from the group consisting of 59444-59449, 59455-59456, 59488-59497, 59568-59583, 59595-59609, and 59612-59621.
[0194] In some embodiments, the disclosure encompasses LNPs that encapsidate a gRNA and an mRNA encoding a dCasX fusion protein linked to a first repressor domain, where the repressor domain is a KRAB domain of any of the embodiments described herein. In some embodiments, the disclosure encompasses LNPs that encapsidate a gRNA and an mRNA encoding a dCasX fusion protein linked to first and second repressor domains, where the first repressor domain is a KRAB domain and the second repressor domain is a DNMT3A catalytic domain. In some embodiments, the disclosure encompasses LNPs that encapsidate a gRNA and an mRNA encoding a dCasX fusion protein linked to first, second, and third repressor domains, where the first repressor domain is a KRAB domain, the second repressor domain is a DNMT3A catalytic domain, and the third domain is a DNMT3L-interacting domain. In some embodiments, the disclosure encompasses LNPs that encapsidate an mRNA encoding a fusion protein of dCasX linked to a gRNA and first, second, third, and fourth repressor domains, where the first repressor domain is a KRAB domain, the second repressor domain is a DNMT3A catalytic domain, the third domain is a DNMT3L-interacting domain, and the fourth domain is a DNMT3A ADD domain. In the foregoing embodiments, the components of the fusion protein can be arranged in alternative configurations, as shown in Figures 7 and 45. In certain embodiments, the disclosure encompasses methods of treating or preventing a disease or disorder in a subject in need thereof by contacting the subject with LNPs encapsulating a dXR:gRNA system of the embodiments described herein, where dXR is the encoding mRNA and the gRNA comprises a targeting sequence complementary to a target nucleic acid in the subject's cells.
[0195] In some embodiments, the present disclosure provides LNPs in which mRNA encoding a gRNA and a dXR are incorporated into a single LNP particle. In certain embodiments, the LNP composition comprises a ratio of gRNA to dXR mRNA of the embodiments described herein, measured by weight, of about 25:1 to about 1:25. In certain embodiments, the LNP formulation comprises a ratio of gRNA to dXR mRNA, such as dXR mRNA, of about 10:1 to about 1:10. In certain embodiments, the LNP formulation comprises a ratio of gRNA to dXR mRNA of about 8:1 to about 1:8. In some embodiments, the LNP formulation comprises a ratio of gRNA to dXR mRNA of about 5:1 to about 1:5. In some embodiments, the ratio ranges are about 3:1 to 1:3, about 2:1 to 1:2, about 5:1 to 1:2, about 5:1 to 1:1, about 3:1 to 1:2, about 3:1 to 1:1, about 3:1, or about 2:1 to 1:1. In some embodiments, the ratio of gRNA to mRNA is about 3:1 or about 2:1. In some embodiments, the ratio of gRNA to dXR mRNA is about 1:1. The ratio can be about 25:1, 10:1, 5:1, 3:1, 1:1, 1:3, 1:5, 1:10, or 1:25.
[0196] In other embodiments, the present disclosure provides LNPs in which the mRNA encoding the gRNA and dXR are incorporated into separate LNP particles, which can be formulated together in various ratios for administration.
[0197] In some embodiments, the optimized mRNA of the present disclosure encoding a CasX protein can be provided in a solution that is mixed with a lipid solution so that the mRNA can be encapsulated in LNPs. A suitable mRNA solution can be any aqueous solution containing the mRNA to be encapsulated at various concentrations. For example, a suitable mRNA solution can contain mRNA at a concentration of about 0.01 mg / ml, 0.05 mg / ml, 0.06 mg / ml, 0.07 mg / ml, 0.08 mg / ml, 0.09 mg / ml, 0.1 mg / ml, 0.15 mg / ml, 0.2 mg / ml, 0.3 mg / ml, 0.4 mg / ml, 0.5 mg / ml, 0.6 mg / ml, 0.7 mg / ml, 0.8 mg / ml, 0.9 mg / ml, 1.0 mg / ml, 1.25 mg / ml, 1.5 mg / ml, 1.75 mg / ml, or 2.0 mg / ml or more. In some embodiments, suitable mRNA solutions include those containing approximately 0.01-2.0 mg / ml, 0.01-1.5 mg / ml, 0.01-1.25 mg / ml, 0.01-1.0 mg / ml, 0.01-0.9 mg / ml, 0.01-0.8 mg / ml, 0.01-0.7 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.4 mg / ml, 0.01-0.3 mg / ml, 0.01-0.2 mg / ml, 0.01-0.1 mg / ml, 0.05-1.0 ...6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.6 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.6 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01 The mRNA may be contained at a concentration in the range of 0.05 to 0.9 mg / ml, 0.05 to 0.8 mg / ml, 0.05 to 0.7 mg / ml, 0.05 to 0.6 mg / ml, 0.05 to 0.5 mg / ml, 0.05 to 0.4 mg / ml, 0.05 to 0.3 mg / ml, 0.05 to 0.2 mg / ml, 0.05 to 0.1 mg / ml, 0.1 to 1.0 mg / ml, 0.2 to 0.9 mg / ml, 0.3 to 0.8 mg / ml, 0.4 to 0.7 mg / ml, or 0.5 to 0.6 mg / ml.In some embodiments, a suitable mRNA solution may contain mRNA at a concentration of up to about 5.0 mg / ml, 4.0 mg / ml, 3.0 mg / ml, 2.0 mg / ml, 1.0 mg / ml, 0.9 mg / ml, 0.8 mg / ml, 0.7 mg / ml, 0.6 mg / ml, 0.5 mg / ml, 0.4 mg / ml, 0.3 mg / ml, 0.2 mg / ml, 0.1 mg / ml, 0.05 mg / ml, 0.04 mg / ml, 0.03 mg / ml, 0.02 mg / ml, 0.01 mg / ml, or 0.05 mg / ml.
[0198] In some embodiments, the gRNA of the present disclosure can be provided in a solution that is mixed with a lipid solution so that the gRNA can be encapsulated in LNPs. Suitable gRNA solutions can be any aqueous solution containing the gRNA to be encapsulated at various concentrations. For example, suitable gRNA solutions can contain gRNA at concentrations of approximately 0.01 mg / ml, 0.05 mg / ml, 0.06 mg / ml, 0.07 mg / ml, 0.08 mg / ml, 0.09 mg / ml, 0.1 mg / ml, 0.15 mg / ml, 0.2 mg / ml, 0.3 mg / ml, 0.4 mg / ml, 0.5 mg / ml, 0.6 mg / ml, 0.7 mg / ml, 0.8 mg / ml, 0.9 mg / ml, 1.0 mg / ml, 1.25 mg / ml, 1.5 mg / ml, 1.75 mg / ml, or 2.0 mg / ml or greater. In some embodiments, a suitable gRNA solution is about 0.01-2.0 mg / ml, 0.01-1.5 mg / ml, 0.01-1.25 mg / ml, 0.01-1.0 mg / ml, 0.01-0.9 mg / ml, 0.01-0.8 mg / ml, 0.01-0.7 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.4 mg / ml, 0.01-0.3 mg / ml, 0.01-0.2 mg / ml, 0.01-0.1 mg / ml, 0.05-1.0 ...6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.6 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01-0.6 mg / ml, 0.01-0.6 mg / ml, 0.01-0.5 mg / ml, 0.01 The gRNA may be contained at a concentration in the range of 0.05-0.9 mg / ml, 0.05-0.8 mg / ml, 0.05-0.7 mg / ml, 0.05-0.6 mg / ml, 0.05-0.5 mg / ml, 0.05-0.4 mg / ml, 0.05-0.3 mg / ml, 0.05-0.2 mg / ml, 0.05-0.1 mg / ml, 0.1-1.0 mg / ml, 0.2-0.9 mg / ml, 0.3-0.8 mg / ml, 0.4-0.7 mg / ml, or 0.5-0.6 mg / ml. In some embodiments, a suitable gRNA solution may contain a gRNA concentration of up to about 5.0 mg / ml, 4.0 mg / ml, 3.0 mg / ml, 2.0 mg / ml, 1.0 mg / ml, 0.9 mg / ml, 0.8 mg / ml, 0.7 mg / ml, 0.6 mg / ml, 0.5 mg / ml, 0.4 mg / ml, 0.3 mg / ml, 0.2 mg / ml, 0.1 mg / ml, 0.05 mg / ml, 0.04 mg / ml, 0.03 mg / ml, 0.02 mg / ml, 0.01 mg / ml, or 0.05 mg / ml.
[0199] Initial formulations of LNPs utilizing permanent cationic lipids resulted in LNPs with a positive surface charge that proved toxic in vivo and were rapidly cleared by phagocytic cells. By changing to ionizable cationic lipids with tertiary or quaternary amines, particularly those with a pKa of less than 7, the resulting LNPs achieve efficient encapsulation of nucleic acid polymers at low pH by electrostatically interacting with the negative charges on the phosphate backbone of mRNA or gRNA. This also results in a primarily neutral system at physiological pH values, thus mitigating the problems associated with permanently charged cationic lipids. As used herein, "ionizable lipid" refers to an amine-containing lipid that can be easily protonated; for example, it can be a lipid whose charge state changes depending on the surrounding pH. An ionizable lipid may be protonated (positively charged) at a pH below the pKa of the cationic lipid and substantially neutral at a pH above the pKa. In one example, an LNP may contain protonated ionizable lipids and / or ionizable lipids that exhibit neutralization. In some embodiments, LNPs have a pKa of 5 to 8,...
Claims
1. A gene repressor system, a) a catalytically inactive Class 2 CRISPR protein; and b) a transcriptional repressor domain comprising an amino acid sequence selected from the group consisting of SEQ ID NOs: 57755, 57750, 57771, and 57779, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto; c) a guide ribonucleic acid (gRNA); Including, i) the transcriptional repressor domain is linked to the catalytically inactive Class 2 CRISPR protein as a fusion protein; ii) the gRNA comprises a targeting sequence complementary to a target nucleic acid sequence of a gene being targeted for repression, silencing, or downregulation; iii) the fusion protein is capable of forming a ribonucleoprotein (RNP) with the gRNA; iv) A gene repressor system, wherein the RNP is capable of binding to the target nucleic acid.
2. 2. The gene repressor system of claim 1, wherein the fusion protein is capable of suppressing expression of a reporter gene to a greater extent than an equivalent fusion protein comprising the sequence of SEQ ID NO: 59626 as a transcriptional repressor domain when assayed in an in vitro cell assay.
3. 3. The gene repressor system of claim 2, wherein the reporter gene is the beta-2-microglobulin (B2M) locus and expression of B2M is repressed by at least about 75%, at least about 80%, at least about 85%, or at least about 90%.
4. 4. The gene repressor system of any one of claims 1 to 3, wherein the transcriptional repressor domain is linked by a linker peptide sequence at or near the C-terminus or at or near the N-terminus of the catalytically inactive Class 2 CRISPR protein.
5. d) a second transcriptional repressor domain; and e) a third transcriptional repressor domain; and 2. The gene repressor system of claim 1, further comprising the catalytically inactive Class 2 CRISPR protein, the first transcriptional repressor domain, the second transcriptional repressor domain, and the third transcriptional repressor domain linked as a fusion protein.
6. 6. The gene repressor system of claim 5, wherein the second transcriptional repressor domain and the third transcriptional repressor domain are each a DNA methyltransferase (DNMT) domain.
7. 7. The gene repressor system of claim 6, wherein the second transcriptional repressor domain is DNMT3A or a subdomain thereof.
8. 8. The gene repressor system of claim 7, wherein the second transcriptional repressor domain is the catalytic domain of DNMT3A (DNMT3A CD).
9. 9. The gene repressor system of claim 8, wherein the DNMT3A CD comprises the sequence of SEQ ID NO: 59450, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
10. The gene repressor system according to any one of claims 5 to 9, wherein the third transcriptional repressor domain is a DNMT3L interaction domain (DNMT3L ID).
11. 11. The gene repressor system of claim 10, wherein the DNMT3L ID comprises the sequence of SEQ ID NO: 59625, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
12. 9. The gene repressor system of claim 8, wherein the fusion protein comprises an ATRX-DNMT3-DNMT3L (ADD) domain linked at its N-terminus to the DNMT3A catalytic domain.
13. 13. The gene repressor system of claim 12, wherein the ADD domain comprises the sequence of SEQ ID NO: 59452, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
14. The gene repressor system of claim 1 , wherein the fusion protein comprises one or more linker peptide sequences.
15. The gene repressor system of claim 1, wherein the catalytically inactive Class 2 CRISPR protein is selected from the group consisting of catalytically inactive Type II, catalytically inactive Type V, or catalytically inactive Type VI proteins.
16. 16. The gene repressor system of claim 15, wherein the catalytically inactive type II protein is a Cas9 protein.
17. The gene repressor system of claim 15, wherein the catalytically inactive V-type protein is selected from the group consisting of catalytically inactive Cas12a (Cpf1), Cas12b (C2c1), Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12f, Cas12g, Cas12h, Cas12i, Cas12j, Cas12k, Cas14, and CasΦ proteins.
18. 18. The gene repressor system of claim 17, wherein the CRISPR protein is a catalytically inactive CasX protein (dCasX), and the dCasX comprises a sequence selected from the group consisting of SEQ ID NOs: 17-36 and 59353-59358, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
19. 19. The gene repressor system of claim 18, wherein the dCasX comprises the sequence of SEQ ID NO: 18, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
20. The gene repressor system of claim 1 , wherein the fusion protein further comprises one or more nuclear localization signals (NLS).
21. a) the one or more NLSs are PKKKRKV (SEQ ID NO: 33289), KRPAATKKAGQAKKKK (SEQ ID NO: 33290), PAAKRVKLD (SEQ ID NO: 33291), RQRRNELKRSP (SEQ ID NO: 33292), NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 33293), RMRIZFKNKGKDTAELRRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 33294), VSRKRPRP (SEQ ID NO: 33295), PPKKARED (SEQ ID NO: 332 96), PQPKKKPL (SEQ ID NO: 166), SALIKKKKKMAP (SEQ ID NO: 33298), DRLRR (SEQ ID NO: 33299), PKQKKRK (SEQ ID NO: 33300), RKLKKKIKKL (SEQ ID NO: 33301), REKKKFLKRR (SEQ ID NO: 33302), KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 33303), RKCLQAGMNLEARKTKK (SEQ ID NO: 33304), PRPRKIPR (SEQ ID NO: 33305), PPRKKRTVV (SEQ ID NO: 33306), NLSKKKKRKREK (SEQ ID NO: 33307), 3307), RRPSRPFRKP (SEQ ID NO: 33308), KRPRSPSSS (SEQ ID NO: 33309), KRGINDRNFWRGENERKTR (SEQ ID NO: 33310), PRPPKMARYDN (SEQ ID NO: 33311), KRSFSKAF (SEQ ID NO: 33312), KLKIKRPVK (SEQ ID NO: 33313), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 33314), PKTRRRPRRSQRKRPPT (SEQ ID NO: 33315), SRRRKANPTKLSENAKKLAKEVEN (SEQ ID NO: 33316), KTRRRPR RSQRKRPPT (SEQ ID NO: 33317), RRKKRRPRRKKRR (SEQ ID NO: 33318), PKKKSRKPKKKSRK (SEQ ID NO: 33319), HKKKHPDASVNFSEFSK (SEQ ID NO: 33320), QRPGPYDRPQRPGPYDRP (SEQ ID NO: 33321), LSPSLSPLLSPSLSPL (SEQ ID NO: 33322), RGKGGKGLGKGGAKRHRK (SEQ ID NO: 33323), PKRGRGRPKRGRGR (SEQ ID NO: 33324), PKKKRKVPPPPAAKRVKLD (SEQ ID NO: 33325),PKKKRKVPPPPPKKKRKV (SEQ ID NO: 33326), PAKRARRGYKC (SEQ ID NO: 33327), KLGPRKATGRW (SEQ ID NO: 33328), PRRKREE (SEQ ID NO: 33329), PYRGRKE (SEQ ID NO: 33330), PLRKRPRR (SEQ ID NO: 33331), PLRKRPRRGSPLRKRPRR (SEQ ID NO: 33332), PAAKRVKLDGGKRTADGSEFESPKKKRKV (SEQ ID NO: 33333), PAAKRVKLDGGKRTADGSEFE SPKKKRKVGIHGVPAA (SEQ ID NO: 33334), PAAKRVKLDGGKRTADGSEFESPKKKRKVAEAAAKEAAAKEAAAKA (SEQ ID NO: 33335), PAAKRVKLDGGKRTADGSEFESPKKKRKVPG (SEQ ID NO: 33336), KRKGSPERGERKRHW (SEQ ID NO: 33337), KRTADSQHSTPKTKRKVEFEPKKKRKV (SEQ ID NO: 33338), and SEQ ID NOs: 37-112; and 21. The gene repressor system of claim 20, wherein the one or more linker peptides comprise a sequence selected from GGPSSGAPPPSGGSPAGSPTSTEEGTSESATPESGPGTSTEPSEGSAPGSPAGSPTSTEEGTSTEPSEGSAPGTSTEPSE (Linker 1, SEQ ID NO: 57621), SSGNSNANSRGPSFSSGLVPLSLRGSH (Linker 2, SEQ ID NO: 57623), GGSGGGS (Linker 3, SEQ ID NO: 57626), and GSGSGGG (Linker 4, SEQ ID NO: 57628).
22. 21. The gene repressor system of claim 20, wherein the one or more NLSs are linked at or near the C-terminus of the fusion protein, or at or near the N-terminus of the fusion protein, or at or near both the N-terminus and C-terminus of the fusion protein.
23. The fusion protein comprises, from the N-terminus to the C-terminus: a) NLS-linker4-DNMT3A CD-linker2-DNMT3L ID-linker1-linker3-dCasX-linker3-first transcriptional repressor domain-NLS; b) NLS-linker3-dCasX-linker3-first transcriptional repressor domain-NLS-linker1-DNMT3A CD-linker2-DNMT3L ID; c) NLS-linker3-dCasX-linker1-DNMT3A CD-linker2-DNMT3L ID-linker3-first transcriptional repressor domain-NLS; d) NLS-first transcriptional repressor domain-linker3-DNMT3A CD-linker2-DNMT3L ID-linker1-dCasX-linker3-NLS; e) NLS-DNMT3A CD-linker2-DNMT3L ID-linker3-first transcriptional repressor domain-linker1-dCasX-linker3-NLS f) NLS-ADD-DNMT3A CD-linker2-DNMT3L ID-linker1-linker3-dCasX-linker3-first transcriptional repressor domain-NLS; g) NLS-linker3-dCasX-linker3-first transcriptional repressor domain-NLS-linker1-ADD-DNMT3A CD-linker2-DNMT3L ID; h) NLS-linker3-dCasX-linker1-ADD-DNMT3A CD-linker2-DNMT3L ID-linker3-first transcriptional repressor domain-NLS; i) NLS-first transcriptional repressor domain-linker3-ADD-DNMT3A CD-linker2-DNMT3L ID-linker1-dCasX-linker3-NLS, or j) The gene repressor system of claim 20, which is composed of NLS-ADD-DNMT3A CD-linker 2-DNMT3L ID-linker 3-first transcriptional repressor domain-linker 1-dCasX-linker 3-NLS.
24. 2. The gene repressor system of claim 1, wherein the gRNA has a scaffold comprising a sequence selected from the group consisting of SEQ ID NOs: 2238-2331, 57544-57589 and 59352, or a sequence having at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity thereto.
25. 25. The gene repressor system of Claim 24, wherein the gRNA scaffold comprises one or more chemical modifications to the sequence.
26. 26. The gene repressor system of claim 25, wherein the chemical modification is the addition of a 2'O-methyl group to one or more nucleotides of the sequence, and / or the chemical modification is the substitution of a phosphorothioate bond between two or more nucleotides of the sequence.
27. 2. The gene repressor system of claim 1, wherein the gRNA comprises a targeting sequence having 15, 16, 17, 18, 19, 20, or 21 nucleotides. (i) a target nucleic acid sequence complementary to the targeting sequence is within 1 kb of a transcription start site (TSS) in a gene; (ii) the target nucleic acid sequence complementary to the targeting sequence is within 1 kb of an enhancer of the gene; (iii) the target nucleic acid sequence complementary to the targeting sequence is in the 3' untranslated region of the gene; or (iv) the target nucleic acid sequence complementary to the targeting sequence is within an exon of the gene.
29. 2. The gene repressor system of claim 1, wherein the RNP is capable of binding to the target nucleic acid but is unable to cleave the target nucleic acid, and upon binding to the target nucleic acid, the gene is epigenetically modified, and after epigenetic modification, transcription of the gene is repressed.
30. 30. The gene repressor system of claim 29, wherein transcription of the gene is suppressed by at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, or at least about 99% compared to the untreated gene when assessed in an in vitro assay.
31. 30. The gene repressor system of claim 29, wherein the repression of transcription of the gene is maintained for at least about 8 hours, at least about 1 day, at least about 7 days, at least about 1 month, or at least about 2 months.
32. A nucleic acid encoding the fusion protein of the gene repressor system of claim 1.
33. 33. The nucleic acid of claim 32, wherein the nucleic acid sequence is mRNA.
34. the mRNA is chemically modified; a) the chemical modification is a substitution of one or more uridine nucleotides of the sequence with N1-methyl-pseudouridine, or b) the nucleic acid of claim 33, wherein the mRNA comprises a sequence selected from the group consisting of SEQ ID NOs: 59584-59585, 59610, 59611, 59622 and 59623.
35. A nucleic acid encoding the gRNA of the gene repressor system of claim 1.
36. 33. A lipid nanoparticle (LNP) comprising the nucleic acid of claim 32, wherein the LNP comprises one or more components selected from the group consisting of an ionizable lipid, a helper phospholipid, a polyethylene glycol (PEG)-modified lipid, and cholesterol.
37. 10. A lipid nanoparticle (LNP) comprising a first nucleic acid encoding the fusion protein of the gene repressor system of claim 1 and a second nucleic acid comprising the gRNA, wherein the LNP comprises one or more components selected from the group consisting of an ionizable lipid, a helper phospholipid, a polyethylene glycol (PEG)-modified lipid, and cholesterol.
38. 10. A lipid nanoparticle (LNP) composition comprising a first population of lipid nanoparticles encapsulating the repressor system of claim 1 and a second population of lipid nanoparticles, wherein the first population comprises lipid nanoparticles encapsulating a first nucleic acid encoding the fusion protein, and the second population of lipid nanoparticles comprises the gRNA, and the LNPs comprise one or more components selected from the group consisting of an ionizable lipid, a helper phospholipid, a polyethylene glycol (PEG)-modified lipid, and cholesterol.
39. A vector comprising the nucleic acid of claim 32 or 35.
40. 40. The vector of claim 39, wherein the vector is selected from the group consisting of a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated viral (AAV) vector, a herpes simplex viral (HSV) vector, a plasmid, a minicircle, a nanoplasmid, and an RNA vector.
41. A method for inhibiting transcription of a target nucleic acid sequence of a gene in a cell population in vitro or ex vivo, said method comprising administering to said cells: a) an RNP comprising the gene repressor system of claim 1; b) a nucleic acid according to claim 32 or 35, c) a vector comprising the nucleic acid of claim 32 or 35; d) lipid nanoparticles according to claim 36 or 37, or e) introducing the lipid nanoparticle composition of claim 38 into the cells; A method wherein, following binding of the introduced or expressed RNP of the gene repressor system to the target nucleic acid, transcription of the gene is repressed in the cell.
42. 42. The method of claim 41, wherein the method mediates a heritable epigenetic change in a gene of the cell.
43. A composition for use as a pharmaceutical in treating a subject having a disorder, said composition comprising: a) an RNP comprising the gene repressor system of claim 1; b) a nucleic acid according to claim 32 or 35, c) a vector comprising the nucleic acid of claim 32 or 35; d) lipid nanoparticles according to claim 36 or 37, or e) a lipid nanoparticle composition according to claim 38.
44. A pharmaceutical composition comprising the gene repressor system of claim 1, the nucleic acid of claim 32 or 35, a vector comprising the nucleic acid of claim 32 or 35, the lipid nanoparticle of claim 36 or 37, or the lipid nanoparticle composition of claim 38, and a pharmaceutically acceptable excipient.