Compositions and methods for identifying regulators of cell type fate determination
By employing neuron-specific transcription factors and CRISPR/Cas9 compositions, the method improves the maturation and conversion of stem cells into neurons, overcoming inefficiencies in current reprogramming methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-04-03
AI Technical Summary
Current methods for reprogramming cell fate, particularly in generating human neuronal subtypes, are inefficient and require additional cofactors or different protocols due to unique differences in mouse and human cell plasticity, limiting the development of high-throughput approaches to systematically profile transcription factors for cell type identity.
Utilizing specific neuron-specific transcription factors and CRISPR/Cas9 compositions to increase the expression of neuron-specific genes in stem cells, including combinations of NEUROG3, SOX4, SOX9, KLF4, and others, with gRNAs targeting these factors to enhance maturation and conversion into neurons.
Enhances the maturation and conversion of stem cells into neurons, addressing the inefficiencies of existing protocols by providing a systematic and efficient method for generating clinically relevant neuronal cell types.
Smart Images

Figure 2026058344000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority based on U.S. Provisional Patent Application No. 62 / 888922, filed on August 19, 2019, U.S. Provisional Patent Application No. 62 / 889361, filed on August 20, 2019, and U.S. Provisional Patent Application No. 62 / 961084, filed on January 14, 2020, each of which is hereby incorporated by reference in its entirety.
[0002] Description of Research and Development by Federal Government Funds This invention was made with government support under grants R21NS103007, DP2OD008586, R01DA036865, F31NS105419, and T32GM008555 awarded by the National Institutes of Health, and grant EFMA - 1830957 awarded by the National Science Foundation. The government has certain rights in this invention. This disclosure relates to DNA targeting compositions, such as CRISPR / Cas9 compositions, and methods for identifying regulators of cell - type fate specification.
Background Art
[0003] The advent of methods for reprogramming cell fate has revolutionized regenerative medicine, disease modeling, and cell therapy. Given the increasing evidence defining specific neuronal subtypes as the origin of neurological diseases, the ability to generate these subtypes in vitro could accelerate the study and treatment of these complex diseases. Some current approaches to cell reprogramming involve overexpressing transcription factors (TFs) to reprogram the transcriptional programs of starting cells. While this approach has been successful in generating clinically relevant cell types, relatively few cell types have been reprogrammed in this manner. Although efforts have been made to catalog the entire set of putative human transcription factors and define their tissue-specific expression, relatively few TFs have been empirically validated for their role in cell fate determination. Furthermore, the selection of fate-determining TFs for cell reprogramming applications often relies on approaches that evaluate small subsets of TFs or use computational models to predict optimal TF combinations. Current strategies for developing novel cell reprogramming protocols using TFs are slow, inadequate, and arduous. While previous studies have been primarily in mice, progress from mouse to human cell reprogramming is crucial. Mouse cells exhibit unique differences in plasticity compared to human cells. Mouse cells generally accept reprogramming more readily, often achieving high conversion efficiency and shorter maturation times. Consequently, human cells often require additional cofactors or entirely different protocols to achieve conversion results comparable to their mouse counterparts. Given that neuronal cell type diversity in the human brain is likely programmed by TF diversity, the continued development of high-throughput approaches to systematically profile the causal role of TFs, particularly those that correlate well with human, in shaping neuronal cell type identity remains crucial. [Overview of the Initiative] [Means for solving the problem]
[0004] In one embodiment, the Disclosure relates to (1) a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (2) a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i) NEUROG3, SOX4, SOX9, KLF4, NR (ii)P RDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii)RUNX3、PRDM1、KLF6、PAX2、RFX3、SOX10、GATA1、KLF5、KLF1、ERF、LHX6、PHOX2B、NANOG、NR5A2、ETV3、NEUROG3、SOX4、SOX9、PAX8、IRF5、CDX4、RARA、BHLHE40、SOX3、KLF4、NR5A1、IRF4、ASCL1、GATA6、SPIB、THRB、FOXH1、NEUROD1、SOX17、CDX2、ZEB2、RARG、INSM1、FOSL1、NEUROG1、SOX1、WT1、PAX5、SOX18、POU5F1、RFX4、KLF7、NKX2-2、OVOL2、FOXJ1、PRDM14、VENTX、LHX8、GFI1、KLF17、OVOL1、OLIG3、HMX3、ZNF521、ONECUT3、OVOL3、ZNF362、AFF1、HMX2、ZNF786、GATA5、TBX3、ZNF385A、ATOH1、PROP1、SOX11、JUN、FOXE3、FERD3L、E2F7;(iv)ZIC2、SPI1、GRHL2、TFAP2C、KLF8、MYB、TCF21、KLF12、TWIST1、SNAI1、RREB1、GCM2、GRHL1、ETS1、BARHL2、GRHL3、ELF3、PTF1A、GSX1、PBX2、NOTO、KLF3、ZNF311、ELMSAN1、ZNF296、PLEK、KMT2A、HES3;(v)HES2、SREBF1、CIC、WHSC1、VDR、HES1、ID2、TCF21、SNAI1、RREB1、GCM2、IRF3、FOXA1、GATA5、GRHL1、SOX5、DMRT1、GCM1、BARHL2、SOX13、ZEB1、PITX2、PTF1A、ZNF282、NPAS2、ZNF160、HES7、ZBED4、SALL4、GLIS3、TBX22、ZNF331、EGR4、ZIC5、ZNF710、ZNF697、ZFP36L2、ELMSAN1、ZNF296、ZNF318、ZNF570、ZNF683、ZFP36L1、HES4、ZNF777、HES5、ZIM2、ZNF579、BMP2、CRAMP1L、TOX3、FEZF2、HES3、ZNF791;(vi)ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 This relates to polynucleotides that can encode a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0005] In a further embodiment, the present disclosure relates to a system for increasing the expression of neuron-specific genes, comprising: (a) a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, Neurog1, SOX18, RFX4, KLF7, SP8, OVOL1, Neurog2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (b) a first gRNA that targets a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i) NE UROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and P LAGL2;(ii)PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SO X2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii)RUNX3、PRDM1、KLF6、PAX2、RFX3、SOX10、GATA1、KLF5、KLF1、ERF、LHX6、PHOX2B、NANOG、NR5A2、ETV3、NEUROG3、SOX4、SOX9、PAX8、IRF5、CDX4、RARA、BHLHE40、SOX3、KLF4、NR5A1、IRF4、ASCL1、GATA6、SPIB、THRB、FOXH1、NEUROD1、SOX17、CDX2、ZEB2、RARG、INSM1、FOSL1、NEUROG1、SOX1、WT1、PAX5、SOX18、POU5F1、RFX4、KLF7、NKX2-2、OVOL2、FOXJ1、PRDM14、VENTX、LHX8、GFI1、KLF17、OVOL1、OLIG3、HMX3、ZNF521、ONECUT3、OVOL3、ZNF362、AFF1、HMX2、ZNF786、GATA5、TBX3、ZNF385A、ATOH1、PROP1、SOX11、JUN、FOXE3、FERD3L、E2F7;(iv)ZIC2、SPI1、GRHL2、TFAP2C、KLF8、MYB、TCF21、KLF12、TWIST1、SNAI1、RREB1、GCM2、GRHL1、ETS1、BARHL2、GRHL3、ELF3、PTF1A、GSX1、PBX2、NOTO、KLF3、ZNF311、ELMSAN1、ZNF296、PLEK、KMT2A、HES3;(v)HES2、SREBF1、CIC、WHSC1、VDR、HES1、ID2、TCF21、SNAI1、RREB1、GCM2、IRF3、FOXA1、GATA5、GRHL1、SOX5、DMRT1、GCM1、BARHL2、SOX13、ZEB1、PITX2、PTF1A、ZNF282、NPAS2、ZNF160、HES7、ZBED4、SALL4、GLIS3、TBX22、ZNF331、EGR4、ZIC5、ZNF710、ZNF697、ZFP36L2、ELMSAN1、ZNF296、ZNF318、ZNF570、ZNF683、ZFP36L1、HES4、ZNF777、HES5、ZIM2、ZNF579、BMP2、CRAMP1L、TOX3、FEZF2、HES3、ZNF791;(vi)ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, K LF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160 The present invention relates to a system that may include a Cas protein or a fusion protein; a second gRNA that targets a second neuron-specific transcription factor selected from ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX; and a Cas protein or a fusion protein. In some embodiments, the fusion protein may include two heterologous polypeptide domains, the first polypeptide domain comprising a Cas protein, a zinc finger protein, or a TALE protein, and the second polypeptide domain having an activity selected from transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, nucleic acid association activity, methylase activity, and demethylase activity.
[0006] In some embodiments, the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, NKX2-2, HES3, and ZFP36L1. In some embodiments, the second neuron-specific transcription factor may be selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, and NKX2-2. In some embodiments, the second neuron-specific transcription factor may be selected from HES3 and ZFP36L1.In some embodiments, the second neuron-specific transcription factor is (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, Neurog1, SOX18, RFX4, KLF7, SP8, OVOL1, Neurog2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1 , LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1 , HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PH OX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, S PIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, N The second polypeptide domain may be selected from KX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7, and the second polypeptide domain has transcriptional activation activity. In some embodiments, the fusion protein is... VP64 dCas9 VP64 Alternatively, it may include dCas9-p300.
[0007] In some embodiments, the second neuron-specific transcription factor is (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TC F21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZB ED4, SALL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, H ES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, M YB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF The second polypeptide domain may be selected from 281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX, and the second polypeptide domain has transcriptional repressive activity. In some embodiments, the fusion protein may contain dCas9-KRAB.In some embodiments, the first gRNA and the second gRNA may each individually comprise a 12-22 base pair complementary polynucleotide sequence of the target DNA sequence, followed by a protospacer flanking motif, and optionally, the gRNA may bind to and target a polynucleotide comprising a sequence selected from SEQ ID NOs. 38-87, and / or include such a sequence, and optionally, the first and / or second gRNA may comprise a crRNA, tracrRNA, or a combination thereof.
[0008] Another aspect of this disclosure provides isolated polynucleotides that can encode the systems detailed herein. Another aspect of this disclosure provides a vector which may comprise isolated polynucleotides as detailed herein. In another aspect, the disclosure relates to cells that may contain isolated polynucleotides or vectors as detailed herein.
[0009] In a further embodiment, the present disclosure relates to a method for increasing the maturation of stem cell-induced neurons. The method comprises (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells, or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof in stem cells; and (i) NEU ROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and P LAGL2;(ii)PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SO X2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P The procedure may include increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0010] Another aspect of this disclosure provides a method for increasing the maturation of stem cell-induced neurons. This method comprises the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; and in stem cells (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX22, ZNF 331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 The procedure may include the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0011] Another aspect of the present disclosure provides a method for increasing the conversion of stem cells into neurons. The method comprises (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells, or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof in stem cells; and (i) NEU ROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and P LAGL2;(ii)PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SO X2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P The procedure may include increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0012] Another aspect of this disclosure provides a method for increasing the conversion of stem cells into neurons. The method comprises the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX22, ZNF 331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 The procedure may include the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0013] Another aspect of this disclosure relates to a method for treating a subject requiring treatment. The method comprises (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, Neurog1, SOX18, RFX4, KLF7, SP8, OVOL1, Neurog2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in the stem cells of the subject, or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the stem cells of the subject, i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1 and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, S OX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P The procedure may include increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0014] Another aspect of this disclosure provides a method for treating subjects requiring treatment. The method comprises the steps of: increasing the levels of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the subject stem cells; and in the subject stem cells: (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, I D2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX22, ZN F331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 The procedure may include the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0015] In some embodiments, the step of increasing the level of a first neuron-specific transcription factor may include at least one of the following: (a) administering a polynucleotide encoding the first neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the first neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the first neuron-specific transcription factor, or a TALE protein that targets the first neuron-specific transcription factor, and the second polypeptide domain has transcriptional activating activity, and if the first polypeptide domain contains a Cas protein, further administering a gRNA that targets the first neuron-specific transcription factor to the stem cells. In some embodiments, the step of increasing the level of the second neuron-specific transcription factor may include at least one of the following: (a) administering a polynucleotide encoding the second neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the second neuron-specific transcription factor to stem cells; and (c) administering a fusion protein containing two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the second neuron-specific transcription factor, or a TALE protein that targets the second neuron-specific transcription factor, and the second polypeptide domain has transcriptional activation activity, and if the first polypeptide domain contains a Cas protein, further administering a gRNA that targets the second neuron-specific transcription factor to stem cells.In some embodiments, the step of reducing the level of a second neuron-specific transcription factor may involve administering a fusion protein comprising two heterologous polypeptide domains to stem cells, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the second neuron-specific transcription factor, or a TALE protein that targets the second neuron-specific transcription factor, and the second polypeptide domain has transcriptional repressive activity; and, if the first polypeptide domain comprises a Cas protein, further administering a gRNA that targets the second neuron-specific transcription factor to the stem cells. In some embodiments, the stem cells may be directly converted to neurons without a pluripotency step. In some embodiments, the stem cells may be pluripotent stem cells, induced pluripotent stem cells, or embryonic stem cells.
[0016] Another aspect of this disclosure provides a system for selecting a polynucleotide for activity as a cell type-specific transcription factor. This system may include a polynucleotide encoding a reporter protein and a cell type marker; a fusion protein comprising two heterologous polypeptide domains, the first polypeptide domain comprising a Cas protein and the second polypeptide domain having transcriptional activation activity; and a library of gRNAs, each of which guide RNAs (gRNAs) target different putative cell type-specific transcription factors. In some embodiments, the cell type-specific transcription factor may be a neuron-specific transcription factor, the cell type marker is a neuron marker, and the neuron marker comprises TUBB3. In some embodiments, the cell type-specific transcription factor may be a muscle-specific transcription factor, the cell type marker is a myogenic marker, and the myogenic marker comprises PAX7. In some embodiments, the cell type-specific transcription factor may be a chondrocyte-specific transcription factor, the cell type marker is a collagen marker, and the collagen marker comprises COL2A1. In some embodiments, the reporter protein may comprise mCherry.
[0017] Another aspect of this disclosure provides isolated polynucleotide sequences capable of encoding the systems detailed herein. Another aspect of this disclosure provides a vector which may comprise an isolated polynucleotide sequence as detailed herein. Another aspect of this disclosure provides cells that may comprise a system, an isolated polynucleotide sequence, or a vector, or a combination thereof, as detailed herein.
[0018] Another aspect of this disclosure provides a method for screening cell type-specific transcription factors. This method may include the steps of: transducing a population of cells in a system detailed herein at an infection multiplicity (MOI) of about 0.2 such that the majority of cells each independently contain one gRNA and target one putative transcription factor; determining the expression level of a reporter protein in each cell; and determining the gRNA level in each cell having high expression of the reporter protein. In some embodiments, high expression of the reporter protein may be defined as being in the top 5% of the cell population; and selecting a putative transcription factor as a cell type-specific transcription factor if the putative transcription factor corresponds to at least two gRNAs enriched in cells having high expression of the reporter protein.
[0019] Another aspect of this disclosure provides a method for screening cell type-specific transcription factor pairs. This method may include the steps of: transducing a population of cells in a system detailed herein at an infection multiplicity (MOI) of about 0.2 such that the majority of cells each independently contain two gRNAs and target two putative transcription factors; determining the expression level of a reporter protein in each cell; and determining the levels of two gRNAs in each cell having high expression of the reporter protein. In some embodiments, high expression of the reporter protein may be defined as being in the top 5% of the cell population; two putative transcription factors are selected as a cell type-specific transcription factor pair if the putative transcription factors correspond to at least two gRNAs enriched in cells having high expression of the reporter protein. In some embodiments, the expression level of the reporter protein in each cell may be determined about 4 days after transduction. In some embodiments, the expression level of the reporter protein in each cell may be determined by flow cytometry. In some embodiments, the gRNA levels in each cell having high expression of the reporter protein may be determined by deep sequencing. In some embodiments, gRNA can increase the expression of reporter proteins in cells by approximately 2–50% compared to untargeted gRNA.
[0020] Another aspect of this disclosure provides polynucleotides encoding muscle-specific transcription factors selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1. Another aspect of the present disclosure provides a system for increasing the expression of muscle-specific genes. This system may comprise (a) a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1; or (b) a fusion protein comprising two heterologous polypeptide domains. In some embodiments, the first polypeptide domain may include a zinc finger protein that targets a muscle-specific transcription factor selected from Cas protein, TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1, or a TALE protein that targets a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1, and the second polypeptide domain has an activity selected from transcriptional activation activity, transcription release factor activity, histone modification activity, nucleic acid association activity, methylase activity, and demethylase activity, and if the first polypeptide domain includes Cas protein, the system further includes a gRNA that targets a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1. In some embodiments, the fusion protein is VP64 dCas9 VP64 Alternatively, it may include dCas9-p300.
[0021] Another aspect of this disclosure provides isolated polynucleotides that can encode the systems detailed herein. Another aspect of this disclosure provides a vector which may comprise isolated polynucleotides as detailed herein. Another aspect of this disclosure provides cells that may contain isolated polynucleotides or vectors as detailed herein.
[0022] Another aspect of this disclosure provides a method for increasing the differentiation of stem cells into myoblasts. This method may include the step of increasing the levels of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in stem cells.
[0023] Another aspect of the present disclosure provides a method for treating a subject requiring treatment. This method may include increasing the levels of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in the stem cells of the subject. In some embodiments, the step of increasing the levels of a muscle-specific transcription factor may include at least one of the following: (a) administering a polynucleotide encoding a muscle-specific transcription factor to stem cells; (b) administering a polypeptide containing a muscle-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells comprising two heterologous polypeptide domains, the first polypeptide domain comprising a Cas protein, a zinc finger protein that targets a muscle-specific transcription factor, or a TALE protein that targets a muscle-specific transcription factor, and the second polypeptide domain having transcriptional activating activity, and further administering a gRNA that targets a muscle-specific transcription factor if the first polypeptide domain comprises a Cas protein.
[0024] This disclosure provides other aspects and embodiments which will become apparent in light of the following detailed description and accompanying drawings. [Brief explanation of the drawing]
[0025] [Figure 1A-1G]High-throughput CRISPRa screening identifies candidate neurogenic transcription factors. (Figure 1A) Schematic diagram of CRISPRa screening for neuronal fate-determining transcription factors in human pluripotent stem cells. VP64dCas9VP64 TUBB3-2A-mCherry reporter cell lines were transduced with a CAS-TF pooled lentivirus library at an MOI of 0.2, and selected for mCherry expression via FACS. gRNA abundance in each cell bin was measured by deep sequencing, and depleted or enriched gRNAs were identified by differential expression analysis. (Figure 1B) The CAS-TF gRNA library was extracted from a previous genome-wide CRISPRa library (Horlbeck, 2016, Compact and highly active next-generation libraries. eLife) and consists of 8505 gRNAs targeting 1496 putative transcription factors. (Figure 1C) TUBB3-2A-mCherry cells were selected based on mCherry signaling to have the highest and lowest 5% expression. Bulk unsorted cell populations were also sampled to establish baseline gRNA distribution. (Figure 1D) Expression difference analysis of normalized gRNA counts between the mCherry high cell population and the unsorted cell population. Red data points indicate FDR < 0.01 by differential DESeq2 analysis (n=3 biological replications). Blue data points indicate a set of 100 scrambled untargeted gRNAs. (Figure 1E) Analysis of TF family types across 17 TFs identified by CAS-TF screening. (Figure 1F) Comparison of mean gene expression across multiple developmental time points and anatomical brain regions for 17 TFs identified by CAS-TF screening and three random sets of 17 TFs. (Figure 1G) Calculation change in gRNA abundance from expression difference analysis between the mCherry high cell population and the mCherry low cell population for all 5 gRNAs from 3 known preneurial TFs compared with random selection of 5 scrambled gRNAs. See also Figures 7A-7D. [Figure 2A-2F]Many candidate factors generate neurons from pluripotent stem cells. (Figure 2A) Validation of 17 factors for TUBB3-2A-mCherry expression 4 days after gRNA transduction (*p<0.05 by global one-way ANOVA with Dunnett's post-hoc test comparing all groups to scramble 1, 1% positive gating for scramble gRNA, n=3 biological replication, error bars represent SEM). (Figure 2B) Relationship between TUBB3-2A-mCherry expression evaluated by individual validation and the multiplicative change in gRNA abundance from library-selected differential expression analysis for all five gRNAs from ATOH1 and NR5A1. (Figure 2C) Validation of 17 factors for induction of panneuronal markers NCAM (top) and MAP2 (bottom) 4 days after gRNA transduction (*p<0.05 by global one-way ANOVA with Dunnett's post-hoc test comparing all groups to scramble 1, n=3 biological replication, error bars represent SEM). (Figure 2D) Immunofluorescence staining of iPSCs to evaluate TUBB3 expression 4 days after transduction with a tetracycline-inducible lentiviral vector containing cDNA encoding the indicated factor, or with an M2rtTA-negative control. Scale bar, 50 μm. (Figure 2E) Immunofluorescence staining of iPSCs to evaluate MAP2 expression with the indicated factor after extended co-culture with astrocytes. Scale bar, 50 μm. (Figure 2F) Immunofluorescence staining of H9 hESCs to evaluate TUBB3 expression 4 days after transduction with the indicated factor. See also Figures 8A-8C, 9A-9D, and 10A-10E. [Figure 3A-3G]Combinatorial gRNA screening identifies cofactors of neuronal differentiation. (Figure 3A) Schematic diagram of combinatorial CRISPRa screening for neuronal fate-determining transcription factors in human pluripotent stem cells. Neurogenic factors and the CAS-TF gRNA library were co-expressed using a dual gRNA expression vector. Two independent screenings were performed on sgASCL1 and sgNGN3. (Figure 3B) Volcano plot of significance (P-value) for fold change in gRNA abundance based on differential DESeq2 analysis between mCherry high cell population and unselected cell population for sgNGN3 paired screening. Red data points indicate FDR < 0.001 (n=3 biological replications). Blue data points indicate a set of 100 scrambled untargeted gRNAs. (Figure 3C) fold change in gRNA abundance for sgASCL1 paired screening compared to sgNGN3 paired screening for all positively enriched gRNAs across both screenings. (Figure 3D) Analysis of TF family type and basal expression levels in pluripotent stem cells for positive hits from both paired screenings. (Figure 3E) Multiplicative changes in gRNA abundance for sets of TFs that are individually inactive but expected to have synergistic activity in sgASCL1 and sgNGN3 paired screenings. (Figure 3F) Verification of TF cofactors for sgNGN3 by TUBB3-2A-mCherry. (Figure 3G) Verification of TF cofactors for sgASCL1 by NCAM staining. (*p<0.05, n=3 biological replications, error bars represent SEM by global one-way ANOVA with Dunnett's post-hoc test comparing all groups to scramble 1). See also Figures 11A-11B and 12A-12D. [Figures 4A-4F]Transcriptional diversity of neurons generated by a single transcription factor. (Figure 4A) Differentially upregulated genes detected in ATOH1 and NEUROG3-induced neurons (FDR < 0.01 and log2 (magnitude change) > 1). (Figure 4B) Enriched gene ontology (GO) terms for a set of 2846 genes shared and upregulated between ATOH1 and NEUROG3. (Figure 4C) Expression levels (log2 (TPM + 1)) of a set of panneuron genes across all analyzed replicate samples. (Figure 4D) Comparison of all detected genes between ATOH1-induced neurons and NEUROG3-induced neurons. Red and blue circles represent genes differentially expressed in NEUROG3 or ATOH1, respectively. (Figure 4E) GO term analysis for markers uniquely upregulated in either NEUROG3 or ATOH1. (Figure 4F) Expression levels (log2 (TPM + 1)) and corresponding z-scores for a set of dopaminergic and glutamatergic markers. [Figure 5A-5N]Transcription and functional maturation of neurons generated by transcription factor pairs. (Figure 5A) Differentially upregulated genes detected in neurons induced from TF pairs (FDR < 0.01 and log2 (magnification change) > 1). (Figure 5B) GO terms enriched in sets of genes differentially upregulated by TF pairs compared to NEUROG3 alone. Upregulation of NTRK3 (Figure 5C) and CDKN1A (Figure 5D) by the addition of RUNX3 or E2F7, respectively. (Figure 5E) SynGO terms for sets of genes differentially upregulated by the addition of LHX8. (Figure 5F) Expression levels for sets of synaptic markers (bottom: log2 (magnification change); top: log2 (TPM + 1)). (Figure 5G) Mean values of membrane properties including (Figure 5H) resting membrane potential (V resting), (Figure 5H) input resistance (Rm), and (Figure 5I) membrane capacitance (Cm) for day 7 neurons generated by NEUROG3 alone or in combination with LHX8. (Figure 5J) Mean values of action potential properties including (Figure 5K) action potential threshold (AP threshold), (Figure 5L) action potential height (AP height), and (Figure 5M) action potential width at half maximum (AP width at half maximum) for day 7 neurons generated by NEUROG3 alone or in combination with LHX8. (Figure 5M) Mean number of generated action potentials against the amplitude of the injected current (*p<0.05 two-way ANOVA). (Figure 5N) Exemplary traces of cells with failed (left), single (center), or multiple (right) action potentials. The corresponding pie charts represent the total percentage of analyzed cells that failed to generate APs (dark shadow), generated a single AP (mid-shadow), or generated multiple APs (faint shadow) in response to a single depolarizing current injection. Regarding Figures 5G to 5L: ns, not significant; *p<0.05. Unpaired t-test (if data passed normality; α=0.05) or Mann-Whitney test (if data failed normality; α=0.05); n=19 cells for NEUROG3 alone; n=22 cells for NEUROG3+LHX8. [Figure 6A-6I]Combinatorial gRNA screening identifies negative regulators of neuronal differentiation. (Figure 6A) Multiplicative changes in gRNA abundance for sgASCL1 paired screening compared to sgNGN3 paired screening for all negatively enriched gRNAs across both screenings. (Figure 6B) Validation of a subset of TFs to assess TUBB3-2A-mCherry-positive cell percentage, and (Figure 6C) expression of the panneuronal marker NCAM (*p<0.05, n=3 biological replications, error bars represent SEM by global one-way ANOVA with Dunnett's post-hoc test comparing all groups to sgNGN3+ scrambled gRNA conditions). (Figure 6D) Validation of the same negative regulators in H9 hESCs. (Figure 6E) Comparison of gRNA effects on neuronal differentiation in iPSCs compared to ESCs. (Figure 6F) Schematic diagram of orthogonal gene activation and repression. (Figure 6G) Relative expression of the top 100 variable genes quantified by z-scores across all three groups tested. (Figure 6H) GO terms enriched by a set of genes differentially expressed in sgNGN3-induced neurons with ZFP36L1 knocked down. (Figure 6I) Exemplary set of genes differentially expressed in relation to neuronal differentiation and morphological development. See also Figures 13A-13C and 14A-14D. [Figures 7A-7D]Generation and characterization of TUBB3-2A-mCherry reporter cell lines. (Figure 7A) Schematic diagram of knock-in of the P2A-mCherry cassette into exon 4 of TUBB3 in human pluripotent stem cell lines using Cas9 nuclease and donor template. (Figure 7B) Targeted activation of endogenous NEUROG2 in pluripotent stem cells by a set of four gRNAs targeting VP64dCas9VP64 and the NEUROG2 promoter. Expression of NCAM (center) and MAP2 (right) by targeted activation of NEUROG2 (n=2 biological replication). (Figure 7C) Flow cytometry of TUBB3-2A-mCherry expression by targeted activation of NEUROG2 by a set of four gRNAs targeting VP64dCas9VP64 and the promoter. (Figure 7D) TUBB3 and MAP2 expression in TUBB3-2A-mCherry cells selected for their highest and lowest mCherry expression levels after NEUROG2 activation by VP64dCas9VP64 and gRNA (n=1 biological replication). [Figures 8A-8C] TF validation using single enriched gRNA. (Figure 8A) Ranking list of the fold change in gRNA abundance between mCherry high-expressing cells and mCherry low-expressing cells in single-factor CAS-TF screening. ASCL1, ATOH7, and ATOH8 all have significantly enriched single gRNAs. Individual validation of sgASCL1, sgATOH7, and sgATOH8 for TUBB3-2A-mCherry expression % (Figure 8B) and MAP2 (left) and NCAM (right) expression (Figure 8C) 4 days after gRNA transduction (*p<0.05, n=3 biological replications by global one-way ANOVA with Dunnett's post-hoc test comparing all groups to scrambled gRNA, *p<0.05, n=3 biological replications, error bars represent SEM). [Figures 9A-9D]Endogenous induction of TFs by VP64dCas9VP64. (Figure 9A) Induction ratios (comparison ratio change to scrambled gRNA, n=2 biological replications) of a subset of 17 TFs enriched by single-factor CAS-TF screening with VP64dCas9VP64 and higher enriched gRNAs. (Figure 9B) Relationship between the induction ratio of each TF and the basal expression of that TF to GAPDH expression. (Figure 9C) Comparison of gRNA enrichment from single-factor CAS-TF screening for two NEUROG2 gRNAs. (Figure 9D) Validation of these two NEUROG2 gRNAs for TF induction and expression of downstream neuron markers (*p<0.05, n=3 biological replications, error bars represent SEM by global one-way ANOVA with Tukey's post-hoc test comparing the two NEUROG2 gRNAs). [Figure 10A-10E] CAS-TF sublibrary gRNA screening. (Figure 10A) Schematic diagram of CRISPRa sublibrary screening for neuron fate-determining transcription factors in human pluripotent stem cells. VP64dCas9VP64 TUBB3-2A-mCherry reporter cell lines were transduced with the CAS-TF pooled lentivirus library at a MOI of 0.2, and sorted for mCherry expression via FACS. gRNA abundance in each cell bin was measured by deep sequencing, and depleted or enriched gRNAs were identified by differential expression analysis. (Figure 10B) The CAS-TF gRNA sublibrary consisted of 3874 gRNAs extracted from several previous genome-wide CRISPRa libraries, targeting 109 putative transcription factors (approximately 33 gRNAs per gene). (Figure 10C) Differential expression analysis of normalized gRNA counts between mCherry-high and mCherry-low cell populations. Red data points indicate FDR < 0.01 by differential DESeq2 analysis (n=3 biological replications). (Figure 10D) Ranking list of enriched gRNA% per gene. (Figure 10E) Examination of 10 factors for TUBB3-2A-mCherry expression 4 days after gRNA transduction (n=2 biological replications). [Figure 11A-11B]Paired gRNA screening using sgASCL1. Volcano plots of significance (P-value) for the multiplicative change in gRNA abundance based on differential DESeq2 analysis between mCherry high vs. mCherry low cell populations for sgASCL1 paired screening: (Figure 11A) mCherry high vs. unselected, and (Figure 11B) mCherry high vs. mCherry low cell populations. Red data points indicate FDR < 0.001 (n=3 biological replications). [Figures 12A-12D] Comparison of single-factor CAS-TF screening and paired CAS-TF screening. (Figures 12A and 12B) sgNGN3 vs single-factor CAS-TF screening for gRNAs that were all positively (Figure 12A) and negatively (Figure 12B) enriched across both screenings, and (Figures 12C and 12D) sgASCL1 vs single-factor CAS-TF screening for gRNAs that were all positively (Figure 12C) and negatively (Figure 12D) enriched across both screenings, showing the fold change in gRNA abundance between mCherry-high and mCherry-low expression cells. [Figure 13A-13C] Gene activation and repression by orthogonal CRISPR systems. (Figure 13A) Targeted expression of ZFP36L1 and HES3 in pluripotent stem cells using dSaCas9KRAB targeting the promoter along with a single gRNA over 7 days (two-sided t-test *p<0.05, n=3 biological replications, error bars represent SEM). Effects of either sgNGN3 (Figure 13B) or sgASLC1 (Figure 13C) on differentiation in ZFP36L1 and HES3 knockdown cell lines (global one-way ANOVA with Dunnett's post-hoc test comparing all groups with either sgNGN3 or sgASLC1 to a control cell line treated with scrambled untargeted Staphylococcus aureus (S. aureus) gRNA *p<0.05, n=3 biological replications, error bars represent SEM). [Figure 14A-14D]Genome-wide expression analysis using orthogonal CRISPR-based gene regulation. (Figure 14A) Differential expression analysis for sgNGN3-induced neurons with HES3 knockdown and (Figure 14B) ZFP36L1 knockdown. Red data points indicate FDR < 0.01 by differential expression analysis using DESeq2 (n=3 biological replications). (Figure 14C) Expression of the Streptococcus pyogenes (S. pyogenes) gRNA target gene, NEUROG3, across the three conditions shown. (Figure 14D) GFP expression in a Streptococcus pyogenes gRNA lentiviral vector was used as a transduction level and gRNA expression surrogate across the three conditions shown. [Figures 15A-15E] Generation and validation of a PAX7-2α-GFP reporter cell line in human ESCs. (Figure 15A) PAX7 gene targeting strategy. A gRNA was designed to target the stop codon of PAX7, and a 2α-GFP donor cassette containing an excisable selection marker was designed for insertion via homologous recombination. (Figure 15B) PCR validation of clones with primers outside the homology arm shows heterozygous insertion of the reporter cassette. (Figure 15C) Sequencing of the 2.6kb product confirms insertion of the 2α-GFP reporter cassette. (Figure 15D) Targeting of the PAX7 promoter in a single clone for CRISPRa-mediated activation demonstrates a shift in GFP. (Figure 15E) The top 15% and bottom 15% of GFP-expressing cells correspond to high and low PAX7 mRNA expression, respectively. [Figures 16A-16E]CRa-TF screening for upstream regulators of PAX7. (Figure 16A) Schematic diagram of CRa-TF screening. H9 Pax7-2a-GFP cells stably expressing VP64dCas9VP64 were transduced with the CRa-TF lentiviral library at an MOI of 0.2. Cells were selected and differentiated for 14 days using the small molecule CHIRON99021 (CHIR) and bFGF. The top 10% and bottom 10% of GFP-expressing cells were selected, and gRNAs were recovered by deep sequencing of DNA. (Figure 16B) Histogram at day 14 of differentiation demonstrates the GFP+ population appearing in three replications of the CRa-TF screening compared to a control without the library. (Figure 16C) MA plot demonstrating significant gRNA hits (p<0.05) in the top 10% compared to unselected cells. (Figure 16D) Verification of individual gRNA hits demonstrating PAX7 induction. (Figure 16E) cDNA delivery of the hit also demonstrates induction of PAX7 (mean ± SEM, n=3). [Figures 17A-17C] Combinatorial CRa-TF screening to identify PAX7 cofactors. (Figure 17A) In a second version of the initial screening, the lentiviral construct was redesigned to include PAX7-targeted gRNA. The lentivirus was transduced at an MOI of 0.2 so that each cell received one copy of PAX7 gRNA and gRNA from the CRa-TF library. (Figure 17B) Histograms at day 7 of differentiation demonstrate a shift in GFP by three replications in the second CRa-TF screening compared to the control without the library. (Figure 17C) Venn diagram showing unique and overlapping significant (p<0.05) hits from both versions of the screening. [Figures 18A-18D]Verification of myogenic lineage induction by CRa-TF hits. (Figure 18A) Schematic diagram of verification by inducible expression of hits. H9 PAX7-2α-GFP expressing TetO-VP64dCasVP64 was transduced with individual gRNA hits and rtTA3. Cells were differentiated in the presence of dox for 28 days. Terminal differentiation was induced by removing dox 14 days before analysis. (Figure 18B) RNA analysis after terminal differentiation demonstrates increased PAX7 expression compared to untargeted gRNA control. (Figure 18C) RNA analysis after terminal differentiation demonstrates increased MYOG expression compared to untargeted gRNA control (mean ± SEM, n=3). (Figure 18D) Image of cells. [Figures 19A-19B] Generation and validation of polyclonal trans-activator strains. (Figure 19A) Schematic diagram of the VP64dCas9VP64-2A-blastosidine expression cassette. (Figure 19B) Activation of endogenous NGN2 after transduction of NGN2. [Figures 20A-20C] TF-targeted gRNA screening to identify regulators of chondrogenesis. (Figure 20A) Schematic diagram of an experiment demonstrating the generation of activator strains in lentiviral packaging of reporter strains and gRNA libraries. After transduction and chondrogenesis of the libraries, high-GFP and low-GFP cells were selected, and gRNAs were recovered from both populations. Next-generation sequencing was used to compare differences in gRNA expression. (Figure 20B) Histogram of GFP fluorescence after library transduction and chondrogenesis. The gate indicates the high-GFP and low-GFP selected populations. (Figure 20C) Volcano plot showing gRNAs significantly enriched in the high-GFP and low-GFP populations (red), as well as gRNAs that did not meet the significance criteria but had high (>3) log2 (magnitude change). See Appendix B for larger volcano plots. [Figure 21A-21D]Validation of SOX9 in directed differentiation. (Figure 21A) Schematic diagram of the experimental design. Differentiation of reporter hiPSCs with SOX9 overexpression into hard segments, followed by flow cytometry at day 6. (Figure 21B) Flow cytometry at day 6 of unmodified strains compared with reporter strains containing (red) and without (black) SOX9 lentivirus. (Figure 21C) Comparison of GFP fluorescence at day 6 and day 21 of differentiation (blue). [Modes for carrying out the invention]
[0026] Cell type-specific transcription factors, as well as methods for using these factors to increase the expression of cell type-specific genes, methods for increasing the maturation of stem cell-induced neurons, methods for increasing the efficiency of stem cell conversion to neurons, and methods for treating subjects requiring treatment are described herein. High-throughput pooled CRISPR activation (CRISPRa) screening for mapping human cell fate regulators and profiling the contribution of putative human transcription factors to neuronal cell fate designation in pluripotent stem cells is further described herein. CRISPRa screening was used as a high-throughput approach to profile thousands of putative transcription factors in the human genome. CRISPR-based gRNA libraries are easier to design and scale than conventional methods and are more readily applicable to testing combinatorial gene interactions and acquiring information from non-coding genomes. Neuronal commitment reporters were used to profile the neurogenic activity of all transcription factors in human pluripotent stem cells. Single-factor screening was performed to identify major regulators of human neuronal fate and to identify many known and previously uncharacterized TFs. Combinatorial screening was performed to identify synergistic and antagonistic TF interactions that enhance or diminish neuronal differentiation. TFs that increase conversion efficiency, influence subtype designation, and improve the maturation of in vitro-induced human neurons were identified.
[0027] In summary, this study highlights the usefulness of DNA targeting systems, such as CRISPR-based technologies, for regulating endogenous gene expression and provides a framework for identifying the causal role of cell fate regulators in defining any cell type of interest. The set of candidate preneural transcription factors selected from the tests detailed herein may serve as a resource for establishing protocols for generating all cell types in the human brain.
[0028] 1.Definition Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. In case of any conflict, the definitions in this document shall prevail. Preferred methods and materials are described below, but similar or equivalent methods and materials may be used in carrying out or testing the present invention. All publications, patent applications, patents and other references referenced herein are incorporated by reference in their entirety. The materials, methods and examples disclosed herein are illustrative and not intended to be limiting.
[0029] As used herein, the terms “comprise(s),” “include(s),” “having,” “has,” “can,” and “contain(s),” and their variations, are intended to be open-ended transitional phrases, terms, or words that do not preclude the possibility of additional actions or structures. The singular forms “a,” “and,” and “the” include plural references unless the context makes it clear that they have a different meaning. This disclosure also considers other embodiments present herein, whether expressly indicated or not, that “comprising,” “consisting of,” and “consisting essentially of.”
[0030] In the enumeration of numerical ranges in this specification, each intervening number having the same precision is explicitly considered. For example, for the range 6 to 9, the numbers 7 and 8 are considered in addition to 6 and 9, and for the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9 and 7.0 are explicitly considered. As used herein, the term “about” applied to one or more values of interest refers to a value similar to the stated reference value. In certain aspects, unless otherwise specified or obvious from the context, the term “about” refers to a range of values that fall within (greater than or less than) 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction of the stated reference value (except where such numbers exceed 100% of the possible value).
[0031] As used interchangeably herein, “adeno-associated virus” or “AAV” refers to a small virus belonging to the genus Dependovirus of the family Parvoviridae that infects humans and some other primate species. AAV is not currently known to cause disease, and as a result, it elicits a very mild immune response.
[0032] As used herein, “amino acid” refers to naturally occurring amino acids and unnaturally synthesized amino acids, as well as amino acid analogs and amino acid mimics that function in a manner similar to naturally occurring amino acids. Naturally occurring amino acids are encoded by the genetic code. Amino acids may be referred herein by their commonly known three-letter symbols or by the single-letter symbols recommended by the IUPAC-IUB Biochemical Nomenclature Committee. Amino acids include side chains and polypeptide backbone portions.
[0033] As used herein, the term "binding region" refers to a region within a nuclease target region that a nuclease recognizes and binds to. As used herein, "coding sequence" or "encoding nucleic acid" means a nucleic acid (RNA or DNA molecule) containing a nucleotide sequence that codes for a protein. The coding sequence may further include start and termination signals operably linked to a regulatory element containing a promoter and a polyadenylation signal, which can direct expression in cells of an individual or mammal to which the nucleic acid has been administered. The coding sequence may be codon-optimized. As used herein, “complementary” or “complementary” refers to Watson-Crick (e.g., AT / U and CG) or Hoogsteen-type base pairing between nucleotides or nucleotide analogs of nucleic acid molecules. “Complementarity” refers to a property shared between two nucleic acid sequences such that, when the two nucleic acid sequences are aligned antiparallel to each other, the nucleotide bases at each position are complementary.
[0034] The terms “control,” “reference level,” and “reference” are used interchangeably herein. A reference level may be a predetermined value or range used as a benchmark to evaluate results measured against it. As used herein, “control group” refers to a group of subjects used as a control. A predetermined level may be a cutoff value from the control group. A predetermined level may be the mean from the control group. The cutoff value (or predetermined cutoff value) may be determined by adaptive indicator model (AIM) methodology. The cutoff value (or predetermined cutoff value) may be determined by receiver operated curve (ROC) analysis from biological samples of a patient group. ROC analysis, commonly known in the biological field, is used to determine the ability of a test to distinguish one condition from another, for example, to determine the performance of each marker in identifying patients with CRC. A description of ROC analysis is provided in PJ Heagerty et al. (Biometrics 2000, 56, 337-44), the entire disclosure of which is incorporated herein by reference. Alternatively, the cutoff value may be determined by quartile analysis of biological samples from the patient group. For example, the cutoff value may be determined by selecting a value corresponding to any value in the 25th–75th percentile range, preferably the 25th, 50th, or 75th percentile, more preferably the 75th percentile. Such statistical analysis may be performed using any method known in the art and can be carried out through any number of commercially available software packages (e.g., Analyse-it Software Ltd., Leeds, UK; StataCorp LP, College Station, TX; SAS Institute Inc., Cary, NC). Healthy or normal levels or ranges for target or protein activity may be defined according to standard practice. Controls may be subjects or cells without the agonists detailed herein. Controls may be subjects or samples thereof whose disease status is known.The subject, or sample thereof, may be health, disease, disease before treatment, disease during treatment, or disease after treatment, or a combination thereof.
[0035] As used herein, the term "fusion protein" refers to a chimeric protein created through the translation of two or more linked genes that originally encoded separate proteins. Translation of a fusion gene yields a single polypeptide possessing functional properties derived from each of the original separate proteins. As used herein, “genetic construct” refers to a DNA or RNA molecule containing a polynucleotide that codes for a protein. The coding sequence includes start and terminate signals operably ligated to a regulatory element containing a promoter and a polyadenylation signal that can direct expression in the cells of an individual to which the nucleic acid molecule has been administered. As used herein, the term “expressible form” refers to a gene construct containing the necessary regulatory elements operably ligated to a protein-coding sequence so that, when present in the cells of an individual, the coding sequence is expressed.
[0036] As used herein, "genome editing" refers to altering a gene. Genome editing may include correcting or restoring a mutated gene. Genome editing may include knocking out a gene, such as a mutated gene or a normal gene. Genome editing can be used to treat diseases or enhance muscle repair by altering a gene of interest.
[0037] As used herein, "identical" or "identity" in the context of two or more nucleic acid or polypeptide sequences means that the sequences have a specified percentage of residues that are the same across a specified region. This percentage can be calculated by optimally aligning the two sequences, comparing them across the specified region, determining the number of positions where identical residues occur in both sequences, obtaining the number of matching positions, dividing the number of matching positions by the total number of positions in the specified region, and multiplying this result by 100 to obtain the percentage of sequence identity. If the two sequences are of different lengths, or if the alignment results in one or more attached ends and the specified region being compared contains only a single sequence, the residues of the single sequence are included in the denominator of the calculation, but not in the numerator. When comparing DNA and RNA, thymine (T) and uracil (U) may be considered equivalent. Identity can be performed manually or by using a computer sequencing algorithm such as BLAST or BLAST 2.0.
[0038] As used interchangeably herein, “mutant gene” or “mutated gene” refers to a gene that has undergone a detectable mutation. A mutant gene has undergone changes such as loss, acquisition, or exchange of genetic material that affect the normal transmission and expression of the gene. As used herein, “disrupted gene” refers to a mutant gene that has a mutation that causes an early stop codon. The disrupted gene product is cleaved compared to the full-length non-disrupted gene product. As used herein, a "normal gene" refers to a gene that has not undergone any alteration, such as loss, acquisition, or exchange of genetic material. Normal genes undergo normal gene transmission and gene expression. For example, a normal gene may be a wild-type gene.
[0039] As used herein, “nucleic acid,” “oligonucleotide,” or “polynucleotide” means at least two nucleotides covalently linked together. A single-stranded description also defines the sequence of the complementary strand. Thus, a polynucleotide also includes the complementary strand of the described single-stranded sequence. Many variants of a polynucleotide can be used for the same purpose as a given polynucleotide. Thus, a polynucleotide also includes substantially identical polynucleotides and their complements. A single strand provides a probe that can hybridize with a target sequence under stringent hybridization conditions. Thus, a polynucleotide also includes probes that hybridize under stringent hybridization conditions. A polynucleotide may be single-stranded or double-stranded, or may contain portions of both double-stranded and single-stranded sequences. Polynucleotides can be nucleic acids, natural or synthetic DNA, genomic DNA, cDNA, RNA, or hybrids, and may contain combinations of deoxyribonucleotides and ribonucleotides, as well as combinations of bases, including, for example, uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, and isoguanine. Polynucleotides can be obtained by chemical synthesis or recombinant methods.
[0040] As used herein, “operably linked” means that the expression of a gene is under the control of a promoter to which the gene is spatially connected. The promoter may be located 5' (upstream) or 3' (downstream) of the gene under its control. The distance between the promoter and the gene is approximately the same as the distance between the promoter and the gene it controls in the gene to which the promoter is induced. As is known in the art, variations in this distance can be considered without loss of promoter function.
[0041] As used herein, "partially functional" refers to a protein encoded by a mutant gene that has biological activity smaller than a functional protein but greater than a non-functional protein.
[0042] A peptide or polypeptide is a linked sequence of two or more amino acids connected by peptide bonds. Polypeptides can be native, synthetic, modified, or a combination of native and synthetic. Peptides and polypeptides include proteins such as binding proteins, receptors, and antibodies. The terms "polypeptide," "protein," and "peptide" are used interchangeably herein. The "primary structure" refers to the amino acid sequence of a particular peptide. The "secondary structure" refers to the locally ordered, three-dimensional structure within a polypeptide. These structures are commonly known as domains, such as enzyme domains, extracellular domains, transmembrane domains, pore domains, and cytoplasmic tail domains. A "domain" is a portion of a polypeptide that forms a compact unit and is typically 15–350 amino acid long. Exemplary domains include those with enzymatic or ligand-binding activity. Typical domains consist of smaller tissue sections, such as β-sheets and α-helix extensions. "Tertiary structure" refers to the complete three-dimensional structure of a polypeptide monomer. "Quaternary structure" refers to a three-dimensional structure formed by the non-covalent association of independent tertiary units. "Motif" is a portion of a polypeptide sequence containing at least two amino acids. Motifs can be 2-20, 2-15, or 2-10 amino acid long. In some embodiments, motifs contain 3, 4, 5, 6, or 7 consecutive amino acids. A domain may consist of a series of motifs of the same type.
[0043] As used interchangeably in this specification, “premature stop codon” or “out-of-frame stop codon” refers to a nonsense mutation in the DNA sequence that results in a stop codon at a location not normally found in wild-type genes. Premature stop codons can cleave proteins or shorten them compared to their full-length version.
[0044] As used herein, “promoter” means a synthetic or naturally occurring molecule that can confer, activate, or enhance the expression of a nucleic acid in a cell. A promoter may include one or more specific transcriptional regulatory sequences for further enhancing nucleic acid expression and / or altering its spatial and / or temporal expression. A promoter may also include distal enhancer or repressor elements, which may be located several thousand base pairs from the transcription start site. Promoters may be derived from sources including viruses, bacteria, fungi, plants, insects, and animals. Promoters may constitutively, differentially, in relation to the cell, tissue, or organ in which expression occurs, in relation to the developmental stage in which expression occurs, or in response to external stimuli such as physiological stress, pathogens, metal ions, or inducers. Typical examples of promoters include the bacteriophage T7 promoter, bacteriophage T3 promoter, SP6 promoter, lac operator-promoter, tac promoter, SV40 late promoter, SV40 early promoter, RSV-LTR promoter, CMV IE promoter, SV40 early promoter or SV40 late promoter, human U6 (hU6) promoter, and CMV IE promoter.
[0045] As used herein, “sample” or “test sample” may mean any sample in which the presence and / or level of a target is detected or determined, or any sample containing a DNA targeting system or its components as detailed herein. A sample may include a liquid, solution, emulsion, or suspension. A sample may include a medical sample. A sample may include any biological fluid or tissue, such as blood, whole blood, blood fractions, such as plasma and serum, muscle, interstitial fluid, sweat, saliva, urine, tears, synovial fluid, bone marrow, cerebrospinal fluid, nasal secretions, sputum, amniotic fluid, bronchoalveolar lavage fluid, gastric lavage, vomit, excrement, lung tissue, peripheral blood mononuclear cells, total leukocytes, lymph node cells, spleen cells, tonsil cells, cancer cells, tumor cells, bile, digestive fluids, skin, or a combination thereof. In some embodiments, the sample includes aliquots. In other embodiments, the sample includes biological fluids. A sample may be obtained by any means known in the art. The sample can be used directly as obtained from the patient, or it can be pretreated by filtration, distillation, extraction, concentration, centrifugation, inactivation of interfering components, addition of reagents, etc., in order to modify the properties of the sample in a manner discussed herein or as known in the art.
[0046] As used interchangeably in this specification, “spacer” and “spacer region” refer to a region within a TALE or zinc finger target region that lies between, but is not part of, the binding regions of two TALE or zinc finger proteins.
[0047] As used herein, “subject” or “patient” may mean an animal that desires or requires the composition or method described herein. The subject may be human or non-human. The subject may be any vertebrate. The subject may be a mammal. The mammal may be a primate or non-primate. The mammal may be a non-primate such as, for example, a dog, cat, horse, cow, pig, mouse, rat, camel, llama, goat, rabbit, sheep, hamster, and guinea pig. The mammal may be a primate such as a human. The mammal may be a non-human primate such as, for example, a monkey, crab-eating macaque, rhesus macaque, chimpanzee, gorilla, orangutan, and gibbon. The subject may be of any age or developmental stage, such as an adult, adolescent, or infant. The subject may be male. The subject may be female. In some embodiments, the subject has certain genetic markers. The subject may be undergoing other forms of treatment.
[0048] "Substantially identical" can mean that the first and second amino acid or polynucleotide sequences are at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% across regions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 1100 amino acids or nucleotides, respectively.
[0049] A "transcription activator-like effector" or "TALE" refers to a protein structure that recognizes and binds to a specific DNA sequence. A "TALE DNA-binding domain" refers to a DNA-binding domain containing an array of tandem 33-35 amino acid repeats, also known as RVD modules, each specifically recognizing a single base pair of DNA. RVD modules can be sequenced in any order to assemble an array that recognizes a defined sequence. The binding specificity of the TALE DNA-binding domain is determined by the RVD array followed by a 20-amino acid single cleavage repeat. A "repeat variable diresidue" or "RVD" refers to a pair of adjacent amino acid residues within a DNA recognition motif (also known as an "RVD module"), containing the 33-35 amino acid of the TALE DNA-binding domain. RVDs determine the nucleotide specificity of the RVD module. RVD modules can be combined to generate an RVD array. As used herein, "RVD array length" refers to the number of RVD modules corresponding to the length of the nucleotide sequence within the TALEN target region recognized by the TALEN; that is, the TALE DNA-binding domain may have 12 to 27 RVD modules, each containing an RVD and recognizing a single base pair of DNA. Specific RVDs recognizing each of the four possible DNA nucleotides (A, T, C, and G) have been identified. Since the TALE DNA-binding domain is modular, repeats recognizing four different DNA nucleotides can be linked together to recognize any specific DNA sequence. These targeted DNA-binding domains can then be combined with catalytic domains to create functional enzymes, including artificial transcription factors, methyltransferases, integrases, nucleases, and recombinases.
[0050] As used herein, "target gene" refers to any nucleotide sequence that encodes a known gene product or a putative gene product. A target gene may be a mutant gene involved in a genetic disorder. In certain embodiments, the target gene is a gene that encodes a transcription factor. As used herein, the term "target region" refers to the region of a target gene designed to be bound by a CRISPR / Cas9-based gene editing system. As used herein, "transgene" refers to a gene or genetic material containing a gene sequence isolated from one organism and introduced into another organism. This non-natural segment of DNA may retain the ability to produce RNA or protein in the transgenic organism, or it may alter the normal function of the genetic code of the transgenic organism. The introduction of a transgene may alter the phenotype of the organism.
[0051] When referring to the protection of a subject from a disease, "treatment" or "treating" means suppressing, inhibiting, improving, or completely eliminating the disease. Preventing a disease involves administering the composition of the present invention to a subject before the onset of the disease. Suppressing a disease involves administering the composition of the present invention to a subject after the induction of the disease but before its clinical manifestation. Inhibiting or improving a disease involves administering the composition of the present invention to a subject after the clinical manifestation of the disease.
[0052] As used herein with respect to polynucleotides, “variant” means (i) a portion or fragment of a reference nucleotide sequence; (ii) a complement of a reference nucleotide sequence or a portion thereof; (iii) a nucleic acid substantially identical to a reference nucleic acid or its complement; or (iv) a nucleic acid that, under stringent conditions, hybridizes with a reference nucleic acid, its complement, or a sequence substantially identical thereto.
[0053] A "variant" of a peptide or polypeptide is one whose amino acid sequence differs due to an insertion, deletion, or conservative substitution of amino acids, but which retains at least one biological activity. A variant can also mean a protein having a substantially identical amino acid sequence to a reference protein, which has an amino acid sequence that retains at least one biological activity. Typical examples of "biologica activity" include the ability to bind to a specific antibody or polypeptide, or the ability to promote an immune response. A variant can mean its functional fragment. A variant can also mean multiple copies of a polypeptide. These multiple copies may be in tandem or separated by a linker. Conservative substitution of amino acids, i.e., replacing an amino acid with a different amino acid that has similar properties (e.g., hydrophilicity, degree and distribution of charged regions), is typically recognized in the art as involving only minor changes. These minor changes can be identified in part by considering the hydropathic index of amino acids as understood in the art. Kyte et al., J. Mol. Biol. 157:105-132 (1982). The hydroxyl index of an amino acid is based on consideration of its hydrophobicity and charge. It is known in the art that amino acids with similar hydroxyl indices can be substituted and still retain protein function. In one embodiment, amino acids with hydroxyl indices of ±2 are substituted. The hydrophilicity of amino acids can also be used to identify substitutions that result in proteins that retain biological function. Considering the hydrophilicity of amino acids in the context of peptides allows for the calculation of the peptide's greatest local mean hydrophilicity. Substitutions can be carried out with amino acids having hydrophilicity values within ±2 of each other. Both the hydroxyl index and hydrophilicity value of an amino acid are influenced by the specific side chain of that amino acid. Consistent with these findings, it is understood that amino acid substitutions of biological function and compatibility depend on the relative similarity of the amino acids, particularly their side chains, as revealed by hydrophobicity, hydrophilicity, charge, size, and other properties.
[0054] As used herein, “vector” means a nucleic acid sequence containing an origin of replication. A vector may be a viral vector, a bacteriophage, a bacterial artificial chromosome, or a yeast artificial chromosome. A vector may be a DNA or RNA vector. A vector may be a self-replicating extrachromosomal vector, preferably a DNA plasmid. For example, a vector may encode the Cas9 protein and at least one gRNA molecule.
[0055] As used herein, "zinc finger" refers to a protein that recognizes and binds to a DNA sequence. The zinc finger domain is the most common DNA-binding motif in the human proteome. A single zinc finger contains approximately 30 amino acids, and this domain typically functions by binding three consecutive base pairs of DNA via the interaction of single amino acid side chains per base pair.
[0056] Unless otherwise defined herein, scientific and technical terms used in connection with this disclosure shall have meanings generally understood by those skilled in the art. For example, any nomenclature and techniques used relating to cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are well known and commonly used in the art. While the meaning and scope of terms should be clear, where there is potential ambiguity, the definitions provided herein shall take precedence over any dictionary or external definitions. Furthermore, unless otherwise required by context, singular terms shall include plural forms, and plural terms shall include singular forms.
[0057] 2. Transcription factors Cell type-specific transcription factors are provided herein. Transcription factors (TFs) are proteins that regulate the rate of transcription of genetic information from DNA to messenger RNA by binding to specific DNA sequences. TFs regulate genes to ensure that genes are expressed in the correct cells, at the correct time, and in the correct amount throughout the lifespan of cells and organisms. TFs transmit complex patterns of endogenous and exogenous signals to dynamic gene expression programs that define cell type identity. Groups of TFs may function in a coordinated manner to, for example, direct cell division, cell proliferation, and cell death throughout life; cell migration and organization (body plan) during embryonic development; and intermittently in response to extracellular signals such as hormones. TFs may act alone or in complex with other proteins, for example, by promoting or blocking the recruitment of RNA polymerase. TFs may be specific to a particular cell type. TFs may be neuron-specific. TFs may be muscle-specific. TFs may be chondrocyte-specific. TF may be specific to any cell type, such as cells of tissues selected from bone marrow, skin, skeletal muscle, adipose tissue, and peripheral blood. The cells may be muscle cells (e.g., smooth muscle cells, skeletal muscle cells, and cardiomyocytes), epithelial cells, endothelial cells, urothelial cells, fibroblasts, hepatocytes, myoblasts, neurons, osteoblasts, osteoclasts, T cells, keratinocytes, hair follicle cells, human umbilical vein endothelial cells (HUVECs), umbilical cord blood cells, neural progenitor cells, chondrocytes, chondrocytes, cholangiocarcinomas, pancreatic islet cells, thyroid cells, parathyroid cells, adrenal cells, hypothalamic cells, pituitary cells, ovarian cells, testicular cells, salivary gland cells, adipocytes, progenitor cells, hematopoietic stem cells (HSCs), adipose mesenchymal stem cells (MSCs), bone marrow mesenchymal stem cells (MSCs), oligodendrocytes, oligodendrocyte progenitor cells, neutrophils, basophils, eosinophils, lymphocytes, monocytes, or cardiomyocytes. A TF may be, for example, a member of the C2H2 ZF, bHLH, or HMG / Sox DNA-binding domain family. A TF may be an activating TF (which activates or increases gene expression), or a TF may be a repressive TF (which suppresses or decreases gene expression).
[0058] TFs can regulate gene expression through various mechanisms. For example, TFs can stabilize or block the binding of RNA polymerase to DNA. TFs can recruit coactivator or corepressor proteins to transcription factor DNA complexes. TFs can directly or indirectly catalyze the acetylation or deacetylation of histone proteins. Histone acetyltransferase (HAT) activity acetylates histone proteins, weakening the association between DNA and histones, which can make DNA more accessible for transcription, thereby upregulating transcription. Histone deacetylase (HDAC) activity deacetylates histone proteins, strengthening the association between DNA and histones, which can make DNA less accessible for transcription, thereby downregulating transcription. TFs can influence the three-dimensional looping of DNA, which in turn can influence gene expression.
[0059] Polynucleotides encoding at least one transcription factor, or the transcription factor polypeptide itself, are provided herein. In some embodiments, the transcription factor is an endogenous transcription factor. Here, “endogenous” refers to a copy of the gene encoding TF at its natural location in the genome of the subject in chromosomal DNA. The transcription factor may direct the expression of a gene in a neuron. The transcription factor may direct the differentiation of a cell into a neuron. In some embodiments, the first transcription factor may act in conjunction with a second transcription factor. The transcription factor may be a putative. The transcription factor may be selected or identified as a neuron-specific transcription factor. Neuron-specific transcription factors may be called neurogenic factors.
[0060] Cell type-specific transcription factors can be activated or repressed. For example, activated or positive neuron-specific transcription factors increase the differentiation of cells into neurons or increase gene expression in neurons. Increased expression of positive neuron-specific transcription factors may improve or increase the differentiation of cells into neurons or increase gene expression in neurons. Repressive or negative neuron-specific transcription factors inhibit the differentiation of cells into neurons or inhibit gene expression in neurons. Knockdown or inhibition of the expression of negative neuron-specific transcription factors may improve or increase the differentiation of cells into neurons or increase gene expression in neurons. Regulation of the expression or protein levels of neuron-specific transcription factors may directly convert stem cells into neurons without the pluripotency stage.
[0061] A first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 is provided herein. Polynucleotides encoding the first neuron-specific transcription factor are further provided. In some embodiments, the first neuron-specific transcription factor is selected from NGN3 and ASCL1, or a combination thereof.
[0062] In some embodiments, a second neuron-specific transcription factor or a polynucleotide encoding a second neuron-specific transcription factor is also provided herein. The first neuron-specific transcription factor may be combined with the second neuron-specific transcription factor. In such embodiments, the first neuron-specific transcription factor may be selected from NGN3 and ASCL1, or a combination thereof. The second neuron-specific transcription factor may be (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, PLAGL2 (listed as "Positive Single Factor CRa-TF" in Table 1). (ii) Selected from "Positive sgNGN3+CRa-TF" in Table 1); (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3 (Selected from "Positive sgNGN3+CRa-TF" in Table 1);(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, C DX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SO X1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, E2F7 (Positive sgASCL1+CRa-TF in Table 1) (iv) Selected from "sgASCL1+CRa-TF"); (iv) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3 (Selected from "Negative Single Factor CRa-TF" in Table 2);(v)HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX22 ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791 (Negative sgNGN3+CRa-TF in Table 2) Selected from sgNGN3+CRa-TF); and (vi)ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, H The following can be selected: ESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, BSX (selected from "Negative sgASCL1+CRa-TF" in Table 2).
[0063] In some embodiments, the second neuron-specific transcription factor is selected from NEUROG3, SOX4, and SOX9. In some embodiments, the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, NKX2-2, HES3, and ZFP36L1. In some embodiments, the second neuron-specific transcription factor is an activating transcription factor selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, and NKX2-2. In some embodiments, the second neuron-specific transcription factor is an inhibitory transcription factor selected from HES3 and ZFP36L1.
[0064] Muscle-specific transcription factors are further provided herein. These may be selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1. Polynucleotides encoding muscle-specific transcription factors are further provided.
[0065] 3. CRISPR / Cas-based gene editing systems The system may be a CRISPR / Cas-based gene editing system. A CRISPR / Cas-based gene editing system includes a nuclease-inactive Cas protein (dCas) or dCas fusion protein, or a promoter or regulatory element of the TF gene or a portion thereof, which can cause activation or suppression of endogenous TF expression. The system may be a CRISPR / Cas9-based gene editing system. As used interchangeably herein, "Clustered Regularly Interspaced Short Palindromic Repeat" and "CRISPR" refer to loci containing multiple short serial repeat sequences found in the genomes of approximately 40% of sequenced bacteria and approximately 90% of sequenced archaea. The CRISPR system is a microbial nuclease system involved in defense against invading phages and plasmids, providing a form of adaptive immunity. CRISPR loci in microbial hosts contain a combination of CRISPR-related (Cas) genes and non-coding RNA elements that can program the specificity of CRISPR-mediated nucleic acid cleavage. Short segments of foreign DNA, called spacers, are incorporated into the genome between CRISPR repeat sequences and serve as "memories" of past exposure. Cas proteins, such as the Cas9 protein, form a complex with the 3' end of sgRNA (also interchangeably referred to herein as "gRNA"), and the protein-RNA pair recognizes its genomic target by complementary base pairing between the 5' end of the sgRNA sequence and a predefined 20 bp DNA sequence known as a protospacer. This complex is directed to homologous loci of pathogen DNA via regions encoded within crRNA, i.e., protospacers, and protospacer-adjacent motifs (PAMs) in the pathogen genome. Non-coding CRISPR arrays are transcribed and cleaved into short crRNAs containing individual spacer sequences within a serial repeat sequence, and these spacer sequences direct Cas nucleases to target sites (protospacers).By simply replacing the 20bp recognition sequence of the expressed sgRNA, the Cas9 nuclease can be directed to a new genomic target. CRISPR spacers are used to recognize and silence exogenous genetic elements, similar to RNAi in eukaryotes.
[0066] Three classes of CRISPR systems (Type I, II, and III effector systems) are known. The Type II effector system uses a single effector enzyme, such as Cas9, to cleave dsDNA by performing targeted double-strand disruption of DNA in four sequential steps. Compared to the Type I and III effector systems, which require multiple distinct effectors acting as a complex, the Type II effector system can function in alternative situations, such as eukaryotic cells. The Type II effector system consists of a long precrRNA transcribed from a spacer-containing CRISPR locus, the Cas9 protein, and tracrRNA involved in precrRNA processing. The tracrRNA hybridizes with a repeating region that separates the spacer in the precrRNA, thus initiating dsRNA cleavage by intrinsic factor RNase III. This cleavage is followed by a second cleavage event within each spacer by Cas9, generating mature crRNA that remains associated with tracrRNA and Cas9, forming the Cas9:crRNA-tracrRNA complex.
[0067] The Cas9:crRNA-tracrRNA complex unwinds the DNA double strand and cleaves it, searching for a sequence that matches the crRNA. Target recognition occurs through the detection of complementarity between the “protospacer” sequence in the target DNA and the remaining spacer sequences in the crRNA. Cas9 mediates the cleavage of target DNA if the precise protospacer adjacency motif (PAM) is also present at the 3' end of the protospacer. For protospacer targeting, the protospacer adjacency motif (PAM), which is a short sequence recognized by the Cas9 nuclease required for DNA cleavage, must follow immediately after the sequence. Different type II systems have different PAM requirements. The Streptococcus pyogenes CRISPR system may have this Cas9 PAM sequence as 5'-NRG-3' (where R is either A or G) (SpCas9), which characterized the specificity of this system in human cells. A unique capability of CRISPR / Cas9-based gene editing systems is their straightforward ability to simultaneously target multiple distinct genomic loci by co-expressing a single Cas9 protein with two or more sgRNAs. For example, while the Streptococcus pyogenes type II system naturally prefers the use of the "NGG" sequence (where "N" can be any nucleotide), the engineered system also accepts other PAM sequences such as "NAG" (Hsu et al., Nature Biotechnology 2013 doi:10.1038 / nbt.2647). Similarly, Cas9 derived from Neisseria meningitidis (NmCas9) typically has the natural PAM NNNNGATT (SEQ ID NO: 12), but exhibits activity across a variety of PAMs, including the highly degenerate NNNNGNNN PAM (SEQ ID NO: 13) (Esvelt et al. Nature Methods 2013 doi:10.1038 / nmeth.2681).
[0068] In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRR(R=A or G) (SEQ ID NO: 8) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRN(R=A or G) (SEQ ID NO: 9) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRT(R=A or G) (SEQ ID NO: 10) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRV(R=A or G) (SEQ ID NO: 11) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In the above embodiment, N may be any nucleotide residue, such as A, G, C, or T. The Cas9 molecule can be manipulated to change its PAM specificity.
[0069] An engineered form of the type II effector system of Streptococcus pyogenes has been shown to function in human cells for genomic manipulation. In this system, the Cas9 protein is directed to a genomic target site by a synthetically reconstituted “guide RNA” (“gRNA,” also used interchangeably herein as chimeric single guide RNA “sgRNA”), which is a crRNA-tracrRNA fusion that typically eliminates the need for RNase III and crRNA processing. CRISPR / Cas9-based manipulative systems for use in genome editing and treatment of genetic diseases are provided herein. CRISPR / Cas9-based manipulative systems can be designed to target any gene, including genes involved in genetic diseases, aging, tissue regeneration, or wound healing. CRISPR / Cas9-based gene editing systems may include the Cas9 protein or Cas9 fusion protein and at least one gRNA. In certain embodiments, the system includes two gRNA molecules. Cas9 fusion proteins may contain domains with activity different from that endogenous to Cas9, such as a transactivation domain.
[0070] The target gene may be involved in any other process in which cell differentiation or gene activation is desired, or it may have mutations such as frameshift mutations or nonsense mutations. In some embodiments, the target or target gene includes a putative transcription factor gene or a portion thereof. The CRISPR / Cas9-based gene editing system may or may not mediate off-target changes to protein-coding regions of the genome. The CRISPR / Cas9-based gene editing system may bind to and recognize target regions.
[0071] a. Cas protein CRISPR / Cas9-based gene editing systems may include Cas9 proteins or Cas fusion proteins. In some embodiments, the Cas protein is a Cas12 protein (also called Cpf1), such as the Cas12a protein. Cas12 proteins may originate from any bacterium or archaeal species, including, but not limited to, Francisella novicida, Acidaminococcus species, Lachnospiraceae species, and Prevotella species. In some embodiments, the Cas protein is a Cas9 protein. The Cas9 protein is an endonuclease that cleaves nucleic acids, is encoded by the CRISPR locus, and is involved in the type II CRISPR system. The Cas9 protein is not limited to these, but is also found in Streptococcus pyogenes, Staphylococcus aureus (S. aureus), Acidovorax avenae, Actinobacillus pleuropneumoniae, Actinobacillus succinogenes, Actinobacillus suis, Actinomyces species, Cycliphilus denitrificans, Aminomonas paucivorans, Bacillus cereus, Bacillus smithii, and Bacillus thuringiensis. Bacteroides species (thuringiensis), Blastopirellula marina, Bradyrhizobium species, Brevibacillus laterosporus, Campylobacter coli, Campylobacter jejuniJejuni), Campylobacter lari, Candidatus Puniceispirillum, Clostridium cellulolyticum, Clostridium perfringens, Corynebacterium accolens, Corynebacterium diphtheria, Corynebacterium matruchotii, Dinoroseobacter shibae, Eubacterium dolichum, Gammaproteobacteria, Gluconacetobacter diazotrophicus, Haemophilus parainfluenza Parainfluenzae), Haemophilus sputorum, Helicobacter canadensis, Helicobacter cinaedi, Helicobacter mustelae, Ilyobacter polytropus, Kingella kingae, Lactobacillus crispatus, Listeria ivanovii, Listeria monocytogenes, Listeriaceae bacteria, Methylocystis species, Methylosinus trichosporium, Mobiluncus murielis Neisseria mulieris, Neisseria bacilliformis, Neisseria cinereacinerea), Neisseria flavescens, Neisseria lactamica, Neisseria species, Neisseria wadsworthii, Nitrosomonas species, Parvibaculum lavamentivorans, Pasteurella multocida, Phascolarctobacterium succinatutens, Ralstonia syzygii, Rhodopseudomonas palustris, Rhodovulum species, Simonsiella muelleri The Cas9 molecule may be derived from any bacterium or archaeal species, including *Sphingomonas* species, *Sporolactobacillus vineae*, *Staphylococcus lugdunensis*, *Streptococcus* species, *Subdoligranulum* species, *Tistrella mobilis*, *Treponema* species, or *Verminephrobacter eiseniae*. In certain embodiments, the Cas9 molecule is a *Streptococcus pyogenes* Cas9 molecule (also known herein as "SpCas9"). In certain embodiments, the Cas9 molecule is a *Staphylococcus aureus* Cas9 molecule (also known herein as "SaCas9").
[0072] A Cas molecule or Cas fusion protein can interact with one or more gRNA molecules and, in cooperation with the gRNA molecules, localize to a target domain and, in certain embodiments, to a site containing a PAM sequence. The ability of a Cas molecule or Cas fusion protein to recognize a PAM sequence can be determined, for example, using transformation assays known in the art.
[0073] In certain embodiments, the ability of a Cas molecule or Cas fusion protein to interact with and cleave a target nucleic acid is dependent on a protospacer-adjacent motif (PAM) sequence. The PAM sequence is a sequence in the target nucleic acid. In certain embodiments, cleavage of the target nucleic acid occurs upstream of the PAM sequence. Cas molecules from different bacterial species can recognize different sequence motifs (e.g., PAM sequences). In certain embodiments, the Cas12 molecule from Francisella nobicida recognizes the sequence motif TTTN (SEQ ID NO: 35). In certain embodiments, the Cas9 molecule from Streptococcus pyogenes recognizes the sequence motif NGG (SEQ ID NO: 1) and directs cleavage of the target nucleic acid sequence 1-10 bp, e.g., 3-5 bp upstream of that sequence. In certain embodiments, the Cas9 molecule of Streptococcus thermophilus recognizes the sequence motif NGGNG (SEQ ID NO: 5) and / or NNAGAAW (W=A or T) (SEQ ID NO: 6) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from these sequences. In certain embodiments, the Cas9 molecule of Streptococcus mutans recognizes the sequence motif NGG (SEQ ID NO: 1) and / or NAAR (R=A or G) (SEQ ID NO: 7) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from this sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRR (R=A or G) (SEQ ID NO: 8) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from this sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRN(R=A or G) (SEQ ID NO: 9) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRT(R=A or G) (SEQ ID NO: 10) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence. In certain embodiments, the Cas9 molecule of Staphylococcus aureus recognizes the sequence motif NNGRRV(R=A or G;V=A or C or G) (SEQ ID NO: 11) and instructs the cleavage of a target nucleic acid sequence 1 to 10 bp upstream, for example, 3 to 5 bp from that sequence.In the above embodiment, N may be any nucleotide residue, such as A, G, C, or T. The Cas9 molecule can be manipulated to change its PAM specificity.
[0074] In certain embodiments, the vector encodes at least one Cas9 molecule that recognizes either NNGRRT (SEQ ID NO: 10) or NNGRRV (SEQ ID NO: 11) protospacer adjacent motif (PAM). In certain embodiments, at least one Cas9 molecule is a Staphylococcus aureus Cas9 molecule. In certain embodiments, at least one Cas9 molecule is a mutant Staphylococcus aureus Cas9 molecule.
[0075] The Cas protein can be mutated to inactivate its nuclease activity. Inactivated Cas9 proteins lacking endonuclease activity (also known as "iCas9" or "dCas9") are targeted by gRNA to genes in bacteria, yeast, and human cells to silence gene expression through steric hindrance. Exemplary mutations associated with the *Streptococcus pyogenes* Cas9 sequence include D10A, E762A, H840A, N854A, N863A, and / or D986A. Exemplary mutations associated with the *Staphylococcus aureus* Cas9 sequence include D10A and N580A. In certain embodiments, the Cas9 molecule is a mutant *Staphylococcus aureus* Cas9 molecule. In some embodiments, dCas9 is a Cas9 molecule containing at least two mutations selected from D10A, E762A, H840A, N854A, N863A, and / or D986A, which are associated with the Streptococcus pyogenes Cas9 sequence. In some embodiments, the Cas protein is the dCas9 protein. In some embodiments, the Cas protein is the dCas12 protein.
[0076] In certain embodiments, the mutant Staphylococcus aureus Cas9 molecule contains the D10A mutation. The nucleotide sequence encoding this mutant Staphylococcus aureus Cas9 is shown in Sequence ID No. 22. In certain embodiments, the mutant Staphylococcus aureus Cas9 molecule contains the N580A mutation. The nucleotide sequence encoding this mutant Staphylococcus aureus Cas9 molecule is shown in Sequence ID No. 23.
[0077] The polynucleotide encoding the Cas9 molecule can be a synthetic polynucleotide. For example, a synthetic polynucleotide can be chemically modified. A synthetic polynucleotide can be codon-optimized, for example, by replacing at least one uncommon or less common codon with a common codon. For example, a synthetic polynucleotide can direct the synthesis of an optimized messenger mRNA, which can be optimized for expression in a mammalian expression system, for example, as described herein. Furthermore, or separately, the nucleic acid encoding the Cas9 molecule or Cas9 polypeptide may include a nuclear localization sequence (NLS). Nuclear localization sequences are known in the art. An exemplary codon-optimized nucleic acid sequence encoding the Cas9 molecule of Streptococcus pyogenes is shown in SEQ ID NO: 14. The corresponding amino acid sequence of the Streptococcus pyogenes Cas9 molecule is shown in SEQ ID NO: 15.
[0078] Exemplary codon-optimized nucleic acid sequences encoding the Staphylococcus aureus Cas9 molecule, which may contain a nuclear localization sequence (NLS), are shown in SEQ ID NOs. 16-20 and 24-25. Another exemplary codon-optimized nucleic acid sequence encoding the Staphylococcus aureus Cas9 molecule is SEQ ID NOs. 27 contains nucleotides 1293-4451. The amino acid sequence of the Staphylococcus aureus Cas9 molecule is shown in SEQ ID NOs. 21. The amino acid sequence of the Staphylococcus aureus Cas9 molecule is shown in SEQ ID NOs. 26.
[0079] b. Fusion protein Alternatively, the CRISPR / Cas-based gene editing system may also include a fusion protein. The fusion protein may include two heterologous polypeptide domains, the first polypeptide domain comprising a DNA-binding protein such as a Cas protein, zinc finger protein, or TALE protein, and the second polypeptide domain having an activity such as transcriptional activation, transcriptional repression, transcriptional release factor activity, histone modification activity, nuclease activity, nucleic acid association activity, methylase activity, or demethylase activity. The fusion protein may include a first polypeptide domain, such as a Cas9 protein or a mutant Cas9 protein, fused to a second polypeptide domain having an activity such as transcriptional activation, transcriptional repression, transcriptional release factor activity, histone modification activity, nuclease activity, nucleic acid association activity, methylase activity, or demethylase activity. In some embodiments, the second polypeptide domain has transcriptional activation activity. In some embodiments, the second polypeptide domain has transcriptional repression activity. In some embodiments, the second polypeptide domain comprises a synthetic transcription factor. The second polypeptide domain may be located at the C-terminus or N-terminus of the first polypeptide domain, or a combination thereof. The fusion protein may contain one second polypeptide domain. The fusion protein may contain two second polypeptide domains. For example, the fusion protein may contain a second polypeptide domain at the N-terminus of the first polypeptide domain and a second polypeptide domain at the C-terminus of the first polypeptide domain. In other embodiments, the fusion protein may contain a single first polypeptide domain and two or more (e.g., two or three) second polypeptide domains in tandem.
[0080] i) Transcriptional activation activity The second polypeptide domain may have transcriptional activation activity, i.e., a transactivation domain. For example, gene expression of endogenous mammalian genes, such as human genes, can be achieved by targeting a mammalian promoter via a gRNA combination with a fusion protein of a first polypeptide domain, such as dCas9 or dCas12, and a transactivation domain. The transactivation domain may include the VP16 protein, multiple VP16 proteins, e.g., the VP48 domain or VP64 domain, the p65 domain of NFκB transcription activator activity, or p300. For example, the fusion protein may be dCas9-VP64. In other embodiments, the Cas9 protein may be VP64-dCas9-VP64 (encoded by the polynucleotides of SEQ ID NO: 36 and SEQ ID NO: 37). In other embodiments, the transcription-activating fusion protein may be dCas9-p300. In some embodiments, p300 may include the polypeptide of SEQ ID NO: 159 or SEQ ID NO: 160.
[0081] ii) Transcriptional repressive activity The second polypeptide domain may possess transcriptional repressive activity. The second polypeptide domain may also have Kruppel association box activity such as a KRAB domain, ERF repressor domain activity, Mxil repressor domain activity, SID4X repressor domain activity, Mad-SID repressor domain activity, or TATA box-binding protein activity. For example, the fusion protein may be dCas9-KRAB. iii) Transcription release factor activity The second polypeptide domain may have transcription release factor activity. The second polypeptide domain may have eukaryotic release factor 1 (ERF1) activity or eukaryotic release factor 3 (ERF3) activity.
[0082] iv) Histone modification activity The second polypeptide domain may have histone modification activity. The second polypeptide domain may have histone deacetylase, histone acetyltransferase, histone demethylase, or histone methyltransferase activity. The histone acetyltransferase may be p300 or a CREB-binding protein (CBP) protein, or a fragment thereof. For example, the fusion protein may be dCas9-p300. In some embodiments, p300 may contain the polypeptide of SEQ ID NO: 159 or SEQ ID NO: 160.
[0083] v) Nuclease activity The second polypeptide domain can have nuclease activity different from that of the Cas9 protein. A nuclease, or protein with nuclease activity, is an enzyme that can cleave phosphodiester bonds between nucleotide subunits of nucleic acids. Nucleases are usually further divided into endonucleases and exonucleases, although some enzymes can fall into both categories. Well-known nucleases include deoxyribonucleases and ribonucleases.
[0084] vi) Nucleic acid associated activity The second polypeptide domain may have nucleic acid association activity or a nucleic acid-binding protein-DNA binding domain (DBD). A DBD is an independently folded protein domain containing at least one motif that recognizes double-stranded or single-stranded DNA. A DBD may be able to recognize a specific DNA sequence (recognition sequence) or may have a general affinity for DNA. The nucleic acid association region may be selected from helix-turn-helix regions, leucine zipper regions, winged helix regions, winged helix-turn-helix regions, helix-loop-helix regions, immunoglobulin folds, B3 domains, zinc fingers, HMG boxes, Wor3 domains, and TAL effector DNA binding domains.
[0085] vii) Methylase activity The second polypeptide domain may have methylase activity involving the transfer of methyl groups to DNA, RNA, proteins, small molecules, cytosine, or adenine. In some embodiments, the second polypeptide domain comprises a DNA methyltransferase.
[0086] viii) Demethylase activity The second polypeptide domain may possess demethylase activity. The second polypeptide domain may include enzymes that remove methyl (CH3-) groups from nucleic acids, proteins (particularly histones), and other molecules. Alternatively, the second polypeptide may convert methyl groups to hydroxymethylcytosine in a mechanism for demethylating DNA. The second polypeptide can catalyze this reaction. For example, the second polypeptide catalyzing this reaction may be Tet1.
[0087] c.gRNA A CRISPR / Cas-based gene editing system includes at least one gRNA molecule. For example, a CRISPR / Cas-based gene editing system may include two gRNA molecules. The gRNA provides targeting for the CRISPR / Cas-based gene editing system. The gRNA is a fusion of two non-coding RNAs: crRNA and tracrRNA. In some embodiments, a polynucleotide comprises crRNA and / or tracrRNA. The sgRNA can target any desired DNA sequence by replacing the sequence encoding a 20bp protospacer that confers targeting specificity through complementary base pairing with the desired DNA target. The gRNA mimics the naturally occurring crRNA:tracrRNA double strand involved in type II effector systems. For example, this double strand may include a 42-nucleotide crRNA and a 75-nucleotide tracrRNA, and acts as a guide for Cas9 to cleave the target nucleic acid. A "target region," "target sequence," or "protospacer" refers to the region of a target gene that a CRISPR / Cas9-based gene editing system targets and binds to. The portion of gRNA in the genome that targets a target sequence may be called a "targeting sequence," "targeting portion," or "targeting domain." A "protospacer" or "gRNA spacer" may refer to the region of a target gene that a CRISPR / Cas9-based gene editing system targets and binds to; a "protospacer" or "gRNA spacer" may also refer to a portion of gRNA in the genome that is complementary to the targeted sequence. The gRNA may include a gRNA scaffold. The gRNA scaffold may promote Cas9 binding to the gRNA and enhance endonuclease activity. The gRNA scaffold is a polynucleotide sequence that follows the portion of the gRNA that corresponds to the sequence targeted by the gRNA. In summary, the gRNA targeting region and the gRNA scaffold form a single polynucleotide.The scaffold may include the polynucleotide sequence of SEQ ID NO: 158. The CRISPR / Cas9-based gene editing system may include at least one gRNA, which targets different DNA sequences. Target DNA sequences may overlap. The target sequence or protospacer is followed by a PAM sequence at the 3' end of the protospacer in the genome. Different type II systems have different PAM requirements. For example, the Streptococcus pyogenes type II system uses the "NGG" sequence (SEQ ID NO: 1) (where "N" can be any nucleotide). In some embodiments, the PAM sequence may be "NGG" (where "N" can be any nucleotide). In some embodiments, the PAM sequence may be NNGRRT (SEQ ID NO: 10) or NNGRRV (SEQ ID NO: 11). At least one gRNA molecule can bind to and recognize the target region.
[0088] The number of gRNA molecules encoded by a gene construct (e.g., an AAV vector) can be at least 1 gRNA, at least 2 different gRNAs, at least 3 different gRNAs, at least 4 different gRNAs, at least 5 different gRNAs, at least 6 different gRNAs, at least 7 different gRNAs, at least 8 different gRNAs, at least 9 different gRNAs, at least 10 different gRNAs, at least 11 different gRNAs, at least 12 different gRNAs, at least 13 different gRNAs, at least 14 different gRNAs, at least 15 different gRNAs, at least 16 different gRNAs, at least 17 different gRNAs, at least 18 different gRNAs, at least 18 different gRNAs, at least 20 different gRNAs, at least 25 different gRNAs, at least 30 different gRNAs, at least 35 different gRNAs, at least 40 different gRNAs, at least 45 different gRNAs, or at least 50 different gRNAs.The number of gRNAs encoded by the vectors disclosed herein is at least 1 to at least 50 different gRNAs, at least 1 to at least 45 different gRNAs, at least 1 to at least 40 different gRNAs, at least 1 to at least 35 different gRNAs, at least 1 to at least 30 different gRNAs, at least 1 to at least 25 different gRNAs, at least 1 to at least 20 different gRNAs, at least 1 to at least 16 different gRNAs, at least 1 to at least 12 different gRNAs, at least 1 to at least 8 different gRNAs, at least 1 to at least 4 different gRNAs, at least 4 to at least 50 different gRNAs, at least 4 to at least 45 different gRNAs, at least 4 to at least 40 different gRNAs, at least 4 to at least 35 different gRNAs, and a small number of other gRNAs. It may be between 4 different gRNAs and at least 30 different gRNAs, at least 4 different gRNAs and at least 25 different gRNAs, at least 4 different gRNAs and at least 20 different gRNAs, at least 4 different gRNAs and at least 16 different gRNAs, at least 4 different gRNAs and at least 12 different gRNAs, at least 4 different gRNAs and at least 8 different gRNAs, at least 8 different gRNAs and at least 50 different gRNAs, at least 8 different gRNAs and at least 45 different gRNAs, at least 8 different gRNAs and at least 40 different gRNAs, at least 8 different gRNAs and at least 35 different gRNAs, 8 different gRNAs and at least 30 different gRNAs, at least 8 different gRNAs and at least 25 different gRNAs, 8 different gRNAs and at least 20 different gRNAs, at least 8 different gRNAs and at least 16 different gRNAs, or 8 different gRNAs and at least 12 different gRNAs.In certain embodiments, a gene construct (e.g., an AAV vector) encodes one gRNA molecule, i.e., a first gRNA molecule, and optionally, a Cas9 molecule. In certain embodiments, a first gene construct (e.g., a first AAV vector) encodes one gRNA molecule, i.e., a first gRNA molecule, and optionally, a Cas9 molecule, and a second gene construct (e.g., a second AAV vector) encodes one gRNA molecule, i.e., a second gRNA molecule, and optionally, a Cas9 molecule.
[0089] A gRNA molecule contains a targeting domain, which is a polynucleotide sequence complementary to the target DNA sequence and the subsequent PAM sequence. The gRNA may contain a "G" at the 5' end of the targeting domain or the complementary polynucleotide sequence. The targeting domain of a gRNA molecule may contain a polynucleotide sequence complementary to the target DNA sequence and the subsequent PAM sequence of at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 30, or at least 35 base pairs. In certain embodiments, the targeting domain of the gRNA molecule has a length of 19 to 25 nucleotides. In certain embodiments, the targeting domain of the gRNA molecule has a length of 20 nucleotides. In certain embodiments, the targeting domain of the gRNA molecule has a length of 21 nucleotides. In certain embodiments, the targeting domain of the gRNA molecule is 22 nucleotides long. In certain embodiments, the targeting domain of the gRNA molecule is 23 nucleotides long.
[0090] gRNAs can target regions within or near genes encoding transcription factors. In certain embodiments, gRNAs can target at least one of the following: exons, introns, promoter regions, enhancer regions, or transcription regions of a gene.
[0091] In some embodiments, the gRNA targets neuron-specific transcription factors. The gRNA may include a targeting domain comprising a polynucleotide sequence corresponding to at least one of SEQ ID NOs. 38-97 shown in Table 3, or its complementary sequence or variant thereof. The gRNA may target a polynucleotide comprising a sequence selected from SEQ ID NOs. 38-97, or its complementary sequence, part, or variant thereof. The gRNA may be encoded by a polynucleotide comprising a sequence selected from SEQ ID NOs. 38-97, or its complementary sequence, part, or variant thereof. The gRNA may include a polynucleotide sequence (e.g., its RNA version), or its complementary sequence, part, or variant thereof, corresponding to at least one of SEQ ID NOs. 38-97.
[0092] JPEG2026058344000002.jpg242165 JPEG2026058344000003.jpg81162
[0093] In some embodiments, the gRNA targets muscle-specific transcription factors. These muscle-specific transcription factors may be selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1. The gRNA may contain a targeting domain comprising a polynucleotide sequence corresponding to at least one of the sequence numbers 98-104 shown in Table 5, or its complementary sequence or variant. The gRNA may target polynucleotides comprising sequences selected from sequence numbers 98-104, or their complementary sequences, parts, or variants. The gRNA may be encoded by polynucleotides comprising sequences selected from sequence numbers 98-104, or their complementary sequences, parts, or variants. The gRNA may contain a polynucleotide sequence (e.g., its RNA version), or its complementary sequence, part, or variant, corresponding to at least one of the sequence numbers 98-104.
[0094] JPEG2026058344000004.jpg52160
[0095] Cells transformed or transcribed in the systems detailed herein may express at least one gRNA. Each cell may independently contain one gRNA and target one putative transcription factor. The intracellular level of at least one gRNA may be determined by any suitable means known in the art, such as deep sequencing. At least one gRNA may be enriched intracellularly. For example, at least one gRNA may be enriched in cells with high expression of a reporter protein. "Enriched" may refer to a statistically significant (p<0.05) increase in the amount of gRNA present in cells with high reporter gene expression. This may be calculated using DESeq2 in the R language, a differential expression analysis package. A cell gRNA, or at least one gRNA, can increase the expression of the reporter protein in the cell by about 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, or 90% compared to a control. The control may be a cell with a non-targeting gRNA. In some embodiments, the gRNA increases the expression of the reporter protein in the cell by about 2–50% compared to a non-targeting gRNA.
[0096] d. Genetic constructs A system for identifying cell type-specific transcription factors or for increasing the expression of a cell type-specific gene, or one or more components thereof, may be encoded by or contained within a gene construct. A gene construct may include polynucleotides such as vectors and plasmids. The construct may be recombinant. In some embodiments, the gene construct includes a promoter operably ligated to a polynucleotide encoding at least one gRNA molecule and / or a Cas molecule or fusion protein. In some embodiments, the gene construct includes a promoter operably ligated to a polynucleotide encoding at least one gRNA molecule and / or a dCas molecule or fusion protein. In some embodiments, the gene construct includes a promoter operably ligated to a polynucleotide encoding at least one gRNA molecule and / or a Cas9 molecule or fusion protein. In some embodiments, the promoter is operably ligated to a polynucleotide encoding a gRNA molecule, a reporter protein, a neuron marker, and / or a Cas9 molecule. In some embodiments, the promoter is operably ligated to a polynucleotide encoding a first gRNA molecule, a second gRNA molecule, a reporter protein, a neuron marker, and / or a Cas9 molecule. Gene constructs may exist intracellularly as functional extrachromosomal molecules. Gene constructs may be kinetochores, linear minichromosomes containing telomeres, or plasmids or cosmids. Gene constructs may be transformed or transduced into cells. Gene constructs may be incorporated into any suitable type of delivery vehicle, including, for example, viral vectors, lentiviral expression, mRNA electroporation, and lipid-mediated transfection. Cells transformed or transduced in the systems or components detailed herein are further provided herein. In some embodiments, the cells are stem cells. The stem cells may be human stem cells. In some embodiments, the cells are embryonic stem cells. The stem cells may be human pluripotent stem cells (iPSCs).Further details provided herein are stem cell-induced neurons, such as neurons derived from iPSCs transformed or transduced with the DNA targeting systems or components thereof.
[0097] Viral delivery systems are further provided herein. Viral delivery systems may include, for example, lentiviruses, retroviruses, mRNA electroporation, or nanoparticles. In some embodiments, the vector is an adeno-associated virus (AAV) vector. AAV vectors are small viruses belonging to the genus Dependvirus of the family Parvoviridae that infect humans and some other primate species. AAV vectors can be used to deliver CRISPR / Cas9-based gene editing systems using various construct configurations. For example, an AAV vector may deliver Cas9 and gRNA expression cassettes in separate vectors or in the same vector. Alternatively, when using small Cas9 proteins derived from species such as Staphylococcus aureus or Neisseria meningitidis, both Cas9 and up to two gRNA expression cassettes can be combined into a single AAV vector within the 4.7 kb packaging limit.
[0098] In some embodiments, the AAV vector is a modified AAV vector. Modified AAV vectors may have enhanced cardiac and / or skeletal muscle tissue tropism. Modified AAV vectors may be able to deliver and express CRISPR / Cas9-based gene editing systems in mammalian cells. For example, a modified AAV vector could be an AAV-SASTG vector (Piacentino et al. Human Gene Therapy 2012, 23, 635-646). Modified AAV vectors may be based on one or more of several capsid types, including AAV1, AAV2, AAV5, AAV6, AAV8, and AAV9. Modified AAV vectors may be based on AAV2 pseudotypes, including alternative myotropic AAV capsids such as AAV2 / 1, AAV2 / 6, AAV2 / 7, AAV2 / 8, AAV2 / 9, AAV2.5, and AAV / SASTG vectors, which efficiently transduce skeletal or cardiac muscle by systemic and local delivery (Seto et al. Current Gene Therapy 2012, 12, 139-151). Modified AAV vectors may also be AAV2i8G9 (Shen et al. J. Biol. Chem. 2013, 288, 28814-28823).
[0099] 4. Systems for increasing neuron-specific transcription of genes A system for increasing neuron-specific transcription of a gene or for increasing the expression of a neuron-specific gene is provided herein. The system may comprise a first gRNA, regulatory region, promoter region, or portion thereof that targets a first neuron-specific transcription factor; and a Cas protein or fusion protein as detailed above. The system may comprise a first gRNA, regulatory region, promoter region, or portion thereof that targets a first neuron-specific transcription factor; and a second gRNA, regulatory region, promoter region, or portion thereof that targets a second neuron-specific transcription factor; and a Cas protein or fusion protein as detailed above. In some embodiments, the second neuron-specific transcription factor is a positive or activating transcription factor, and the second polypeptide domain of the fusion protein has transcriptional activating activity. In some embodiments, the second neuron-specific transcription factor is a negative or repressing transcription factor, and the second polypeptide domain of the fusion protein has transcriptional repressing activity.
[0100] 5. Systems for identifying cell type-specific transcription factors For example, compositions and methods for selecting or identifying cell type-specific transcription factors, such as neuron-specific transcription factors, muscle-specific transcription factors, or chondrocyte-specific transcription factors, are provided herein. The system comprises a reporter protein and a polynucleotide encoding a cell type marker; a Cas protein or fusion protein as detailed above; and a library of gRNAs that target the putative transcription factors. Cell type-specific transcription factors, or polynucleotide sequences encoding cell type-specific transcription factors, or polynucleotide sequences encoding gRNAs that target cell type-specific transcription factors, which are selected or identified by the compositions and methods detailed herein, are further provided herein.
[0101] a. Reporter protein Polynucleotides can encode reporter proteins. Reporter proteins are encoded by a reporter gene and, in a recombinant system, trigger some determinable or detectable properties simultaneously with the expression of another gene, thereby indicating the expression of that other gene. Reporter proteins can generate detectable signals. A variety of reporter proteins can be used, differing in the physical nature of the signaling (e.g., fluorescence, electrochemistry, nuclear magnetic resonance (NMR), and electron paramagnetic resonance (EPR)) and the chemical properties of the reporter protein itself. In some embodiments, the signal from the reporter protein is a fluorescent signal.
[0102] In some embodiments, the reporter protein is a fluorescent protein. Examples of fluorescent proteins include luciferase, enhanced blue fluorescent protein (EBFP), enhanced blue fluorescent protein-2 (EBFP2), mKATE, iRFP (infrared fluorescent protein), enhanced yellow fluorescent protein (EYFP), yellow fluorescent protein (YFP), Katushka, Ds-Red express, red fluorescent protein, red fluorescent protein turbo, TurboRFP, TagRFP, green fluorescent protein (GFP), blue fluorescent protein (BFP), cyan fluorescent protein (CFP), enhanced green fluorescent protein (EGFP), AcGFP, TurboGFP, emerald, thistle green, Zs green, sapphire, T-sapphire, enhanced cyan fluorescent protein (ECFP), mCFP, cerulean, CyPet, AmCyanl, green ishi cyan, mTFPl (Teal), topaz, Venus, m This includes Citrine, YPet, PhiYFP, ZsYellowl, mBanana, Kusabira Orange, mOrange, dTomato, dTomato-Tandem, DsRed, DsRed2, DsRed-Express(Tl), DsRed-Monomer, mTangerine, mStrawberry, AsRed2, mRFPl, JRed, mCherry, HcRedl, mRaspberry, HcRedl, HcRed-Tandem, mPlum, and AQ143, or combinations thereof. In some embodiments, the reporter protein comprises mCherry. mCherry may comprise a polypeptide having the amino acid sequence of SEQ ID NO: 28 and may be encoded by polynucleotides comprising SEQ ID NO: 29. In some embodiments, the reporter protein is any polypeptide that can be identified by immunohistochemistry or antibody staining.
[0103] Cells transfected or transformed with polynucleotides may express a reporter protein. For example, the intracellular expression level of the reporter protein can be determined. The expression level of the reporter protein can be determined at various points in time after transfection of cells by the systems detailed herein. For example, the intracellular expression level of the reporter protein can be determined about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 days after transduction. In some embodiments, the intracellular expression level of the reporter protein is determined about 4 days after transduction. The fluorescent protein can be assayed by any suitable means known in the art, for example, by FACS, flow cytometry, or fluorescence microscopy. In some embodiments, cells transfected or transformed with polynucleotides have higher expression of the reporter protein compared to a control. The control may be another cell or cells transfected or transformed with a polynucleotide containing a different gRNA. "High expression" of the reporter protein can be defined as being in the top 5% expression level in the cell population.
[0104] b. Cell type markers Polynucleotides can encode markers that exhibit expression in specific cell types, states, or stages. For example, polynucleotides can encode neuronal markers. Neuronal markers are genes that are expressed only or predominantly in neuronal cells. Neuronal markers can be subtype-specific markers that are expressed only in certain subtypes of neurons. Neuronal markers can also be panneuronal markers. Panneuronal markers are genes that are expressed only or predominantly in neuronal cells and most neuronal cells. Panneuronal markers may also be called neuronal lineage markers. Neuronal markers can be expressed at any point in neurogenesis and in cells differentiated into neurons. Neuronal markers can be selected from, for example, TUBB3, Neurod1, Neurog1, Neurog2, ASCL1, SYN1, NCAM, and MAP2. In some embodiments, the panneuronal marker is TUBB3. TUBB3 is a gene that encodes polypeptide β-3 tubulin (also called β-tubulin III), a microtubule element of the tubulin family found almost exclusively in neurons. In some embodiments, the cell type-specific transcription factor is a neuron-specific transcription factor, the cell type marker is a neuron marker, and the neuron marker includes TUBB3.
[0105] In other embodiments, the cell type marker is a muscle or myogenic marker. The muscle or myogenic marker is a gene expressed only or predominantly in muscle cells. The muscle or myogenic marker may be a subtype-specific marker expressed only in certain subtypes of muscle cells. The muscle or myogenic marker may be a panmuscular or panmyogenic marker. The panmuscular or panmyogenic marker is a gene expressed only or predominantly in muscle cells and most muscle cells. The myogenic marker may include PAX7. In some embodiments, the cell type-specific transcription factor is a muscle-specific transcription factor, the cell type marker is a myogenic marker, and the myogenic marker includes PAX7.
[0106] In other embodiments, the cell type marker is a collagen marker. The collagen marker is a gene expressed only or predominantly in chondrocytes. The collagen marker may be a subtype-specific marker expressed only in certain subtypes of chondrocytes. The collagen marker may be a pancollagen marker. A pancollagen marker is a gene expressed only or predominantly in chondrocytes and most chondrocytes. The collagen marker may include COL2A1. In some embodiments, the cell type-specific transcription factor is a chondrocyte-specific transcription factor, the cell type marker is a collagen marker, and the collagen marker includes COL2A1.
[0107] The polynucleotide encoding the reporter protein can be operably ligated to the polynucleotide encoding the cell type marker, as detailed below. The polynucleotide encoding the reporter protein may be within the same reading frame as the polynucleotide encoding the cell type marker. Therefore, the reporter protein can serve as an expression or translation reporter for the cell type marker.
[0108] Cells transfected or transformed with polynucleotides may express cell type markers. For example, the expression level of intracellular cell type markers can be determined. The expression level of cell type markers can be determined at various time points after transfection of cells by the systems detailed herein. For example, the expression level of intracellular cell type markers can be determined approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 days after transduction. Cell type markers can be assayed by any suitable means known in the art, for example, by immunohistochemistry, qRT-PCR, and RNA sequencing.
[0109] c.gRNA library A system for selecting or identifying transcription factors may further include a library of gRNAs. The gRNA library may target putative transcription factors. For example, a gRNA may target the promoter of a gene encoding a transcription factor. Each gRNA may be different. The gRNA library may contain multiple gRNAs, each targeting a putative transcription factor. In some embodiments, each gRNA targets a different putative transcription factor. Some gRNAs may target the same putative transcription factor, with each gRNA targeting a different region of the gene encoding the transcription factor. In some embodiments, the different regions may overlap. In some embodiments, the gRNA library may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 gRNAs for each transcription start site of a transcription factor. A gRNA library may contain at least approximately 1000, at least approximately 2000, at least approximately 3000, at least approximately 4000, at least approximately 5000, at least approximately 6000, at least approximately 7000, at least approximately 8000, or at least approximately 9000 gRNAs.
[0110] 6. Pharmaceutical Compositions Pharmaceutical compositions comprising the above-described gene constructs or systems are further provided herein. The systems detailed herein, or at least one component thereof, can be formulated into pharmaceutical compositions according to standard techniques well known to those skilled in the pharmaceutical art. The pharmaceutical compositions can be formulated according to the mode of administration used. If the pharmaceutical compositions are for injection, they are sterile, pyrogen-free, and particle-free. Preferably, isotonic formulations are used. Generally, additives for isotonicity may include sodium chloride, dextrose, mannitol, sorbitol, and lactose. In some cases, isotonic solutions such as phosphate-buffered saline are preferred. Stabilizers include gelatin and albumin. In some embodiments, vasoconstrictors are added to the formulation.
[0111] The composition may further contain pharmaceutically acceptable excipients. pharmaceutically acceptable excipients may be functional molecules acting as vehicles, adjuvants, carriers, or diluents. The term "pharmaceutically acceptable carrier" can refer to any type of non-toxic, inert solid, semi-solid, or liquid filler, diluent, encapsulating material, or formulation aid. Examples of pharmaceutically acceptable carriers include diluents, lubricants, binders, disintegrants, colorants, flavorings, sweeteners, antioxidants, preservatives, lubricants, solvents, suspending agents, wetting agents, surfactants, emollients, sprays, water-retaining agents, powders, pH adjusters, and combinations thereof. Pharmacoherent excipients may include immunostimulatory complexes (ISCOMS), Freund's incomplete adjuvants, LPS analogs including monophosphoryl lipid A, muramil peptides, quinone analogs, surfactants such as squalene and vesicles, hyaluronic acid, lipids, liposomes, calcium ions, viral proteins, polyanions, polycations, or nanoparticles, or other known transfection accelerators.
[0112] The transfection promoter may be a polyanion, polycation, or lipid containing poly-L-glutamic acid (LGS). The transfection promoter is poly-L-glutamic acid, and more preferably, poly-L-glutamic acid is present in the composition for genome editing in skeletal muscle or cardiac muscle at a concentration of less than 6 mg / mL. The transfection promoter may also include surfactants such as immunostimulatory complexes (ISCOMS), Freund's incomplete adjuvants, LPS analogs containing monophosphoryl lipid A, muramyl peptides, quinone analogs, and squalene and squalene vesicles, and hyaluronic acid may also be used and administered in conjunction with the gene construct. In some embodiments, the DNA vector encoding the composition may also include liposomes containing lipids, lecithin liposomes or DNA-liposome mixtures (see, for example, International Patent Application Publication No. 9324640), other liposomes known in the art, calcium ions, viral proteins, polyanions, polycations, or nanoparticles, or other known transcription promoters. In some embodiments, the transfection accelerator is a polyanion, polycation, or lipid containing poly-L-glutamic acid (LGS).
[0113] 7. Administration The systems, or at least one component thereof, or pharmaceutical compositions containing them, as detailed herein, may be administered to a subject. Such compositions may be administered by dosage and techniques well known to those skilled in the pharmaceutical art, taking into account factors such as the age, sex, weight, and condition of a particular subject, as well as the route of administration. The systems, or at least one component thereof, gene constructs, or compositions containing them disclosed herein may be administered to a subject by different routes, including oral, parenteral, sublingual, transdermal, rectal, transmucosal, topical, intranasal, vaginal, inhalation, buccal administration, intrapleural, intravenous, intra-arterial, intraperitoneal, subcutaneous, intradermal, epidermal, intramuscular, intranasal, intrathecal, intracranial, and intra-articular, or combinations thereof. In certain embodiments, the systems, gene constructs, or compositions containing them are administered to a subject intramuscularly, intravenously, or in combination thereof. For veterinary use, DNA targeting systems, gene constructs, or compositions containing them may be administered as appropriately acceptable formulations in accordance with normal veterinary practice. Veterinarians can easily determine the most appropriate administration regimen and route for a particular animal. Systems, gene constructs, or compositions containing them may be administered by traditional syringes, needle-free injectors, "microprojectile bombardment gone guns," or other physical methods such as electroporation ("EP"), "hydrodynamic methods," or ultrasound.
[0114] Systems, gene constructs, or compositions containing these may be delivered to a target by several techniques, including, with and without, in vivo electroporation, liposome-mediated delivery, nanoparticle-enhanced delivery, and recombinant vector delivery, such as recombinant lentiviruses, recombinant adenoviruses, and recombinant adeno-associated viruses, including DNA injection (also known as DNA vaccination). The compositions may be injected into the brain or other components of the central nervous system.
[0115] 8. Method a. Methods to increase neuronal maturation of stem cells Methods for increasing neuronal maturation of stem cells, or for increasing the maturation of stem cell-induced neurons, are provided herein. These methods may comprise (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells, and increasing the level of a second neuron-specific transcription factor, which is an activated or positively positive neuron-specific transcription factor, in stem cells. In other embodiments, the method may include the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; and decreasing the level of a second neuron-specific transcription factor, which is a repressor or negative neuron-specific transcription factor, in stem cells.
[0116] In some embodiments, the step of increasing the level of a first neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the first neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the first neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the first neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity.
[0117] In some embodiments, the step of increasing the level of a second neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the second neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the second neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the second neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity.
[0118] In some embodiments, the step of reducing the level of a second neuron-specific transcription factor includes administering to stem cells a fusion protein comprising a second neuron-specific transcription factor, a regulatory region, a promoter region, or a gRNA that targets a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional repressive activity.
[0119] b. Methods to increase the conversion of stem cells into neurons A method for increasing the conversion of stem cells into neurons is provided herein. This method may comprise (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells, and increasing the level of a second neuron-specific transcription factor, which is an activated or positively positive neuron-specific transcription factor, in stem cells. In other embodiments, the method may include the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; and decreasing the level of a second neuron-specific transcription factor, which is a repressor or negative neuron-specific transcription factor, in stem cells.
[0120] In some embodiments, the step of increasing the level of a first neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the first neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the first neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the first neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity.
[0121] In some embodiments, the step of increasing the level of a second neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the second neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the second neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the second neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity. In some embodiments, the step of reducing the level of a second neuron-specific transcription factor includes administering to stem cells a fusion protein comprising a second neuron-specific transcription factor, a regulatory region, a promoter region, or a gRNA that targets a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional repressive activity.
[0122] c. Method of treating the target A method for treating subjects requiring treatment is provided herein. This method may include (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells of interest; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells of interest, and increasing the level of a second neuron-specific transcription factor, which is an activated or positively positive neuron-specific transcription factor, in stem cells of interest. In other embodiments, the method may include the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the target stem cell; and decreasing the level of a second neuron-specific transcription factor, which is a repressive or negative neuron-specific transcription factor, in the target stem cell.
[0123] In some embodiments, the step of increasing the level of a first neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the first neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the first neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the first neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity.
[0124] In some embodiments, the step of increasing the level of a second neuron-specific transcription factor includes at least one of the following: (a) administering a polynucleotide encoding the second neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing the second neuron-specific transcription factor to stem cells; and (c) administering a fusion protein to stem cells that targets the second neuron-specific transcription factor, a regulatory region, a promoter region, or a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional activation activity. In some embodiments, the step of reducing the level of a second neuron-specific transcription factor includes administering to stem cells a fusion protein comprising a second neuron-specific transcription factor, a regulatory region, a promoter region, or a gRNA that targets a portion thereof, and a first polypeptide domain comprising a DNA-binding protein such as a Cas protein, a zinc finger protein, or a TALE protein, and a second polypeptide domain comprising two heterologous polypeptide domains having transcriptional repressive activity.
[0125] d. Methods for screening neuron-specific transcription factors A method for screening neuron-specific transcription factors is provided herein. The method may include: transducing a population of cells in the system described in any one of claims 1 to 3 at an infection multiplicity (MOI) of about 0.2 such that the majority of cells independently contain one gRNA and target one putative transcription factor; determining the expression level of a reporter protein in each cell; determining the gRNA level in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; and selecting a putative transcription factor as a neuron-specific transcription factor if the putative transcription factor corresponds to at least two gRNAs enriched in cells having high expression of the reporter protein. "Enriched" may mean a statistically significant (p<0.05) increase in gRNA abundance in cells having high reporter gene expression.
[0126] In some embodiments, the expression level of the reporter protein in each cell is determined approximately 4 days after transduction. In some embodiments, the expression level of the reporter protein in each cell is determined by flow cytometry. In some embodiments, the gRNA level in each cell with high reporter protein expression is determined by deep sequencing. In some embodiments, the gRNA increases reporter protein expression in the cell by approximately 2–50% compared to untargeted gRNA.
[0127] e. Methods for screening pairs of neuron-specific transcription factors A method for screening pairs of neuron-specific transcription factors is provided herein. The method may include: transducing a population of cells in the system described in any one of claims 1 to 3 at an infection multiplicity (MOI) of about 0.2 such that the majority of cells each independently contain two gRNAs and target two putative transcription factors; determining the expression level of a reporter protein in each cell; determining the levels of two gRNAs in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; and selecting two putative transcription factors as a pair of neuron-specific transcription factors if the putative transcription factors correspond to at least two gRNAs enriched in cells having high expression of the reporter protein.
[0128] In some embodiments, the expression level of the reporter protein in each cell is determined approximately 4 days after transduction. In some embodiments, the expression level of the reporter protein in each cell is determined by flow cytometry. In some embodiments, the gRNA level in each cell with high reporter protein expression is determined by deep sequencing. In some embodiments, the gRNA increases reporter protein expression in the cell by approximately 2–50% compared to untargeted gRNA. [Examples]
[0129] 9. Examples (Example 1) material and method Construction of a TUBB3-2A-mCherry pluripotent stem cell line. A TUBB3-2A-mCherry reporter line was constructed using human iPS cell lines (RVR-iPSCs). RVR-iPSCs were reprogrammed from BJ fibroblasts with retrovirus and characterized as previously done (Lee et al. Cell 2012, 51, 547-558). 3 × 10⁶ cells were used to generate the TUBB3-2A-mCherry reporter line. 6Individual cells were separated with Accutase (Stemcell Tech, 7920) and electroporated with 6 μg of the gRNA-Cas9 expression vector and 3 μg of the TUBB3 targeting vector using the P3 Primary Cell 4D-Nucleofector Kit (Lonza, V4XP-3032). Transfected cells were seeded onto Matrigel (Corning, 354230)-coated 10-cm dishes in complete mTesR (Stemcell Tech, 85850) supplemented with 10 μM Rock inhibitor (Y-27632, Stemcell Tech, 72304). Twenty-four hours after transfection, positive selection was initiated with 1 μg / mL puromycin for 7 days. After selection, the cells were transfected with the CMV-CRE recombinase expression vector to remove the floxed puromycin selection cassette. Transfected cells were expanded and seeded at low density (180 cells / cm 2 ) for clonal isolation. The resulting clones were mechanically picked, expanded, and genomic DNA (gDNA) was extracted using QuickExtract DNA Extraction Solution (Lucigen, QE09050) for PCR screening of targeted vector integration. VP64 dCas9 VP64 After lentiviral transduction of dCas9, a second round of clonal isolation was performed using the same protocol.
[0130] Plasmid construction. The lentiviral VP64 dCas9 VP64 plasmid was generated by modifying Addgene plasmid number 59791 to replace GFP with the BSD blasticidin resistance gene. The lentiviral dSaCas9 KRABPlasmids were constructed. A gRNA expression plasmid for single CAS-TF screening was constructed by modifying Addgene plasmid number 83925 to contain a puromycin resistance gene instead of Bsr, along with an optimized gRNA scaffold (Chen et al. Cell 2013, 155, 1479-149). A gRNA expression plasmid for paired CAS-TF screening was constructed by further modifying the single gRNA expression plasmid to contain an additional gRNA cassette expressing either sgNGN3 or sgASCL1 under the control of the mU6 Pol III promoter, along with the previously described (Adamson et al. Cell 2016, 167, 1867-1882 e1821) modified gRNA scaffold. Individual gRNAs were aligned as oligonucleotides (Integrated DNA Technologies), phosphorylated, hybridized, and cloned into gRNA expression plasmids using the BsmBI site. The protospacers used for individual gRNA cloning are listed in Table 3 above.
[0131] A TUBB3-targeting vector was cloned by inserting an approximately 700 bp homology arm (surrounding the TUBB3 stop codon), and the P2A-mCherry sequence was surrounded by a flox puromycin-resistant cassette. This was then amplified by PCR from the genomic DNA of RVR-iPS cells.
[0132] The cDNA encoding TF was synthesized from a cDNA pool by PCR amplification or as gBlocks (Integrative DNA Technologies) and cloned into Addgene plasmid number 52047 using EcoRI and XbaI restriction sites. TetO gene expression was achieved by co-delivery of M2rtTA (Addgene number 20342).
[0133] Lentivirus production and titer measurement. HEK293T cells were obtained from the American Tissue Collection Center (ATCC) and purchased through the Duke University Cell Culture Facility. Cells were maintained in high-glucose DMEM supplemented with 10% FBS and 1% penicillin-streptomycin and cultured at 37°C and 5% CO2. gRNA library, VP64 dCas9 VP64 and dSaCas9 KRAB For lentivirus production, the calcium phosphate precipitation method (Salmon and Trono, 2007 Curr. Protoc. Hum. Genet. Chapter 12, Unit 12 10) was used to obtain 4.5 × 10⁻⁶ cells. 5 Each cell was transfected with 6 μg of pMD2.G (Addgene No. 12259), 15 μg of psPAX2 (Addgene No. 12260), and 20 μg of transduction vector. The medium was changed 12–14 hours after transfection, and the viral supernatant was collected 24 and 48 hours after this medium change. The viral supernatant was pooled, centrifuged at 600 g for 10 minutes, passed through a PVDF 0.45 μm filter (Millipore, SLHV033RB), and concentrated 50-fold in 1 × PBS using a Lenti-X Concentrator (Clontech, 631232) according to the manufacturer's protocol.
[0134] To produce lentiviruses for gRNA and cDNA validation, use Lipofectamine 3000 (Invitrogen, L3000008) according to the manufacturer's instructions, producing 0.4 × 10⁶ molecules. 6Cells were transfected with 200 ng of pMD2.G, 600 ng of psPAX2, and 200 ng of transfection vector. The medium was changed 12–14 hours after transfection, and the viral supernatant was collected 24 and 48 hours after this medium change. The viral supernatant was pooled, centrifuged at 600 g for 10 minutes, and concentrated 50-fold in 1 × PBS using a Lenti-X Concentrator (Clontech, 631232) according to the manufacturer's protocol. Lentivirus serial dilution 6 × 10 4 The titers of lentiviral gRNA library pools for single or paired CAS-TF libraries were determined by transduction into individual cells and measurement of GFP expression % four days after transduction using an Accuri C6 flow cytometer (BD). All lentiviral titers were measured in the TUBB3-2A-mCherry cell line used for CAS-TF single and paired gRNA screening.
[0135] CAS-TF gRNA library design and cloning. Putative TFs were selected from a previous catalog of human transcription factors (Vaquerizas et al. Nat. Rev. Genet. 2009, 10, 252-263). A gRNA library consisting of 5 gRNAs per TSS targeting 1496 TFs was extracted from a previous genome-wide CRISPRa library (Horlbeck, 2016 Compact and highly active next-generation libraries. eLife). This library contained a set of 100 scrambled untargeted gRNAs extracted from the same genome-wide library for a total of 8505 gRNAs. Oligonucleotide pools (Custom Arrays) were PCR amplified and cloned using Gibson assembly into single gRNA expression plasmids for single CAS-TF screening or dual gRNA expression plasmids for paired CAS-TF screening using sgASCL or sgNGN3.
[0136] A sublibrary was designed by extracting additional gRNAs from several previously published CRISPRa genome-wide libraries (Gilbert et al. Cell 2014, 159, 647-66; Horlbeck, 2016 Compact and highly active next-generation libraries. eLife; Konermann et al. Nature 2015, 517, 583-588; Sanson et al. Nat. Commun. 2018, 9, 5416) to obtain an average of 33 gRNAs per gene targeting 109 TFs. This library contained a set of 300 scrambled non-targeting gRNAs for a total of 3874 gRNAs. Oligonucleotide pools (Twist Bioscience) were PCR amplified and cloned into single-gRNA expression plasmids, as was done with the original CAS-TF library.
[0137] Single and paired CAS-TF neuron differentiation screening. Each CAS-TF screening was performed in triplicate via independent transduction. For each replication, 24 × 10⁻¹⁶ neurons were examined. 6 TUBB3-2A-mCherry VP64 dCas9 VP64iPSCs were isolated using Accutase (Stemcell Tech, 7920) and transduced in suspension across five Matrigel-coated 15 cm dishes in mTesR (Stemcell Tech 85850) supplemented with 10 μM Rock inhibitor (Y-27632, Stemcell Tech, 72304). Transduction to cells at a MOI of 0.2 yielded one gRNA per cell and approximately 550-fold coverage of the CAS-TF gRNA library. 18–20 hours after transduction, the medium was replaced with fresh mTesR without Rock inhibitor. Antibiotic selection was initiated 30 hours after transduction by directly adding 1 μg / mL puromycin (Sigma, P8833) to the plate without changing the medium. 48 hours after transduction, the culture medium was replaced with neurogenic medium supplemented with 1 μg / mL puromycin (DMEM / F-12 Nutrient Mix (Gibco, 11320), 1×B-27 serum-free supplement (Gibco, 17504), 1×N-2 supplement (Gibco, 17502), and 25 μg / mL gentamicin (Sigma, G1397)), and the medium was replaced daily for the remainder of the experiment.
[0138] Cells were collected for sorting 5 days after transduction of the gRNA library for single-factor CAS-TF screening and sgASCL1 pair screening. Cells were collected 4 days after transduction for sgNGN3 pair screening. Cells were washed once with 1×PBS, separated using Accutase, filtered through a 30 μm CellTrics filter (Sysmex, 04-004-2326), and resuspended in FACS buffer (0.5% BSA (Sigma, A7906), 2 mM EDTA (Sigma, E7889) in PBS). Before sorting, 4.8×10⁶ cells were selected. 6 Aliquots of individual cells were taken to form an unsorted bulk population. The highest and lowest 5% of cells were selected based on mCherry expression, resulting in 4.8 × 10⁶ cells. 6Individual cells were sorted into separate vials. Sorting was performed using an SH800 FACS Cell Sorter (Sony Biotechnology). After sorting, genomic DNA was collected using the DNeasy Blood and Tissue Kit (Qiagen, 69506).
[0139] Sublibrary screening. CAS-TF sublibrary screening was performed in three consecutive sequences using independent transduction. For each replication, 9.6 × 10⁻⁶ 6 TUBB3-2A-mCherry VP64 dCas9 VP64 iPSCs were isolated using Accutase (Stemcell Tech, 7920) and transduced in suspension across two Matrigel-coated 15 cm dishes in mTesR (Stemcell Tech 85850) supplemented with 10 μM Rock inhibitor (Y-27632, Stemcell Tech, 72304). Transduction to cells at a MOI of 0.2 yielded approximately 495-fold coverage of one gRNA per cell and the CAS-TF gRNA sublibrary. 18–20 hours after transduction, the medium was replaced with fresh mTesR without Rock inhibitor. Antibiotic selection was initiated 30 hours after transduction by directly adding 1 μg / mL puromycin (Sigma, P8833) to the plate without changing the medium. 48 hours after transduction, the culture medium was replaced with neurogenic medium supplemented with 1 μg / mL puromycin (DMEM / F-12 Nutrient Mix (Gibco, 11320), 1×B-27 serum-free supplement (Gibco, 17504), 1×N-2 supplement (Gibco, 17502), and 25 μg / mL gentamicin (Sigma, G1397)), and the medium was replaced daily for the remainder of the experiment.
[0140] Cells were collected for sorting 5 days after transduction of the gRNA library. Cells were washed once with 1×PBS, separated using Accutase, filtered through a 30 μm CellTrics filter (Sysmex, 04-004-2326), and resuspended in FACS buffer (0.5% BSA (Sigma, A7906) and 2 mM EDTA (Sigma, E7889) in PBS). Before sorting, 2×10 6 Aliquots of individual cells were taken to form an unsorted bulk population. The highest and lowest 5% of cells were selected based on mCherry expression, resulting in 2 × 10⁶ cells. 6 Individual cells were sorted into separate vials. Sorting was performed using an SH800 FACS Cell Sorter (Sony Biotechnology). After sorting, genomic DNA was collected using the DNeasy Blood and Tissue Kit (Qiagen, 69506). gRNA library sequencing. gRNA libraries were amplified from each genomic DNA sample over 100 μL PCR reactions using Q5 hot-start polymerase (NEB, M0493) with 1 μg of genomic DNA per reaction. PCR amplification was performed using the following primers, with 25 cycles at an annealing temperature of 60°C, according to the manufacturer's instructions: Fwd:5'-AATGATACGGCGACCACCGAGATCTACACAATTTCTTGGGTAGTTTGCAGTT Rev:5'-CAAGCAGAAGACGGCATACGAGAT-(6-bp index sequence)-GACTCGGTGCCACTTTTTCAA
[0141] The amplified libraries were purified using Agincourt AMPure XP beads (Beckman Coulter, A63881) with dual size selection of 0.65x and then 1x to obtain 282bp amplicons. Each sample was purified and quantified using the Qubit dsDNA High Sensitivity Assay Kit (Thermo Fisher, Q32854). The samples were pooled and sequenced using MiSeq (Illumina) with 20bp paired-end sequencing using the following custom reads and index primers: Lead 1:5'-GATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCG (Sequence ID 32). Index: 5'-GCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTC (Sequence ID 33). Lead 2:5'-GTTGATAACGGACTAGCCTTATTTAAACTTGCTATGCTGTTTCCAGCATAGCTCTTAAAC (Sequence ID 34).
[0142] Data processing and enrichment analysis. FASTQ files were aligned to a custom index of 8505 protospacers (generated from the bowtie2 build function) using Bowtie2 (Langmead and Salzberg Nat. Methods 2012, 9, 357-359). Counts of each gRNA were extracted and used for further analysis. All enrichment analyses were performed in R. Individual gRNA enrichment was determined using the DESeq2 package (Love et al. Genome Biol. 2014, 15, 550) to compare gRNA abundances between high and low conditions, unselected and low conditions, or unselected and high conditions for each screening. TFs were selected as hits if two or more gRNAs were significantly enriched in the mCherry high cell bin compared to both the unselected cell bin and the mCherry low cell bin (FDR < 0.01).
[0143] In vivo expression comparison. RNA sequencing data generated as part of the Brainspan Developmental Transcriptome Atlas was downloaded (Miller et al. Nature 2014, 508, 199-206). Mean expression for 17 TFs identified by single-factor CAS-TF screening was calculated for each listed developmental time and anatomical region between 8 and 13 weeks post-conception. Random sets of the 17 TFs were similarly analyzed, and representative comparisons are shown in Figure 1F.
[0144] gRNA and cDNA validation. Top enriched gRNAs from screening were cloned into appropriate gRNA expression vectors as previously described. Transduction was performed in 24-well plates, and gRNA validation was performed in the same manner as in screening, except that the virus was delivered with a high MOI. Cells were collected for flow cytometry or qRT-PCR 4 days after gRNA transduction.
[0145] For immunofluorescence staining experiments, the cDNA encoding the top-enriched TF was PCR-amplified as previously described and cloned into a doxycycline-inducible expression vector. Cells were co-transduced in suspension with TFs shown together with a separate lentivirus encoding M2rtTA (Addgene number 20342) in mTesR supplemented with 10 μM Rock inhibitor. Unmodified iPSCs were used in these experiments to allow staining with red fluorophores without interference from the mCherry reporter. 18–20 hours after transduction, the medium was replaced with neurogenic medium supplemented with 0.1 μg / mL doxycycline (Sigma, D9891). Staining was performed 4 days after transduction as previously described. For subsets of TFs, the TUBB3-2A-mCherry cell line was used to select the best mCherry-expressing cells 3 days after transduction. The cells were re-seed into a pre-established monolayer of human astrocytes (Lonza, CC-2565) and cultured for an additional 8 days in neurogenic medium before staining. gRNA and cDNA validation in H9 human embryonic stem cells was performed in the same manner as described for iPSCs. PolyclonalVP64 dCas9 VP64 The H9 ESC strain was established via lentiviral transduction, and gRNA was delivered using a separate lentivirus.
[0146] Quantitative RT-PCR. Cells were separated with Accutase (StemCell Tech, 7920) and centrifuged at 300g for 5 minutes. Total RNA was isolated using RNeasy Plus (Qiagen, 74136) and QIAshredder kit (Qiagen, 79656). Reverse transcription was performed using the SuperScript VILO Reverse Transcription Kit (Invitrogen, 11754) with 0.1 μg of total RNA per sample in a 10 μL reaction. 1.0 μL of cDNA per PCR reaction was used with Perfecta SYBR Green Fastmix (Quanta BioSciences, 95072) using the CFX96 Real-Time PCR Detection System (Bio-Rad). Amplification efficiency across the appropriate dynamic range of all primers was optimized using dilutions of purified amplicons. All amplicon products were validated by gel electrophoresis and melting curve analysis. All qRT-PCR results are presented as the quantification change of RNA normalized to GAPDH expression. The primers used in this test can be found in Table 4.
[0147] JPEG2026058344000005.jpg247161
[0148] Immunofluorescence staining. Cells were briefly washed with PBS and then fixed with 4% paraformaldehyde (Santa Cruz, sc-281692) for 20 minutes at room temperature. Cells were washed twice with PBS and then incubated with blocking buffer (10% goat serum (Sigma, G6767) and 2% BSA (Sigma, A7906) in PBS) for 30 minutes at room temperature. Cells were permeabilized with 0.2% Triton-X 100 (Sigma, T8787) for 10 minutes at room temperature. The following primary antibodies were used in a 2-hour incubation at room temperature: mouse anti-TUBB3 (1:1000 dilution, BioLegend, 801201); rabbit anti-MAP2 (1:500 dilution, Sigma, AB5622). Cells were washed three times with PBS and then incubated with secondary antibody and DAPI (Invitrogen, D3571) in blocking solution for 1 hour at room temperature. The following secondary antibodies were used: Alexa Fluor 488 goat anti-mouse (1:500 dilution, Invitrogen, A-11001); Alexa Fluor 594 goat anti-rabbit (1:500 dilution, Invitrogen, A-11012). Cells were washed three times with PBS and imaged with a Zeiss 780 upright confocal microscope.
[0149] For NCAM staining of live cells for gRNA validation, cells were isolated with Accutase (StemCell Tech, 7920), centrifuged at 300g for 5 minutes, and then 10 × 10 6 Cells / mL were resuspended in staining buffer (0.5% BSA (Sigma, A7906) and 2 mM EDTA (Sigma, E7889) in PBS). Mouse anti-CD56 (NCAM, Invitrogen, 12-0567) was used to stain 1 × 10⁶ cells. 6 The solution was added at a dose of 0.6 μg per cell and incubated at 4°C for 30 minutes. The cells were washed with 1 mL of staining buffer, centrifuged at 300 g for 5 minutes, and resuspended in staining buffer for analysis using an SH800 FACS Cell Sorter (Sony Biotechnology).
[0150] RNA sequencing by tetO cDNA expression. TUBB3-2A-mCherry iPSCs were co-transduced with M2rtTA and a lentivirus encoding the indicated tetO-cDNA. Cells were transduced in mTesR containing 10 μM Rock inhibitor. The following day, the medium was replaced with neurogenic medium supplemented with 0.1 μg / mL doxycycline (DMEM / F-12 Nutrient Mix (Gibco, 11320), 1×B-27 serum-free supplement (Gibco, 17504), 1×N-2 supplement (Gibco, 17502), and 25 μg / mL gentamicin (Sigma, G1397)). Cells were sorted after 2 or 3 days of transgene expression using an SH800 FACS Cell Sorter in semi-purity mode. The selected cells were re-seed in Matrigel-coated 24-well plates and cultured in neurogenic medium supplemented with 10 ng / mL of BDNF, GDNF, and NT-3 (PeproTech), respectively, until collection at 6 or 7 days later.
[0151] Total RNA was extracted using the RNeasy Mini Kit (Qiagen), and an RNA-seq library was developed using 100 ng of RNA. The RNA sequencing library was prepared using the Trueq Stranded mRNA Kit (Illumina) according to the manufacturer's protocol. The library was sequenced using a NextSeq 500 in High Output Mode with 75 bp paired-end reads. Reads were first trimmed using Trimmomatic v0.32 to remove adapters, and then aligned to GRCh38 using a STAR aligner (Langmead et al. Nat. Methods 2012, 9, 357-359). Gene counts were obtained from the subread package (version 1.4.6-p4) using featureCounts with comprehensive gene annotation in Gencode v22. Differential expression analysis was performed using DESeq2, where gene counts were fitted to a generalized linear model (GLM) of a negative binomial distribution, and Wald statistics determined significant hits. Genes were included in the analysis if at least three samples had a TPM > 1 across all tested conditions. Gene ontology analysis was performed using the Gene Ontology Consortium database (Ashburner et al., 2000, The Gene Ontology Consortium, 2017) and the Synaptic Gene Ontology Consortium database (Koopmans et al. Neuron 2019, 103, 217-234 e214).
[0152] Electrophysiology. TUBB3-2A-mCherry iPSCs were co-transduced with lentiviruses encoding M2rtTA and tetO-NEUROG3, either alone or in combination with tetO-LHX8. Cells were transduced in mTesR containing 10 μM Rock inhibitor. The following day, the medium was replaced with neurogenic medium supplemented with 0.1 μg / mL doxycycline. Cells were sorted after 3 days of transgene expression using an SH800 FACS Cell Sorter in semi-purity mode. Sorted cells were re-seeded on matrixel-coated coverslips, and the remainder of the experiment were cultured in neurogenic medium supplemented with 10 ng / mL BDNF, GDNF, and NT-3 (PeproTech), respectively.
[0153] Whole-cell patch-clamp recordings were performed on cultured cells 7 days after induction of transgene expression under a Zeiss Axio Examiner.D1 microscope. To avoid osmotic shock, the culture medium was gradually replaced with artificial CSF (aCSF) over approximately 5 minutes, and then the coverslip was transferred to the recording chamber. The aCSF contained 124 mM NaCl, 26 mM NaHCO3, 10 mM D-glucose, 2 mM CaCl2, 3 mM KCl, 1.3 mM MgSO4, and 1.25 mM NaH2PO4 (310 mOsm / L), which was continuously bubbling at room temperature with 95% O2 and 5% CO2. Cells were examined under a 20x water immersion objective lens using infrared irradiation and differential interference contrast (IR-DIC). Experimenters were blinded to the conditions, and the most morphologically complex neurons were selected for recording. Electrodes (4–7 MΩ) were drawn out of borosilicate glass capillaries using a P-97 puller (Sutter Instrument) and filled with an intracellular solution containing 135 mM K-methanesulfonate, 8 mM NaCl, 10 mM HEPES, 0.3 mM EGTA, 4 mM MgATP, and 0.3 mM Na2GTP (adjusted to pH 7.3 with KOH and 295 mOsm / L with sucrose). After rupturing the gigaohm seal, the membrane resistance was measured in voltage clamp mode using a short hyperpolarization pulse, and the membrane capacitance was estimated from the amplifier's capacitance compensation circuit. The resting membrane potential was then recorded in current clamp mode. Finally, an input-output curve was generated by applying a small holding current to adjust the membrane potential to approximately -60 mV and injecting an increasing amount of current. Data were recorded using a Multiclamp 700B amplifier (Molecular Devices) and digitized at 50 kHz using a Digidata 1550 (Molecular Devices). Action potential characteristics were calculated based on the first action potential generated using a custom MATLAB script. Action potentials were visually counted if they had a characteristic two-component depolarization phase, regardless of peak amplitude. All experiments were analyzed blinded to the conditions, and only recordings that remained stable throughout the entire data acquisition period were used.
[0154] Orthogonal CRISPR-based gene regulation. TUBB3-2A-mCherry VP64 dCas9 VP64 All-in-one dSaCas9 containing iPSCs with either ZFP36L1, HES3, or scrambled Staphylococcus aureus gRNA. KRAB Transduction was performed using lentivirus (Thankore et al. Nat. Commun. 2018, 9, 1674). Two days later, antibiotic selection was initiated with 0.5 μg / mL puromycin, and the cells were cultured in mTesR for a further 7 days. dSaCas9 KRAB Nine days after transduction with Staphylococcus aureus gRNA, the cells were transduced with a lentivirus encoding either sgNGN3 or sgASCL1, and then switched to neurogenic medium. Cells were collected for mRNA sequencing three days after gRNA transduction and for flow cytometry four days after gRNA transduction.
[0155] Total RNA was isolated using RNeasy Plus (Qiagen, 74136) and the QIAshredder kit (Qiagen, 79656). Libraries were prepared and sequenced using Genewiz on Illumina Hiseq with 2×150bp paired-end reads. The mean quality score for sequencing was 39.03, with 94.48% of reads being ≥30. The average number of reads per sample was approximately 50,000,000. mRNA sequencing analysis was performed as previously described for the tetO cDNA experiment. GFP transgene expression was quantified by aligning trimmed reads to a custom GFP index generated with the bowtie2 build function using bowtie2. Raw counts were normalized with respect to sequencing depth and presented as relative counts across the three conditions analyzed. Statistical methods. Statistical analysis was performed using GraphPad Prism 7. For details of the specific statistical tests performed for each experiment, please refer to the legend in the diagram. Statistical significance is indicated by stars ( *This is represented by ) and shows a calculated p-value < 0.05.
[0156] (Example 2) Generation of human pluripotent stem cell lines for CRISPRa screening of neuronal cell fate To enable enrichment of neuronal cells within the CRISPRa screening framework, a 2A-mCherry sequence was inserted into exon 4 of the panneuronal marker TUBB3 in human pluripotent stem cell lines (Figure 7A). TUBB3 is expressed almost exclusively in neurons and is induced early in in vitro differentiation and reprogramming of cells into neurons. 2A-mediated ribosome skipping ensures that mCherry serves as a translation reporter for TUBB3 while mitigating interference with endogenous TUBB3 function resulting from direct protein fusion.
[0157] To enable efficient and robust targeted gene activation in the TUBB3-P2A-mCherry cell line, a lentiviral vector was used to fused dCas9 ( ) to the VP64 transactivation domain at both the N-terminus and C-terminus, under the control of the human ubiquitin C promoter (Kabadi et al. Nucleic Acids Res. 2014, 42, e147). VP64 dCas9 VP64 We established a clonal cell line that expresses ). VP64 dCas9 VP64 This has been used for some time to achieve sufficiently robust endogenous gene activation for cell fate reprogramming. VP64 dCas9 VP64To evaluate the CRISPRa approach for neuronal differentiation in the TUBB3-2A-mCherry cell line, we delivered a pool of four lentiviral gRNAs that, when ectopically overexpressed or endogenously activated by CRISPRa, target the proximal promoter of NEUROG2, a major neurogenesis regulator sufficient to generate neurons from pluripotent stem cells (Chavez et al. Nat. Methods 2015, 12, 326-328; Zhang et al. Neuron 2013, 78, 785-798). After 5 days of gRNA expression, we detected upregulation of the target gene NEUROG2, as well as the early panneuronal markers NCAM and MAP2 (Figure 7B). VP64 dCas9 VP64 Targeted gene activation was achieved only when both NEUROG2 gRNA and the specified gene were co-expressed (Figure 7B).
[0158] Following NEUROG2 gRNA delivery, 15% of mCherry-positive cells were detected 6 days after transduction compared to untreated control cells (Figure 7C). To assess the applicability of the TUBB3-2A-mCherry reporter cell line as a surrogate for the neuronal phenotype, fluorescence-activated cell sorting (FACS) was used to isolate cells with the highest and lowest 10% mCherry expression. mCherry-high cells also had high mRNA expression levels of the mCherry-tagged gene TUBB3, as well as MAP2 (Figure 7D). TUBB3-2A-mCherry cells and the CRISPRa approach were used for all screenings described in this study.
[0159] (Example 3) CRISPRa screening for major regulators of neuronal cell fate To identify a set of neuronal cell fate regulators in an unbiased manner, we performed a CRISPRa pooled gRNA screening in the TUBB3-2A-mCherry cell line (Figure 1A). The gRNA library consisted of gRNAs targeting a set of putative human TFs (Vaquerizas et al. Nat. Rev. Genet. 2009, 10, 252-263). TFs are essential for cell fate designation and are widely applied for cell reprogramming and directed differentiation applications. We selected a set of 1496 TFs and constructed a targeting gRNA library of five gRNAs for each transcription start site, extracted from a genome-wide library of optimized CRISPRa gRNAs (Horlbeck, 2016, Compact and highly active next-generation libraries. eLife) (Figure 1B).
[0160] A CRISPRa-TF gRNA lentivirus library (referred to as CRISPR-activated screening TF or CAS-TF) was transduced at an infection multiplicity (MOI) of 0.2 and library coverage of 550-fold to ensure that most cells activated a single TF, illustrating the stochastic and often inefficient nature of in vitro cell differentiation (Figure 1A). After 5 days of gRNA expression, the top and bottom 5% of mCherry-expressing cells were isolated using FACS (Figure 1C), and gRNA abundance was quantified by differential expression analysis after deep sequencing of protospacers in each sorting bin. The 5% tail portion of the mCherry distribution was recovered to allow for the identification of subtle changes in TUBB3 expression. Cells were sorted on day 5 posttransduction to limit postmittal neuron loss during extended culture time or passage, while allowing sufficient time for TF expression and reporter gene induction. Compared to the bulk, unsorted population of cells, gRNAs were significantly enriched in the mCherry-high-expression cell bin (FDR<0.01; Figure 1D). Similar results were observed when mCherry-high-expression cells were compared to mCherry-low-expression cells (Figure 8A). A set of 100 scrambled untargeted gRNAs remained unchanged between different cell bins (Figure 1D).
[0161] The degree of transcriptional activation achieved by dCas9-based activators can vary across the set of gRNAs for a given target gene. Consequently, a mixture of active and inactive gRNAs was expected for most target genes. Furthermore, off-target gRNA activity could facilitate false positives by modulating reporter gene expression independent of the predicted TF target. To ensure avoidance of overinterpretation of single-gRNA results, TFs were selected as high-confidence hits if they contained at least two gRNAs significantly enriched in the mCherry high-expression cell bin compared to both the unselected cell bin and the mCherry low-cell bin (FDR < 0.01). This approach yielded a list of 17 TFs as candidate neurogenic factors (Figure 1E). The majority of these TFs belonged to one of three of the most abundant families across all human transcription factors: C2H2 ZF, bHLH, or HMG / Sox DNA-binding domain families (Figure 1E).
[0162] The expression of 17 candidate neurogenic factors was analyzed using publicly available gene expression data from developing human brains, selected as part of BrainSpan (Miller et al. Nature 2014, 508, 199-206) (http: / / brainspan.org). The mean expression of the 17 factors, calculated across several anatomical regions and developmental time points of the human brain (see Example 1), was observed to be higher than that of a randomly generated set of the 17 TFs (Figure 1F).
[0163] As further demonstration of the fidelity of CAS-TF screening, it was observed that three well-characterized preneuronal factors, Neurod1, Neurog1, and Neurog2, each possessed several gRNAs enriched in mCherry-highly expressing cells, while a random set of five scrambled untargeted gRNAs remained unchanged (Figure 1G). A fourth gene with expected preneuronal activity, ASCL1, was not selected as a high-confidence hit based on stringent selection criteria. However, a single ASCL1 gRNA was enriched in mCherry-highly expressing cells (Figure 8A), and this gRNA was sufficient to generate mCherry-positive cells expressing NCAM and MAP2 (Figures 8B and 8C).
[0164] (Example 4) Verification of candidate neurogenic transcription factors To validate the activity of candidate neurogenic TFs, the most enriched gRNAs for 17 TFs identified by CAS-TF screening were individually tested. These gRNAs were transduced into TUBB3-2A-mCherry cell lines with a high MOI, and reporter expression was evaluated after 4 days (Figure 2A). All tested gRNAs increased the number of mCherry-positive cells to varying degrees (approximately 2% to 50%) compared to delivery of scrambled untargeted gRNAs, but only a subset of 10 factors showed statistically significant increases (Figure 2A; α=0.05). To validate CRISPRa activity, it was confirmed that all TFs were upregulated in response to the expression of the appropriate gRNA (Figure 9A). The degree of TF induction correlated directly with the basal expression level of the target gene, consistent with previous reports (Konerman Nature 2015, 517, 583-588) (Figure 9B).
[0165] Further validation of all five gRNAs represented in the CAS-TF library for ATOH1 and NR5A1 revealed a direct correlation between enrichment calculated from pooled screening when gRNAs were tested individually and the degree of differentiation assessed by reporter gene expression (Figure 2B). In some cases, gRNAs that were not significantly enriched in screening were still capable of moderate gene activation and neuronal induction (Figures 9C and 9D). For example, NEUROG2 gRNA was parallelized by NCAM and MAP2 induction, but was sufficient to upregulate NEUROG2 that was not enriched in CAS-TF screening (Figures 9C and 9D).
[0166] Given the reliance on a single reporter gene as a surrogate for the neuronal phenotype, it was expected that the TFs enriched by CAS-TF screening would include both major regulators of neuronal fate sufficient to initiate differentiation, and cofactors or downstream effectors that only modulate one or a subset of neuronal genes. To clarify these differences within the set of candidate factors, we first evaluated the expression of two other neuronal markers, NCAM and MAP2, four days after gRNA delivery. Some TFs upregulated one or both of these markers, while others produced no change or even downregulated them (Figure 2C). For example, SOX4 induced one of the largest increases in mCherry expression % at an average of 34%, but produced no detectable effect on NCAM and MAP2 expression (Figures 2A and 2C).
[0167] Immunofluorescence staining was used to assess the presence of neuronal morphology by the expression of a subset of TFs identified by CAS-CF screening (Figure 2D). Each TF-encoding cDNA was overexpressed to ensure robust TF expression and control differential gRNA activity. Several factors, including NEUROG3 and NEUROD1, generated cells with complex dendrites positively stained for TUBB3 within 4 days of expression (Figure 2D). In contrast, many TFs upregulated TUBB3 as expected but failed to generate cells with neuronal morphology. We hypothesized that the lack of morphological development in these cells could be due to a slower differentiation kinetics. Other neuronal reprogramming paradigms often require long-term culture to achieve morphological maturity. To illustrate this, cells were cultured with primordial astrocytes for an additional 11 days, and it was found that, over the extended culture period, ATOH1, ATOH7, and ASCL1 were sufficient to generate cells with complex neuronal morphology positively stained for MAP2 (Figure 2E). No similar morphological maturation was observed in KLF7, NR5A1, and OVOL1 during long-term culture. To explain the variability in the expression of these TFs across various pluripotent stem cell lines, and to confirm whether the lack of complete neuronal differentiation for several factors is a cell line-specific phenomenon, KLF7, NR5A1, and OVOL1 were also tested in H9 embryonic stem cells. A clear upregulation of TUBB3 without neuronal morphological development was similarly observed (Figure 2F). As expected, NEUROG3 was able to induce rapid differentiation with clear neuronal morphological development.
[0168] While the 17 high-confidence TF hits had a high validation rate, many preneuronal TFs, like ASCL1, were suspected of not meeting the stringent cutoff criteria. In fact, there were 109 other TFs that contained a single gRNA significantly enriched in at least mCherry-high-expressing cells, but were not called hits. To further investigate these TFs, we first focused on those that shared a subfamily with one of the 17 high-confidence hits. For example, ATOH1 was a high-confidence hit with several enriched gRNAs, while ATOH7 and ATOH8 both had only a single enriched gRNA (Figure 8A). When these gRNAs were tested individually, both ATOH7 and ATOH8 were sufficient to generate mCherry-positive cells expressing NCAM and / or MAP2 (Figures 8B and 8C), indicating that many hits with only a single enriched gRNA represent true positivity by this cutoff.
[0169] To more comprehensively examine these 109 activities, a secondary sublibrary screening was performed targeting only these TFs (Figures 10A-10E). This screening was conducted in the same manner as the primary CAS-TF screening (Figure 10A), but the new sublibrary consisted of an average of 33 gRNAs per TF (Figure 10B). This screening revealed additional gRNAs enriched in mCherry-high cells (Figure 10C). However, most of the genes in the sublibrary had relatively few enriched gRNAs, similar to the pool of scrambled untargeted gRNAs (Figure 10D). Only a few genes had gRNAs enriched in more than 40% of mCherry-high cells. However, individual examination of these gRNAs revealed almost subtle effects on the mCherry reporter (Figure 10E). This analysis indicates a robust CRISPRa screening design and confirms that our screening design is successful in identifying the most robust neurogenic factors.
[0170] (Example 5) Combinatorial gRNA screening identifies neuronal cofactors. TFs often function cooperatively to organize gene expression programs. Similarly, TF-mediated cell reprogramming often benefits from the co-expression of combinations of TFs, improving conversion efficiency, maturation, and subtype designation. The mechanisms underlying the improvements observed with co-expressed TFs are often unknown, and since effective cofactors may have minimal activity when expressed individually, predicting effective TF cocktails can be challenging. To address this challenge, we performed pooled screening of gRNA pairs to identify novel combinations of regulators that govern neuronal differentiation in human pluripotent stem cells.
[0171] The inventors hypothesized that some co-regulators of neuronal differentiation, when expressed alone, lack detectable activity and therefore would not be identified by initial single-factor CAS-TF screening. Rather, these co-regulators may require pairing with other neurogenic factors to reveal their activity. To enable the identification of such TFs, we chose to perform a screening of the remaining CAS-TF library to pair validated neurogenic TFs identified from single-factor screening (Figure 3A). Two such independent screenings were performed on single gRNAs for either NEUROG3 (sgNGN3) or ASCL1 (sgASCL1) (Figure 3A). The gRNA pairs were co-expressed in single lentiviral vectors from two independent RNA polymerase III promoters in a format adapted from a previous study (Adamson et al. Cell 2016, 167, 1867-1882 e1821). NEUROG3 and ASCL1 were selected due to their potent neurogenic activity, but they exhibited different differentiation kinetics (Figures 2D and 2E). Pair screening was performed, here so that each cell received a single pair of gRNAs, as described for single-factor screening.
[0172] Due to the constitutive presence of validated neurogenic factors in each cell, a clear population of mCherry-positive cells emerged. This basal neurogenic stimulation allowed for the detection of novel positive cofactors of differentiation, as well as the easy detection of negative regulators in mCherry-low-expressing cells (Figures 3B, 11A, and 11B).
[0173] Effective cofactors that enhance conversion efficiency are often shared across different neuronal reprogramming paradigms, but can contribute to subtype designation in context-dependent ways. Similarly, we hypothesized that many cofactors are shared between NEUROG3 and ASCL1. Consistent with this hypothesis, we found that the majority of positive regulators were shared between the two screenings (Figure 3C). However, there were several factors that were uniquely enriched when combined with either NEUROG3 or ASCL1 (Figure 3C). For example, FEV was positively enriched only with NEUROG3, while NKX2.2 was positively enriched only with ASCL1. Importantly, both the sgNGN3 and sgASCL1 screenings identified novel TFs not observed in single-factor CAS-TF screenings (Figures 12A-12D). Many of these TFs, including LHX6, LHX8, and HMX2, are involved in neuronal development and subtype designation, but have not been broadly characterized for neuronal in vitro generation. A list of all candidate neurogenic factors identified across all three screenings can be found in Table 1.
[0174] JPEG2026058344000006.jpg226166 JPEG2026058344000007.jpg173163
[0175] Positive hits from the two paired CAS-TF screenings encompassed a diverse set of TF families (Figure 3D). While the majority of these TFs were not expressed or were low in pluripotent stem cells, several factors were more highly expressed (Consortium. Nature 2012, 489, 57-74) (Figure 3D). A set of eight TFs was selected for further validation. These TFs were expected to have minimal activity on their own, but enhanced neurogenic activity when co-expressed with NEUROG3 and / or ASCL1 (Figure 3E). This subset of eight TFs was selected for further characterization, but numerous other candidate factors, revealed by CRISPRa paired screening, exist that may be subject to future testing (Table 1). All tested TFs, when paired with sgNGN3, improved the conversion efficiency to mCherry-positive cells by up to 3-fold compared to sgNGN3 co-expressed with scrambled gRNA (Figure 3F). Since sgASCL1 only increased the mCherry reporter to a moderate level, we chose to use NCAM staining for gRNA validation for pairing with this gRNA. Only E2F7 and HMX2 had a moderate effect on NCAM expression on their own (Figure 3G). However, some TFs significantly increased the neurogenic activity of ASCL1, including up to 8-fold for E2F7 (Figure 3G). Consistent with the results expected from the screening, NKX2.2 had a significant effect only with ASCL1 and no significant effect with NEUROG3 (Figures 3E, 3F, and 3G).
[0176] (Example 6) Neurogenic transcription factors regulate subtype specificity and maturation. Neuronal subtype identity and the degree of synaptic maturation are key features that define the usefulness of in vitro induced neurons for disease modeling and cell therapy applications. Consequently, the development of protocols to improve maturation kinetics and the purity of subtype designation has been a major focus in this field. Given the diversity of neurogenic TFs identified through CRISPRa screening and the range of conversion efficiencies observed through validation experiments, we hypothesized that many of these TFs appear to influence subtype identity and maturation in different ways. To begin addressing this question, we performed bulk mRNA sequencing to more comprehensively assess the degree of neuronal conversion and compare transcriptional diversity in neuronal populations generated by different TFs.
[0177] We began by analyzing neurons induced from single TFs. While TF combinations often enhance the specificity of subtype generation and improve conversion efficiency and maturation kinetics, a single TF may be sufficient to generate functional neurons with subtype tendencies. Initially, we chose to perform mRNA sequencing on neurons induced from ATOH1 or NEUROG3 overexpression (Figures 4A–4F). These TFs possess some of the highest conversion efficiencies determined through validation experiments (Figures 2A–2F) and facilitate the isolation of sufficient material for sequencing. Furthermore, while the neurogenic activity of both ATOH1 and NEUROG3 has been previously confirmed, our understanding of their roles in in vitro neuronal differentiation remains incomplete.
[0178] TUBB3-mCherry-positive cells were purified using FACS after overexpression of cDNA encoding either ATOH1 or NEUROG3, and mRNA sequencing was performed after 7 days of transgene expression. Both neuronal populations had over 3000 genes that were upregulated compared to the undifferentiated pluripotent stem cell initiation population (Figure 4A). A set of shared genes was enriched in gene ontology (GO) terms related to neuronal differentiation and development (Figure 4B). Importantly, the set of panneuronal genes was highly enriched across all replications for ATOH1 (3 replications) and NEUROG3 (2 replications) compared to pluripotent stem cells (Figure 4C).
[0179] Surprisingly, strong correlations were observed across all detectable genes between ATOH1-induced and NEUROG3-induced neurons, demonstrating remarkable consistency in the induction of core neuronal programs and the suppression of pluripotency networks (Figure 4D). However, a subset of genes were more highly expressed in either ATOH1 or NEUROG3 (Figure 4D). These genes were enriched with GO terms associated with glutamatergic activity for NEUROG3 and dopaminergic activity for ATOH1 (Figure 4E). Indeed, examining the expected sets of markers for the two neuronal subtypes revealed a clear enrichment of dopaminergic markers for ATOH1 and glutamatergic markers for NEUROG3 (Figure 4F). While certain canonical markers of dopaminergic neurons, such as tyrosine hydroxylase (TH), remained low in expression, many TFs associated with dopaminergic designation, such as LMX1A, were more highly expressed in ATOH1-induced neurons (Figure 4F).
[0180] In many cases, combinations of TFs can aid in the accuracy of neuronal subtype designation or enhance conversion efficiency and maturation. We hypothesized that cofactors identified by paired gRNA screening, when combined with neurogenic factors identified by single-factor screening, would serve as primary candidates for modulating subtype identity and maturation. As a result, we chose to perform mRNA sequencing on neurons induced from NEUROG3, either alone or in combination with E2F7, RUNX3, or LHX8. These three cofactors were preferred due to their substantial impact on differentiation efficiency as assessed through gRNA validation (Figures 3A–3G). NEUROG3 was selected due to its predetermined preference for generating glutamatergic neurons, often considered the default subtype. cDNA encoding NEUROG3 was overexpressed either alone or in combination with E2F7, RUNX3, or LHX8, followed by mRNA sequencing after 6 days of transgene expression. Similar to the comparison between ATOH1 and NEUROG3, all TF pairs shared a core set of upregulatory genes (Figure 5A). However, the genes uniquely upregulated in each TF pair compared to NEUROG3 alone were enriched in GO terms related to neuronal differentiation and development, consistent with the previously measured increase in TUBB3 expression and the improved conversion efficiency in the expression of these neuronal cofactors (Figure 5B).
[0181] Importantly, each TF pair uniquely upregulated genes associated with the designation and maturation of specific neuronal subtypes. For example, the addition of RUNX3 led to increased expression of NTRK3, which encodes the TrkC neurotrophin-3 receptor associated with the development of dorsal root ganglion neurons (Figure 5C). The addition of E2F7 led to increased expression of CDKN1A, which encodes the p21 cell cycle regulator involved in neuronal fate commitment and morphogenesis (Figure 5D). A subset of genes more highly expressed with the addition of LHX8 was enriched in the synaptic gene ontology (SynGO) term, a hallmark of neuronal maturation, associated with synaptic development (Figure 5E). Consistent with GO term analysis, a set of genes associated with synaptic development, regulation, and function was clearly upregulated with the addition of LHX8 (Figure 5F).
[0182] To evaluate whether the addition of LHX8 affected the electrophysiological maturation of NEUROG3-induced neurons, patch-clamp recordings of TUBB3-2A-mCherry-positive cells were performed 7 days after transgene induction. No differences were observed in the resulting membrane potentials (Figure 5G), but with the addition of LHX8, a decrease in membrane resistance (Figure 5H) and an increase in membrane volume (Figure 5I) were observed compared to NEUROG3 alone. LHX8 improved several metrics of action potential maturation, including a decrease in firing threshold (Figure 5J), an increase in action potential height (Figure 5K), and a decrease in action potential half-width (Figure 5L). Furthermore, neurons with LHX8 fired action potentials more frequently for a given step depolarization by current injection (Figure 5M) and had a higher proportion of recording cells firing multiple action potentials (Figure 5N). Cells generated by NEUROG3 alone more frequently failed to fire or fired only a single low-amplitude action potential (Figure 5N).
[0183] (Example 7) Combinatorial gRNA screening identifies negative regulators of neuronal fate. The conversion efficiency achieved in cell reprogramming and differentiation protocols often varies depending on the starting and ending cell types. Generally, more distantly related cell types or older cell lines are less receptive to conversion. For example, reprogramming astrocytes into neurons is often more efficient than reprogramming fibroblasts into neurons, and this efficiency is further reduced in adult fibroblasts compared to embryonic fibroblasts. These discrepancies in reprogramming outcomes can be explained in part by variations in gene expression profiles and the epigenetic landscape of cells of different types or developmental ages. Consequently, this cellular situation can create barriers that hinder proper TF activity, reducing conversion efficiency and fidelity.
[0184] High-throughput loss-of-function RNAi screening has helped identify molecular barriers that hinder cell type reprogramming and affect conversion efficiency. Importantly, removing such barriers often results in significant improvements in reprogramming outcomes. Through paired CRISPRa screening, TFs whose activation impedes neuronal differentiation were identified (Figures 3B, 11A, and 11B). Negative regulators of these candidates included several members of the HES gene family of downstream canonical neuron repressors of Notch signaling, in addition to many other uncharacterized TFs. A list of all candidate negative regulators identified across all three screenings can be found in Table 2.
[0185] JPEG2026058344000008.jpg233163 JPEG2026058344000009.jpg131160
[0186] Interestingly, the majority of negative regulators were shared across sgNGN3 and sgASCL1 screenings (Figure 6A). These consisted of a diverse set of TFs across many TF families with broad basal expression in embryonic stem cells. When individually tested with single gRNAs co-expressed with NEUROG3 gRNA, several TFs, including HES1 and DMRT1, reduced the percentage of mCherry-positive cells to regulated levels (Figure 6B). To demonstrate that this repression was not limited to the reporter gene alone, we also demonstrated up to an 8-fold reduction in NCAM expression by seven of the eight repressors tested (Figure 6C). Similar suppression of neuronal differentiation was observed when these factors were tested in H9 human embryonic stem cells (Figure 6D). In fact, there was a significant correlation between the relative effects of these negative regulators in iPSCs compared to ESCs (Figure 6E), highlighting the robustness of these effects across multiple pluripotent stem cell lines.
[0187] The inventors hypothesized that some of these identified negative regulators, which are basally expressed in pluripotent stem cells, may act as barriers to neuronal conversion, and therefore, inhibiting them may improve differentiation efficiency. Cas9 proteins from various bacterial species can be programmed for orthogonal gene regulation and epigenetic modification. Therefore, orthogonal dSaCas9 based on the Cas9 protein from Staphylococcus aureus. KRAB Using (Thakore et al. Nat. Commun. 2018, 9, 1674), we chose to target the promoters of two negative regulators basally expressed in pluripotent stem cells, ZFP36L1 and HES3 (Figure 6F). dSaCas9 KRAB Targeting the promoters of these genes resulted in 10-fold and 4-fold transcriptional repression of ZFP36L1 and HES3, respectively (Figure 13A).
[0188] dSaCas9 for targeted gene expression KRAB The use of orthogonal for the simultaneous activation of neurogenic factors VP64 dSpCas9VP64 Co-expression of TUBB3-2A-mCherry becomes possible (Figure 6F). P64 dSpCas9 VP64 iPSCs initially co-express ZFP36L1, HES3, or scrambled Staphylococcus aureus gRNA with dSaCas9. KRAB Cells were transduced using lentiviruses. Nine days after transduction with Staphylococcus aureus gRNA, cells were transduced with lentiviruses encoding either sgNGN3 or sgASCL1 from Streptococcus pyogenes, and analyzed four days after this final transduction. Knockdown of ZFP36L1 resulted in a twofold increase in the percentage of mCherry-positive cells obtained with sgNGN3 compared to a control cell line expressing scrambled Staphylococcus aureus gRNA (Figure 13B). Similarly, knockdown of ZFP36L1 resulted in a 1.2-fold increase in mCherry reporter gene expression levels in the NCAM-positive population of differentiated cells obtained with sgASCL1 (Figure 13C).
[0189] To identify the genome-wide effects of this orthogonal CRISPR-based regulation, mRNA sequencing was performed in neurons induced from NGN3 activation concurrently with the suppression of ZFP36L1 or HES3. HES3 knockdown resulted in only very subtle changes in gene expression compared to cells treated with scrambled Staphylococcus aureus gRNA (Figure 14A), while ZFP36L1 knockdown led to significant changes in the overall gene expression profile compared to NGN3 activation alone (Figures 6G and 14B). Subtle increases in the expression of NEUROG3 and Streptococcus pyogenes gRNA, quantified by GFP transgene expression in gRNA vectors, were also observed in ZFP36L1 knockdown cells (Figures 14C and 14D). Genes upregulated in neurons by ZFP36L1 knockdown were enriched in GO terms related to neuronal differentiation and morphological development (Figure 6H). In contrast, genes downregulated by ZFP36L1 knockdown were enriched in GO terms related to cell cycle development and progression (Figure 6H). Examples of genes upregulated by ZFP36L1 knockdown include the neuronal transcription factors NEUROD4, INSM1, and OLIG2, as well as genes involved in neuronal morphogenesis, including NEFL, NGEF, and NTN1 (Figure 6I).
[0190] (Example 8) Consideration As detailed herein, 1496 putative human transcription factors were systematically profiled for their role in regulating neuronal differentiation in pluripotent stem cells through single and combinatorial CRISPR screening. This study highlights the usefulness of CRISPR-based technologies for disrupting gene expression in a high-throughput manner and emphasizes the robust nature of dCas9-based gene activation for testing the causal role of gene expression in complex cellular phenotypes.
[0191] The use of early panneuronal markers such as TUBB3 as a surrogate for the neuronal phenotype enabled the identification of a broad set of TFs with diverse neurogenic activity. For example, NEUROG3 was sufficient to rapidly generate neurons within 4 days of expression, while ATOH7 and ASCL1 required longer culture times to achieve similar phenotypes (Figures 2D and 2E). The addition of cofactors, such as those identified by combinatorial gRNA screening, may improve the efficiency and kinetics of differentiation, as can be seen in other cell reprogramming studies (Pang et al. Nature 2011, 476, 220-223). Furthermore, several TFs, including KLF7, NR5A1, and OVOL1, induced TUBB3 expression but failed to generate neurons (Figure 2D). These TFs may serve as cofactors or downstream regulators requiring co-expression of other neurogenic factors to achieve more complete differentiation. In fact, many of the TFs identified by single-factor screening were also hits in paired gRNA screening (Table 1).
[0192] Several TFs with apparent neurogenic activity, including ASCL1 and ATOH7, were found to possess only single gRNAs enriched by CAS-TF screening (Figure 8). Since single enriched gRNAs may be the result of off-target activity or noise, accurately classifying these gRNAs can be challenging. The use of multiple gRNAs per gene or next-generation dCas9-based activator platforms may help more accurately define true positive effects. Indeed, sublibrary screening using multiple gRNAs per gene revealed several additional candidate hits (Figure 10). Further improvements in gRNA design and screening analysis can continue to make CRISPR-based screening more robust and scalable to more complex phenotypes.
[0193] Through the use of paired gRNA screening, a set of TFs that improve neuronal differentiation efficiency, maturation, and subtype designation was identified. Interestingly, most of these TFs did not possess neurogenic activity on their own, as assessed by single-factor CAS-TF screening. This observation highlights the importance of synergistic TF interactions governing cell differentiation and supports the use of unbiased methods for identifying these TFs. E2F7 was identified as improving neuronal conversion efficiency, likely due to its known role in inhibiting cell proliferation as a critical switch in the conversion of proliferative pluripotent stem cells to postmittal neurons (Figures 3F and 3G). Furthermore, RUNX3 was found to uniquely induce subtype-specific receptor gene expression (Figure 5C), thus potentially being a useful additive to differentiation protocols for more accurately guiding neuronal subtype identity. The neuronal cofactor LHX8 had a profound impact on markers of neuronal maturation, as seen in enrichment of many synapse-related genes and a clear improvement in electrophysiological maturation (Figure 5). Functional synapse formation is an essential phenotype for in vitro induced neurons and is often the rate-limiting step. Improving synaptic maturation through TF programming may help facilitate the development of useful neuronal models for disease modeling and drug screening.
[0194] Future studies may utilize advanced screening platforms to further characterize cell lineage identifiers. A more comprehensive list of neuronal TFs may have been identified by performing screenings that rely on multiple neuronal markers or using maturation or subtype identity markers. Alternatively, instead of assaying a small number of separate markers, these screenings could be performed with single-cell RNA sequencing (scRNA-seq) output to more accurately define the diversity of neuronal phenotypes obtained with various TF combinations and evaluate these results against a growing atlas of scRNA-seq data from human brain samples. TFs identified from the screenings detailed herein may serve as primary candidates for sublibraries to be tested in these alternative approaches, which may be more limited in scale by library size.
[0195] Paired gRNA screening also identified negative regulators of neuronal differentiation. Knockdown of one of these TFs, ZFP36L1, was sufficient to improve differentiation and resulted in an overall change in gene expression toward a more differentiated neuronal phenotype (Figure 6G, Figure 6H, Figure 6I). While the effect on differentiation was somewhat modest in this example, more dramatic improvements may be seen in less convertible cell types such as adult fibroblasts. Importantly, many of the negative regulators identified in the screening are expressed in other cell types used in reprogramming studies, such as fibroblasts and astrocytes. Additional CRISPRa screening targeting epigenetic modifiers or other gene subsets besides TF may help further elucidate the extent to which gene activation can regulate neuronal cell fate. Continued development of synthetic systems for programmable regulation of endogenous gene expression and chromatin state, as well as their application to more complex in vitro and in vivo models, may enable testing to more comprehensively define the gene networks and epigenetic mechanisms governing cell fate determination.
[0196] Overall, as detailed herein, a broad set of transcription factors controlling neuronal fate designation in human cells has been identified. This catalog of factors may serve as a basis for developing protocols for generating diverse neuronal cell types with high efficiency and fidelity for applications in regenerative medicine and disease modeling. Finally, the CRISPRa screening platform detailed herein can be extended to other cell reprogramming paradigms to facilitate the in vitro generation of many clinically relevant cell types.
[0197] (Example 9) High-throughput CRISPR activation screening to identify novel driving factors of myogenic progenitor cell fate Skeletal muscle regeneration is a complex process mediated by muscle satellite cells. While the cascade of events driving proper myogenic differentiation from muscle satellite cells is well-characterized, the upstream events that determine satellite cell fate during embryonic development are not fully understood. The transcription factor PAX7 plays a crucial role in the designation and maintenance of satellite cells, and its overexpression may designate myogenic progenitor cell fate in human pluripotent stem cells. To investigate novel drivers of satellite cell fate, we generated a PAX7-2a-GFP cell line in human H9 embryonic stem cells. We systematically identified independent drivers of PAX7 expression by applying a gRNA library targeted by the promoters of all human transcription factors and co-delivering CRISPR / Cas9-based transcription activators. Subsequently, we investigated cofactors of PAX7 by performing a secondary screening, applying the gRNA library together with PAX7 promoter-targeted gRNAs. This secondary screening identified a distinct set of transcription factors, bringing the total number of identified transcription factors to 21. Individual validations demonstrated the induction of PAX7 expression and the adoption of myogenic cell fate for some of the hits. The data generated from this study can be used to identify potential therapeutic targets for skeletal muscle regeneration in cell therapy and gene therapy contexts.
[0198] Generation of PAX7-2a-GFP cell lines. Human H9 ESCs (obtained from WiCell Stem Cell Bank) were used in these studies, maintained in mTeSR (Stem Cell Technologies), and seeded on tissue culture-treated plates coated with ES-certified Matrigel (Corning). H9 ESCs were co-transfected with a Cas9-gRNA plasmid targeting the PAX7 isoform A stop codon and a donor plasmid having homology arms complementary to exon 8 and 3'UTR of PAX7 isoform A. Transfection was performed in a 4 mm cuvette at 250 V, 750 μF, and infinite resistance using GenePulser Xcell (Bio-Rad). The donor plasmid also contained a PGK-PuroR cassette surrounded by a loxP site to enable selective cell proliferation by donor plasmid incorporation. After two weeks of puromycin selection (1 μg / mL), clones were harvested and screened by PCR for integration of the donor cassette at the correct genomic locus. Selective-positive clones were transfected with Cre recombinase plasmid to remove the large PGK-PuroR cassette. Cells were seeded at low density, clones were harvested, and screened for correct integration using primers outside the donor template. The resulting PCR bands were confirmed by Sanger sequencing.
[0199] Generation of a CRISPR-activated transcription factor (CRa-TF) gRNA library. Putative human transcription factors were selected based on a previously curated list. Corresponding gRNAs available for the gene list were extracted from the human subpool CRISPRa library. 100 scrambled untargeted gRNAs were also extracted from this library. The custom library consists of 1496 unique genes with 5 targeted gRNAs per transcription start site and 100 scrambled untargeted gRNAs for the total library size of 8505 gRNAs. The oligonucleotide pool (Custom Array) was PCR amplified and cloned using PAX7 promoter-targeted gRNAs, either into single gRNA expression plasmids for single CRa-TF screening or into dual gRNA expression plasmids for paired CRa-TF screening using Gibson assembly.
[0200] Lentivirus generation. HEK293T cells were obtained from the American Tissue Collection Center (ATCC), purchased through the Duke University Cancer Center facility, and cultured at 37°C and 5% CO2 in Dulbecco's Modified Eagle Medium (Invitrogen) supplemented with 10% FBS (Sigma) and 1% penicillin / streptomycin (Invitrogen). Approximately 3.5 million cells were seeded per 10 cm TCPS dish. After 24 hours, cells were transfected with expression plasmids, pMD2.G envelope plasmid (Addgene number 12259), and psPAX2 second-generation packaging plasmid (Addgene number 12260) using calcium phosphate precipitation. The medium was changed 12 hours after transfection, and the viral supernatant was collected 24 and 48 hours after this medium change. The viral supernatant was pooled, centrifuged at 500 g for 5 minutes, passed through a 0.45 μm filter, and concentrated 20-fold using a Lenti-X Concentrator (Clontech) according to the manufacturer's protocol. The lentiviral gRNA library was titrated by flow cytometry.
[0201] High-throughput CRa-TF screening for upstream regulators of PAX7. VP64 dCas9 VP64 Undifferentiated H9 PAX7-2a-GFP cells that stably express PAX7-2a were isolated, and 22.5 × 10⁴ cells were measured. 6 Individual cells were transduced with the CRa-TF lentivirus library at an MOI of 0.2 per replication (3.1 × 10⁻¹⁰). 4 cells / cm 2 The goal was to achieve a library coverage of 500-fold per replication. Cells were selected with 1 μg / mL puromycin for 6 days. For differentiation, hESCs were isolated into single cells with Accutase (Stem Cell Technologies) and seeded on Matrigel-coated plates in mTeSR medium supplemented with 10 μM Y27632 (Stem Cell Technologies) (3.6 × 10⁶). 4 cells / cm 2 The following day, the mTeSR medium was replaced with E6 medium supplemented with 10 μM CHIR99021 (Sigma) to initiate mesoderm differentiation. Two days later, CHIR99021 was removed, and the cells were maintained in E6 medium supplemented daily with 10 ng / mL FGF2 (Sigma). The cells were not passaged during the differentiation duration, which was 2 weeks for version 1 screening and 1 week for version 2 pre-analysis screening.
[0202] One or two weeks after differentiation induction, cells were isolated with 0.2% collagenase II (ThermoFisher) and washed with neutralizing medium (10% FBS in DMEM / F12). Cells were pelleted by centrifugation and resuspended in fluid medium (5% FBS in PBS). Cells were gated for positive mCherry expression, and the top 10% and bottom 10% of GFP-expressing cells were sorted into separate tubes using a SONY SH800 flow cytometer. The sorted cells were pelleted, and genomic DNA was extracted using the Qiagen DNeasy kit. Unsorted cells were also set aside for genomic DNA isolation and used as input controls. gRNA sequences were recovered from genomic DNA by PCR. Sequence determination was performed on Illumina Miseq using 21 bp paired-end sequencing with custom reads and index primers.
[0203] Data processing and enrichment analysis. FASTQ files were aligned to a custom index (generated from the bowtie2 build function) using Bowtie with the options -p32--end-to-end --very-sensitive -3 1 -I 0 -X 200. Counts for each gRNA were extracted and used for further analysis. All enrichment analyses were performed using the R language. For individual gRNA enrichment analyses, the DESeq2 package was used to compare each screening between high and low conditions, unselected and low conditions, or unselected and high conditions. Individual gRNA validation. Protospacers from the top enriched gRNAs found in each screening were ordered as oligonucleotides from the IDT and cloned into lentiviral gRNA expression vectors as described above. The same H9 PAX7-2a-GFP cell line used for pooled CRa-TF screening was used for individual gRNA validation. The cells were transduced with the individual gRNAs and subjected to the same puromycin selection and differentiation protocols as the original screening, albeit on a smaller scale.
[0204] RNA was isolated using the RNeasy Plus RNA Isolation Kit (Qiagen). cDNA was synthesized using the SuperScript VILO cDNA Synthesis Kit (Invitrogen). Real-time PCR using PerfeCTa SYBR Green FastMix (Quanta Biosciences) was performed on the CFX96 Real-time PCR Detection System (Bio-Rad). The results were analyzed using ΔΔC. t The method is used to express the target gene as a multiplier increase in expression normalized to GAPDH expression.
[0205] Immunofluorescence staining of cultured cells. For differentiation, cells were grown to confluence and differentiated in 24-well tissue culture plates coated with Matrigel. Immunofluorescence staining was performed directly in the wells. Cells were fixed with 4% PFA for 15 minutes and permeabilized at room temperature for 1 hour with blocking buffer (PBS supplemented with 3% BSA and 0.2% Triton X-100). Samples were incubated overnight at 4°C with PAX7 (1:20, Developmental Studies Hybridoma Bank) and myosin heavy chain MF20 (1:200, Developmental Studies Hybridoma Bank). Samples were washed with PBS for 15 minutes and incubated at room temperature for 1 hour with a 1:500 diluted compatible secondary antibody and DAPI from Invitrogen. Samples were washed three times with PBS for 5 minutes each, the wells were maintained in PBS, and imaged using conventional fluorescence microscopy.
[0206] Results: Generation of a PAX7 reporter strain in human ESCs. PAX7 may be important for satellite cell designation, function, and maintenance. Since adult satellite cells are also identified by the unique expression of PAX7, we decided to use this gene to generate a satellite cell reporter strain. Three gRNAs designed to cleave near the stop codon of PAX7 were tested in H9 ESCs, and SURVEYOR analysis revealed that gRNA 1 had the best cleavage activity. A donor template containing a homology arm and a P2A-eGFP sequence for insertion downstream of the last exon of PAX7 was designed (Figure 15A). H9 ESCs were cotransfected with a donor vector containing a CRISPR / Cas9 plasmid and a loxP-adjacent PGK-PuroR cassette to allow selection of the recombination event. Resistant clones were molecularly validated, and the selection cassette was excised by Cre recombination. The resulting clones were further validated by PCR using primers designed to pan outside the homology arm (Figure 15B). Larger embedded bands from multiple clones were validated by Sanger sequencing to ensure the in-frame position of the reporter cassette (Figure 15C). Smaller wild-type bands were also sequenced to ensure that no indels were generated in the non-reporter allele. One clone was selected and used for subsequent testing.
[0207] In cells, VP64 dCas9 VP64 Reporter activity was validated by transduction with a lentiviral vector encoding a gRNA targeted by the PAX promoter, thereby activating endogenous gene expression. Flow cytometry analysis showed a clear shift in GFP expression in the clonal population compared to untransduced cells (Figure 15D). The top 15% and bottom 15% of GFP-expressing cells were selected, and RNA was extracted for qRT-PCR, demonstrating a positive correlation between GFP and PAX7 expression (Figure 15E). CRa-TF screening to identify novel regulators of PAX7 expression. To systematically identify TFs acting upstream of PAX7, gRNA libraries targeting the promoters of all putative TFs were generated based on a previously selected list. Corresponding gRNAs available for the list of genes were extracted from a previously generated human subpool CRISPRa library. The custom CRISPRa-TF (CRa-TF) library generated for testing contained 5 targeted gRNAs per transcription start site for 1496 unique genes and 100 scrambled untargeted gRNAs for the total library size of 8505 gRNAs.
[0208] Since PAX7 is expressed in the neural crest induced from the ectoderm during embryogenesis, screening was paired with a mesoderm differentiation protocol to facilitate myogenic lineage designation. Differentiation of hPSCs into mesoderm cells can be initiated by the addition of the small molecule CHIR99021, a GSK3 inhibitor. Before differentiation, VP64 dCas9 VP64 The cell line was transduced to stably express FGF2. Next, the CRa-TF library was transduced at an MOI of 0.2, selection was applied, and the cells were allowed to differentiate for 2 weeks in serum-free medium in the presence of FGF2 (Figure 16A). The inventors previously determined that 2 weeks of mesoderm differentiation alone is not sufficient to induce GFP expression.
[0209] CRa-TF library and differentiation, GFP +A identifiable population of cells emerged, and the top 10% and bottom 10% of GFP-expressing cells were selected by FACS (Figure 16B). Next-generation sequencing (NGS) was performed to identify gRNAs enriched in either group. When low-GFP-expressing cells were compared to unselected cells, no hits were found, indicating that this population of cells completely lacked PAX7 expression. When high-GFP-expressing cells were compared to unselected cells, 10 unique genes (not containing PAX7 gRNA) were identified as significant (Figure 16C). These gRNAs were individually cloned into lentiviral vectors and validated in the same cells using a two-week differentiation protocol (Figure 16D). Equivalent cDNAs were also cloned into lentiviral constructs, and it was determined that protein delivery, to varying degrees, could lead to PAX7 activation (Figure 16E).
[0210] Combinatorial CRa-TF screening to identify synergistic TFs with PAX7. While mesodermal differentiation by small molecules has been shown to generate myogenic cells, this also leads to differentiation into heterogeneous cell types, including neurons. Mesodermal differentiation by CHIR99021 is also used for the differentiation of pluripotent cells into cardiac and renal lineages. It has been previously demonstrated that PAX7 cDNA expression during differentiation time can influence cells to adopt a myogenic cell fate more readily than alternative lineages.
[0211] A secondary screening was performed by adding an mU6-PAX7 promoter-targeted gRNA cassette to a lentiviral CRa-TF library (Figure 17A). This screening also had the potential to identify TFs that synergistically work with PAX7 to enhance myogenic progenitor cell designation. Anticipating rapid upregulation of PAX7, the screening was performed as previously described, except that differentiation was reduced from two weeks to one week. After one week of differentiation, a clear shift in the GFP population was observed, and the top 10% and bottom 10% of GFP-expressing cells were selected (Figure 17B). This secondary screening revealed 13 TFs that, when co-expressed with PAX7, have an additive effect on PAX7 expression. Overall, both screenings yielded a list of 21 TFs that upregulate PAX7 at the mesodermal differentiation stage (Figure 17C).
[0212] Verification of hit TFs that promote myogenic differentiation. Next, we wanted to determine whether TFs could not only upregulate PAX7 expression but also produce myogenic cells. Each of the 21 TF gRNA hits was cloned into a lentiviral vector expressing rtTA3 and used a tetracycline-inducible promoter. VP64 dCas9 VP64The expression of [specific gene] was driven. Both constructs were transduced into H9 PAX7-2a-GFP cell lines, and the cells were differentiated for 28 days in the presence of doxycycline (dox) using a passage step on day 14. After 28 days, dox was removed to allow downregulation of PAX7, which in turn allowed upregulation of downstream myogenic genes to induce terminal differentiation of myogenic progenitor cells into myocytes (Figure 18A). qRT-PCR analysis showed slightly upregulated PAX7 expression under most conditions after 2 weeks of terminal differentiation compared to scrambled gRNA controls. Surprisingly, three TFs, MYOD, DMRT1, and PAX3, demonstrated higher PAX7 expression compared to PAX7 gRNA expression controls (Figure 18B). When the expression of the downstream myogenic marker, MYOG, was examined, it was found to be highly expressed in 8 of the 21 novel TF gRNA hits (Figure 18C). Finally, immunofluorescence staining was performed on fixed differentiated cells to check for the presence of myosin heavy chain (MHC)-positive muscle fibers (Figure 18D). PAX7 was also stained to determine whether any of the novel hits could generate cell types capable of maintaining the PAX7+ satellite cell phenotype. Many of the putative hits expressing MYOG also showed the presence of MHC+ muscle fibers. DMRT1 showed the highest number of PAX7+ nuclei and generated muscle fibers most robustly.
[0213] Discussion. This study screened all human TFs for myogenic progenitor cell fate designation using an unbiased systematic approach. Using PAX7 expression as a surrogate for satellite cell designation, we generated a PAX7-2α-GFP human embryonic stem cell line to reveal novel upstream regulators of PAX7 during myogenic differentiation. Using individual and combinatorial CRISPRa screening, we generated a list of 21 putative TFs demonstrating PAX7 activation. A subset of these TFs also demonstrated the ability to differentiate ESCs into muscle fibers. Hits such as TWIST1 and PAX3 were not surprising, as their importance for paraxial mesoderm development has been previously characterized. PAX3, in particular, is a paralog of PAX7, and these have overlapping functions as upstream regulators of myogenesis. MYOD and MYOG were interesting hits, as they are understood to be downstream of PAX7 expression during myogenesis. A possible explanation is that the overexpression of these myogenic factors propels embryonic stem cells toward a myogenic program, generating major muscle fibers in the sarcomeres, which then creates a positive feedback loop that generates further PAX7-induced embryonic myoblasts. In two versions of the CRISPRa screening performed in this study, SOX9 and SOX10 were the only TFs that appeared as hits in both. SOX9 and SOX10 are both important TFs in development, and SOX factors are generally involved in cell fate determination. The involvement of SOX9 ranges from chondrogenesis to central nervous system development and has also been shown to enhance the differentiation of ESCs into progenitor cells of all three germ layers. Like SOX9 and PAX7, SOX10 also plays a crucial role in neural crest development. Unlike PAX7, SOX10 is not expressed in the mesoderm; SOX10-deficient embryos show a significant decrease in PAX7+ muscle progenitor cells and reduced sarcomereogenesis. The combination of previous studies linking SOX9 and SOX10 to differentiation and appropriate myogenesis, and the appearance of these TFs in CRa-TF screening, solidifies their importance in designating myogenic progenitor cells.
[0214] Of all the hits analyzed, one TF, DMRT1 in particular, showed an exciting ability to generate a large number of PAX7+ cells in abundant muscle fibers in vitro. DMRT1 is a particularly unexpected hit, as it is primarily recognized as a sex-determining gene. This gene is predominantly expressed in Sertoli cells and is required for testicular maturation. Interestingly, PAX7 was recently identified as a marker for a rare subpopulation of mouse spermatogonial cells with stem cell-like characteristics. While there is no defined association between DMRT1 and PAX7 in either spermatogenesis or myogenesis context, the results suggest that DMRT1 acts upstream of PAX7, activating its expression and giving cells a stem cell phenotype. In the mesodermal differentiation context used for screening, this results in myogenic progenitor cell and myofibrillation. While this process may not be a naturally occurring phenomenon, DMRT1 overexpression could be utilized to generate robust myogenic progenitor cells for cell therapy.
[0215] In conclusion, a robust CRISPRa screening was performed on all human TFs, revealing hits that were a combination of expected, interesting, and surprising findings. These results shed light on the understanding of satellite cell development and PAX7 upstream regulators, which may be useful for manipulating myogenic progenitor cells. The approach developed in this study has broad utility for discovering novel TFs and enhancing the manipulation of other cell lines.
[0216] (Example 10) Identification of transcription factors that regulate cartilage formation Novel driving factors for chondrocyte-specific gene expression were identified using a high-throughput CRISPR activation screening similar to that detailed in Example 9. Genes specifically expressed in collagen were used as chondrocyte-specific markers. Chondrocyte-specific transcription factors were identified. Generation of TF-targeted CRISPR activation library. gRNAs targeting the annotated TFs described in the previous example were extracted from the library, and a library consisting of 8435 gRNAs (approximately 5 gRNAs per TF) was obtained. The library was amplified and cloned into a modified lentiviral CRISPR construct containing the mCherry-2A-Puro® expression cassette using Gibson Assembly.
[0217] Lentivirus production and titer measurement. Lentiviral packaging of the gRNA library and the VP64-dCas9-VP64 expression vector was performed by transfecting 3E6 HEK 293T with the pooled gRNA library plasmid or the VP64-dCas9-VP64 plasmid (20 μg), pMD2.G (Addgene, 12259, 6 μg), and psPAX2 (Addgene, 12260, 15 μg) using calcium phosphate precipitation. After 16 hours, the medium was changed. Viral supernatants were collected after both 24 hours and 48 hours and concentrated using the Lenti-X concentrator system (Clonetech) according to the manufacturer's instructions.
[0218] Eight hours after seeding, transduction efficiency of the lentivirus containing the gRNA library was measured by transducing COL2A1-2A-GFP;VP64-dCas9-VP64 hiPSCs in 24-well plates at 60K cells / cm 2 A 10-fold serial dilution of the concentrated lentivirus ranging from 5E-5 to 5 μL was added to the medium. The medium was changed 16 hours after transduction, and mCherry fluorescence was measured using a BD Accuri C6 cytometer to determine the transduction efficiency at D3. Generation and verification of CRISPR activator hiPSC lines. COL2A1-2A-GFP reporter hiPSCs were transduced with a lentivirus having an expression cassette of dCas9 fused with the VP64 transactivation domain at the N-terminus and C-terminus as described above. Cells were selected with 100 μg / mL blasticidin for 5 days. The obtained polyclonal lines were verified by transduction of NGN2-targeting gRNA. Three days later, the cells were lysed and NGN2 expression was evaluated by qRT-PCR.
[0219] Gene expression. Monolayer and pellet cells were rinsed with DPBS. Monolayer cells were lysed with 350 μl of buffer RL (Norgen Biotek, Thorold, Canada). RNA was isolated using a total RNA purification kit according to the manufacturer's recommendation (Norgen Biotek). Reverse transcription was performed using SuperScript™ VILO™ Master Mix (Thermo Fisher) according to the manufacturer's instructions. Quantitative RT-PCR was performed using Fast SYBR™ Green Master Mix (Thermo Fisher) on a QuantStudio3 (Thermo Fisher) and a CFX96 Real Time System (Biorad, Hercules, CA) according to the manufacturer's protocol. ΔΔC T method was used to calculate the fold change. The following primer pairs were used to evaluate the gene expression of NGN2: F: 5’-CAGGCCAAAGTCACAGCAAC-3’ (SEQ ID NO: 151) R: 5’-CGATCCGAGCAGCACTAACA-3’ (SEQ ID NO: 152)
[0220] Lentiviral gRNA screening of TF-targeting libraries. To maintain a coverage of the library over 500-fold, 4.5×10 8Transduced cells were transduced with a lentiviral gRNA library in 25 mL of complete mTeSR at an MOI of 0.2 into 5 15 cm matrigel-coated dishes, ensuring that most cells contained 0 or 1 gRNA. Transduced cells were selected with 0.5 μg / mL puromycin for 3 days and then transferred to 4 15 cm matrigel-coated dishes at 10 K / cm². 2 The plants were passed through at a density of 5 × 10 6 Individual cell samples were taken and used as input controls for each replication. 24 hours after seeding, cells were further selected with puromycin for 2 days to ensure complete selection. Cells were differentiated into bone progenitor cells for 21 days as described in 2.4.3. At this point, the upper / lower 5th percentiles were collected and added to the unselected population. After selection, input, unselected, and GFP were collected. 高 , and GFP 低 The population was collected for genomic DNA purification (Qiagen).
[0221] gRNA library sequencing. gRNA libraries were amplified from each population by amplifying 12 μg of gRNA divided into 12 100 μL PCR reactions using Q5 hot-start polymerase (NEB, M0493L). The following PCR conditions were used: 60°C annealing temperature, 20'' extension time, 25 cycles. The following primers were used: F:5'AATGATACGGCGACCACCGAGATCTACACAATTTCTTGGGTAGTTTGCAGTT-3'(Sequence ID 153) R:5'-CAAGCAGAAGACGGCATACGAGAT(NNNNNN)GACTCGGTGCCACTTTTTCAA-3'(Sequence ID 154) (NNNNNN represents a 6bp barcode sequence).
[0222] The PCR amplification libraries were purified using Agencourt AMPure XP beads (Beckman Coulter) with double selection by first adding a 0.65× PCR volume, and then a 1× original PCR volume of beads, to remove large fragments and primer dimers. After resuspending in water, the library concentration of each sample was determined using the Qubit dsDNA High Sensitivity Kit (ThermoFisher). The samples were pooled and 21bp paired-end sequencing was performed on Illumina Miseq using the following reads and index primers: Lead 1: 5'-GATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCG-3' (Sequence ID 155) Lead 2:5'-GTTGATAACGGACTAGCCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC-3'(Sequence ID 156) Index: 5'-GCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTC-3' (Sequence ID 157)
[0223] Analysis of differential gRNA enrichment. Using Bowtie2 with options -p32 --end-to-end --very-sensitive -3 2 -I 0 -X 200, FASTQ files generated by MiSeq sequencing were aligned to a custom index. A summary table was then created for the number of reads for each gRNA in each sequencing population. Significant enrichment of each gRNA was assessed using the DESeq2 package in the R language. Unselected and GFP 高 Unselected and GFP 低 , and GFP 高 and GFP 低 We compared them; here, GFP 高 vs GFP 低 Only data on this topic is shown.
[0224] Validation of candidate TFs. Reporter hiPSCs were transduced with a lentivirus containing SOX9 cDNA, as described in 4.4.3, alongside a non-transduced control. After a 2-day recovery period, cells were differentiated according to the chondrogenesis protocol described in 2.4.2, but were collected at the sclerotid stage (D6). At this point, chondrogenic differentiation was evaluated using flow cytometry with an Accuri C6 cytometer. Identification of candidate regulators of hiPSC chondrogenesis. To evaluate the effect of activated TF on chondrogenic differentiation, we generated a cell line that stably expresses dCas9 fused to both the N-terminal and C-terminal VP64 transactivation domains (VP64-dCas9-VP64) in a COL2A1-2A-GFP background (Figure 19A). Transduced cells were selected to generate a polyclonal activator cell line. This polyclonal cell line robustly activated endogenous neurogenin 2 (NGN2) after transduction of gRNA targeting its promoter (Figure 19B).
[0225] To generate a TF-targeted CRISPR activation library, TF-targeted gRNAs were extracted from a previously described, publicly available, genome-scale activation library, similarly detailed in Example 9. The gRNA library was cloned into a lenti CRISPR construct with an mCherry-2a-Puro® expression cassette to allow selection of transduced cells (Figure 20A). Transduction of the lenti CRISPR library with a low MOI of infection into the activator / reporter strain ensured one gRNA per cell and maintained sufficient library coverage (over 500x). Transduced cells were then differentiated (Figure 20A). Transduction of the gRNA library appeared to eliminate the bimodal distribution of GFP on day 21; nevertheless, GFP 高 / 低 The population was selected (Figure 20B). Significant differential enrichment (adjusted p-value < 0.05) of 36 gRNAs was observed (Figure 20C).
[0226] In particular, two gRNAs that target SOX9 are GFP 高Significant enrichment was observed in the population. Strong enrichment was also observed for two gRNAs targeting SOX10, another transcription factor known to be involved in limb blastogenesis. The roles of SOX15 and TBR1 have not been investigated or defined. Interestingly, several further gRNAs were found to be GFP. 低 The population was enriched. As expected, gRNAs that target TFs strongly expressed in the pluripotent state, such as PRDM14 and NR5A2, were enriched in this population. However, other commonly cited pluripotent TFs, such as NANOG and OCT4, were not enriched in this population. Surprisingly, gRNAs that target TFs induced during chondrogenesis, such as PITX1, HES1, ID4, SP9, and SIX6, were also enriched in GFP. 低 The gRNAs were enriched in the population. gRNAs that were more than 3-fold enriched in any population but did not meet the significance criteria are colored blue (Figure 20C).
[0227] Preliminary validation of screening results using SOX9 overexpression. SOX9 is a known chondrogenic transcription factor that directly binds to the promoter and enhancer elements of genes encoding cartilage matrix proteins, but the effect of SOX9 activation in stepwise differentiation was unknown. Gene expression data from time-course experiments suggested that SOX9 activation occurs at D12 of this differentiation protocol. To determine the effect of SOX9 overexpression on chondrogenesis in the differentiation scheme, a lentivirus encoding SOX9 cDNA was transduced into reporter hiPSCs, and reporter fluorescence was evaluated after 6 days of differentiation (Figure 21A). At this stage, the cells are not exposed to chondrogenic growth factor BMP-4, but it would be useful to establish a protocol that avoids the need for redundant (6-15 day) prechondrogenic differentiation to monolayer. In fact, much of the variability observed in the chondrogenic differentiation protocol occurs at this differentiation stage. Approximately 2-3% of the entire population GFP after 6 days of differentiation induced by SOX9 overexpression and before BMP-4 treatment. 高The population was observed (Figure 21B). SOX9 transduction also appeared to broaden the distribution of reporter fluorescence to the left. The fluorescence intensity of this population generated by SOX9 overexpression was comparable to that of reporter cells on day 21 of differentiation, but the proportion of these cells was quite low (Figure 21C).
[0228] Discussion. Here, we show a high-throughput screening of all TFs for their ability to regulate chondrogenesis. GFP 高 SOX9, which was predicted to be enriched in the population, served as an internal control. Other factors known to be involved in chondrogenesis, such as SOX10, were also enriched in the GFP 高 population. SOX10 has been shown to be involved in limb bud chondrogenesis and, together with SOX9 and SOX8, may prepare the chondrogenic program and be involved in promoting hypertrophic differentiation of chondrocytes. The potential roles of TBR1 and SOX15 in chondrogenesis may not be as clear; SOX15 has been associated with muscle regeneration, and TBR1 is known to be expressed in glutamatergic neurons. The screening generated many more hits that were enriched in the GFP 低 population. Strong activation of most TFs could interfere with chondrogenic specification at various differentiation stages. The gRNA most significantly enriched in this population targets PRDM14, a regulator of naive pluripotency. gRNAs targeting NR5A2, which is similarly pluripotent and highly expressed, are also enriched in this population. Notably, gRNAs targeting TFs involved in chondrogenesis and activated during chondrogenesis, such as PITX1, are also enriched in the GFP 低 population.
[0229] In a validation experiment to test SOX9 overexpression in a differentiation context, after 6 days of differentiation and before the addition of BMP-4, the appearance of the GFP 高 population was observed, suggesting that exogenous delivery of the TF could avoid the pre-chondrogenic stage of differentiation. hiPSC-derived nodules appear to be primed to appropriately activate COL2A1 in response to SOX9. A detailed analysis of the histogram shown in Figure 21B revealed that overexpression of SOX9 led to an increase in the GFP 高In addition to generating SOX9, it appears to increase the height of the left tail portion of the histogram, suggesting that SOX9 overexpression may also inhibit chondrogenesis in a subset of cells.
[0230] In summary, we demonstrate the usefulness of a high-throughput hiPSC chondrogenesis platform using a COL2A1 knock-in reporter for screening chondrogenesis-promoting TFs. The screening successfully enriched gRNAs targeting the known chondrogenesis TF SOX9 and generated several other interesting hits. The TFs discovered herein may improve techniques for generating hiPSC-induced cartilage or for specifying various chondrocyte subtypes (e.g., cartilage vs. growth plate).
[0231] The foregoing descriptions of specific embodiments will fully illustrate the general nature of the invention that others can readily modify such specific embodiments and / or adapt such specific embodiments to various uses without diverting from the general concepts of the disclosure and without excessive experimentation, by applying knowledge within the scope of the skills possessed by those skilled in the art. Therefore, such adaptations and modifications are intended to be within the meaning and scope of equivalents of the disclosed embodiments, based on the teachings and guidance presented herein. Since the expressions and terms herein are for illustrative purposes only and not for limitation, it should be understood that the terms and terms herein should be interpreted by those skilled in the art in light of the teachings and guidance.
[0232] The scope and width of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined solely in accordance with the following claims and their equivalents. All publications, patents, patent applications, and / or other documents cited in this application are incorporated by reference in whole for all purposes to the same extent that each individual publication, patent, patent application, and / or other document is individually indicated as being incorporated by reference for all purposes. For the sake of completeness, various aspects of the present invention are shown in the following numbered sections:
[0233] Item 1. (1) A first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (2) A first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i)NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC 1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, Z NF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6 , PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEURO D1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KL F17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, E2F7;(iv)ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PB X2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (v) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IR F3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX2 2, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (vi) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, H A polynucleotide encoding a second neuron-specific transcription factor selected from ESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0234] Item 2. A system for increasing the expression of neuron-specific genes, comprising: (a) a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (b) a first gRNA that targets a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAG L2;(ii)PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX 2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii)RUNX3、PRDM1、KLF6、PAX2、RFX3、SOX10、GATA1、KLF5、KLF1、ERF、LHX6、PHOX2B、NANOG、NR5A2、ETV3、NEUROG3、SOX4、SOX9、PAX8、IRF5、CDX4、RARA、BHLHE40、SOX3、KLF4、NR5A1、IRF4、ASCL1、GATA6、SPIB、THRB、FOXH1、NEUROD1、SOX17、CDX2、ZEB2、RARG、INSM1、FOSL1、NEUROG1、SOX1、WT1、PAX5、SOX18、POU5F1、RFX4、KLF7、NKX2-2、OVOL2、FOXJ1、PRDM14、VENTX、LHX8、GFI1、KLF17、OVOL1、OLIG3、HMX3、ZNF521、ONECUT3、OVOL3、ZNF362、AFF1、HMX2、ZNF786、GATA5、TBX3、ZNF385A、ATOH1、PROP1、SOX11、JUN、FOXE3、FERD3L、E2F7;(iv)ZIC2、SPI1、GRHL2、TFAP2C、KLF8、MYB、TCF21、KLF12、TWIST1、SNAI1、RREB1、GCM2、GRHL1、ETS1、BARHL2、GRHL3、ELF3、PTF1A、GSX1、PBX2、NOTO、KLF3、ZNF311、ELMSAN1、ZNF296、PLEK、KMT2A、HES3;(v)HES2、SREBF1、CIC、WHSC1、VDR、HES1、ID2、TCF21、SNAI1、RREB1、GCM2、IRF3、FOXA1、GATA5、GRHL1、SOX5、DMRT1、GCM1、BARHL2、SOX13、ZEB1、PITX2、PTF1A、ZNF282、NPAS2、ZNF160、HES7、ZBED4、SALL4、GLIS3、TBX22、ZNF331、EGR4、ZIC5、ZNF710、ZNF697、ZFP36L2、ELMSAN1、ZNF296、ZNF318、ZNF570、ZNF683、ZFP36L1、HES4、ZNF777、HES5、ZIM2、ZNF579、BMP2、CRAMP1L、TOX3、FEZF2、HES3、ZNF791;(vi) A second gRNA targeting a second neuron-specific transcription factor selected from ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX; and a system comprising a Cas protein or a fusion protein, the fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein, or a TALE protein, and the second polypeptide domain has an activity selected from transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, nucleic acid association activity, methylase activity, and demethylase activity.
[0235] Item 3. The polynucleotide according to Item 1 or the system according to Item 2, wherein the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, NKX2-2, HES3, and ZFP36L1. Item 4. The polynucleotide or system according to Item 3, wherein the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, and NKX2-2. Item 5. The polynucleotide or system according to Item 3, wherein the second neuron-specific transcription factor is selected from HES3 and ZFP36L1.
[0236] Section 6. The second neuron-specific transcription factor is (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEU ROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MA Z, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NAN OG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THR B, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, O The system described in item 2, selected from VOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7, wherein the second polypeptide domain has transcriptional activating activity.
[0237] Section 7. Fusion protein VP64 dCas9 VP64 Or the system described in item 6, including dCas9-p300. Section 8. The second neuron-specific transcription factor is (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1 , RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4 , GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2 , ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12 , VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF 3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX, wherein the second polypeptide domain has transcriptional repressive activity, as described in item 2.
[0238] Item 9. The system described in Item 8, wherein the fusion protein contains dCas9-KRAB. Item 10. A system according to any one of items 2 to 9, wherein the first gRNA and the second gRNA each individually comprise a 12-22 base pair complementary polynucleotide sequence of the target DNA sequence, followed by a protospacer flanking motif, and optionally, the gRNA binds to and targets a polynucleotide comprising a sequence selected from SEQ ID NOs. 38-87, and / or comprises the same, and optionally, the first and / or second gRNA comprise a crRNA, tracrRNA, or a combination thereof.
[0239] Item 11. An isolated polynucleotide encoding the system described in any one of items 2 to 10. Item 12. A vector containing the isolated polynucleotide described in Item 11. Item 13. Cells containing isolated polynucleotides as described in Item 11 or vectors as described in Item 12.
[0240] Item 14. A method for increasing the maturation of stem cell-induced neurons, comprising: (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; In cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL 1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1 , SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P A method comprising the step of increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0241] Item 15. A method for increasing the maturation of stem cell-induced neurons, comprising the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, V DR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX 22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF6 83, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 A method comprising the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0242] Item 16. A method for increasing the conversion of stem cells to neurons, comprising: (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in stem cells; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; In cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL 1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1 , SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P A method comprising the step of increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0243] Item 17. A method for increasing the conversion of stem cells to neurons, comprising the steps of: increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in stem cells; (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, V DR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX 22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF6 83, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 A method comprising the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0244] Section 18. A method for treating a subject requiring treatment, comprising: (a) increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2 in the subject stem cells; or (b) increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the subject stem cells; In stem cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, Neurod1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3;(iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8, IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FOSL1, NEUROG1, SOX1, WT1, P A method comprising the step of increasing the level of a second neuron-specific transcription factor selected from AX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7.
[0245] Item 19. A method for treating a subject requiring treatment, comprising the steps of increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the target stem cells; and in the target stem cells, (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRHL1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SALL4, GLIS3, TBX 22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF570, ZNF6 83, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791;(iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1 , IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1 A method comprising the step of reducing the level of a second neuron-specific transcription factor selected from A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX.
[0246] 20. The method according to any one of items 14 to 19, wherein the step of increasing the level of a first neuron-specific transcription factor comprises at least one of the following: (a) administering a polynucleotide encoding the first neuron-specific transcription factor to stem cells; (b) administering a polypeptide comprising the first neuron-specific transcription factor to stem cells; and (c) administering a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the first neuron-specific transcription factor, or a TALE protein that targets the first neuron-specific transcription factor, and the second polypeptide domain has transcriptional activation activity, and if the first polypeptide domain comprises a Cas protein, further administering a gRNA that targets the first neuron-specific transcription factor to stem cells.
[0247] 21. The method according to any one of items 14, 16, and 18, wherein the step of increasing the level of a second neuron-specific transcription factor comprises at least one of the following: (a) administering a polynucleotide encoding a second neuron-specific transcription factor to stem cells; (b) administering a polypeptide containing a second neuron-specific transcription factor to stem cells; and (c) administering a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets a second neuron-specific transcription factor, or a TALE protein that targets a second neuron-specific transcription factor, and the second polypeptide domain has transcriptional activating activity, and if the first polypeptide domain contains a Cas protein, further administering a gRNA that targets a second neuron-specific transcription factor to stem cells.
[0248] The method according to any one of items 15, 17, and 19, wherein the step of reducing the level of a second neuron-specific transcription factor comprises administering to stem cells a fusion protein comprising two heterologous polypeptide domains, the first polypeptide domain comprising a Cas protein, a zinc finger protein that targets a second neuron-specific transcription factor, or a TALE protein that targets a second neuron-specific transcription factor, and the second polypeptide domain having transcriptional repressive activity, and further administering to stem cells a gRNA that targets a second neuron-specific transcription factor if the first polypeptide domain comprises a Cas protein. Item 23. The method according to any one of items 14-22, wherein stem cells are directly converted to neurons without a pluripotency stage.
[0249] Item 24. The cells described in Item 13 or the method described in any one of Items 14-23, wherein the stem cells are pluripotent stem cells, induced pluripotent stem cells, or embryonic stem cells. Item 25. A system for selecting polynucleotides for activity as cell type-specific transcription factors, comprising: a polynucleotide encoding a reporter protein and a cell type marker; a fusion protein containing two heterologous polypeptide domains, the first polypeptide domain containing a Cas protein and the second polypeptide domain having transcriptional activation activity; and a library of gRNAs, each of which guide RNAs (gRNAs) target different putative cell type-specific transcription factors.
[0250] Item 26. The system described in Item 25, wherein the cell type-specific transcription factor is a neuron-specific transcription factor, the cell type marker is a neuron marker, and the neuron marker includes TUBB3. Item 27. The system described in Item 25, wherein the cell type-specific transcription factor is a muscle-specific transcription factor, the cell type marker is a myogenic marker, and the myogenic marker contains PAX7. Item 28. The system described in Item 25, wherein the cell type-specific transcription factor is a chondrocyte-specific transcription factor, the cell type marker is a collagen marker, and the collagen marker includes COL2A1. Item 29. A system described in any one of items 25-28, wherein the reporter protein contains mCherry. Item 30. An isolated polynucleotide sequence encoding the system described in any one of items 25-29.
[0251] Item 31. A vector containing the isolated polynucleotide sequence described in Item 30. Cells containing the system described in any one of sections 25-29, the isolated polynucleotide sequence described in section 30, or the vector described in section 31, or a combination thereof. Item 33. A method for screening cell type-specific transcription factors, comprising: transducing a population of cells in the system described in any one of items 25-29 at an infection multiplicity (MOI) of about 0.2 such that the majority of cells each independently contain one gRNA and target one putative transcription factor; determining the expression level of a reporter protein in each cell; determining the gRNA level in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; and selecting a putative transcription factor as a cell type-specific transcription factor if the putative transcription factor corresponds to at least two gRNAs enriched in cells having high expression of the reporter protein.
[0252] Item 34. A method for screening cell type-specific transcription factor pairs, comprising the steps of: transducing a population of cells in any one of items 25-29 at an infection multiplicity (MOI) of about 0.2 such that the majority of cells each independently contain two gRNAs and target two putative transcription factors; determining the expression level of a reporter protein in each cell; determining the levels of two gRNAs in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; and selecting two putative transcription factors as a cell type-specific transcription factor pair if the putative transcription factors correspond to at least two gRNAs enriched in cells having high expression of the reporter protein. Item 35. The method according to item 33 or 34, wherein the expression level of the reporter protein in each cell is determined approximately 4 days after transduction.
[0253] Item 36. The method according to any one of items 33-35, wherein the expression level of the reporter protein in each cell is determined by flow cytometry. Item 37. The method according to any one of items 33-36, wherein the gRNA level in each cell having high expression of the reporter protein is determined by deep sequencing. Item 38. The method according to any one of items 33-37, wherein the gRNA increases the expression of the reporter protein in cells by approximately 2-50% compared to the untargeted gRNA. Item 39. A polynucleotide encoding a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1. Item 40. A system for increasing the expression of muscle-specific genes, comprising: (a) a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1; or (b) a zinc finger protein in which the first polypeptide domain targets a muscle-specific transcription factor selected from Cas protein, TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1, or selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 A system comprising a fusion protein containing a TALE protein that targets a selected muscle-specific transcription factor, wherein the second polypeptide domain comprises two heterologous polypeptide domains having an activity selected from transcriptional activation activity, transcription release factor activity, histone modification activity, nucleic acid association activity, methylase activity, and demethylase activity, and further comprising a gRNA that targets a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 if the first polypeptide domain contains a Cas protein.
[0254] Item 41. Fusion protein VP64 dCas9 VP64 Or the system described in item 40, including dCas9-p300.
[0255] Item 42. An isolated polynucleotide encoding the system described in any one of items 40-41. Item 43. A vector containing the isolated polynucleotide described in Item 42. Cells containing an isolated polynucleotide as described in item 42 or a vector as described in item 43. Item 45. A method for increasing the differentiation of stem cells into myoblasts, comprising the step of increasing the level of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in the stem cells.
[0256] Item 46. A method for treating a subject requiring treatment, comprising the step of increasing the level of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in the stem cells of the subject. The method according to item 47. The method according to item 45 or 46, wherein the step of increasing the level of muscle-specific transcription factors comprises at least one of the following: (a) administering a polynucleotide encoding a muscle-specific transcription factor to stem cells; (b) administering a polypeptide containing a muscle-specific transcription factor to stem cells; and (c) administering a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets a muscle-specific transcription factor, or a TALE protein that targets a muscle-specific transcription factor, and the second polypeptide domain has transcriptional activation activity, and if the first polypeptide domain contains a Cas protein, further administering a gRNA that targets a muscle-specific transcription factor.
[0257] array Sequence ID 1 NGG (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 2 NGA (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 3 NGAN (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 4 NGNG (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 5 NGGNG (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 6 NNAGAAW (W=A or T; N is any nucleotide residue, e.g., A, G, C, or T) Sequence ID 7 NAAR (R = A or G; N is any nucleotide residue, e.g., A, G, C, or T) Sequence ID 8 NNGRR (R = A or G; N is any nucleotide residue, e.g., A, G, C, or T) Sequence ID 9 NNGRRN (R = A or G; N is any nucleotide residue, e.g., A, G, C, or T) Sequence ID 10 NNGRRT (R = A or G; N is any nucleotide residue, e.g., A, G, C, or T)
[0258] Sequence ID 11 NNGRRV (R = A or G; N is any nucleotide residue, e.g., A, G, C, or T) Sequence ID 12 NNNNGATT (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 13 NNNNGNNN (where N is any nucleotide residue, such as A, G, C, or T) Sequence ID 14 Codon-optimized polynucleotide encoding Streptococcus pyogenes Cas9
[0259] JPEG2026058344000010.jpg251154 JPEG2026058344000011.jpg102156
[0260] Sequence ID 15 Amino acid sequence of codon-optimized polynucleotide encoding Streptococcus pyogenes Cas9
[0261] Sequence ID 16 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9 JPEG2026058344000012.jpg246156 JPEG2026058344000013.jpg31154
[0262] Sequence ID 17 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9 JPEG2026058344000014.jpg246156 JPEG2026058344000015.jpg33155
[0263] Sequence ID 18 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9 JPEG2026058344000016.jpg244155 JPEG2026058344000017.jpg30157
[0264] Sequence ID 19 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9
[0265] Sequence ID 20 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9 JPEG2026058344000018.jpg246156 JPEG2026058344000019.jpg36155
[0266] Sequence ID 21 Amino acid sequence of the codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9
[0267] Sequence ID 22 Polynucleotide sequence of the D10A mutant of Staphylococcus aureus Cas9 JPEG2026058344000020.jpg246155 JPEG2026058344000021.jpg29154
[0268] Sequence ID 23 Polynucleotide sequence of the N580A mutant of Staphylococcus aureus Cas9 JPEG2026058344000022.jpg245156 JPEG2026058344000023.jpg31155
[0269] Sequence ID 24 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9
[0270] Sequence ID 25 Codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9
[0271] Sequence ID 26 Amino acid sequence of the codon-optimized nucleic acid sequence encoding Staphylococcus aureus Cas9
[0272] Sequence ID 27 A vector (pDO242) encoding a codon-optimized nucleic acid sequence that codes for Staphylococcus aureus Cas9.
[0273] Sequence ID 28 mCherry polypeptide MVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGFKWERVMNFEDGGVVTVTQDSSLQDGEFIYK VKLRGTNFPSDGPVMQKKTMGWEASSERMYPEDGALKGEIKQRLKLKDGGHYDAEVKTTYKAKKPVQLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMDELYKPKKKRKVGGPKKKRKV
[0274] Sequence ID 29 mCherry polynucleotide atggtgagcaagggcgaggaggataacatggccatcatcaaggagttcatgcgcttcaaggtgcacatggagggctccgtgaacggccacgagttcgagatcgagggcgagggcgagggccgcccctacgagggcacccagaccgccaagctgaaggtgaccaagggtggccccctgcccttcgcctgggacatcctgtcccctcagttcatgtacggctccaaggcctacgtgaagcaccccgccgacatccccgactacttgaagctgtccttccccgagggcttcaagtgggagcgcgtgatgaacttcgaggacggcggcgtggtgaccgtgacccaggactcctccctgcaggacggcgagttcatctacaaggtgaagctgcgcggcaccaacttcccctccgacggccccgtaatgcagaagaagaccatgggctgggaggcctcctccgagcggatgtaccccgaggacggcgccctgaagggcgagatcaagcagaggctgaagctgaaggacggcggccactacgacgctgaggtcaagaccacctacaaggccaagaagcccgtgcagctgcccggcgcctacaacgtcaacatcaagttggacatcacctcccacaacgaggactacaccatcgtggaacagtacgaacgcgccgagggccgccactccaccggcggcatggacgagctgtacaagcccaagaagaagaggaaggtgggtggccctaagaaaaagagaaaggtgtga
[0275] SEQ ID NO: 30 Fwd: 5’-AATGATACGGCGACCACCGAGATCTACACAATTTCTTGGGTAGTTTGCAGTT SEQ ID NO: 31 Rev: 5’-CAAGCAGAAGACGGCATACGAGAT-(6-bp index sequence)-GACTCGGTGCCACTTTTTCAA SEQ ID NO: 32 Lead 1:5'-GATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACACCG
[0276] Sequence ID 33 Index: 5'-GCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTC Sequence ID 34 Lead 2:5'-GTTGATAACGGACTAGCCTTATTTAAACTTGCTATGCTGTTTCCAGCATAGCTCTTAAAC Sequence ID 35 tttn(N can be any nucleotide residue, e.g., A, G, C, or T)
[0277] Sequence ID 36 VP64-dCas9-VP64 protein
[0278] Sequence ID 37 VP64-dCas9-VP64 DNA
[0279] Sequence ID 159 Human p300 protein (with L553M mutation)
[0280] Sequence ID 160 Human p300 core effector protein (sequence number 134, aa1048~1664) IFKPEELRQALMPTLEALYRQDPESLPFRQPVDPQLLGIPDYFDIVKSPMDLSTIKRKLDTGQYQEPWQYVDDIWLMFNNAWLYNRKTSRVYKYCSKLSEVFEQEIDPVMQSLGYCCGRKLEFSPQTLCCYGKQLCTIPRDATYYSYQNRYHFC EKCFNEIQGESVSLGDDPSQPQTTINKEQFSKRKNDTLDPELFVECTECGRKMHQICVLHHEIIWPAGFVCDGCLKKSARTRKENKFSAKRLPSTRLGTFLENRVNDFLRRQNHPESGEVTVRVVHASDKTVEVKPGMKARFVDSGEMAESFPY RTKALFAFEEIDGVDLCFFGMHVQEYGSDCPPPNQRRVYISYLDSVHFFRPKCLRTAVYHEILIGYLEYVKKLGYTTGHIWACPPSEGDDYIFHCHPPDQKIPKPKRLQEWYKKMLDKAVSERIVHDYKDIFKQATEDRLTSAKELPYFEGDFW PNVLEESIKELEQEEEERKREENTSNESTDVTKGDSKNAKKKNNKKTSKNKSSLSRGNKKKPGMPNVSNDLSQKLYATMEKHKEVFFVIRLIAGPAANSLPPIVDPDPLIPCDLMDGRDAFLTLARDKHLEFSSLRRAQWSTMCMLVELHTQSQD
[0281] Sequence ID 158 Polynucleotide sequences of gRNA scaffolds gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgctttttt
Claims
1. (1) A first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (2) A first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii)RUNX3、PRDM1、KLF6、PAX2、RFX3、SOX10、GATA1、KLF5、KLF1、ERF、LHX6、PHOX2B、NANOG、NR5A2、ETV3、NEUROG3、SOX4、SOX9、PAX8、IRF5、CDX4、RARA、BHLHE40、SOX3、KLF4、NR5A1、IRF4、ASCL1、GATA6、SPIB、THRB、FOXH1、NEUROD1、SOX17、CDX2、ZEB2、RARG、INSM1、FOSL1、NEUROG1、SOX1、WT1、PAX5、SOX18、POU5F1、RFX4、KLF7、NKX2-2、OVOL2、FOXJ1、PRDM14、VENTX、LHX8、GFI1、KLF17、OVOL1、OLIG3、HMX3、ZNF521、ONECUT3、OVOL3、ZNF362、AFF1、HMX2、ZNF786、GATA5、TBX3、ZNF385A、ATOH1、PROP1、SOX11、JUN、FOXE3、FERD3L、E2F7; (iv)ZIC2、SPI1、GRHL2、TFAP2C、KLF8、MYB、TCF21、KLF12、TWIST1、SNAI1、RREB1、GCM2、GRHL1、ETS1、BARHL2、GRHL3、ELF3、PTF1A、GSX1、PBX2、NOTO、KLF3、ZNF311、ELMSAN1、ZNF296、PLEK、KMT2A、HES3; (v)HES2、SREBF1、CIC、WHSC1、VDR、HES1、ID2、TCF21、SNAI1、RREB1、GCM2、IRF3、FOXA1、GATA5、GRHL1、SOX5、DMRT1、GCM1、BARHL2、SOX13、ZEB1、PITX2、PTF1A、ZNF282、NPAS2、ZNF160、HES7、ZBED4、SALL4、GLIS3、TBX22、ZNF331、EGR4、ZIC5、ZNF710、ZNF697、ZFP36L2、ELMSAN1、ZNF296、ZNF318、ZNF570、ZNF683、ZFP36L1、HES4、ZNF777、HES5、ZIM2、ZNF579、BMP2、CRAMP1L、TOX3、FEZF2、HES3、ZNF791; (vi) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, H ESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36 L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX A second neuron-specific transcription factor selected from A polynucleotide that codes for something.
2. A system for increasing the expression of neuron-specific genes, (a) A first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; or (b) A first gRNA that targets a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof; and (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii)PRDM1、LHX6、NEUROG3、PAX8、SOX3、KLF4、FLI1、FOXH1、FEV、SOX17、FOS、INSM1、SOX2、WT1、SOX18、ZNF670、LHX8、OVOL1、E2F7、AFF1、HMX2、MAZ、RARA、PROP1、FOSL1、PAX5、KLF3; (iii)RUNX3、PRDM1、KLF6、PAX2、RFX3、SOX10、GATA1、KLF5、KLF1、ERF、LHX6、PHOX2B、NANOG、NR5A2、ETV3、NEUROG3、SOX4、SOX9、PAX8、IRF5、CDX4、RARA、BHLHE40、SOX3、KLF4、NR5A1、IRF4、ASCL1、GATA6、SPIB、THRB、FOXH1、NEUROD1、SOX17、CDX2、ZEB2、RARG、INSM1、FOSL1、NEUROG1、SOX1、WT1、PAX5、SOX18、POU5F1、RFX4、KLF7、NKX2-2、OVOL2、FOXJ1、PRDM14、VENTX、LHX8、GFI1、KLF17、OVOL1、OLIG3、HMX3、ZNF521、ONECUT3、OVOL3、ZNF362、AFF1、HMX2、ZNF786、GATA5、TBX3、ZNF385A、ATOH1、PROP1、SOX11、JUN、FOXE3、FERD3L、E2F7; (iv)ZIC2、SPI1、GRHL2、TFAP2C、KLF8、MYB、TCF21、KLF12、TWIST1、SNAI1、RREB1、GCM2、GRHL1、ETS1、BARHL2、GRHL3、ELF3、PTF1A、GSX1、PBX2、NOTO、KLF3、ZNF311、ELMSAN1、ZNF296、PLEK、KMT2A、HES3; (v) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRH L1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SA LL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF57 0, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (vi) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1, RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, H ESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP36 L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX A second gRNA that targets a second neuron-specific transcription factor selected from; and Cas protein or fusion protein Includes, The fusion protein comprises two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein, or a TALE protein, and the second polypeptide domain has an activity selected from transcriptional activation activity, transcriptional repression activity, transcription release factor activity, histone modification activity, nuclease activity, nucleic acid association activity, methylase activity, and demethylase activity.
3. The polynucleotide according to claim 1 or the system according to claim 2, wherein the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, NKX2-2, HES3, and ZFP36L1.
4. The polynucleotide or system according to claim 3, wherein the second neuron-specific transcription factor is selected from LHX8, LHX6, E2F7, RUNX3, FOXH1, SOX2, HMX2, and NKX2-2.
5. The polynucleotide or system according to claim 3, wherein the second neuron-specific transcription factor is selected from HES3 and ZFP36L1.
6. The second neuron-specific transcription factor described above is (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8 , IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FO SL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7 Selected from, The system according to claim 2, wherein the second polypeptide domain has transcriptional activation activity.
7. The aforementioned fusion protein VP64 dCas9 VP64 The system according to claim 6, or comprising dCas9-p300.
8. The second neuron-specific transcription factor described above is (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRH L1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SA LL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF57 0, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1 , RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP3 6L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX Selected from, The system according to claim 2, wherein the second polypeptide domain has transcriptional repressive activity.
9. The system according to claim 8, wherein the fusion protein comprises dCas9-KRAB.
10. The system according to any one of claims 2 to 9, wherein the first gRNA and the second gRNA each individually comprise a 12-22 base pair complementary polynucleotide sequence of the target DNA sequence, followed by a protospacer adjacent motif, and optionally, the gRNA binds to and targets a polynucleotide comprising a sequence selected from SEQ ID NOs. 38-97, and / or comprises the same, and optionally, the first and / or second gRNA comprises crRNA, tracrRNA, or a combination thereof.
11. An isolated polynucleotide encoding the system according to any one of claims 2 to 10.
12. A vector comprising an isolated polynucleotide as described in claim 11.
13. A cell comprising an isolated polynucleotide according to claim 11 or a vector according to claim 12.
14. A method for increasing the maturation of stem cell-induced neurons, (a) In the stem cells, increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2, or (b) The step of increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the stem cells; In the aforementioned stem cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8 , IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FO SL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7 The steps include increasing the level of a second neuron-specific transcription factor selected from and A method that includes this.
15. A method for increasing the maturation of stem cell-induced neurons, The steps include increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the stem cells; In the aforementioned stem cells, (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRH L1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SA LL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF57 0, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1 , RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP3 6L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX The steps include reducing the level of a second neuron-specific transcription factor selected from and A method that includes this.
16. A method for increasing the conversion of stem cells into neurons, (a) In the stem cells, increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2, or (b) The step of increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the stem cells; In the aforementioned stem cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8 , IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FO SL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7 The steps include increasing the level of a second neuron-specific transcription factor selected from and A method that includes this.
17. A method for increasing the conversion of stem cells into neurons, The steps include increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the stem cells; In the aforementioned stem cells, (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRH L1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SA LL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF57 0, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1 , RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP3 6L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX The steps include reducing the level of a second neuron-specific transcription factor selected from and A method that includes this.
18. A method of treating an object that requires treatment, (a) In the target stem cells, increasing the level of a first neuron-specific transcription factor selected from NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4, KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2, or (b) The step of increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the target stem cells; In the aforementioned target stem cells, (i) NEUROG3, SOX4, SOX9, KLF4, NR5A1, NEUROD1, SOX17, SMAD1, ATOH1, INSM1, NEUROG1, SOX18, RFX4 , KLF7, SP8, OVOL1, NEUROG2, ERF, PRDM1, OLIG3, HIC1, SOX3, FOXJ1, SOX10, KLF6, ASCL1, and PLAGL2; (ii) PRDM1, LHX6, NEUROG3, PAX8, SOX3, KLF4, FLI1, FOXH1, FEV, SOX17, FOS, INSM1, SOX2, WT1, SOX18, ZNF670, LHX8, OVOL1, E2F7, AFF1, HMX2, MAZ, RARA, PROP1, FOSL1, PAX5, KLF3; (iii) RUNX3, PRDM1, KLF6, PAX2, RFX3, SOX10, GATA1, KLF5, KLF1, ERF, LHX6, PHOX2B, NANOG, NR5A2, ETV3, NEUROG3, SOX4, SOX9, PAX8 , IRF5, CDX4, RARA, BHLHE40, SOX3, KLF4, NR5A1, IRF4, ASCL1, GATA6, SPIB, THRB, FOXH1, NEUROD1, SOX17, CDX2, ZEB2, RARG, INSM1, FO SL1, NEUROG1, SOX1, WT1, PAX5, SOX18, POU5F1, RFX4, KLF7, NKX2-2, OVOL2, FOXJ1, PRDM14, VENTX, LHX8, GFI1, KLF17, OVOL1, OLIG3, HMX3, ZNF521, ONECUT3, OVOL3, ZNF362, AFF1, HMX2, ZNF786, GATA5, TBX3, ZNF385A, ATOH1, PROP1, SOX11, JUN, FOXE3, FERD3L, and E2F7 The steps include increasing the level of a second neuron-specific transcription factor selected from and A method that includes this.
19. A method of treating an object that requires treatment, The steps include increasing the level of a first neuron-specific transcription factor selected from NGN3 and ASCL1, or a combination thereof, in the target stem cells; In the aforementioned target stem cells, (i) ZIC2, SPI1, GRHL2, TFAP2C, KLF8, MYB, TCF21, KLF12, TWIST1, SNAI1, RREB1, GCM2, GRHL1, ETS1, BARHL2, GRHL3, ELF3, PTF1A, GSX1, PBX2, NOTO, KLF3, ZNF311, ELMSAN1, ZNF296, PLEK, KMT2A, HES3; (ii) HES2, SREBF1, CIC, WHSC1, VDR, HES1, ID2, TCF21, SNAI1, RREB1, GCM2, IRF3, FOXA1, GATA5, GRH L1, SOX5, DMRT1, GCM1, BARHL2, SOX13, ZEB1, PITX2, PTF1A, ZNF282, NPAS2, ZNF160, HES7, ZBED4, SA LL4, GLIS3, TBX22, ZNF331, EGR4, ZIC5, ZNF710, ZNF697, ZFP36L2, ELMSAN1, ZNF296, ZNF318, ZNF57 0, ZNF683, ZFP36L1, HES4, ZNF777, HES5, ZIM2, ZNF579, BMP2, CRAMP1L, TOX3, FEZF2, HES3, ZNF791; (iii) ETV1, ZIC2, GSC2, CIC, GRHL2, REST, TFAP2C, SALL1, NFKB1, ELF2, HES1, MYB, KLF12, VSX2, NFE2, SNAI1, TRERF1 , RREB1, IRF1, IRF3, KLF2, MYOD1, SOX15, BARX1, GRHL1, SOX5, ETS1, SKIL, BARHL2, SOX13, ERG, GRHL3, ZNF281, ELF3, HESX1, KLF15, PITX2, PTF1A, GSX1, ZNF160, ETV5, MYBL1, NOTO, DPF1, MECOM, GLIS3, KLF3, TBX22, ESX1, ZNF337, ZFP3 6L2, ELMSAN1, ZNF618, ZNF296, ZNF318, ZNF570, ZNF497, ZFP36L1, HES5, BMP2, CRAMP1L, ZNF821, KMT2A, HES3, and BSX The steps include reducing the level of a second neuron-specific transcription factor selected from and A method that includes this.
20. The step of increasing the level of the first neuron-specific transcription factor is (a) Administering the stem cells a polynucleotide encoding the first neuron-specific transcription factor; (b) administering the polypeptide containing the first neuron-specific transcription factor to the stem cells; and (c) Administering a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the first neuron-specific transcription factor, or a TALE protein that targets the first neuron-specific transcription factor, and the second polypeptide domain having transcriptional activation activity, to the stem cells, and if the first polypeptide domain comprises a Cas protein, further administering a gRNA that targets the first neuron-specific transcription factor to the stem cells. The method according to any one of claims 14 to 19, comprising at least one of the following.
21. The step of increasing the level of the second neuron-specific transcription factor is (a) Administering the stem cells a polynucleotide encoding the second neuron-specific transcription factor; (b) administering the polypeptide containing the second neuron-specific transcription factor to the stem cells; and (c) A fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the second neuron-specific transcription factor, or a TALE protein that targets the second neuron-specific transcription factor, and the second polypeptide domain has transcriptional activation activity, is administered to the stem cells, and if the first polypeptide domain comprises a Cas protein, a gRNA that targets the second neuron-specific transcription factor is further administered to the stem cells. The method according to any one of claims 14, 16, and 18, comprising at least one of the following.
22. The method according to any one of claims 15, 17, and 19, wherein the step of reducing the level of the second neuron-specific transcription factor is administered to the stem cells a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the second neuron-specific transcription factor, or a TALE protein that targets the second neuron-specific transcription factor, and the second polypeptide domain has transcriptional repressive activity; and if the first polypeptide domain comprises a Cas protein, further administering to the stem cells a gRNA that targets the second neuron-specific transcription factor.
23. The method according to any one of claims 14 to 22, wherein the stem cells are directly converted into neurons without a pluripotency stage.
24. The cell according to claim 13 or the method according to any one of claims 14 to 23, wherein the stem cells are pluripotent stem cells, induced pluripotent stem cells, or embryonic stem cells.
25. A system for selecting polynucleotides for activity as cell type-specific transcription factors, Polynucleotides encoding reporter proteins and cell type markers; A fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein and the second polypeptide domain has transcriptional activation activity; and A library of gRNAs, each of which targets a different putative cell type-specific transcription factor. A system that includes this.
26. The system according to claim 25, wherein the cell type-specific transcription factor is a neuron-specific transcription factor, the cell type marker is a neuron marker, and the neuron marker includes TUBB3.
27. The system according to claim 25, wherein the cell type-specific transcription factor is a muscle-specific transcription factor, the cell type marker is a myogenic marker, and the myogenic marker includes PAX7.
28. The system according to claim 25, wherein the cell type-specific transcription factor is a chondrocyte-specific transcription factor, the cell type marker is a collagen marker, and the collagen marker includes COL2A1.
29. The system according to any one of claims 25 to 28, wherein the reporter protein comprises mCherry.
30. An isolated polynucleotide sequence encoding the system according to any one of claims 25 to 29.
31. A vector comprising the isolated polynucleotide sequence described in claim 30.
32. A cell comprising the system according to any one of claims 25 to 29, the isolated polynucleotide sequence according to claim 30, or the vector according to claim 31, or a combination thereof.
33. A method for screening cell type-specific transcription factors, Transduction into a population of cells in the system described in any one of claims 25 to 29, at an infection multiplicity (MOI) of about 0.2, such that the majority of the cells each independently contain one gRNA and target one putative transcription factor; The steps include determining the expression level of the reporter protein in each cell; A step of determining the gRNA level in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; If the putative transcription factor corresponds to at least two gRNAs enriched in the cell having high expression of the reporter protein, the step of selecting the putative transcription factor as a cell type-specific transcription factor is as follows: A method that includes this.
34. A method for screening cell type-specific transcription factor pairs, Transduction into a population of cells in the system described in any one of claims 25 to 29, at an infection multiplicity (MOI) of about 0.2, such that the majority of the cells each independently contain two gRNAs and target two putative transcription factors; The steps include determining the expression level of the reporter protein in each cell; A step of determining the two gRNA levels in each cell having high expression of the reporter protein, wherein high expression of the reporter protein is defined as being in the top 5% of the cell population; If the putative transcription factors correspond to at least two gRNAs enriched in the cells having high expression of the reporter protein, the step is to select the two putative transcription factors as a pair of cell type-specific transcription factors. A method that includes this.
35. The method according to claim 33 or 34, wherein the expression level of the reporter protein in each cell is determined approximately four days after transduction.
36. The method according to any one of claims 33 to 35, wherein the expression level of the reporter protein in each cell is determined by flow cytometry.
37. The method according to any one of claims 33 to 36, wherein the gRNA level in each cell having high expression of the reporter protein is determined by deep sequencing.
38. The method according to any one of claims 33 to 37, wherein the gRNA increases the expression of the reporter protein in the cell by about 2 to 50% compared to the untargeted gRNA.
39. A polynucleotide encoding a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1.
40. A system for increasing the expression of muscle-specific genes, (a) a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1; or (b) A fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a zinc finger protein that targets a muscle-specific transcription factor selected from Cas protein, TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1, or a TALE protein that targets a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1, and the second polypeptide domain is a fusion protein having an activity selected from transcriptional activation activity, transcription release factor activity, histone modification activity, nucleic acid association activity, methylase activity, and demethylase activity. A system comprising, if the first polypeptide domain contains a Cas protein, further comprising a gRNA that targets a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1.
41. The aforementioned fusion protein VP64 dCas9 VP64 The system according to claim 40, or comprising dCas9-p300.
42. An isolated polynucleotide encoding the system according to any one of claims 40 to 41.
43. A vector comprising an isolated polynucleotide as described in claim 42.
44. A cell comprising the isolated polynucleotide described in claim 42 or the vector described in claim 43.
45. A method for increasing the differentiation of stem cells into myoblasts, The step of increasing the level of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in the aforementioned stem cells. A method that includes this.
46. A method of treating an object that requires treatment, The step of increasing the level of a muscle-specific transcription factor selected from TWIST1, PAX3, MYOD, MYOG, SOX9, SOX10, and DMRT1 in the target stem cells. A method that includes this.
47. The step of increasing the level of the muscle-specific transcription factor is, (a) Administering the polynucleotide encoding the muscle-specific transcription factor to the stem cells; (b) administering the polypeptide containing the muscle-specific transcription factor to the stem cells; and (c) Administering a fusion protein comprising two heterologous polypeptide domains, wherein the first polypeptide domain comprises a Cas protein, a zinc finger protein that targets the muscle-specific transcription factor, or a TALE protein that targets the muscle-specific transcription factor, and the second polypeptide domain has transcriptional activation activity, to the stem cells, and if the first polypeptide domain comprises a Cas protein, further administering a gRNA that targets the muscle-specific transcription factor. The method according to claim 45 or 46, comprising at least one of the following.