Cell Reprogramming

A predictive framework using gene expression and network data identifies transcription factors for transdifferentiation, enhancing cell conversion efficiency and reducing immune rejection and tumorigenicity in therapeutic applications.

JP7806009B2Active Publication Date: 2026-01-26MOGRIFY LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023215034
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2015-12-23
Filing Date
2023-12-20
Publication Date
2026-01-26
Estimated Expiration
2036-12-23

AI Technical Summary

Technical Problem

Current methods for transdifferentiation of cells are inefficient and lack a systematic approach to identify necessary factors for converting one cell type to another, and there is a need for immunologically matched cells for therapeutic applications that avoid immune rejection and tumorigenicity.

Method used

A predictive framework combining gene expression data with regulatory network information to identify transcription factors required for transdifferentiation, using methods to determine and rank TFs based on differential gene expression and network scores, and optionally removing redundant TFs.

Benefits of technology

Accurately predicts transcription factors for transdifferentiation, achieving high conversion efficiency to target cell characteristics, with cells showing low tumorigenicity and reduced immune rejection risk.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007806009000044
    Figure 0007806009000044
  • Figure 0007806009000045
    Figure 0007806009000045
  • Figure 0007806009000046
    Figure 0007806009000046
Patent Text Reader

Abstract

To provide methods and compositions for converting one cell type to another cell type.SOLUTION: A method for determining transcription factors required for conversion of a source cell to a cell exhibiting at least one characteristic of a target cell type comprises the steps of: determining differential expression of genes in the source and target cell types; determining a network score for each transcription factor (TF) in each of the source and target cell types based on the differential gene expression over at least one network, where the network contains information on interactions that affect gene expression; ranking the TFs based on an informational combination of the differential gene expression and the network scores, thereby identifying the set of transcription factors for the conversion from the source cell to the cell exhibiting the at least one characteristic of the target cell type.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application claims priority from Australian Provisional Patent Application No. 2015905349, the entire disclosure of which is incorporated herein in its entirety.

[0002] The present invention relates to methods and compositions for converting one cell type to another, particularly to the transdifferentiation of cells into different cell types. [Background technology]

[0003] Cell-based regenerative therapies require the generation of specific cell types to replenish tissues damaged by injury, disease, or age. Embryonic stem cells (ESCs) have the potential to differentiate into all cell types from the human body and are therefore being extensively investigated as a source of replacement therapy. However, because ESCs are established from cultured blastocysts, they cannot be derived patient-specifically. Therefore, immune rejection and ethical issues remain major barriers preventing the translation of ESC technology, particularly human ESC technology, into clinical applications.

[0004] Cell replacement therapy has the potential to rapidly generate a variety of therapeutically important cell types directly from easily accessible tissues, such as the skin or blood. Such immunologically matched cells would also have a lower risk of rejection after transplantation. Furthermore, because these cells are well differentiated, they would exhibit lower tumorigenicity.

[0005] Transdifferentiation, the process of converting one cell type into another without going through a pluripotent state, may hold great promise in regenerative medicine but has not yet been applied with high confidence. While it may be possible to change the phenotype of one somatic cell type to another, the elements required for conversion are difficult to identify and are often unknown. Identifying factors that directly reprogram cell type identity is currently limited, among other things, by the expense of comprehensive experimental testing of a set of relevant factors, an approach that is both inefficient and unmasterable. Summary of the Invention [Problem to be solved by the invention]

[0006] There is a need for new and / or improved methods to identify the factors necessary to convert one cell type to another. There is also a need for cells and cell populations that can be used for therapeutic applications.

[0007] The reference in the specification to any prior art is not an admission or suggestion that this prior art forms part of the common general knowledge in any jurisdiction or that this prior art could reasonably be expected to be understood, regarded as relevant, and / or combined with other prior art by a person skilled in the art. [Means for solving the problem]

[0008] The present invention relates to a predictive framework that combines gene expression data with regulatory network information to predict the reprogramming factors required to induce cell transformation (i.e., transforming a source cell into a cell exhibiting characteristics of a target cell type). This framework accurately predicts transcription factors used in known transdifferentiation, as well as experimentally validated previously unknown transcription factors involved in transdifferentiation. The present invention also relates to methods and compositions for the direct reprogramming of source cells into cells with characteristics of a target cell type (i.e., transdifferentiation or cell reprogramming).

[0009] The present invention provides a method for determining transcription factors required to convert a source cell into a cell having at least one characteristic of a target cell type, the method comprising: - determining differential expression of genes in the source cell type and the target cell type; - determining a network score for each transcription factor (TF) in each of the source cell type and the target cell type across at least one network based on the differential gene expression, the network containing information on interactions that affect gene expression; - ranking the TFs based on a combination of information from the network scores and differential gene expression, thereby identifying a set of transcription factors for converting the source cell into a cell having at least one characteristic of the target cell type.

[0010] The present invention provides a method for determining transcription factors required to convert a source cell into a cell having at least one characteristic of a target cell type, the method comprising: - determining a gene score for each differentially expressed gene in the source cell type and the target cell type; - determining a network score for each transcription factor (TF) in each of the source cell type and the target cell type by performing a weighted sum of each gene score across at least one network, the network containing information on interactions that affect gene expression; - ranking TFs based on a combination of gene scores and network scores; - based on a comparison of the ranked lists for each cell type, identifying a set of transcription factors for conversion of the source cell into a cell having at least one characteristic of the target cell type.

[0011] Preferably, the gene score is a combination of the log fold change of differential expression and the adjusted P-value. The gene score can be calculated based on tree-based methods or Bayesian clustering.

[0012] Preferably, the network contains information on protein-DNA interactions, protein-DNA, and protein-RNA interactions. Typically, the network contains information on interactions between transcription factors and regulatory regions of genes. Typically, the regulatory regions are promoter regions of genes.

[0013] Preferably, the method further comprises the step of collecting expression data for each gene before determining the gene score.

[0014] Preferably, the method further comprises the step of removing transcriptionally redundant TFs from the ranked list from each cell type.

[0015] The present invention provides a method for determining transcription factors required to convert a source cell into a cell having at least one characteristic of a target cell type, the method comprising: - collecting expression data for each gene in the source cell type and the target cell type; - calculating tree-based differential expression versus background for each gene in each sample, and then combining the log fold change and adjusted P value with the gene score; - calculating a network score for each TF by performing a weighted sum of gene scores across at least one sub-network centered on each TF; - ranking the TFs based on a combination of gene scores and network scores. Tep and; - calculating a set of transcription factors for conversion between any two cell types based on a comparison of the ranked lists from each cell type; and optionally - removing transcriptionally redundant TFs from the list, Thereby, the transcription factors required for conversion of the source cell type to the target cell type are determined.

[0016] The present invention provides a method for determining transcription factors required to convert a source cell into a cell having at least one characteristic of a target cell type, the method comprising: - collecting expression data for each gene (x) in each sample (s); - calculating the tree-based differential expression relative to background for each gene in each sample, followed by the log fold change

number

number

number

number

number

[0017] Preferably, the identified set of transcription factors affects the expression of at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% of the genes expressed in the target cell type.

[0018] The source and target cell types can be any cell type listed in the FANTOM5 dataset or any cell type listed herein, including in Table 4. obtain.

[0019] Typically, the sub-network is generated from gene expression data to which MARA has been applied, or from the STRING database (referred to herein as

number

[0020] Preferably, the method further comprises forming a cell transformation landscape by arranging cell types on a 2D surface based on their required TFs and adding height based on the average coverage of required genes directly regulated by the selected TFs.

[0021] Preferably, any of the methods described herein further comprise the step of forming a cell transformation landscape by arranging cell types on a 2D surface based on their required TFs and adding height based on the average coverage of required genes directly regulated by the selected TFs.

[0022] In any of the methods of the invention described above, the method further comprises increasing the amount of a transcription factor in the source cell type determined to be necessary for conversion of the source cell type to the target cell type.

[0023] The present invention provides a method for identifying an agent useful for promoting conversion of a source cell type to a target cell type, the method comprising: - determining one or more transcription factors required for conversion of a source cell type into a target cell type by any method described herein; - screening one or more candidate agents for their ability to increase the amount of one or more transcription factors required for conversion of a source cell type to a target cell type; Including, Agents that increase the amount of one or more transcription factors are useful agents for promoting the conversion of a source cell type to a target cell type.

[0024] Preferably, the candidate agent can be any compound one desires to test, including, but not limited to, a protein (such as an antibody or fragment thereof or an antibody mimetic), a peptide, a nucleic acid (including RNA, DNA, antisense oligonucleotides, peptide nucleic acids), a carbohydrate, an organic compound, a small molecule, a natural product, a library extract, a biological fluid, etc. The candidate compound can also be part of a library, e.g., a collection of compounds including alterations or modifications.

[0025] The present invention provides a method of reprogramming a source cell, the method comprising increasing protein expression of one or a transcription factor, or a variant thereof, in a source cell, wherein the source cell is reprogrammed to have at least one characteristic of a target cell: the source cells are selected from the group consisting of skin fibroblasts, epidermal keratinocytes, embryonic stem cells, pluripotent stem cells, mesenchymal stem cells, monocytes or cardiac fibroblasts; the target cells are selected from the group consisting of chondrocytes, hair follicles, CD4+ T cells, CD8+ T cells, NK-cells, hemopoietic stem cells (HSCs), mesenchymal stem cells (MSCs), adipose MSCs, bone marrow MSCs, oligodendrocytes, oligodendrocyte precursors, skeletal muscle cells, smooth muscle cells, fetal cardiomyocytes, epithelial cells, endothelial cells, keratinocytes and astrocytes; - the transcription factor is one or more of those listed in Table 4.

[0026] The present invention provides a method for generating cells having at least one characteristic of a target cell from a source cell, the method comprising: -increasing the amount of one or more transcription factors, or variants thereof, in the source cell; - culturing the source cells for a time and under conditions sufficient to differentiate them into target cells; thereby generating cells from the source cells having at least one characteristic of the target cells; the source cells are selected from the group consisting of skin fibroblasts, epidermal keratinocytes, embryonic stem cells, pluripotent stem cells, mesenchymal stem cells, monocytes or cardiac fibroblasts; the target cells are selected from the group consisting of chondrocytes, hair follicles, CD4+ T cells, CD8+ T cells, NK (natural killer) cells, hemopoietic stem cells (HSCs), mesenchymal stem cells (MSCs), adipose MSCs, bone marrow MSCs, oligodendrocytes, oligodendrocyte precursors, skeletal muscle cells, smooth muscle cells, fetal cardiomyocytes, epithelial cells, endothelial cells, keratinocytes and astrocytes; - the transcription factor is one or more of those listed in Table 4.

[0027] The present invention also provides a method of reprogramming a source cell listed in Table 4, the method comprising increasing protein expression of a transcription factor of Table 4 or a variant thereof in a source cell, wherein the source is reprogrammed to have at least one characteristic of a target cell.

[0028] The present invention provides a method for reprogramming a source cell into a cell having at least one characteristic of a target cell, the method comprising: i) providing a source cell, or a cell population comprising source cells; ii) transfecting the source cell with one or more nucleic acids comprising nucleotide sequences encoding one or more transcription factors; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of a target cell: the source cells are selected from the group consisting of skin fibroblasts, epidermal keratinocytes, embryonic stem cells, pluripotent stem cells, mesenchymal stem cells, monocytes or cardiac fibroblasts; the target cells are selected from the group consisting of chondrocytes, hair follicles, CD4+ T cells, CD8+ T cells, NK-cells, hemopoietic stem cells (HSCs), mesenchymal stem cells (MSCs), adipose MSCs, bone marrow MSCs, oligodendrocytes, oligodendrocyte precursors, skeletal muscle cells, smooth muscle cells, fetal cardiomyocytes, epithelial cells, endothelial cells, keratinocytes and astrocytes; - the transcription factor is one or more of those listed in Table 4.

[0029] In any of the methods of the invention described herein, the source cells are fibroblasts; (a) the target cell is a chondrocyte, and the transcription factor is any one or more of BARX1, PITX1, SMAD6, FOXC1, SIX2, and AHR; (b) the target cell is a hair follicle and the transcription factor is any one or more of ZIC1, PRRX2, RARB, VDR, FOXD1, and CREB3; (c) the target cell is a CD4+ T cell and the transcription factor is any one or more of RORA, LEF1, JUN, FOS, and BACH2; (d) the target cell is a CD8+ T cell and the transcription factor is any one or more of RORA, FOS, SMAD7, JUN, and RUNX3; (e) the target cell is a NK cell and the transcription factor is any one or more of RORA, SMAD7, FOS, JUN, and NFATC2; (f) the target cell is an HSC and the transcription factor is any one or more of MYB, GATA1, GFI1, and GFI1B; (g) the target cells are adipose MSCs and the transcription factors are any one or more of NOTCH3, HIC1, ID1, ESRRA, IR1, SIX5, SREBF1, and SNAI2; (h) the target cells are bone marrow MSCs and the transcription factors are any one or more of SIX1, ID1, HOXA7, FOXC2, HOXA9, MAFB, and IRX5; (i) the target cells are oligodendrocyte precursors and the transcription factor is any one or more of NKX2-1, ANKRD1, FOXA2, CDH1, ZFP42, IGF1, ICAM1, and FOS; (j) the target cells are skeletal muscle cells and the transcription factors are MYOG, HIC1 MYOD1, FOXD1, PITX3, SIX2, HOXA7, and JUNB; (k) the target cell is a smooth muscle cell and the transcription factor is any one or more of GATA6, LIF, JUNB, CREB3, MEIS1, and PBX1; (l) the target cells are fetal cardiomyocytes and the transcription factors are any one or more of BMP10, GATA6, TBX5, FHL2, NKX2-5, HAND2, GATA4, and PPARGC1A; (m) the target cell is an astrocyte and the transcription factor is any one or more of SOX2, SOX9, ARNT2, E2F5, PBX1, SMAD1, and RUNX2; (n) the target cell is an epithelial cell and the transcription factor is any one or more of FOS, DBP, HES1, FOXA2, ESRRA, CDH1, FOXQ1, and PAX6; (o) the target cell is an endothelial cell and the transcription factor is any one or more of SOX17, SMAD1, TAL1, IRF1, TCF7L1, MXD4, and JUNB; or (p) The target cells are keratinocytes and the transcription factor is any one or more of FOXQ1, SOX9, MAFB, CDH1, FOS, and REL.

[0030] All of the transcription factors described in (a) to (p) immediately above can be used. Preferably, the fibroblasts are skin fibroblasts.

[0031] In any of the methods of the invention described herein, the source cells are keratinocytes; (a) the target cells are chondrocytes and the transcription factors are any one or more of BARX1, PITX1, SMAD6, TGFB3, FOXC1, and SIX2; (b) the target cell is a hair follicle and the transcription factor is any one or more of RUNX1T1, ZIC1, PRRX1, MSX1, EBF1, FOXD1, and RUNX2; (c) the target cell is a CD4+ T cell and the transcription factor is any one or more of RORA, LEF1, JUN, FOS, and NR3C1; (d) the target cell is a CD8+ T cell and the transcription factor is any one or more of RORA, FOS, SMAD7, JUN, and RUNX3; (e) the target cell is a NK cell and the transcription factor is any one or more of RORA, SMAD7, FOS, JUN, NFATC2, and RUNX3; (f) the target cell is an HSC and the transcription factor is any one or more of MYB, GATA1, GFI1, and GFI1B; (g) the target cells are adipose MSCs and the transcription factors are any one or more of TWIST1, HIC1, ID1, MSX1, IRF1, HOXB7, SNAI2, and E2F1; (h) the target cells are bone marrow MSCs and the transcription factors are any one or more of SIX1, TWSIT1, ID1, HMOX1, FOXC2, and HOXA7; (i) the target cells are oligodendrocyte precursor cells and the transcription factor is any one or more of NKX2-1, ANKRD1, ZFP42, FOS, IGF1, ICAM1, FOXA2, and CDH1; (j) the target cell is a skeletal muscle cell and the transcription factor is any one or more of MYOG, MYOD1, RF1, PITX3, HOXA7, FOXD1, and SOX8; (k) the target cell is a smooth muscle cell and the transcription factor is any one or more of IRF1, GATA6, LIF, and MEIS1; (l) the target cell is an endothelial cell and the transcription factor is any one or more of SOX17, TAL1, SMAD1, IRF1, and TCF7L1; or (m) The target cell is an epithelial cell and the transcription factor is any one or more of NOTCH1, HR, DBP, OTX1, ESRRA, FOXQ1, PAX6, and IRX5.

[0032] In the above (a) to (m), all of the transcription factors can be used. Preferably, the keratinocytes are epithelial keratinocytes. More preferably, the keratinocytes are oral mucosal keratinocytes. More preferably, when the source cells are mucosal keratinocytes, the target cells are corneal epithelial cells.

[0033] In any of the methods of the invention described herein, the source cells are embryonic stem cells; (a) the target cell is a chondrocyte, and the transcription factor is any one or more of BARX1, PITX1, SMAD6, and NFKB1; (b) the target cell is a hair follicle and the transcription factor is any one or more of TWIST1, ZIC1, NR2F2, PRRX1, NFKB1, and AHR; (c) the target cell is a CD4+ T cell and the transcription factor is any one or more of RORA, LEF1, JUN, FOS, and BACH2; (d) the target cell is a CD8+ T cell and the transcription factor is any one or more of RORA, FOS, SMAD7, and JUN; (e) the target cell is a NK cell and the transcription factor is any one or more of RORA, SMAD7, FOS, JUN, and NFATC2; (f) the target cell is an HSC and the transcription factor is any one or more of MYB, IL1B, KLF1, GATA1, GFI1, GFI1B, and NFE2; (g) the target cells are adipose MSCs and the transcription factor is any one or more of TWIST1, SNAI2, IRF1, MXD4, NFKB1, MSX1, HOXB7, and ESRRA; (h) the target cells are bone marrow MSCs and the transcription factors are any one or more of IRF1, RUNX1, CEBPB, AHR, FOXC2, and HOXA9; (i) the target cells are oligodendrocyte precursor cells and the transcription factor is any one or more of NKX2-1, ANKRD1, FOXA2, LMO3, FOS, IGF1, ICAM1, and CDH1; (j) the target cell is a skeletal muscle cell and the transcription factor is any one or more of MYOG, IRF1, MYOD1, FOXD1, NFKB1, JUNB, and HOXA7; (k) the target cell is a smooth muscle cell and the transcription factor is any one or more of IRF1, NFKB1, JUNB, FOSL2, GATA6, and MEIS1; (l) the target cell is an endothelial cell and the transcription factor is any one or more of SOX17, TAL1, SMAD1, HOXB7, JUNB.IRF1, and NFKB1; (m) the target cell is an astrocyte and the transcription factor is any one or more of IRF1, SOX9, ARNT2, PAX6, SNAI2, SOX5, and RUNX2; (n) the target cell is a keratinocyte and the transcription factor is any one or more of SOX9, NFKB1, MYC, NR2F2, AHR, FOSL1, and FOSL2; or (o) The target cell is an epithelial cell, and the transcription factor is any one or more of MYC, IL1B, FOS, NFKB1, ESRRA, FOXQ1, IRF1, and PAX6.

[0034] All of the transcription factors listed in (a) to (o) immediately above can be used. Preferably, the embryonic stem cells are human embryonic stem cells.

[0035] In any of the methods of the invention described herein, the source cells are monocytic cells, the target cells are HSCs, and the transcription factor is any one or more of MYB, IL1B, GATA1, GFI1, and GFI1B. Preferably, all of the listed transcription factors can be used. Cut.

[0036] In any of the methods of the invention described herein, the source cells are cardiac fibroblasts, the target cells are fetal cardiomyocytes, and the transcription factor is any one or more of BMP10, GATA6, TBX5, ANKRD1, HAND1, PPARGC1A, NKX2-5, and GATA4. Preferably, all of the listed transcription factors can be used.

[0037] In any of the methods of the present invention described herein, the source cells are pluripotent cells, the target cells are endothelial cells, and the transcription factors are any one or more of SOX17, TAL1, HOXB7, NFKB1, IRF1, JUNB, and SMAD1. Preferably, all of the listed transcription factors can be used. Preferably, the pluripotent cells are induced pluripotent stem cells (iPSCs).

[0038] In any of the methods of the present invention described herein, the source cells are pluripotent cells, the target cells are astrocytes, and the transcription factors are any one or more of PAX6, POU3F2, SNAI2, RUNX2, SOX5, E2F5, and HMGB2. Preferably, all of the listed transcription factors can be used. Preferably, the pluripotent cells are induced pluripotent stem cells (iPSCs).

[0039] In any of the methods of the invention described herein, the source cells are pluripotent cells, the target cells are keratinocytes, and the transcription factors are any one or more of TP63, TFAP2A, MYC, NFKBIA, SOX9, and NFKB1. Preferably, all of the listed transcription factors can be used. Preferably, the pluripotent cells are induced pluripotent stem cells (iPSCs).

[0040] In any of the methods of the invention described herein, the source cells are bone marrow stem cells, the target cells are astrocytes, and the transcription factors are any one or more of SOX2, SOX9, ARNT2, MYBL2, POU3F2, E2F1, and HMGB2. Preferably, all of the listed transcription factors can be used.

[0041] Preferably, the at least one characteristic of the target cell is the upregulation of any one or more target cell markers and / or a change in cell morphology. Relevant markers are described herein and known in the art. Exemplary markers of target cells include: - Chondrocytes: CD49, CD10, CD9, CD95, integrin α10β1,105, and production of sulfated glycosaminoglycans (GAGs); - Hair follicles: CD200, PHLDA1 and follistatin; -CD4+T- cells: CD3, CD4; -CD8+T- cells: CD3, CD8; -NK- cells: CD56, CD2; -HSC: CD45, CD19 / 20, CD14 / 15, CD34, CD90; - Adipose MSCs: CD13, CD29, CD90, CD105, CD10, CD45 and differentiation towards osteoblasts, adipocytes and chondrocytes in vitro; - MSCs from bone marrow: CD13, CD29, CD90, CD105, CD10, and differentiation towards osteoblasts, adipocytes and chondrocytes in vitro; -oligodendrocytes and oligodendrocyte precursors; QPCR for NG2 and PDGFRα, Olig2 and Nkx2.2; -Skeletal muscle cells: MyoD, myogenin and desmin; -Smooth muscle cells: myocardin, smooth muscle alpha actin and smooth muscle myosin heavy chain; -Fetal cardiomyocytes: MEF2C, MYH6, ACTN1, CDH2 and GJA1; -Endothelial cells: PeCAM (CD31), VE-cadherin and VEGFR2; -Keratinocytes: keratin 1, keratin 14, pankeratin and involucrin; Astrocytes: GFAP, S100B and ALDH1L1; and -Epithelial cells: cytokeratin 15 (CK15), cytokeratin 3 (CK3), involucrin and connexin 4 Includes.

[0042] Typically, suitable conditions for target cell differentiation include culturing the cells for a sufficient time in a suitable medium. The sufficient time for culturing may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. The suitable medium may be one set forth in Table 9.

[0043] The present invention also provides cells having at least one characteristic of a target cell produced by the methods described herein.

[0044] In any of the methods described herein, the method can further include enriching cells having at least one characteristic of the target cell type to increase the proportion of cells in the population having at least one characteristic of the target cell type. The enrichment step can be performed in culture for a time and under conditions sufficient to produce a population of cells as described below.

[0045] In any of the methods described herein, the method may further comprise administering to the individual a cell, or a cell population comprising cells, having at least one characteristic of the target cell type.

[0046] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of a target cell, the cells being generated by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of a target cell.

[0047] The present invention also relates to kits for producing cells having at least one characteristic of a target cell disclosed herein. In some embodiments, the kit comprises one or more nucleic acids having one or more nucleic acid sequences encoding a transcription factor or variants thereof described herein. Preferably, the kit can be used to produce cells having at least one characteristic of a target cell recited in Table 4. Preferably, the kit can be used with a source cell recited in Table 4. In some embodiments, the kit further comprises instructions for reprogramming a source cell into a cell having at least one characteristic of a target cell according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0048] The present invention relates to a composition comprising at least one target cell and at least one agent that increases protein expression of one or more transcription factors in the target cell. Preferably, the target cell is one of those described herein. Furthermore, the transcription factor is any one of those described herein. Preferably, the target cell and transcription factor are those described in Table 4.

[0049] The present invention provides a method for reprogramming fibroblasts, the method comprising increasing protein expression of any one or more of FOXQ1, SOX9, MAFB, CDH1, FOS and REL, or variants thereof, in fibroblasts, wherein the fibroblasts are reprogrammed to produce fewer keratinocytes. It has been reprogrammed to have at least one characteristic.

[0050] The present invention provides a method for producing cells having at least one characteristic of keratinocytes from fibroblasts, the method comprising: - Increasing the amount of any one or more of FOXQ1, SOX9, MAFB, CDH1, FOS and REL, or variants thereof, in fibroblasts; - culturing the fibroblasts for a time and under conditions sufficient for keratinocyte differentiation; thereby producing cells from the fibroblasts having at least one characteristic of keratinocytes.

[0051] The present invention provides a method for reprogramming fibroblasts into cells having at least one characteristic of a keratinocyte, the method comprising: i) providing a fibroblast, or a cell population comprising a fibroblast; ii) transfecting the fibroblast with one or more nucleic acids comprising a nucleotide sequence encoding the polypeptides FOXQ1, SOX9, MAFB, CDH1, FOS, and REL; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of a keratinocyte.

[0052] Preferably, the at least one characteristic of a keratinocyte is an upregulation of any one or more keratinocyte markers, including keratin 1, keratin 14, and involucrin, and / or a change in cell morphology, wherein the cell morphology is a cobblestone appearance.

[0053] Typically, suitable conditions for keratinocyte differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0054] The present invention also provides cells having at least one characteristic of a keratinocyte produced by the methods described herein.

[0055] In any of the methods described herein, the method can further include enriching cells having at least one characteristic of the target cell type to increase the proportion of cells in the population having at least one characteristic of the target cell type. The enrichment step can be performed in culture for a time and under conditions sufficient to produce a population of cells as described below.

[0056] In any of the methods described herein, the method may further comprise administering to the individual a cell, or a cell population comprising cells, having at least one characteristic of a keratinocyte.

[0057] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of a keratinocyte, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of a keratinocyte.

[0058] The present invention also relates to kits for producing cells having at least one characteristic of a keratinocyte, as disclosed herein. In some embodiments, the kits include a nucleic acid sequence encoding (i) a FOXQ1 polypeptide or a variant thereof, and (ii) a SOX9 polypeptide or a variant thereof, and (iii) a MAFB polypeptide or a variant thereof, and (iv) a CDH1 polypeptide or a variant thereof. and (v) a nucleic acid sequence encoding a FOS polypeptide or variant thereof, and (vi) a nucleic acid sequence encoding a REL polypeptide or variant thereof. In some embodiments, the kit further comprises instructions for reprogramming fibroblasts into cells having at least one characteristic of a keratinocyte according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0059] The present invention relates to a composition comprising at least one fibroblast and at least one agent that increases protein expression of any one or more of FOXQ1, SOX9, MAFB, CDH1, FOS, and REL in the fibroblast.

[0060] The present invention provides a method for reprogramming fibroblasts, the method comprising increasing protein expression of any one or more of SOX17, SMAD1, TAL1, IRF1, TCF7L1, MXD4 and JUNB, or variants thereof, in fibroblasts, wherein the fibroblasts are reprogrammed to have at least one characteristic of an endothelial cell.

[0061] The present invention provides a method for generating cells having at least one characteristic of endothelial cells from fibroblasts, the method comprising: - increasing the amount of any one or more of SOX17, SMAD1, TAL1, IRF1, TCF7L1, MXD4 and JUNB, or variants thereof, in fibroblasts; - culturing the fibroblasts for a time and under conditions sufficient for endothelial differentiation; thereby generating cells from the fibroblasts having at least one characteristic of endothelial cells.

[0062] The present invention provides a method for reprogramming fibroblasts into cells having at least one characteristic of an endothelial cell, the method comprising: i) providing a fibroblast, or a cell population comprising a fibroblast; ii) transfecting the fibroblast with one or more nucleic acids comprising a nucleotide sequence encoding the polypeptides SOX17, SMAD1, TAL1, IRF1, TCF7L1, MXD4 and JUNB; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of an endothelial cell.

[0063] Preferably, at least one characteristic of endothelial cells is upregulation of any one or more endothelial markers, including CD31 (Pe-CAM), VE-cadherin, and VEGFR2, and / or a change in cell morphology, and the cell morphology may be capillary-like structures.

[0064] Typically, suitable conditions for endothelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0065] The present invention also provides cells having at least one characteristic of an endothelial cell produced by the methods described herein.

[0066] In any of the methods described herein, the method may further comprise administering to the individual a cell, or a cell population comprising cells, having at least one characteristic of an endothelial cell.

[0067] The present invention also provides a method for the preparation of a cell line comprising the steps of: (a) providing a cell line comprising: a cell line; Populations are provided, the cells being generated by the methods described herein, wherein preferably at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of an endothelial cell.

[0068] The present invention also relates to kits for producing cells having at least one characteristic of an endothelial cell, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX17 polypeptide or a variant thereof; and (ii) a nucleic acid sequence encoding a SMAD1 polypeptide or a variant thereof; and (iii) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof, (iv) a nucleic acid sequence encoding a TCF7L1 polypeptide or a variant thereof, (v) a nucleic acid sequence encoding an MXD4 polypeptide or a variant thereof, (vi) a nucleic acid sequence encoding a TAL1 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a JUNB polypeptide or a variant thereof. In some embodiments, the kits further include instructions for reprogramming fibroblasts into cells having at least one characteristic of an endothelial cell according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0069] The present invention relates to a composition comprising at least one fibroblast and at least one agent that increases protein expression of any one or more of SOX17, SMAD1, TAL1, IRF1, TCF7L1, MXD4, and JUNB in ​​the fibroblast.

[0070] The present invention provides a method for reprogramming fibroblasts, the method comprising increasing protein expression of any one or more of SOX2, SOX9, ARNT2, E2F5, PXB1, SMAD1, and RUNX2, or variants thereof, in fibroblasts, wherein the fibroblasts are reprogrammed to have at least one characteristic of an astrocyte.

[0071] The present invention provides a method for generating cells having at least one characteristic of an astrocyte from a fibroblast, the method comprising: - Increasing the amount of any one or more of SOX2, SOX9, ARNT2, E2F5, PXB1, SMAD1, and RUNX2 or variants thereof in fibroblasts; - culturing the fibroblasts for a time and under conditions sufficient for astrocyte differentiation; thereby generating cells from the fibroblasts having at least one characteristic of astrocytes.

[0072] The present invention provides methods for reprogramming fibroblasts into cells having at least one characteristic of an astrocyte, the method comprising: i) providing a fibroblast, or a cell population comprising a fibroblast; ii) transfecting the fibroblast with one or more nucleic acids comprising a nucleotide sequence encoding the polypeptides SOX2, SOX9, ARNT2, E2F5, PXB1, SMAD1, and RUNX2; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of an astrocyte.

[0073] Preferably, at least one characteristic of astrocytes is the upregulation of any one or more astrocyte markers and / or a change in cell morphology. Astrocyte markers include GFAP, S100B, and ALDH1L1. Preferably, the marker used is GFAP. Preferably, the morphology observed is the presence of astrocytic processes. Typically, suitable conditions for astrocyte differentiation include culturing the cells in a suitable medium for a sufficient time. A sufficient time of culturing is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 , 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 days. Suitable media may be those shown in Table 9.

[0074] The present invention also provides cells having at least one characteristic of an astrocyte produced by the methods described herein.

[0075] In any of the methods described herein, the method may further comprise administering to the individual a cell, or a cell population comprising cells, having at least one characteristic of an astrocyte.

[0076] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of astrocytes, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of astrocytes.

[0077] The present invention also relates to kits for producing cells having at least one characteristic of astrocytes, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX2 polypeptide or a variant thereof; and (ii) a nucleic acid sequence encoding a SOX9 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding an ARNT2 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an E2F5 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a PXB1 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a SMAD1 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a RUNX2 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for reprogramming fibroblasts into cells having at least one characteristic of astrocytes according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0078] The present invention relates to a composition comprising at least one fibroblast and at least one agent that increases protein expression of any one or more of SOX2, SOX9, ARNT2, E2F5, PXB1, SMAD1, and RUNX2 in the fibroblast.

[0079] The present invention provides a method for reprogramming fibroblasts, the method comprising increasing protein expression of any one or more of FOS, DBP, HES1, FOXA2, ESRRA, CDH1, FOXQ1 and PAX6 or variants thereof in fibroblasts, wherein the fibroblasts are reprogrammed to have at least one characteristic of an epithelial cell.

[0080] The present invention provides a method for generating cells having at least one characteristic of epithelial cells from fibroblasts, the method comprising: - increasing the amount of any one or more of FOS, DBP, HES1, FOXA2, ESRRA, CDH1, FOXQ1 and PAX6 or variants thereof in fibroblasts; - culturing the fibroblasts for a time and under conditions sufficient for epithelial differentiation; thereby generating cells from the fibroblasts having at least one characteristic of epithelial cells.

[0081] The present invention provides a method for reprogramming fibroblasts into cells having at least one characteristic of an epithelial cell, the method comprising: i) providing a fibroblast, or a cell population comprising a fibroblast; ii) administering to the patient a polypeptide selected from the group consisting of FOS, DBP, HES1, FOXA2, ES2, and the like. iii) transfecting the fibroblasts with one or more nucleic acids comprising nucleotide sequences encoding RRA, CDH1, FOXQ1 and PAX6; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an epithelial cell.

[0082] Preferably, at least one characteristic of epithelial cells is upregulation of any one or more epithelial markers and / or a change in cell morphology. Epithelial markers include cytokeratin 15 (CK15), cytokeratin 3 (CK3), involucrin, and connexin 4. Preferably, the observed morphology is a cobblestone appearance.

[0083] Typically, suitable conditions for epithelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0084] The present invention also provides a cell having at least one characteristic of an epithelial cell produced by the methods described herein.

[0085] In any of the methods described herein, the method may further comprise administering to the individual a cell, or a cell population comprising cells, having at least one characteristic of an epithelial cell.

[0086] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of an epithelial cell, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of an epithelial cell.

[0087] The present invention also relates to kits for producing cells having at least one characteristic of epithelial cells, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a FOS polypeptide or a variant thereof; and (ii) a nucleic acid sequence encoding a DBP polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a nuclear FOXA2 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an ESRRA polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a CDH1 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a FOXQ1 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a PAX6 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for reprogramming fibroblasts into cells having at least one characteristic of epithelial cells according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0088] The present invention relates to a composition comprising at least one fibroblast and at least one agent that increases protein expression of any one or more of FOS, DBP, HES1, FOXA2, ESRRA, CDH1, FOXQ1 and PAX6 in the fibroblast.

[0089] The present invention provides a method for reprogramming keratinocytes, the method comprising increasing protein expression of any one or more of SOX17, TAL1, SMAD1, IRF1, HOXB7 and TCF7L1 in keratinocytes, wherein the keratinocytes are reprogrammed to have at least one characteristic of an endothelial cell.

[0090] The present invention provides a method for producing cells having at least one characteristic of endothelial cells from keratinocytes, the method comprising: Increasing the amount of any one or more of SOX17, TAL1, SMAD1, IRF1, HOXB7 and TCF7L1, or variants thereof, in keratinocytes; Culturing the keratinocytes for a time and under conditions sufficient for endothelial differentiation; thereby producing cells from the keratinocytes having at least one characteristic of endothelial cells.

[0091] The present invention provides a method for reprogramming keratinocytes into cells having at least one characteristic of an endothelial cell, the method comprising: i) providing a keratinocyte, or a cell population comprising a keratinocyte; ii) transfecting the keratinocyte with one or more nucleic acids comprising a nucleotide sequence encoding the polypeptides SOX17, TAL1, SMAD1, IRF1, HOXB7, and TCF7L1; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of an endothelial cell.

[0092] Preferably, in any aspect of the present invention, the endothelial cells are microvascular endothelial cells.

[0093] Preferably, the at least one characteristic of endothelial cells is the upregulation of any one or more endothelial markers and / or a change in cell morphology. Endothelial markers include CD31, VE-cadherin and VEGFR2.

[0094] Typically, suitable conditions for endothelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0095] The present invention also provides cells having at least one characteristic of a microvascular endothelial cell produced by the methods described herein.

[0096] The present invention also provides cell populations in which at least 5% of the cells have at least one characteristic of endothelial cells, preferably microvascular endothelial cells, and the cells are generated by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of endothelial cells, preferably microvascular endothelial cells.

[0097] The present invention also relates to kits for producing cells having at least one characteristic of an endothelial cell, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX17 polypeptide or a variant thereof; and (ii) a nucleic acid sequence encoding a TAL1 polypeptide or a variant thereof; and (iii) a nucleic acid sequence encoding a SMAD1 polypeptide or a variant thereof; and (iv) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof, (v) a nucleic acid sequence encoding a TCF7L1 polypeptide or a variant thereof; and (vi) a nucleic acid sequence encoding a HOXB7 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for reprogramming keratinocytes into cells having at least one characteristic of an endothelial cell according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0098] The present invention relates to a composition comprising at least one keratinocyte and at least one agent that increases protein expression of any one or more of SOX17, TAL1, SMAD1, IRF1, HOXB7, and TCF7L1 in the keratinocyte.

[0099] The present invention provides a method for reprogramming keratinocytes, the method comprising increasing protein expression of any one or more of NOTCH1, HR, DBP, OTX1, ESRRA, FOXQ1, PAX6 and IRX5 in keratinocytes, wherein the keratinocytes are reprogrammed to have at least one characteristic of an epithelial cell.

[0100] The present invention provides a method for producing cells having at least one characteristic of epithelial cells from keratinocytes, the method comprising: Increasing the amount of any one or more of NOTCH1, HR, DBP, OTX1, ESRRA, FOXQ1, PAX6 and IRX5 or variants thereof in keratinocytes; Culturing the keratinocytes for a time and under conditions sufficient for epithelial differentiation; thereby producing cells from the keratinocytes having at least one characteristic of epithelial cells.

[0101] The present invention provides a method for reprogramming keratinocytes into cells having at least one characteristic of an epithelial cell, the method comprising: i) providing a keratinocyte, or a cell population comprising keratinocytes; ii) transfecting the keratinocytes with one or more nucleic acids comprising nucleotide sequences encoding the polypeptides NOTCH1, HR, DBP, OTX1, ESRRA, FOXQ1, PAX6 and IRX5; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of an epithelial cell.

[0102] Preferably, in any aspect of the present invention, the epithelial cells are corneal epithelial cells.

[0103] Preferably, at least one characteristic of epithelial cells is upregulation of any one or more epithelial markers, including cytokeratin 15 (CK15), cytokeratin 3 (CK3), involucrin, and connexin 4, and / or a change in cell morphology.

[0104] Typically, suitable conditions for endothelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0105] The present invention also provides cells having at least one characteristic of an epithelial cell, preferably a corneal epithelial cell, produced by the methods described herein.

[0106] The present invention also provides cell populations in which at least 5% of the cells have at least one characteristic of epithelial cells, preferably corneal epithelial cells, and the cells are generated by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of epithelial cells, preferably corneal epithelial cells.

[0107] The present invention also relates to kits for producing cells having at least one characteristic of an epithelial cell, as disclosed herein. In some embodiments, the kits include: (i) a NOTCH The kit may comprise any one or more of (i) a nucleic acid sequence encoding a keratinocyte reprogramming kit (1) or a variant thereof; and (ii) a nucleic acid sequence encoding an HR polypeptide or a variant thereof; and (iii) a nucleic acid sequence encoding a DBP polypeptide or a variant thereof, and (iv) a nucleic acid sequence encoding an OTX1 polypeptide or a variant thereof, (v) a nucleic acid sequence encoding an ESRRA polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a FOXQ1 polypeptide or a variant thereof; (vii) a nucleic acid sequence encoding a PAX6 polypeptide or a variant thereof; and (viii) a nucleic acid sequence encoding an IRX5 polypeptide or a variant thereof. In some embodiments, the kit further comprises instructions for reprogramming a keratinocyte into a cell having at least one characteristic of an epithelial cell according to the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0108] The present invention relates to a composition comprising at least one keratinocyte and at least one agent that increases protein expression of any one or more of NOTCH1, HR, DBP, OTX1, ESRRA, FOXQ1, PAX6 and IRX5 in the keratinocyte.

[0109] The present invention provides a method for differentiating embryonic stem cells, the method comprising increasing protein expression of any one or more of SOX17, TAL1, SMAD1, HOXB7, JUNB, IRF1 and NFKB1 in embryonic stem cells, wherein the embryonic stem cells differentiate to have at least one characteristic of an endothelial cell.

[0110] The present invention provides a method for generating cells having at least one characteristic of an endothelial cell from embryonic stem cells, the method comprising: Increasing the amount of any one or more of SOX17, TAL1, SMAD1, HOXB7, JUNB, IRF1 and NFKB1 or a mutant thereof in embryonic stem cells; Culturing the embryonic stem cells for a time and under conditions sufficient for endothelial differentiation; thereby producing cells from the embryonic stem cells that have at least one characteristic of endothelial cells.

[0111] The present invention provides a method for differentiating embryonic stem cells into cells having at least one characteristic of an endothelial cell, the method comprising: i) providing embryonic stem cells, or a cell population comprising embryonic stem cells; ii) transfecting the embryonic stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides SOX17, TAL1, SMAD1, HOXB7, JUNB, IRF1 and NFKB1; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an endothelial cell.

[0112] Preferably, in any aspect of the present invention, the endothelial cells are microvascular endothelial cells.

[0113] Preferably, the at least one characteristic of endothelial cells is the upregulation of any one or more endothelial markers and / or a change in cell morphology. Endothelial markers include CD31, VE-cadherin and VEGFR2.

[0114] Typically, suitable conditions for endothelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0115] The present invention also provides cells having at least one characteristic of a microvascular endothelial cell produced by the methods described herein.

[0116] The present invention also provides cell populations in which at least 5% of the cells have at least one characteristic of endothelial cells, preferably microvascular endothelial cells, and the cells are generated by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of endothelial cells, preferably microvascular endothelial cells.

[0117] The present invention also relates to kits for producing cells having at least one characteristic of an endothelial cell, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX17 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a TAL1 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a SMAD1 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a HOXB7 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a JUNB polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating embryonic stem cells into cells having at least one characteristic of an endothelial cell by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0118] The present invention relates to a composition comprising at least one embryonic stem cell and at least one agent that increases protein expression of any one or more of SOX17, TAL1, SMAD1, HOXB7, JUNB, IRF1 and NFKB1 in the embryonic stem cell.

[0119] The present invention provides a method for differentiating embryonic stem cells, the method comprising increasing protein expression of IRF1, SOX9, ARNT2, PAX6, SNAI2, SOX5 and RUNX2 in embryonic stem cells, wherein the embryonic stem cells differentiate to have at least one characteristic of an astrocyte.

[0120] The present invention provides a method for producing or generating cells having at least one characteristic of an astrocyte from embryonic stem cells, the method comprising: Increasing the amount of any one or more of IRF1, SOX9, ARNT2, PAX6, SNAI2, SOX5 and RUNX2, or mutants thereof, in embryonic stem cells; Culturing the embryonic stem cells for a time and under conditions sufficient for astrocyte differentiation; thereby producing cells from the embryonic stem cells having at least one characteristic of an astrocyte.

[0121] The present invention provides methods for differentiating embryonic stem cells into cells having at least one characteristic of an astrocyte, the method comprising: i) providing embryonic stem cells, or a cell population comprising embryonic stem cells; ii) transfecting the embryonic stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the following polypeptides: IRF1, SOX9, ARNT2, PAX6, SNAI2, SOX5, and RUNX2; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an astrocyte.

[0122] Preferably, at least one characteristic of astrocytes is the upregulation of any one or more astrocyte markers and / or a change in cell morphology. Astrocyte markers include GFAP, S100B, and ALDH1L1. Preferably, the marker used is GFAP. Preferably, the observed morphology is the presence of astrocytic processes.

[0123] Typically, suitable conditions for astrocyte differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. A suitable medium may be one set forth in Table 9.

[0124] The present invention also provides cells having at least one characteristic of an astrocyte produced by the methods described herein.

[0125] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of astrocytes, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of astrocytes.

[0126] The present invention also relates to kits for producing cells having at least one characteristic of astrocytes, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a SOX9 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding an ARNT2 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding a PAX6 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a SNAI2 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a RUNX2 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a SOX5 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating embryonic stem cells into cells having at least one characteristic of astrocytes by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0127] The present invention relates to a composition comprising at least one embryonic stem cell and at least one agent that increases the protein expression of IRF1, SOX9, ARNT2, PAX6, SNAI2, SOX5 and RUNX2 in the embryonic stem cell.

[0128] The present invention provides a method for differentiation of embryonic stem cells, the method comprising increasing protein expression of any one or more of SOX9, NFKB1, MYC, NR2F2, FOSL1, AHR and FOSL2 in embryonic stem cells, wherein the embryonic stem cells differentiate to have at least one characteristic of a keratinocyte.

[0129] The present invention provides a method for generating cells having at least one characteristic of a keratinocyte from an embryonic stem cell, the method comprising: Increasing the amount of any one or more of SOX9, NFKB1, MYC, NR2F2, FOSL1, AHR and FOSL2, or variants thereof, in embryonic stem cells; Culturing the embryonic stem cells for a time and under conditions sufficient for keratinocyte differentiation; thereby producing cells from the embryonic stem cells having at least one characteristic of a keratinocyte.

[0130] The present invention provides a method for differentiating embryonic stem cells into cells having at least one characteristic of a keratinocyte, the method comprising: i) providing an embryonic stem cell, or a cell population comprising embryonic stem cells; ii) transfecting the embryonic stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides SOX9, NFKB1, MYC, NR2F2, FOSL1, AHR, and FOSL2; and iii) culturing the cell or cell population and, optionally, monitoring the cell or cell population for at least one characteristic of a keratinocyte. This includes:

[0131] Preferably, at least one characteristic of a keratinocyte is an upregulation of any one or more keratinocyte markers, including pankeratin, keratin 1, keratin 14, and involucrin, and / or a change in cell morphology, and the cell morphology is a cobblestone appearance.

[0132] Typically, suitable conditions for keratinocyte differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0133] The present invention also provides cells having at least one characteristic of a keratinocyte produced by the methods described herein.

[0134] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of a keratinocyte, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of a keratinocyte.

[0135] The present invention also relates to kits for producing cells having at least one characteristic of a keratinocyte, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX9 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a MYC polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding a FOSL2 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding an NR2F2 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a FOSL1 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a nuclear AHR polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating embryonic stem cells into cells having at least one characteristic of a keratinocyte by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0136] The present invention relates to a composition comprising at least one embryonic stem cell and at least one agent that increases protein expression of SOX9, NFKB1, MYC, NR2F2, FOSL1, AHR, and FOSL2 in the embryonic stem cell.

[0137] The present invention provides a method for differentiating embryonic stem cells, the method comprising increasing protein expression of any one or more of MYC, IL1B, FOS, NFKB1, ESRRA, FOXQ1, IRF1 and PAX6 in embryonic stem cells, and differentiating the embryonic stem cells to have at least one characteristic of an epithelial cell, preferably a corneal epithelial cell.

[0138] The present invention provides a method for generating cells having at least one characteristic of an epithelial cell from embryonic stem cells, the method comprising: Increasing the amount of any one or more of MYC, IL1B, FOS, NFKB1, ESRRA, FOXQ1, IRF1 and PAX6, or mutants thereof, in embryonic stem cells; Culturing the embryonic stem cells for a time and under conditions sufficient for epithelial differentiation; thereby producing cells from the embryonic stem cells that have at least one characteristic of an epithelial cell.

[0139] The present invention provides a method for differentiating embryonic stem cells into cells having at least one characteristic of an epithelial cell, the method comprising: i) providing embryonic stem cells, or a cell population comprising embryonic stem cells; ii) transfecting the embryonic stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides MYC, IL1B, FOS, NFKB1, ESRRA, FOXQ1, IRF1 and PAX6; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of an epithelial cell.

[0140] Preferably, at least one characteristic of epithelial cells is upregulation of any one or more epithelial markers, including cytokeratin 15 (CK15), cytokeratin 3 (CK3), involucrin, and connexin 4, and / or a change in cell morphology, and the cell morphology may be a cobblestone appearance.

[0141] Typically, suitable conditions for epithelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0142] The present invention also provides a cell having at least one characteristic of an epithelial cell produced by the methods described herein.

[0143] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of an epithelial cell, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of an epithelial cell.

[0144] The present invention also relates to kits for producing cells having at least one characteristic of an epithelial cell, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a MYC polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding an IL1B polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a FOS polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding an ESRRA polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a FOXQ1 polypeptide or a variant thereof; (vii) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof; and (viii) a nucleic acid sequence encoding a PAX6 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating embryonic stem cells into cells having at least one characteristic of an epithelial cell by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0145] The present invention relates to a composition comprising at least one embryonic stem cell and at least one agent that increases protein expression of any one or more of MYC, IL1B, FOS, NFKB1, ESRRA, FOXQ1, IRF1 and PAX6 in the embryonic stem cell.

[0146] The present invention provides a method for generating endothelial cells from pluripotent stem cells, comprising differentiating the pluripotent stem cells, the method comprising increasing the protein expression of any one or more of SOX17, TAL1, NFKB1, IRF1, HOXB7, JUNB, and SMAD1 in the pluripotent stem cells. The method comprises: causing the pluripotent stem cells to proliferate and differentiate to have at least one characteristic of an endothelial cell.

[0147] In any aspect of the present invention (including any method or composition), the pluripotent stem cells may be induced pluripotent stem cells (iPSCs).

[0148] The present invention provides a method for generating cells having at least one characteristic of an endothelial cell from pluripotent stem cells, the method comprising: Increasing the amount of any one or more of SOX17, TAL1, NFKB1, HOXB7, JUNB, IRF1 and SMAD1, or mutants thereof, in pluripotent stem cells; Culturing the pluripotent stem cells for a time and under conditions sufficient for endothelial differentiation; thereby generating cells from the pluripotent stem cells having at least one characteristic of endothelial cells.

[0149] The present invention provides a method for differentiating pluripotent stem cells into cells having at least one characteristic of an endothelial cell, the method comprising: i) providing pluripotent stem cells, or a cell population comprising pluripotent stem cells; ii) transfecting the pluripotent stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides SOX17, TAL1, NFKB1, HOXB7, JUNB, IRF1 and SMAD1; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an endothelial cell.

[0150] Preferably, at least one characteristic of endothelial cells is upregulation of any one or more endothelial markers and / or a change in cell morphology. Endothelial markers include pan-CD31, VE-cadherin and VEGFR2.

[0151] Typically, suitable conditions for endothelial differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0152] The present invention also provides cells having at least one characteristic of an endothelial cell produced by the methods described herein.

[0153] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of an endothelial cell, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of an endothelial cell.

[0154] The present invention also relates to kits for producing cells having at least one characteristic of an endothelial cell, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX17 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a TAL1 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an IRF1 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a SMAD1 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a HOXB7 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding a JUNB polypeptide or a variant thereof. In some embodiments, the kits further include a method for producing at least one endothelial cell by the methods disclosed herein. and instructions for differentiation of pluripotent stem cells into cells having the characteristic. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0155] The present invention relates to a composition comprising at least one pluripotent stem cell and at least one agent that increases protein expression of any one or more of SOX17, TAL1, NFKB1, IRF1, HOXB7, JUNB, and SMAD1 in the pluripotent stem cell.

[0156] The present invention provides a method of generating astrocytes from pluripotent stem cells comprising differentiating the pluripotent stem cells, the method comprising increasing protein expression of any one or more of PAX6, SNAI2, POU3F2, SOX5, E2F5, RUNX2, and HMGB2 in the pluripotent stem cells, wherein the pluripotent stem cells differentiate to have at least one characteristic of an astrocyte.

[0157] The present invention provides a method for generating cells having at least one characteristic of an astrocyte from pluripotent stem cells, the method comprising: Increasing the amount of any one or more of PAX6, SNAI2, POU3F2, SOX5, E2F5, RUNX2, and HMGB2, or variants thereof, in pluripotent stem cells; Culturing the pluripotent stem cells for a time and under conditions sufficient for astrocyte differentiation; thereby generating cells from the pluripotent stem cells having at least one characteristic of an astrocyte.

[0158] The present invention provides methods for differentiating pluripotent stem cells into cells having at least one characteristic of an astrocyte, the method comprising: i) providing pluripotent stem cells, or a cell population comprising pluripotent stem cells; ii) transfecting the pluripotent stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides PAX6, SNAI2, POU3F2, SOX5, E2F5, RUNX2, and HMGB2; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an astrocyte.

[0159] Preferably, at least one characteristic of astrocytes is the upregulation of any one or more astrocyte markers and / or a change in cell morphology. Astrocyte markers include GFAP, S100B, and ALDH1L1. Preferably, the marker used is GFAP. Preferably, the observed morphology is the presence of astrocytic processes.

[0160] Typically, suitable conditions for astrocyte differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. A suitable medium may be one set forth in Table 9.

[0161] The present invention also provides cells having at least one characteristic of an astrocyte produced by the methods described herein.

[0162] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of astrocytes, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of astrocytes.

[0163] The present invention also provides a cell having at least one characteristic of an astrocyte as disclosed herein. The present invention relates to kits for manufacturing. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a PAX6 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a SNAI2 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a RUNX2 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an HMGB2 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a POU3F2 polypeptide or a variant thereof; (vi) a nucleic acid sequence encoding a SOX5 polypeptide or a variant thereof; and (vii) a nucleic acid sequence encoding an E2F5 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating pluripotent stem cells into cells having at least one characteristic of an astrocyte by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0164] The present invention relates to a composition comprising at least one pluripotent stem cell and at least one agent that increases protein expression of any one or more of PAX6, SNAI2, POU3F2, SOX5, E2F5, RUNX2, and HMGB2 in the pluripotent stem cell.

[0165] The present invention provides a method for generating keratinocytes from pluripotent stem cells, comprising differentiating the pluripotent stem cells, the method comprising increasing protein expression of any one or more of TFAP2A, MYC, SOX9, TP63, NFKBIA and NFKB1 in the pluripotent stem cells, wherein the pluripotent stem cells differentiate to have at least one characteristic of a keratinocyte.

[0166] The present invention provides a method for generating cells having at least one characteristic of a keratinocyte from a pluripotent stem cell, the method comprising: Increasing the amount of any one or more of TFAP2A, MYC, SOX9, TP63, NFKBIA, and NFKB1 or a mutant thereof in pluripotent stem cells; Culturing the pluripotent stem cells for a time and under conditions sufficient for keratinocyte differentiation; thereby generating cells having at least one characteristic of keratinocytes from the pluripotent stem cells.

[0167] The present invention provides a method for differentiating pluripotent stem cells into cells having at least one characteristic of a keratinocyte, the method comprising: i) providing pluripotent stem cells, or a cell population comprising pluripotent stem cells; ii) transfecting the pluripotent stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides TFAP2A, MYC, SOX9, TP63, NFKBIA, and NFKB1; and iii) culturing the cell or cell population and optionally monitoring the cell or cell population for at least one characteristic of a keratinocyte.

[0168] Preferably, the at least one characteristic of a keratinocyte is an upregulation of any one or more keratinocyte markers, including keratin 1, keratin 14, and involucrin, and / or a change in cell morphology, wherein the cell morphology is a cobblestone appearance.

[0169] Typically, suitable conditions for keratinocyte differentiation include culturing the cells for a sufficient time in a suitable medium, which may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 days. Suitable media may be those set forth in Table 9.

[0170] The present invention also provides cells having at least one characteristic of a keratinocyte produced by the methods described herein.

[0171] The present invention also provides a cell population in which at least 5% of the cells have at least one characteristic of a keratinocyte, the cells being produced by the methods described herein. , at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of a keratinocyte.

[0172] The present invention also relates to kits for producing cells having at least one characteristic of a keratinocyte, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a TFAP2A polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a MYC polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding a SOX9 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding a TP63 polypeptide or a variant thereof; or (vi) a nucleic acid sequence encoding an NFKB1 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating pluripotent stem cells into cells having at least one characteristic of a keratinocyte by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0173] The present invention relates to a composition comprising at least one pluripotent stem cell and at least one agent that increases protein expression of any one or more of TFAP2A, MYC, SOX9, TP63, NFKBIA, and NFKB1 in the pluripotent stem cell.

[0174] The present invention provides a method for generating astrocytes from bone marrow stem cells, comprising differentiating the bone marrow stem cells, the method comprising increasing protein expression of any one or more of SOX2, SOX9, ARNT2, MYBL2, POU3F2, E2F1 and HMGB2 in the bone marrow stem cells, wherein the bone marrow stem cells differentiate to have at least one characteristic of an astrocyte.

[0175] The present invention provides a method for generating cells having at least one characteristic of astrocytes from bone marrow stem cells, the method comprising: Increasing the amount of any one or more of SOX2, SOX9, ARNT2, MYBL2, POU3F2, E2F1 and HMGB2 or variants thereof in bone marrow stem cells; Culturing the bone marrow stem cells for a time and under conditions sufficient for astrocyte differentiation; thereby generating cells from the bone marrow stem cells having at least one characteristic of astrocytes.

[0176] The present invention provides methods for differentiating bone marrow stem cells into cells having at least one characteristic of an astrocyte, the method comprising: i) providing bone marrow stem cells, or a cell population comprising bone marrow stem cells; ii) transfecting the bone marrow stem cells with one or more nucleic acids comprising a nucleotide sequence encoding any one or more of the polypeptides SOX2, SOX9, ARNT2, MYBL2, POU3F2, E2F1 and HMGB2; and iii) culturing the cells or cell population and optionally monitoring the cells or cell population for at least one characteristic of an astrocyte.

[0177] Preferably, at least one characteristic of astrocytes is the upregulation of any one or more astrocyte markers and / or a change in cell morphology. Astrocyte markers include GFAP, S100B, and ALDH1L1. Preferably, the marker used is GFAP. Preferably, the observed morphology is the presence of astrocytic processes.

[0178] Typically, suitable conditions for astrocyte differentiation include culturing the cells in a suitable medium for a sufficient period of time. The incubation period may be 3, 24, 25, 26, 27, 28, 29 or 30 days. Suitable media may be those shown in Table 9.

[0179] The present invention also provides cells having at least one characteristic of an astrocyte produced by the methods described herein.

[0180] The invention also provides cell populations in which at least 5% of the cells have at least one characteristic of astrocytes, the cells being produced by the methods described herein. Preferably, at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% of the cells in the population have at least one characteristic of astrocytes.

[0181] The present invention also relates to kits for producing cells having at least one characteristic of astrocytes, as disclosed herein. In some embodiments, the kits include any one or more of: (i) a nucleic acid sequence encoding a SOX2 polypeptide or a variant thereof; (ii) a nucleic acid sequence encoding a SOX9 polypeptide or a variant thereof; (iii) a nucleic acid sequence encoding an ARNT2 polypeptide or a variant thereof; (iv) a nucleic acid sequence encoding an E2F1 polypeptide or a variant thereof; (v) a nucleic acid sequence encoding an HMGB2 polypeptide or a variant thereof; or (vi) a nucleic acid sequence encoding a POU3F2 polypeptide or a variant thereof. In some embodiments, the kits further include instructions for differentiating bone marrow stem cells into cells having at least one characteristic of astrocytes by the methods disclosed herein. Preferably, the present invention provides kits for use in the methods of the present invention described herein.

[0182] The present invention relates to a composition comprising at least one bone marrow stem cell and at least one agent that increases protein expression of any one or more of SOX2, SOX9, ARNT2, MYBL2, E2F1, POU3F2 and HMGB2 in the bone marrow stem cell.

[0183] Typically, protein expression or amount of a transcription factor described herein is increased by contacting the cell with an agent that increases expression of the transcription factor. Preferably, the agent is selected from the group consisting of: nucleotide sequences, proteins, aptamers and small molecules, ribosomes, RNAi agents and peptide-nucleic acids (PNAs) and analogs or variants thereof. Preferably, the agent is exogenous.

[0184] Typically, protein expression or amount of a transcription factor described herein is increased by introducing into a cell at least one nucleic acid comprising a nucleotide sequence encoding the transcription factor or encoding a functional fragment of the transcription factor. Preferably, the nucleotide sequence encoding the transcription factor is at least 70%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical to a sequence having an accession number listed in Table 3.

[0185] Preferably, the nucleic acid further comprises a heterologous promoter. Preferably, the nucleic acid is in a vector, such as a viral vector or a non-viral vector. Preferably, the vector is a viral vector comprising a genome that does not integrate into the host cell genome. The viral vector may be a retroviral vector or a lentiviral vector.

[0186] Any of the methods described herein may have one or more, or all, steps performed in vitro, ex vivo, or in vivo.

[0187] As used herein, unless the context otherwise requires, the term "comprise" and variations of that term, such as "comprises," "comprises," and "comprised of," are not intended to exclude further additives, components, integers, or steps.

[0188] Further aspects of the invention and further embodiments of the aspects described in the preceding paragraphs will become apparent from the following description, given by way of example, and with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0189] [Figure 1-1] The Mogrify algorithm predicts TFs for cellular transformation. This works as follows: (A) Mogrify aims to find TFs that are not only differentially expressed but also appear to be involved in the regulation of a large number of differentially expressed genes in a given cell type. (B) To select an appropriate background for DESeq (Anders, S. & Huber, W. (2010) Genome Biol. 2010;11(10):R106), we use the cell type ontology tree created as part of the FANTOM5 consortium (Forrest, A.R. et al. Nature 507, 462-470 (2014)). We calculate adjusted p-values ​​and log-fold changes for genes in samples. [Figure 1-2] (C) For each TF, we build a local network neighborhood of influence, weighting downstream effects on the gene by its connection distance and parental out-degree. (D) We maximize regulatory coverage by removing TFs whose influence overlaps with other factors. [Figure 2] Mogrify predictions for several known transdifferentiation events have been published in the literature. TFs that Mogrify accurately identifies from the published list are highlighted. Samples are classified using the FANTOM cell ontology (Forrest, A.R. et al., supra). For each publication, transcription factors in the initial maximum coverage set are shown in green, and the entire predicted Mogrify set is shown in orange. For example, transdifferentiation between fibroblasts and myoblasts (Lattanzi, L. et al. J. Clin. Invest. 101, 2119-28 (1998)) requires only MYOD, which was identified by Mogrify. [Figure 3-1]Experimental validation of a novel transition predicted by Mogrify: induction of keratinocytes from dermal fibroblasts. (A) Transcription factor network predicted by Mogrify as involved in keratinocyte transdifferentiation from dermal fibroblasts. [Figure 3-2] (B) Schematic of the method used for the transdifferentiation assay. (C) qPCR analysis of the indicated markers in cells harvested on days 12–16 during transdifferentiation. All values ​​are experimental replicates and refer to gene expression in dermal fibroblasts (n=3, error bars indicate s.e.m.). (D) Brightfield and GFP images at day 24 show the cobblestone morphology of transdifferentiated cells (upper panel) and GFP+ control cells (lower panel). [Figure 4-1] Experimental validation of a novel transformation predicted by Mogrify: induction of microvascular endothelial cells from keratinocytes. (A) Schematic of the transcription factor network predicted by Mogrify as involved in microvascular endothelial cell transdifferentiation from keratinocytes. (B) Overview of the method used for the transdifferentiation assay. (C) Flow cytometry analysis of CD31 expression at days 0, 14, and 18 of transdifferentiation. [Figure 4-2] (D) qPCR analysis of the indicated expression markers in CD31+ cells harvested on day 18 of transdifferentiation. All values ​​are experimental replicates and refer to gene expression in keratinocytes (n = 3, error bars indicate sem). (E) Immunofluorescence analysis of endothelial markers CD31 and VE-cadherin on day 18 for vector-free control cells (a) and transdifferentiated cells (b-f). Scale bar = 50 μm. [Figure 5] Comparison to published conversions. As more transcription factors are added to the list, the additional coverage values ​​for each conversion show that coverage always approaches 100% within eight transcription factors. [Figure 6]Benchmarking against Existing Cellular Transduction TF Technologies. To demonstrate how Mogrify's performance compares to other published methods in deriving sets of TFs for cellular transduction, we report two statistics. First (top panel), the recovery rate for each technique; 100% recovery means that the technique also found the entire set of TFs used in the published transduction. Consequently, if that technique had been used to configure the experiment, the known transduction set would have been discovered in the first iteration. For Mogrify, this represents 6 / 10 of the published transduction cases; for CellNet and D'Allessio et al., this is true for only 1 / 10 of the published transduction cases. Second (bottom panel), the average rank of the recovered TFs is plotted. Ignoring the TFs missed by each technique, this test shows how well each technique performs at somehow prioritizing the necessary TFs. With the exception of the transduction between fibroblasts and cardiac (cardiomyocyte) cells, Mogrify performed best in all cases. If none of the correct TFs were predicted, the average rank is not shown. This corresponds to four transitions in CellNet and one transition in D'Alessio et al. [Figure 7] Reprogramming landscape of human cell types. Samples are classified using the cell ontology terms provided by Forrest et al., supra. Expression profiles of ontology terms, including repeats, are arranged in the XY plane using multidimensional scaling to provide cell types, with cell types with similar expression profiles residing in close proximity to each other. Mogrify then calculates the height on the landscape according to the normalized cumulative coverage of the top 8 TFs, such that a transformation in which the top-ranked TF regulates all of the required genes has a height of 1, and vice versa. [Figure 8]Experimental validation of a novel transformation predicted by Mogrify: induction of endothelial cells from dermal fibroblasts. A: Immunofluorescence analysis of endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation. Scale bar, 25 μm. B: qPCR analysis showing the expression levels of endothelial-associated genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation. [Figure 9] Experimental validation of a novel transformation predicted by Mogrify: induction of endothelial cells from hESCs. A: Immunofluorescence analysis of endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation. Scale bar, 25 μm. B: qPCR analysis showing the expression levels of endothelial-associated genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation. [Figure 10] Induction of endothelial cells from hESCs. A: Flow cytometry analysis of PeCAM expression on days 12 and 18 of transdifferentiation. FSC, forward scatter. B: Quantification of PeCAM-positive cells on day 18 of transdifferentiation. N=3. [Figure 11] Experimental validation of a novel transformation predicted by Mogrify: induction of endothelial cells from hiPSCs. A: Immunofluorescence analysis of endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation. Scale bar, 25 μm. B: qPCR analysis showing the expression levels of endothelial-associated genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation. [Figure 12] Induction of endothelial cells from hiPSCs. A: Flow cytometry analysis of PeCAM expression on days 12 and 18 of transdifferentiation. FSC, forward scatter. B: Quantification of PeCAM-positive cells on day 18 of transdifferentiation. N=3. [Figure 13] Experimental validation of a novel transformation predicted by Mogrify: astrocyte induction from fibroblasts. Immunofluorescence analysis of the astrocyte marker GFAP at day 21 of transdifferentiation. Scale bar, 25 μm. [Figure 14] Experimental validation of a novel transformation predicted by Mogrify: astrocyte induction from hESCs. Immunofluorescence analysis of the astrocyte marker GFAP at day 21 of transdifferentiation. Scale bar, 25 μm. [Figure 15] Experimental validation of a novel transformation predicted by Mogrify: astrocyte induction from hiPSCs. Immunofluorescence analysis of the astrocyte marker GFAP at day 21 of transdifferentiation. 25 μm. [Figure 16] Experimental validation of a novel transformation predicted by Mogrify: astrocyte induction from BM-MSCs. Immunofluorescence analysis of the astrocyte marker GFAP at day 21 of transdifferentiation. Scale bar, 25 μm. [Figure 17] Experimental validation of a novel transformation predicted by Mogrify: keratinocyte derivation from hESCs. Immunofluorescence analysis of the keratinocyte marker pankeratin at day 21 of transdifferentiation. Scale bar, 25 μm. [Figure 18] Experimental validation of a novel transformation predicted by Mogrify: keratinocyte derivation from hiPSCs. A: Immunofluorescence analysis of the keratinocyte marker keratin 14 (KRT14) at day 21 of transdifferentiation. Scale bar, 25 μm. B and C: qPCR analysis showing the expression levels of the keratinocyte-associated genes keratin 14 (B) and keratin 1 (C) at day 21 of transdifferentiation. DETAILED DESCRIPTION OF THE INVENTION

[0190] It will be understood that the invention disclosed and defined herein extends to all alternative combinations of two or more of the individual features mentioned or apparent from the text or drawings, all of these different combinations constituting various alternative aspects of the invention.

[0191] Reference will now be made in detail to certain embodiments of the invention. While the invention will be described in conjunction with the embodiments, it will be understood that the intention is not to limit the invention to those embodiments. On the contrary, the invention is intended to cover all alternatives, modifications, and equivalents, which may be included within the scope of the present invention as defined by the claims.

[0192] Those skilled in the art will recognize many methods and materials similar or equivalent to those described herein, which could be used in the practice of this invention. The invention is not limited in any way to the methods and materials described. It will be understood that the invention disclosed and defined herein extends to all alternative combinations of two or more of the individual features mentioned or apparent from the text or drawings. All of these different combinations constitute various alternative aspects of the invention.

[0193] For purposes of interpretation of this specification, terms used in the singular will also include the plural and vice versa.

[0194] The present invention provides a practical and efficient mechanism for systematically performing cell conversion, facilitating the generalization of human cell reprogramming. The present invention combines gene expression data with regulatory network information to achieve what neither of these alone is sufficient to reliably and accurately identify the transcription factors necessary to convert a source cell type into a cell possessing at least one characteristic of a target cell type. Furthermore, in some embodiments, the present invention provides a set of transcription factors for conversion, rather than a ranked list of all transcription factors.

[0195] Expression data for each gene in a sample can be determined by any known method, including those described herein. The data can be generated de novo or derived from an existing database.

[0196] Differential expression can be determined using DESeq, edgeR, baySeq23, BBSeq24, NOISeq25, or QuasiSeq protocols, or any other process known to those skilled in the art that determines differential expression relative to background in one or more samples, or pairwise comparisons. It can be calculated using the following formula.

[0197] The tree-based background approach cited in various methods of the present invention is based on the principle of excluding cell types whose ontology is very similar while including other cell types that are close to the background in the tree. This can be achieved by choosing a point near the top of the tree that serves as a cutoff point. Samples that are in the same clade as the analyzed cell type may be removed, and samples that are not in the same clade but are still below this point may be included. The result is a set of samples that is broad enough to provide robust results, but narrow enough to maintain a manageable level of statistical power.

[0198] An alternative to tree-based approaches is Bayesian clustering, specifically the DGEclust approach described in Vavoulis et al. Genome Biology. 2015, 16:39.

[0199] To calculate the network-based influence sphere of a transcription factor, any network or subnetwork containing a source of network information regarding the interactions of transcription factors that affect gene expression can be used. Typically, this is information regarding the interactions of transcription factors with other biological molecules, such as DNA, RNA, or proteins. For example, any network information regarding protein-DNA interactions between transcription factors and known binding sites in the promoters or regulatory regions of genes. An example of a source of such network information is Motif Activity Response Analysis (MARA) (FANTOM Consortium, Suzuki et al. 2009. Nat Genet 41:553-5620). A further example of a source of network information is a database of protein-protein, protein-DNA, protein-RNA, and / or biological pathway interactions. An example of a source of such network information is the STRING (Search Tool for Retrieval of Interacting Genes / Protein) database. Exemplary databases and methods for calculating the network-based influence sphere of a transcription factor are described in the Examples. Preferably, when a MARA-derived network is used to generate a network score, any technique that identifies transcription start sites, such as cap analysis gene expression (CAGE), is used to generate the gene score.

[0200] A weighted sum of gene influences can be calculated across one or more networks to generate one or more influence lists. Preferably, at least two influence lists, such as those described herein, are generated. Weightings that can be applied include weightings that allow genes that are increasingly distant from direct regulation to have relatively little impact on the network score (referred to herein as distance weighting), and weightings that compensate for highly ubiquitous transcription factors and prevent them from receiving artificially high scores due to the regulation of a large number of marginally differentially expressed genes (referred to herein as edge weighting).

[0201] As used herein, "Mogrify" refers to the methods described herein for determining transcription factors required for the conversion of a source cell to a cell having at least one characteristic of a target cell. In any embodiment of the present invention, the Mogrify method can be implemented on a variety of computer processing systems, e.g., laptop computers, netbook computers, tablet computers, smartphones, desktop computers, and server computers. In one embodiment, the computer system includes a processor and a data storage device, wherein the data storage device stores a set of computer-readable media. In one aspect, the computer system further includes an algorithm for comparing expression profiles between source and / or target cells. In one embodiment, the computer readable medium stores an expression profile or a series of expression profiles from different cell types. In a further embodiment, the computer readable medium stores details of transcription factors involved in regulating a network of genes. It will be appreciated that the particular type of computer processing system will determine the appropriate hardware and architecture to be used.

[0202] Determination of gene and network scores may be by any method described herein, including the Examples.

[0203] The ranking of transcription factors may be by any method described herein, taking into account the gene scores and network scores described herein, or the influence lists described herein, etc. Preferably, scores or influence lists based on differential expression analysis and / or scores, or influence lists based on the interactions of transcription factors that directly and / or indirectly affect gene expression, are used to rank the transcription factors.

[0204] To identify a set of TFs for a given transformation, the ranked lists of source and target cell types are compared: if a TF for a target cell type is already expressed in the source target, it may be removed from the list.

[0205] Removal of transcriptionally redundant TFs from the ranked list for each cell type may be by any method described herein, including by comparing the list of genes that each TF directly regulates. For a given TF, if there is a higher-ranked TF that regulates more than 98% of the genes that the TF regulates, it may be removed. Thus, the resulting predictions include TFs that are diverse in their regulatory influence.

[0206] The process of cellular reprogramming changes the type of progeny a cell can give rise to and includes the distinct processes of forward programming and transdifferentiation. In some embodiments, forward programming of a multipotent or pluripotent cell provides a cell with at least one characteristic of a cell type that has a more differentiated phenotype than the multipotent or pluripotent cell. In other embodiments, transdifferentiation of one somatic cell provides a cell with at least one characteristic of another somatic cell type.

[0207] The present invention provides compositions and methods for directly reprogramming or transdifferentiating source cells into target cells without an intermediate induced pluripotent stem cell (iPS) stage before the source cells become target cells. Compared to iPS cell technology, transdifferentiation is highly efficient and carries a significantly lower risk of teratoma formation in downstream applications. Furthermore, transdifferentiation can be used to directly convert one cell type into another in vivo, whereas iPS cell technology cannot.

[0208] The source cells can be any cell type described herein, including somatic cells or disease cells. Somatic cells can be adult cells, or cells derived from an adult that exhibit one or more detectable characteristics of an adult, or non-embryonic cells. Disease cells can also be cells that exhibit one or more detectable characteristics of a disease or condition, e.g., disease cells can be cancer cells that exhibit one or more clinical or biochemical markers of cancer. Examples of source cells include hematopoietic cells, e.g., lymphocytes, myeloid cells; buccal mucosa cells, epithelial cells, mesenchymal cells, keratinocytes, and hepatocytes. Examples of source cells are listed in Table 4.

[0209] As used herein, the term "somatic cell" refers to any cell that forms the body of an organism, as opposed to a germ cell. In mammals, germ cells (also known as "gametes") are sperm. The somatic cells are the spermatozoa and the egg, which fuse during fertilization to produce a cell called a zygote, from which the entire mammalian embryo develops. All other cell types in a mammal's body—except for the sperm and eggs, the cells from which they are created (gametocytes), and undifferentiated stem cells—are somatic cells: internal organs, skin, bone, blood, and connective tissue are all composed of somatic cells. In some embodiments, the somatic cells are "non-embryonic somatic cells," meaning somatic cells that are not present in or obtained from an embryo or that do not result from the propagation of such cells in vitro. In some embodiments, the somatic cells are "adult somatic cells," meaning cells that are present in or obtained from an organism other than an embryo or fetus, or that result from the propagation of such cells in vitro. Somatic cells can be immortalized to provide an unlimited supply of cells, for example, by increasing the level of telomerase reverse transcriptase (TERT). For example, the level of TERT can be increased by increasing TERT transcription from an endogenous gene or by introducing a transgene via any gene delivery method or system.

[0210] Unless otherwise indicated, methods for reprogramming somatic cells can be performed both in vivo and in vitro (wherein in vivo is performed when the somatic cells are present within a subject, and in vitro is performed using isolated somatic cells maintained in culture).

[0211] Embryonic cells, such as embryonic stem cells, can be cells derived from an embryonic cell line and not directly derived from an embryo or fetus. Alternatively, embryonic cells can be cells derived from an embryo or fetus, but the cells are obtained or isolated without disrupting or having any negative effect on the development of the embryo or fetus.

[0212] Differentiated somatic cells, including cells from humans, including fetal, neonatal, juvenile, or adult primates, are suitable source cells for the methods of the present invention. Suitable somatic cells include, but are not limited to, bone marrow cells, epithelial cells, endothelial cells, fibroblasts, hematopoietic cells, keratinocytes, liver cells, intestinal cells, mesenchymal cells, myeloid progenitor cells, and spleen cells. Alternatively, somatic cells can be cells capable of self-proliferation and differentiation into other cell types, including blood stem cells, muscle / bone stem cells, brain stem cells, and liver stem cells. Suitable somatic cells are competent to incorporate transcription factors, including genetic material encoding the transcription factors, or can be made competent using methods commonly known in the scientific literature. Methods for improving uptake can vary depending on the cell type and expression system. Exemplary conditions used to prepare recipient somatic cells with suitable transduction efficiency are well known to those of skill in the art. Starting somatic cells can have a doubling time of approximately 24 hours.

[0213] As used herein, the term "isolated cell" refers to a cell that has been removed from the organism in which it was originally found, or the progeny of such a cell. Optionally, the cell has been cultured in vitro, for example, in the presence of other cells. Optionally, the cell is later introduced into a second organism, or reintroduced into the organism from which it was isolated (or from which it is a descendant).

[0214] As used herein, the term "isolated population," in reference to an isolated cell population, refers to a population of cells that have been removed and separated from a mixed or heterogeneous population of cells. In some embodiments, an isolated population is a substantially pure population of cells compared to the heterogeneous population from which the cells are isolated or enriched.

[0215] The term "substantially pure" in reference to a particular cell population refers to a population of cells that is at least about 75%, preferably at least about 85%, more preferably at least about 90%, and most preferably at least about 95% pure, with respect to the cells that make up the total cell population. In other words, the term "substantially pure" or "essentially purified" in reference to a population of target cells refers to an approximately Less than 20%, more preferably less than about 15%, 10%, 8%, 7%, and most preferably less than about 5%, 4%, 3%, 2%, 1%, or less than 1% refers to a cell population that is not a target cell or their progeny as defined by that term herein.

[0216] As used herein, the term "cancer" refers to cells capable of autonomous growth, i.e., an abnormal state or pathology characterized by rapidly proliferating cell growth. The term is intended to include all types of cancerous growths or oncogenic processes, metastatic tissues, or malignantly cancerous cells, tissues, or organs, regardless of histopathological type or stage of invasiveness. The term "cancer" includes malignant tumors of various organ systems, such as those affecting the lung, breast, thyroid, lymphatic system, gastrointestinal, and genitourinary tract, as well as adenocarcinomas, including most colon cancers, renal cell carcinoma, prostate cancer, and / or testicular tumors, non-small cell carcinoma of the lung, small intestine cancer, and esophageal cancer. The term "carcinoma" is recognized in the art and refers to malignant tumors of epithelial or endocrine tissues, including respiratory system carcinoma, gastrointestinal system carcinoma, genitourinary system carcinoma, testicular carcinoma, breast carcinoma, prostate carcinoma, endocrine system carcinoma, and melanoma. Exemplary carcinomas include those forming from tissue of the cervix, lung, prostate, breast, head and neck, colon, and ovary. The term "carcinoma" also includes carcinosarcomas, which include, for example, malignant tumors composed of carcinomatous and sarcomatous tissue. "Adenocarcinoma" refers to a carcinoma derived from glandular tissue or in which the tumor cells form recognizable glandular structures. The term "sarcoma" is recognized in the art as and refers to a malignant tumor of mesenchymal derivation.

[0217] As used herein, a reference to a "target cell" may be a reference to any one or more of the cells referred to herein as target cells or target cell types, such as those in the top row of Table 4.

[0218] A source cell is determined to have been converted into a target cell or become a target-like cell by the methods of the present invention if the source cell exhibits at least one characteristic of the target cell type. For example, a human fibroblast would be confirmed to have been converted into a keratinocyte-like cell if the cell exhibits at least one characteristic of the target cell type. Typically, the cell will exhibit one, two, three, four, five, six, seven, eight, or more characteristics of the target cell type. For example, if the target cell is a keratinocyte, the cell is confirmed or determined to be a keratinocyte-like cell if upregulation of any one or more keratinocyte markers and / or a change in cell morphology is detectable. Preferably, the keratinocyte markers include keratin 1, keratin 14, and involucrin, and the cell morphology exhibits a cobblestone appearance. In any embodiment of the present invention, the characteristics of the target cell can be determined by analysis of cell morphology, gene expression profile, activity assay, protein expression profile, surface marker profile, or differentiation potential. Examples of characteristics or markers include those described herein and known to those of skill in the art. Other examples of relevant markers include, for example, in the case of conversion of keratinocytes to hematopoietic stem cells (HSCs): CD45 (pan hematopoietic marker), CD19 / 20 (B cell marker), CD14 / 15 (myeloid), CD34 (progenitor / SC marker), CD90 (SC), and α-integrin (keratinocyte marker not expressed by HSCs); in the case of conversion of human embryonic stem cells to hematopoietic stem cells: Runx1 (GFP), CD45 (pan hematopoietic marker), CD19 / 20 (B cell marker), CD14 / 15 (myeloid), CD34 (progenitor / SC marker). Examples of markers for many of the transformations described herein include CD40 (a marker for ESCs), CD90 (SCs), and Tra-1-160 (an ESC marker not expressed in HSCs); and for blastogenesis of aged or adult HSCs: comparison between the transcriptional signatures of young and aged human HSCs (e.g., using RNA-seq), and functional characterization of "blastogenic HSCs" by assessing blastogenic cells 1, 3, and 6 months after transplantation into animals to determine myeloid bias (where loss of myeloid bias indicates "blastogenic" HSCs). Examples of markers for many of the transformations described herein are shown in Table 1 below.

[0219] [Table 1]

[0220] Transcription factors cited herein are cited by their HUGO Gene Nomenclature Committee (HGNC) Symbol. Exemplary nucleotide sequences for each transcription factor are provided in Tables 2 and 3 below. The nucleotide sequences are derived from the Ensembl database (Flicek et al. (2014). Nucleic Acid Research Volume 42, Issue D1. Pp. D749-D755) version 83. Any homologs, orthologs, or paralogs of the transcription factors cited herein are also contemplated for use in the present invention.

[0221] Table 2a and b below: Ensembl gene accession numbers (nucleotide sequences; NT-seq) and exemplary transcription factors (TFs) that can be used according to the methods described herein. The source cell type is shown in the left-most column, and the target cell type is shown in the top row. Transcription factors that can be used to convert the source cell type into a cell having at least one characteristic of the target cell type are shown.

[0222] [Table 2]

[0223] [Table 3]

[0224] [Table 4]

[0225] [Table 5]

[0226] [Table 6]

[0227] The term "variant" refers to a polypeptide that is at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identical to the full-length polypeptide. The present invention contemplates the use of variants of the transcription factors described herein, including the sequences listed in Tables 2a and 2b. Variants may be fragments of the full-length polypeptides or naturally occurring splice variants. A variant may be a polypeptide at least 70%, 80%, 85%, 90%, 95%, 98%, or 99% identical to a fragment of the polypeptide, where the fragment has a length at least 50%, 60%, 70%, 80%, 85%, 90%, 95%, 98%, or 99% of the full length of the wild-type polypeptide, or a domain thereof that has a functional activity of interest, such as the ability to promote conversion of a source cell type to a target cell type. In some embodiments, the domain is at least 100, 200, 300, or 400 amino acids in length, starting at any amino acid position in the sequence and extending toward the C-terminus. It is preferable to avoid modifications known in the art that eliminate or substantially reduce protein activity. In some embodiments, the variant lacks N- and / or C-terminal portions of the full-length polypeptide, e.g., up to 10, 20, or 50 amino acids from either end. In some embodiments, a polypeptide refers to a polypeptide having the sequence of a mature (full-length) polypeptide, which has had one or more portions, such as a signal peptide, removed during normal intracellular proteolytic processing (e.g., during co-translational or post-translational processing). In some embodiments, where a protein is produced by a method other than purifying the protein from cells that naturally express the protein, the protein is a chimeric polypeptide, meaning that the protein contains portions from two or more different species. In some embodiments, where a protein is produced by a method other than purifying the protein from cells that naturally express the protein, the protein is a derivative, meaning that the protein contains additional sequence not related to the protein, so long as the sequence does not substantially reduce the biological activity of the protein. Those skilled in the art will know, or can readily ascertain, whether a particular polypeptide variant, fragment, or derivative is functional using assays known in the art. For example, the ability of a transcription factor variant to convert a source cell to a target cell type can be assessed using the assays disclosed in the Examples herein.Another convenient assay involves measuring the ability to activate transcription of a reporter construct containing a transcription factor binding site operably linked to a nucleic acid sequence encoding a detectable marker, such as luciferase. In certain embodiments of the invention, functional variants or fragments have at least 50%, 60%, 70%, 80%, 90%, 95% or more of the activity of the full-length wild-type polypeptide.

[0228] The term "increasing the amount of" in relation to increasing the amount of a transcription factor refers to increasing the amount of the transcription factor in a cell of interest (e.g., a source cell such as a fibroblast or keratinocyte). In some embodiments, the amount of a transcription factor is "increased" in a cell of interest (e.g., a cell into which an expression cassette directing the expression of one or more transcription factor-encoding polynucleotides has been introduced) if the amount of the transcription factor is at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or more greater than in a subject (e.g., a fibroblast or keratinocyte into which none of the expression cassettes described below have been introduced). However, any method of increasing the amount of a transcription factor is contemplated, including any method of increasing the amount, rate, or efficiency of transcription, translation, stability, or activity of the transcription factor (or pre-mRNA or mRNA encoding it). Additionally, downregulating or interfering with negative regulators of transcription expression to increase the efficiency of existing transcription (e.g., SINEUP) is also contemplated.

[0229] The term "agent," as used herein, means any compound or substance, including, but not limited to, a small molecule, a nucleic acid, a polypeptide, a peptide, a drug, an ion, etc. An "agent" can be any chemical compound, entity, or moiety, including, but not limited to, synthetic and naturally occurring proteinaceous and non-proteinaceous entities. In some embodiments, an agent is a nucleic acid, a nucleic acid analog, a protein, an antibody, a peptide, an aptamer, an oligomer of nucleic acid, an amino acid, or a carbohydrate, including, but not limited to, a protein, an oligonucleotide, a ribozyme, a DNAzyme, a glycoprotein, an siRNA, a lipoprotein, an aptamer, modifications and combinations thereof, etc. In certain embodiments, the agent is a small molecule having a chemical moiety. For example, the chemical moiety includes an unsubstituted or substituted alkyl, aromatic, or heterocyclyl moiety, including macrolides, leptomycin, and its related natural products or analogs. The compound may be known to have the desired activity and / or properties, or may be selected from a library of diverse compounds.

[0230] The term "exogenous," when used with reference to a protein, gene, nucleic acid, or polynucleotide within a cell or organism, refers to a protein, gene, nucleic acid, or polynucleotide that has been introduced into the cell or organism by artificial or natural means; or, when used with reference to a cell, refers to a cell that has been isolated and then introduced into another cell or organism by artificial or natural means. An exogenous nucleic acid may be from a different organism or cell, or may be one or more additional copies of a nucleic acid that occurs naturally within the organism or cell. An exogenous cell may be from a different organism or from the same organism. As a non-limiting example, an exogenous nucleic acid is one that is in a chromosomal location that is different from that of the native cell, or is flanked by nucleic acid sequences that are different from those found in nature. An exogenous nucleic acid may also be extrachromosomal, such as an episomal vector.

[0231] Screening one or more candidate agents for their ability to increase the amount of one or more transcription factors required for conversion of a source cell type to a target cell type can include contacting a system that allows for the production or expression of the transcription factor with the candidate agent and determining whether the amount of the transcription factor is increased. The system can be in vivo, e.g., a tissue or cell within an organism, or in vitro, e.g., a cell isolated from an organism, or an in vitro transcription assay, or ex vivo in a cell or tissue. The amount of the transcription factor can be measured directly or indirectly by determining the amount of protein or RNA (e.g., mRNA or pre-mRNA). The candidate agent functions to increase the amount of the transcription factor by increasing any step in the transcription of the gene encoding the transcription factor or by increasing translation of the corresponding mRNA. Alternatively, the candidate agent can reduce the inhibitory activity of a repressor of transcription of the gene encoding the transcription factor, or the activity of a molecule that leads to degradation of the mRNA encoding the transcription factor or the transcription factor protein itself.

[0232] Suitable detection means include the use of labels such as radionucleotides, enzymes, coenzymes, fluorescers, chemiluminescers, chromogens, enzyme substrates or cofactors, enzyme inhibitors, prosthetic group complexes, free radicals, particles, dyes, and the like. Such labeled reagents may be used in a variety of well-known assays, e.g., radioimmunoassays, enzyme immunoassays such as ELISAs, fluorescence immunoassays, etc. See, e.g., U.S. Patent Nos. 3,766,162; 3,791,932; 3,817,837; and 4,233,402.

[0233] The methods of the present invention include high-throughput screening applications. For example, high-throughput screening assays, including any of the assays according to the present invention, can be used, in which an aliquot of a system allowing for the production or expression of a transcription factor is exposed to multiple candidate agents in different wells of a multi-well plate. Furthermore, high-throughput screening assays according to the present disclosure include aliquots of a system allowing for the production or expression of a transcription factor exposed to multiple candidate agents in different wells of a multi-well plate in any type of miniaturized assay system.

[0234] The disclosed methods may be "miniaturized" in assay systems by any acceptable method of miniaturization, including but not limited to multi-well plates, microchips, or slides, such as 24, 48, 96, or 384 wells per plate. Assays are advantageously reduced in size to be performed on microchip supports containing smaller amounts of reagents and other materials. Any miniaturization of the process that is conducive to high throughput screening is within the scope of the present invention.

[0235] In any of the methods of the invention, the target cells can be transferred to the same mammal from which the source cells were obtained. In other words, the source cells used in the methods of the invention can be autologous, i.e., obtained from the same individual to whom the target cells are to be administered. Alternatively, the target cells can be allogeneically transferred to another individual. Preferably, the cells are autologous to the subject in the method of treatment or prevention of the individual's medical condition.

[0236] The term "cell culture medium" (also referred to herein as "culture medium" or "culture medium"), as referred to herein, is a medium for culturing cells, containing nutrients that maintain cell viability and support growth. Cell culture media can include any of the following in appropriate combinations: salts, buffers, amino acids, glucose or other sugars, antibiotics, serum or serum replacement, and other components such as peptide growth factors. Cell culture media commonly used for particular cell types are known to those of skill in the art. Exemplary cell culture media for use in the methods of the present invention are shown in Table 9.

[0237] [Table 7]

[0238] [Table 8]

[0239] [Table 9]

[0240] The nucleic acids described herein, or vectors containing the nucleic acids, can include one or more of the sequences cited in Table 3 above, or a sequence encoding any one or more of the amino acid sequences listed in Table 3.

[0241] The term "expression" refers to the cellular processes involved in the production of RNA and proteins, and, where appropriate, protein secretion, including, but not limited to, transcription, translation, folding, modification, and processing, as applicable.

[0242] The terms "isolated" or "partially purified," as used herein, refer to a nucleic acid or polypeptide that has been separated from at least one other component (e.g., nucleic acid or polypeptide) that is present with the nucleic acid or polypeptide as found in its natural source, and / or that would be present with the nucleic acid or polypeptide when expressed or, in the case of a secreted polypeptide, when secreted by a cell. Chemically synthesized nucleic acids or polypeptides or those synthesized using in vitro transcription / translation are considered "isolated."

[0243] The term "vector" refers to a carrier DNA molecule into which a DNA sequence can be inserted for introduction into a host or source cell. Preferred vectors are capable of autonomous replication and / or expression of nucleic acids to which they are linked. Vectors capable of directing the expression of an operably linked gene are referred to herein as "expression vectors." Thus, an "expression vector" is a specialized vector that contains the necessary regulatory regions required for the expression of a gene of interest in a host cell. In some embodiments, the gene of interest is operably linked to other sequences within the vector. Vectors can be viral or non-viral. When a viral vector is used, the viral vector is preferably replication-deficient, which can be achieved by removing all viral nucleic acid encoded for replication. Replication-deficient viral vectors remain infectious and enter cells in a manner similar to replicating adenoviral vectors, but do not replicate or propagate after entering the cell. Vectors also encompass liposomes and nanoparticles, as well as other means of delivering DNA molecules to cells.

[0244] The term "operably linked" refers to regulatory sequences required for expression of a coding sequence that are positioned within a DNA molecule in the appropriate position relative to the coding sequence to effect expression of the coding sequence. This same definition is sometimes applied to the arrangement of coding sequences and transcription control elements (e.g., promoters, enhancers, and termination elements) within an expression vector. The term "operably linked" includes having an appropriate initiation signal (e.g., ATG) in front of the polynucleotide sequence to be expressed, and maintaining the correct reading frame to permit expression of the polynucleotide sequence under the control of the expression control sequences and production of the desired polypeptide encoded by the polynucleotide sequence.

[0245] The term "viral vector" refers to the use of a virus or virus-associated vector as a carrier of a nucleic acid construct into a cell. The construct may be integrated into and packaged into a non-replicating, defective viral genome, such as adenovirus, adeno-associated virus (AAV), or herpes simplex virus (HSV), or others, including retroviral and lentiviral vectors, for infection or transduction into the cell. The vector may or may not integrate into the genome of the cell. The construct may also include viral sequences for transfection, if desired. Alternatively, the construct may be incorporated into a vector capable of episomal replication, such as EPV and EBV vectors.

[0246] As used herein, the term "adenovirus" refers to a virus of the Adenoviridae family. Adenoviruses are medium-sized (90-100 nm), non-enveloped (naked), icosahedral viruses composed of a nucleocapsid and a double-stranded linear DNA genome.

[0247] As used herein, the term "non-integrating viral vector" refers to a viral vector that does not integrate into the host genome; expression of genes delivered by the viral vector is transient. Because there is little to no integration into the host genome, non-integrating viral vectors have the advantage of not causing DNA mutations due to insertion at random locations in the genome. For example, non-integrating viral vectors remain extrachromosomal and do not insert their genes into the host genome, which could disrupt the expression of endogenous genes. Non-integrating viral vectors can include, but are not limited to, adenoviruses, alphaviruses, picornaviruses, and vaccinia viruses. These viral vectors are "non-integrating" viral vectors as the term is used herein, even though any of them may, in certain rare circumstances, integrate viral nucleic acid into the genome of a host cell. Importantly, the viral vectors used in the methods described herein do not integrate their nucleic acid into the genome of a host cell, generally, or as a major part of their life cycle under the conditions used.

[0248] The vectors described herein can be constructed and engineered using methods generally known in the scientific literature to enhance their safety for use in therapy, to include selection and enrichment markers, if desired, and to optimize expression of the nucleotide sequences they contain. Vectors must contain structural components that allow them to self-replicate within the source cell type. For example, the known Epstein-Barr oriP / Nuclear Antigen-1 (EBNA-I) combination (see, e.g., Lindner, SE and, incorporated by reference in its entirety as if set forth herein) can be used to generate vectors that can be used in therapeutic applications. (See B. Sugden, The plasmid replicon of Epstein-Barr virus: mechanistic insights into efficient, licensed, extrachromosomal replication in human cells, Plasmid 58:1 (2007)) is sufficient to support the autonomous replication of the vector, and other combinations known to function in mammalian, particularly primate, cells may also be used. Standard techniques for constructing suitable expression vectors are well known to those skilled in the art and can be found in publications such as Sambrook J, et al., "Molecular cloning: a laboratory manual," (3rd ed. Cold Spring Harbor Press, Cold Spring Harbor, NY 2001), which is incorporated by reference in its entirety as if set forth herein.

[0249] In the methods of the invention, genetic material encoding the relevant transcription factors required for conversion is delivered into the source cells by one or more reprogramming vectors. Each transcription factor may be introduced into the source cells as a polynucleotide transgene encoding the transcription factor operably linked to a heterologous promoter capable of driving expression of the polynucleotide in the source cells.

[0250] Suitable reprogramming vectors are any of those described herein, including episomal vectors such as plasmids that do not encode all or part of a viral genome sufficient to generate infectious or replication-competent virus, although vectors may contain structural elements derived from more than one virus. One or more reprogramming vectors can be introduced into a single source cell. One or more transgenes can be provided on a single reprogramming vector. A single strong constitutive transcription promoter can provide transcriptional control for multiple transgenes, which can be provided as expression cassettes. Separate expression cassettes on a vector can be under the transcriptional control of separate strong constitutive promoters, which may be copies of the same promoter or different promoters. A variety of heterologous promoters are known in the art and can be used depending on factors such as the desired expression level of transcription factors. As exemplified below, it may be advantageous to use different promoters with different strengths to control the transcription of separate expression cassettes in the source cell. Another consideration in selecting a transcription promoter is the rate at which the promoter is silenced. Those skilled in the art will recognize that it may be advantageous to reduce expression of one or more transgenes or transgene expression cassettes after the gene's product has completed or substantially completed its role in the reprogramming method. Exemplary promoters are the human EF1α elongation factor promoter, the CMV cytomegalovirus immediate early promoter, and the CAG chicken albumin promoter, as well as corresponding homologous promoters from other species. In human somatic cells, both EF1α and CMV are strong promoters, but the CMV promoter is silenced more efficiently than the EF1α promoter, and thus transgene expression under its control is silenced earlier than transgenes under its control. Transcription factors can be expressed in source cells in relative ratios that can be varied to regulate reprogramming efficiency. Preferably, when multiple transgenes are encoded on a single transcript, an internal ribosome entry site is provided upstream of the transgenes in the direction of the transcription promoter.The relative ratios of factors may vary depending on the factors being delivered, but one of skill in the art in possession of this disclosure will be able to determine the optimal ratios of factors.

[0251] Those skilled in the art will recognize that the advantageous efficiency of introducing all factors via a single vector rather than multiple vectors increases the difficulty of vector introduction as the total vector size increases. Those skilled in the art will also recognize that the location of transcription factors on the vector can affect their temporal expression and the resulting reprogramming efficiency. Therefore, applicants have used various combinations of factors in combination with vectors. Several such combinations are shown herein to support reprogramming.

[0252] After introduction of the reprogramming vector, and while the source cell is being reprogrammed, the vector may persist in the target cell while the introduced transgene is transcribed and translated. Transgene expression is advantageously maintained in cells reprogrammed to the target cell type. The reprogramming vector may be downregulated or shut down. The reprogramming vector may remain extrachromosomal. In very low efficiency cases, the vector may integrate into the cell's genome. The following examples are intended to illustrate and in no way limit the invention.

[0253] Methods suitable for nucleic acid delivery for transformation of cells, tissues, or organisms as used herein are intended to include virtually any method described herein or known to those of skill in the art by which nucleic acid (e.g., DNA) can be introduced into a cell, tissue, or organism (e.g., Stadtfeld, and Hochedlinger, Nature Methods 6(5):329-330 (2009); Yusa et al., Nat. Methods 6:363-369 (2009); Woltjen, et al., Nature 458,766-770 (9 Apr. 2009)). Such methods include, but are not limited to, direct delivery of DNA, such as by ex vivo transfection (Wilson et al., Science, 244:1344-1346, 1989; Nabel and Baltimore, Nature 326:711-713, 1987), optionally using lipid-based transfection reagents such as Fugene 6 (Roche) or Lipofectamine (Invitrogen), microinjection (Harland and Weintraub, J. Cell 2009, 1989, incorporated herein by reference). Biol., 101:1094-1099, 1985; U.S. Patent No. 5,789,215), including by injection (U.S. Patent Nos. 5,994,624, 5,981,274, 5,945,100, 5,780,448, 5,736,524, 5,702,932, 5,656,610, 5,589,466, and 5,580,859, each of which is incorporated herein by reference); by electroporation (U.S. Patent No. 5,384,253, incorporated herein by reference; Tur-Kaspa et al., Mol. Cell Biol., 6:716-718, 1986; Potter et al., Proc. Nat'l Acad. Sci. USA, 81:7161-7165, 1984); by calcium phosphate precipitation (Graham and Van Der Eb, Virology, 52:456-467, 1973; Chen and Okayama, Mol. Cell Biol., 7(8):2745-2752, 1987; Rippe et al., Mol. Cell Biol., 10:689-695, 1990); by using DEAE-dextran followed by polyethylene glycol (Gopal, Mol. Cell Biol., 5:1188-1190, 1985); by direct sonic loading (Fechheimer et al., Proc. Nat'l Acad. Sci. USA, 84:8463-8467, 1987); by liposome-mediated transfection (Nicolau and Sene, Biochim. Biophys. Acta, 721:185-190, 1982; Fraley et al., Proc. Nat'l Acad. Sci. USA, 76:3348-3352, 1979; Nicolau et al., Methods Enzymol., 149:157-176, 1987; Wong et al., Gene, 10:87-94, 1980; Kaneda et al., Science, 243:375-378, 1989; Kato et al. al., J. Biol. Chem., 266:3361-3364, 1991) and receptor-mediated transfection (Wu and Wu, Biochemistry, 27:887-892, 1988; Wu and Wu, J. Biol. Chem., 262:4429-4432, 1987); and any combination of these methods, each of which is incorporated herein by reference.

[0254] A number of polypeptides capable of mediating the introduction of relevant molecules into cells have been previously described and can be applied to the present invention. See, e.g., Langel (2002) Cell Penetrating Peptides: Processes and Applications tions, CRC Press, Pharmacology and Toxicology Series. Examples of polypeptide sequences that improve transport across membranes include, but are not limited to, the Drosophila homeoprotein Antennapedia transcription protein (AntHD) (Joliot et al., New Biol. 3:1121-34, 1991; Joliot et al., Proc. Natl. Acad. Sci. USA, 88:1864-8, 1991; Le Roux et al., Proc. Natl. Acad. Sci. USA, 90:9120-4, 1993), the herpes simplex virus structural protein VP22 (Elliott and O'Hare, Cell 88:223-33, 1997); the HIV-1 transcriptional activator TAT protein (Green and Loewenstein, Cell 55:1179-1188, 1988; Frankel and Pabo, Cell 55:1 289-1193, 1988); Kaposi FGF signal sequence (kFGF); protein transduction domain-4 (PTD4); penetratin, M918, transportan-10; nuclear localization sequences, PEP-I peptide; amphipathic peptides (e.g., MPG peptide); delivery-enhancing transporters such as those described in U.S. Pat. No. 6,730,293 (including, but not limited to, peptide sequences containing at least 5-25 or more consecutive arginines or 5-25 or more arginines in a contiguous set of 30, 40, or 50 amino acids; including, but not limited to, peptides having sufficient, e.g., at least 5, guanidino or amidino moieties); and the commercially available Penetratin™ 1 peptide, and Diatos Peptide Vectors (“DPVs”) of the Vectocell® platform available from Daitos SA, Paris, France. See also WO 2005 / 084158 and WO 2007 / 123667 and the additional transporters described therein. Not only are these proteins able to cross the plasma membrane, but the binding of other proteins, such as the transcription factors described herein, is sufficient to stimulate cellular uptake of these complexes.

[0255] Table 4a. Exemplary conversions of the present invention. Source cell types are shown in the leftmost column and target cell types are shown in the top row. Transcription factors required for conversion of a source cell type to a cell having at least one characteristic of the target cell type are shown.

[0256] [Table 10]

[0257] [Table 11]

[0258] Table 4b. Further exemplary conversions of the present invention. Source cell types are shown in the leftmost column and target cell types are shown in the top row. Transcription factors required for conversion of a source cell type to a cell having at least one characteristic of the target cell type are shown.

[0259] [Table 12]

[0260] The present invention includes the following non-limiting examples. [Example]

[0261] Example 1 To predict the set of TFs required for each cell transformation, we identify TFs that are not only differentially expressed between cell types but also exert regulatory influence on other differentially expressed genes in the local network (see Figure 1a). A single score capturing differential expression across all genes in all cells is defined by combining the log fold change and adjusted p-value. The regulatory impact of each TF in each cell is calculated by performing a weighted sum of the differential expression scores across the known interactome (defined by STRING and MARA, see Figure 1c). This sum is weighted by two factors: (1) the directness of regulation, i.e., how many intermediate interactions there are between the TF and downstream genes, and (2) the specificity, i.e., the number of other genes that upstream TFs also regulate. This weighted sum ranks TFs within each cell type according to their influence. The final step is to select the optimal set of TFs with the greatest combined influence across differentially expressed genes in the target cell type compared to the source. This is done by adding TFs to the set in ranked order by differential impact, excluding those whose combined impact does not increase the impact of the set, until it reaches 98% of the expressed target cell genes (see Figure 1d and methodology below). Biologically speaking, Mogrify identifies TFs that control the parts of the regulatory network that are most responsible for the uniqueness of the target cell type.

[0262] Mogrify may include one or more steps, which are outlined below and described in more detail in the following sections.

[0263] 1. Collect expression data for each gene (x) in each sample (s). 2. Calculate the tree-based differential expression of each gene in each sample relative to background, followed by the log fold change

number

number

number

number

number

[0264] Step 1: Expression data obtained from the FANTOM5 dataset Mogrify uses 700 libraries of clustered CAGE tags, which provide TSS positions. These are mapped to their corresponding genes (provided by the FANTOM5 consortium (Forrest, A.R. et al. Nature 507, 462-470 (2014)). This data is used to generate tag counts for each gene in each library. In total, all genes expressed at least 20 TPM (tags per million) in at least one sample are included. There are 5,878 different genes (1408 of which are TFs).

[0265] Step 2: Tree-based differential expression Calculating differential expression is a common problem when analyzing biological data, and several techniques exist for doing so. We chose to use DESeq for this work because it performs well in benchmark evaluations, allows for the analysis of several non-repeated datasets, and has a short runtime. To calculate differential expression, two groups must be identified: the set of samples for which differential expression is desired and the background to which it is compared. The issue of selecting the correct background is important. Too many unrelated samples can reduce the statistical power of the test. Too few or too few samples in the background makes it impossible to distinguish which genes are actually differentially expressed. One solution is to perform an exhaustive calculation of pairwise tests between each of the cell types. This approach has two problems: first, it is very computationally expensive, and second, it does not reveal differentially expressed genes between a sample and the average background, but rather between the two samples in detail. In the case of Mogrify, we are interested in genes that are important for a given cell type in all circumstances, and therefore oppose sample collection. To do this, we implemented a tree-based background selection method based on the FANTOM5 cell ontology (Figure 1B). The principle of this approach is to exclude cell types whose ontology is very similar, while including other cell types that are close to the background in the tree. This was achieved by choosing a point near the top of the tree to serve as a cutoff point. We removed samples that were in the same clade as the analyzed cell type, and included samples that were not in the same clade but still below this point. The result was a set of samples that was broad enough to provide robust results, yet narrow enough to maintain a manageable level of statistical power.

[0266] This tree-based background selection for DEseq is performed on the entire FANTOM5 library (sorted by replicates) generating the log fold change and FDR-adjusted p-value for each gene in each sample. Due to the presence of a non-uniform background, the results of each differential expression calculation cannot be directly compared; therefore, in the remaining steps, these numbers are used to rank the genes in each sample, and it is the rankings that are compared.

[0267] Because we are only interested in identifying TFs with high levels of influence, we used the following equation to calculate the log fold change and FDR-adjusted p-value as a single positive score:

number

number

number

number

[0268] This formula gives very high guarantees to genes with high log fold change and low adjusted p-value scores, and vice versa. This is applied to all genes in each sample, forming a 15,878 gene matrix of differential expression with 700 samples.

[0269] Step 3: Calculating the network-based influence range of TFs To assess the importance of each TF, its effect on the TF's local neighborhood is calculated using two sources of network information: the STRING database and Motif Activity Response Analysis (MARA). These two techniques, described below, involve different types of interactions. MARA provides protein-DNA interactions with known binding sites within the promoter regions of genes. This represents a low-level, directed regulatory network of interactions. STRING is a meta-database of interactions, including various types of interactions, including protein-protein, protein-DNA, protein-RNA, and biological pathways. This provides an overview of the interactions that take place to affect gene expression both directly and indirectly.

[0270] To calculate influence, a weighted sum of gene influences (from step 2) is performed over the local network neighborhood of the transcription factor. This local network is constrained to a maximum of three edges, and the effect of each node decreases with distance from the seed TF where it is located, depending on its out-degree from its parent (Figure 1C). Distance weighting is used to ensure that genes increasingly distant from direct regulation have relatively little influence on the score. Edge weighting is used to compensate for highly ubiquitous transcription factors and prevent them from receiving artificially high scores due to regulation of a large number of marginally differentially expressed genes. We

number

number

[0271] The equation for this weighted sum is:

number

[0272] We do this across both the MARA and STRING networks, and generate two TF-influence lists.

number

[0273] Step 4: Rank the TFs based on the results of steps 2 and 3 The results of steps 3 and 4 are:

number

[0274] Step 5: Compute all pairwise experimental comparisons to form predictions To predict the set of TFs for a given transformation, the ranked lists from the source and target cell types are compared: If a TF from the target cell type list is already expressed in the source target (>20 TPM), it is removed from the list.

[0275] Step 6: Removal of transcriptionally redundant TFs After the final ranking is completed, regulatory redundancy is removed. This is achieved by comparing the list of genes directly regulated by each TF. For a given TF, if there is a higher-ranked TF that regulates more than 98% of the genes it regulates, it is removed. This means that the resulting predictions include TFs that are diverse in their regulatory sphere of influence. This cutoff was empirically chosen to minimize the number of predicted factors while maximizing network coverage (Figure 5).

[0276] Step 7: Forming a cellular reprogramming landscape based on steps 1–6 To create the reprogramming landscape, we calculated the X and Y coordinates independently of the Z coordinate. To reduce the complexity of the landscape, we averaged the gene expression profiles of individual samples classified by the cell ontology provided by FANTOM5. The result is a set of 314 ontologies containing at least three samples, from which we obtain the average gene expression. The X and Y coordinates were calculated by performing multidimensional scaling (MDS) of these profiles. The result of MDS is a projection of the data in which the distance between points is maintained in a two-dimensional reduction from the multidimensional reality. As a result, two points close to each other in the XY plane of the landscape have similar expression profiles and therefore represent similar cell types. The Z-axis of the landscape was calculated by considering the regulatory coverage of the top eight Mogrify-predicted TFs. For the entire transformation, we focused on the set of genes expressed in the ontology, some of which are directly regulated by each TF. We calculated the area under the curve of cumulative coverage for the top 8 TFs, normalized by the maximum possible AUC, and retrieved a height value between 0 and 1 for each ontology. Thus, a height of 1 represents an ontology in which all of the required genes are directly regulated by the top-ranked TF, whereas a height of 0 represents an ontology in which none of the top 8 TFs directly regulate any of the required genes. The X, Y, and Z values ​​were then used in the R package plot3D to generate a landscape using the image2D and persp3D packages. The highest-ranking distinct stem cell was found with a gene set enrichment score of 0.41 and a p-value of 0.011.

[0277] Example 2 To evaluate Mogrify's predictive power, we first focused on those involving human cells and determined how it performed against well-known, previously published direct cell transformations. These should not be considered absolute, complete combinations, but rather a reference point of positive examples useful for comparison. As shown in Figure 2, in almost all cases, Mogrify predicts the complete set of TFs previously shown to function, but may include upstream TFs in place of published factors. For example, it is known that human fibroblasts can be converted into iPS cells by introducing OCT4 (also known as POU5F1), SOX2, KLF4, and MYC, or OCT4, SOX2, NANOG, and LIN28. Mogrify predicted NANOG, OCT4, and SOX2 as the top three TFs for this transformation, a combination that has been experimentally validated. Previous work has shown that the conversion of B cells and fibroblasts into macrophage-like cells was possible through the expression of CEBPa and PU.1 (also known as SPI1) (Xie, H., Ye, M., Feng, R. & Graf, T. Cell 117, 663-676 (2004); Rapino, F. et al. Cell Rep. 3, 1153-63 (2013)), which is perfectly predicted by Mogrify. For the conversion of human dermal fibroblasts into cardiomyocytes, we chose not to use the FANTOM5 set of data because it lacks many important cardiomyocyte genes (indicating deficiencies in the origin of the sample). Despite the heterogeneity of the cells and the use of a less-than-ideal cardiac sample, the Mogrify prediction list includes four of the five TFs (or closely related factors) used in human conversion (Fu, J.-D. et al. Stem cell reports 1, 235-47 (2013)). There are numerous reports in the literature of transdifferentiation of various cell types into neurons in both mice and humans (Table 5).

[0278] [Table 13]

[0279] Although the set of TFs used varies, likely due to the heterogeneity and complexity of neurons, factors common to all experiments are predicted by Mogrify (Table 6).

[0280] [Table 14]

[0281] Finally, between human fibroblasts and hepatocytes, Mogrify predicts highly similar TF combinations required for transformation and maturation (Figure 2). Using the transformations shown in Figure 2, we evaluated the ability of Mogrify, CellNet, and the entropy-based approach from D'Alessio et al. (Stem Cell Reports, Volume 5, Issue 5, 10 November 2015, Pages 763-775) to recover these known factors. The average recovery of published transcription factors for Mogrify was 84%, compared with 31% for CellNet and 51% for D'Alessio et al. (Figure 6). In six of the ten transformations in Figure 2, Mogrify recovered 100% of the required TFs, indicating that if Mogrify had been used to provide the TF sets for these transformations, the experiment would have been successful from the start. In contrast, CellNet and D'Allesio et al. recovered all factors for only 1 in 10 transformations.

[0282] While Mogrify's primary goal is to predict TFs for cellular transformations, it also maps the landscape of human cell types in terms of naturally occurring states and the transitions between them, capturing a core control set of TFs that describe each individual cell type. This may in itself help researchers clarify the roles of different TFs in their preferred cell types. Indeed, Mogrify offers a significant advantage over current laboratory strategies for cell reprogramming, helping to predict TFs whose overexpression will induce directed cellular transformations. Mogrify has been pre-calculated for transformations between all possible combinations of 307 FANTOM5 tissues / cell types, yielding 93,942 directed transformations. Mogrify can be applied to numerous other cell types not included in FANTOM5, provided their expression signatures (e.g., RNAseq or CAGE) are known. Mogrify provides a starting point and a systematic means to explore novel transformations in humans. Because Mogrify incorporates a TF redundancy step, it can provide a finite set of TFs as predictors for cellular transformations, which is more useful than simply ranking all TFs.

[0283] To compare the performance of Mogrify with other methods, benchmarking experiments were performed. First, the effect on performance of using the full Mogrify algorithm was evaluated compared to using only MARA, STRING, and differential expression components. Next, a comparison was made with CellNet and D'Allesio et al., which are currently the only other techniques that provide a means to calculate transcription factor sets for a wide variety of cell types. To perform the comparison, the set of transcription factors from a published transduction shown in Figure 2 was used as true positives. The benchmark consisted of evaluating the performance of each technique in recovering these TFs using the following steps:

[0284] 1) For each transition, we identify the number of transcription factors considered: Mogrify is the only method that provides a set of TFs, not a ranked list of all TFs, and since the goal is to compare other methods with Mogrify, the information generated by Mogrify regarding the number of factors used was shared with the other methods. That is, no method can use more factors than another method. For example, for the transition between B cells and macrophages, Mogrify predicts that 8 TFs should be sufficient, so the top 8 TFs from all methods are used for comparison.

[0285] 2) For each method, check whether the correct transcription factors are predicted: For each published set of transcription factors, we compare the predictions by each method and extract two statistics: first, the recovery rate of published transcription factors (i.e., 100% if all factors were included in the predicted set), and second, the average rank of the published factors (i.e., for each correctly identified TF, sum the rank and divide by the total number of correctly identified TFs).

[0286] The results from these two steps can be found in Tables 7 and 8, and a summary of Mogrify's comparison to CellNet and D'Allesio et al. can be found in Figure 6.

[0287] To extract results for CellNet, we used publicly available datasets for fibroblasts (GSE14897) and B cells (GSE65136) as a starting point and used a web interface to CellNet (cellnet.hms.harvard.edu) to provide predictions for each transition in Figure 2. D'Allessio et al. These ranked lists were used for comparison.

[0288] Table 7: Benchmarking results comparing the performance of Mogrify, CellNet, and D'Alessio et al. For each transformation in Figure 2, predictions from each technology are shown. The ranked lists from CellNet and D'Alessio et al. were cut off by the size of the Mogrify set. To compare these sets, we extract the average rank and the overall recovery efficiency from the published set. These statistics provide a guide to the performance each technology would have achieved for these transformations. Failure to identify a published transcription factor does not necessarily mean that the predicted transcription factor from each technology cannot transform cells; this benchmark is designed to evaluate performance based only on available data. For CellNet's myoblast predictions, we used skeletal muscle GRN.

[0289] [Table 15]

[0290] [Table 16]

[0291] Table 8: Benchmarking results comparing the performance of Mogrify and its individual components (MARA, STRING, and differential expression). For each transformation in Figure 2, predictions for Mogrify and each of Mogrify's individual components are shown. The ranked lists of MARA, STRING, and differential expression components are cut off at the size of the set predicted by Mogrify. To compare these sets, the mean rank and overall recovery efficiency from the published set are extracted. These statistics provide a guide to the performance each technology would have achieved for these transformations. Failure to identify published transcription factors does not necessarily mean that the predicted transcription factors from each technology cannot transform cells; This benchmark is designed to evaluate performance based solely on available data.

[0292] [Table 17]

[0293] [Table 18]

[0294] Example 3 To experimentally demonstrate the predictive capabilities of Mogrify, we performed 11 novel cell transformations using human cells: fibroblasts to keratinocytes (results in Example 4); keratinocytes to endothelial cells (results in Example 5); fibroblasts to endothelial cells (results in Example 6); Endothelial cells from embryonic stem cells (results in Example 7); Endothelial cells from induced pluripotent stem cells (results in Example 8); fibroblasts to astrocytes (results in Example 9); Astrocytes from embryonic stem cells (results in Example 10); Astrocytes from induced pluripotent stem cells (results in Example 11); Astrocytes from bone mesenchymal stem cells (results in Example 12); Keratinocytes from embryonic stem cells (results in Example 13); and -Keratinocytes from induced pluripotent stem cells (results in Example 14).

[0295] In this example, materials and methods are described.

[0296] Lentivirus generation For lentivirus production, 293T human embryonic kidney (HEK; Sigma) cells were cultured in T-75 flasks. After reaching 90-95% confluence, they were transfected with the -lv165 vector expressing the relevant transcription factors (e.g., CDH1, FOS, FOXQ1, HOXB6, IRF1, MAFB, REL, SMAD1, SOX9, SOX17, TAL1, TCF7L1, MXD4, NFKB1, SOX2, ARNT2, RUNX2, PAX6, SNAI2, HMGB2, E2F1, MYC, FOSL2, or TFAP2A) driven by the EF1α promoter and IRES2-eGFP (GeneCopoeia) using LTX Lipofectamine (Invitrogen) transfection agent. Viral supernatants were collected 24 and 36 hours after transfection and concentrated using ultracentrifugal filters (Millipore). The viral concentrates were then stored at -80°C. Titration was based on eGFP expression measured by flow cytometry. The cell lines used in these experiments were negative for mycoplasma contamination.

[0297] cell culture Before use in experiments, human adult epidermal keratinocytes (HEKa; GIBCO) and human dermal fibroblasts (HDFs; GIBCO) were cultured at 2.5 × 10 3 cells / cm 2 The cells were expanded to 100 μg / ml and passaged at least three times. HEKa cells were cultured in keratinocyte serum-free medium (KSFM; GIBCO) containing 10% HKGS (GIBCO) and 1% Pen / Strep (GIBCO). Meanwhile, HDFs were cultured in medium 106 (GIBCO) containing 10% LSGS (GIBCO) and 1% Pen / Strep. The cells were then frozen in liquid nitrogen for later use. For endothelial cell transdifferentiation from keratinocytes, cells were thawed and cultured at 2.5 × 10 cells / ml until they reached 90% confluence. 3 cells / cm 2 They were then seeded in KSFM medium at 5.0 × 10 3 cells / cm 2The cells were then seeded again for 2 days and then infected with concentrated lentiviral particles of HOXB6, IRF1, SMAD1, SOX17, TAL1, and TCF7L1 in KSFM medium in the presence of polybrene (Millipore). After virus addition (12–24 h), the medium was replaced with fresh KSFM medium. On day 4, the medium was replaced with human endothelial serum-free medium (GIBCO) with 1% Pen / Strep containing human VEGF (50 ng / μl; PeproTech), human BMP4 (20 ng / μl; PeproTech), and human FGF2 (20 ng / μl; PeproTech). For keratinocyte transdifferentiation from fibroblasts, cells were cultured at 2.5 × 10 cells / ml until they reached 90% confluence. 3 cells / cm 2 They were then seeded at 2.5 × 10 in mouse fibroblast medium (MEFM). 3 cells / cm 2 The cells were then reseeded in 1% PEG / MSC for 24 hours and subsequently transduced with lentiviral particles of CDH1, FOS, FOXQ1, MAFB, REL, and SOX9 in MEFM in the presence of polybrene for 24 hours. The medium was replaced with KSFM containing n / Strep, retinoic acid, and human BMP4 (R&D). Throughout the entire experiment, fresh medium was added at least once every two days. Each of these experiments was repeated 3-4 times.

[0298] [Table 19]

[0299] Flow cytometry At various time points, transdifferentiated cells were detached using 0.25% trypsin-EDTA (GIBCO) for 3 minutes at 37°C. Cells were then prepared for flow cytometry analysis or sorting. Cells were incubated with anti-human CD31-APC (17-0319-41, eBioscience) for 15 minutes at 4°C, washed with DPBS (GIBCO), centrifuged at 1000 rpm for 7 minutes, and resuspended in medium containing propidium iodide (Sigma-Aldrich). An LSR-II analyzer (BD Bioscience) and an Influx cell sorter (BD Biosciences) were used for data analysis and sorting, respectively.

[0300] qPCR Total RNA was extracted using the RNeasy Micro kit (Qiagen) according to the manufacturer's instructions. The extracted RNA was reverse transcribed into cDNA using the Superscript III kit (Invitrogen). Real-time quantitative PCR reactions were performed in triplicate using Brilliant II SYBR Green QPCR Master Mix (Stratagene) on a 7500 Real-Time PCR System. The qPCR primer sequences were: F-CD31:CCTTCTGCTCTGTTCAAGCC R-CD31:GGGTTCAGGTTCTTCCCATTT F-VE:ATGAGAATGACAATGCCCCG R-VE:TGTCTATTGCGGAGATCTGCAG F-VEGFR2:GGCCCAATAATCAGAGTGGCA R-VEGFR2:CCAGTGTCATTTCCGATCACTTT F-keratin 1: AGAGTGGACCAACTGAAGAGT R-keratin 1: ATTCTCTGCATTTGTCCGCTT F-keratin 14:AGACCAAAGGTCGCTACTGC R-keratin 14:AGGAGAACTGGGAGGAGGAG F-involucrin: CTGCCTCAGCCTTACTGTGA R-involucrin: GGAGGAGGAACAGTCTTGAGG F-β-ACTIN:CATGTACGTTGCTATCCAGGC R-β-ACTIN:CTCCTTAATGTCACGCACGAT is.

[0301] Immunofluorescence Cells were fixed with 4% paraformaldehyde in DPBS for 10 minutes at room temperature. Because the markers of interest were expressed on the cell surface, there was no need to permeabilize the cells. Cells were blocked with 5% donkey serum in DPBS for 30 minutes and then incubated with primary antibodies (goat polyclonal anti-CD31, sc-1506; Santa Cruz; and rabbit polyclonal anti-VE-cadherin, ab33168; Abcam) overnight at 4°C. The next day, cells were incubated with secondary antibodies (donkey anti-goat Alexa Flour-555; Invitrogen; and donkey anti-rabbit Alexa Flour-647; Invitrogen) for 2 hours at room temperature. Finally, cells were covered with 4',6-diamidino-2-phenylindole (DAPI; Life Technologies) for 1 minute. All images were captured using an inverted Nikon Eclipse Ti epifluorescence microscope with a Nikon Digital Sight DS-U2 camera and processed and analyzed using FIJI software.

[0302] Example 4 - Conversion of human fibroblasts to keratinocytes (iKer) For this conversion, cells were transduced with FOXQ1, SOX9, MAFB, CDH1, FOS, and REL, as predicted by Mogrify (Figure 3A and Table 10).

[0303] [Table 20]

[0304] By day 16 after transduction, keratinocyte-associated markers, keratin 1, keratin 14, and involucrin, were significantly upregulated in the transdifferentiated cells (Figure 3C). Furthermore, by week 3, the majority of transduced cells exhibited a cobblestone morphology, a typical feature exhibited by keratinocytes. Adjacent untransduced GFP-negative cells or control cells transduced with a GFP-only virus maintained their fibroblast morphology (arrows in Figure 3D). The morphological and molecular characteristics of these reprogrammed cells demonstrate that Mogrify successfully predicts the TFs required to induce the transformation of human fibroblasts into keratinocyte-like cells.

[0305] Example 5 - From adult human keratinocytes (HEKa) to microvascular endothelial cells (iECs) For this conversion, we chose to use SOX17, TAL1, SMAD1, IRF1, and TCF7L1 from the six TFs proposed by Mogrify (Figure 4 and Table 11).

[0306] [Table 21]

[0307] These five TFs are predicted to regulate up to 92% of the genes required for iECs. After these TFs were overexpressed in HEKa cells, we found that the cells needed to be maintained in their culture medium for up to four days (Figure 4B). We used FACS to examine the kinetics of cell reprogramming using the stable endothelial marker CD31 (Figure 4C). By day 14 after transduction, we observed that more than 2% of infected cells had upregulated CD31, and by day 18, nearly 10% had upregulated CD31. At this point, we isolated these CD31 cells and evaluated the expression of endothelial-related genes (CD31, VE-cadherin, and VEGFR2) by qPCR, which showed clear reactivation of all genes evaluated (Figure 4D). Finally, we performed immunofluorescence (IF) to verify the morphology and expression of the transdifferentiated cells. As shown in Figure 4E, only cells transduced with the predicted TFs—but not the control cells—presented the correct morphology and expressed CD31 and VE-cadherin on their surface. The morphological and molecular characteristics of these reprogrammed cells demonstrate the successful transition of human keratinocytes into human endothelial-like cells.

[0308] Example 6 - From fibroblasts to endothelial cells Transcription factors used: SOX17, SMAD1, TAL1, IRF1, TCF7L1 and MXD4. (Mogrify also identified the factor JUNB, but this was not used).

[0309] Transdifferentiation strategy: 24 hours before viral transduction of transcription factors, human dermal fibroblasts were plated onto well plates at 5k cells / cm in medium 10 with LSGS (Life Technologies). 2The cells were seeded at 1000 x g for 1 hour. The next day, lentiviral particles encoding transcription factors were transduced into the cells in Medium 106 with Polybrene (Merck Millipore). Immediately after transduction, the well plates were centrifuged at 1900 rpm for 60 minutes. On day 5, the medium was replaced with endothelial medium (Medium 131, Life Technologies) supplemented with VEGF (50 ng / ml, Miltenyi Biotec), FGF2 (20 ng / ml, Miltenyi Biotec), and BMP4 (20 ng / ml, Miltenyi Biotec). The medium was changed every 2 days throughout the experiment.

[0310] Immunofluorescence analysis showed evidence of expression of the endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation (FIG. 8A).

[0311] qPCR analysis also showed expression levels of endothelial-related genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation ( Fig. 8B ).

[0312] Example 7 - From embryonic stem cells to endothelial cells Transcription factors used: SOX17, SMAD1, TAL1, NFKB1 and IRF1. (Mogrify also identified the factors HOXB7 and JUNB, but these were not used).

[0313] Transdifferentiation strategy: 24 hours before viral transduction of transcription factors, human embryonic stem cells (H9) were plated in Essential 8 medium (Life Technologies) at 5k cells / cm on Matrigel-coated (BD Falcon) well plates. 2The cells were seeded at 100°C for 1 hour. The next day, lentiviral particles encoding transcription factors were transduced into the cells in Essential 8 medium with polybrene (Merck Millipore). Immediately after transduction, the well plates were centrifuged at 1900 rpm for 60 minutes. On day 5, the medium was replaced with endothelial medium (Medium 131, Life Technologies) supplemented with VEGF (50 ng / ml, Miltenyi Biotec), FGF2 (20 ng / ml, Miltenyi Biotec), and BMP4 (20 ng / ml, Miltenyi Biotec). The medium was changed every 2 days throughout the experiment.

[0314] Immunofluorescence analysis showed evidence of expression of the endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation (FIG. 9A).

[0315] qPCR analysis showed expression levels of endothelial-related genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation (Fig. 9B).

[0316] FIG. 10 shows the results of flow cytometry analysis of PeCAM expression on days 12 and 18 of transdifferentiation and quantification of PeCAM-positive cells on day 18 of transdifferentiation.

[0317] Example 8 - From pluripotent stem cells to endothelial cells Transcription factors used: SOX17, TAL1, NFKB1, IRF1, and SMAD1. (Mogrify also identified the factors HOXB7 and JUNB, but these were not used.)

[0318] Transdifferentiation strategy: 24 h prior to viral transduction of transcription factors, human induced pluripotent stem cells (32F donor) were cultured in Essential 8 Medium (Life Technologies). 5k cells / cm on Matrigel-coated (BD Falcon) well plates in 2The cells were seeded at 100°C for 1 hour. The next day, lentiviral particles encoding transcription factors were transduced into the cells in Essential 8 medium with polybrene (Merck Millipore). Immediately after transduction, the well plates were centrifuged at 1900 rpm for 60 minutes. On day 5, the medium was replaced with endothelial medium (Medium 131, Life Technologies) supplemented with VEGF (50 ng / ml, Miltenyi Biotec), FGF2 (20 ng / ml, Miltenyi Biotec), and BMP4 (20 ng / ml, Miltenyi Biotec). The medium was changed every 2 days throughout the experiment.

[0319] Immunofluorescence analysis shows the expression of the endothelial markers PeCAM and VE-cadherin at day 18 of transdifferentiation (FIG. 11A).

[0320] qPCR analysis shows the expression levels of the endothelial-related genes VEGFR2 and VE-cadherin at day 18 of transdifferentiation (FIG. 11B).

[0321] Figure 12 shows flow cytometry analysis of PeCAM expression on days 12 and 18 of transdifferentiation. FSC, forward scatter and quantification of PeCAM positive cells on day 18 of transdifferentiation.

[0322] Example 9 - From fibroblasts to astrocytes Transcription factors used: SOX2, SOX9 ARNT2, SMAD1 and RUNX2. (Mogrify also identified factors E2F5 and PBX1, but these were not used).

[0323] Transdifferentiation strategy: 24 hours before viral transduction of transcription factors, human dermal fibroblasts were plated onto well plates at 5k cells / cm in medium 10 with LSGS (Life Technologies). 2The cells were seeded at 1000 x g for 1 hour. The next day, lentiviral particles encoding transcription factors were transduced into the cells in a medium containing polybrene (Merck Millipore). Immediately after transduction, the well plates were then centrifuged at 1900 rpm for 60 minutes. On day 5, the medium was replaced with astrocyte medium (Life Technologies) supplemented with IL1β (10 ng / ml, Sigma-Aldrich). On day 7, the medium was replaced with astrocyte medium. The medium was changed every 2 days throughout the experiment.

[0324] Immunofluorescence analysis shows the expression of the astrocytic marker GFAP at day 21 of transdifferentiation (FIG. 13).

[0325] Example 10 - Embryonic Stem Cells (H9) to Astrocytes Transcription factors used: IRF1, SOX9, ARNT2, PAX6, SNAI2, RUNX2. (Mogrify also predicted the factor SOX5, but this was not used.)

[0326] Transdifferentiation strategy: Human embryonic stem cells (H9) were plated onto Matrigel-coated (BD Falcon) well plates at 5k cells / cm in Essential 8 medium (Life Technologies) 24 h prior to viral transduction of transcription factors. 2 The cells were seeded at 1000 x g for 10 min. The next day, lentiviral particles encoding transcription factors were transduced into the cells in Essential 8 medium with polybrene (Merck Millipore). Immediately after transduction, the well plates were then centrifuged at 1900 rpm for 60 min. On day 2, the medium was replaced with N2 medium containing B27 supplement (Life Technologies) and 0.6 μM CHIR99021 (Miltenyi Biotec). On day 6, the medium was replaced with astrocyte medium (Life Technologies) supplemented with IL1β (10 ng / ml, Sigma-Aldrich). On day 8, the medium was replaced with astrocyte medium. The medium was replaced every 2 days throughout the experiment.

[0327] Immunofluorescence analysis shows the expression of the astrocytic marker GFAP at day 21 of transdifferentiation (FIG. 14).

[0328] Example 11 - From pluripotent stem cells to astrocytes Transcription factors used: PAX6, SNAI2, RUNX2, HMGB2. (Mogrify also predicted factors POU3F2.E2F5 and SOX5, but these were not used.)

[0329] Transdifferentiation strategy: Human induced pluripotent stem cells (32F donor) were plated onto Matrigel-coated (BD Falcon) well plates at 5k cells / cm in Essential 8 medium (Life Technologies) 24 h prior to viral transduction of transcription factors. 2 The cells were seeded at 1000 x g for 10 min. The next day, lentiviral particles encoding transcription factors were transduced into the cells in Essential 8 medium with polybrene (Merck Millipore). Immediately after transduction, the well plates were then centrifuged at 1900 rpm for 60 min. On day 2, the medium was replaced with N2 medium containing B27 supplement (Life Technologies) and 0.6 μM CHIR99021 (Miltenyi Biotec). On day 6, the medium was replaced with astrocyte medium (Life Technologies) supplemented with IL1β (10 ng / ml, Sigma-Aldrich). On day 8, the medium was replaced with astrocyte medium. The medium was replaced every 2 days throughout the experiment.

[0330] Immunofluorescence analysis showed expression of the astrocytic marker GFAP at day 21 of transdifferentiation (FIG. 15).

[0331] Example 12 - From mesenchymal stem cells to astrocytes Transcription factors: SOX2, SOX9, ARNT2, MYBL2, E2F1, HMGB2. (Mogrify also identified the factors HOXB7 and JUNB, but these were not used.)

[0332] Transdifferentiation strategy: 24 h before viral transduction of transcription factors, bone marrow mesenchymal stem cells (donor 7081) were plated onto well plates at 5k cells / cm in MSC medium (α-MEM with 15% FBS, glutamine, penicillin, and streptomycin; Life Technologies). 2 The cells were seeded at 1000 x g for 10 min. The next day, lentiviral particles encoding transcription factors were transduced into the cells in MSC medium containing polybrene (Merck Millipore). Immediately after transduction, the well plates were then centrifuged at 1900 rpm for 60 min. On day 5, the medium was replaced with astrocyte medium (Life Technologies) supplemented with IL1β (10 ng / ml, Sigma-Aldrich). On day 7, the medium was replaced with astrocyte medium. The medium was changed every 2 days throughout the experiment.

[0333] Immunofluorescence analysis showed expression of the astrocytic marker GFAP at day 21 of transdifferentiation (FIG. 16).

[0334] Example 13 - From embryonic stem cells to keratinocytes Transcription factors used: SOX9, NFKB1, MYC, FOSL2. (Mogrify also predicted factors NR2F2, FOSL1 and AHR, but these were not used.)

[0335] Transdifferentiation strategy: Human embryonic stem cells (H9) were plated onto Matrigel-coated (BD Falcon) well plates at 5k cells / cm in Essential 8 medium (Life Technologies) 24 h prior to viral transduction of transcription factors. 2 The next day, lentiviral particles encoding the transcription factors were inoculated into the cells using polybrene (Merck Millippo Cells were transduced in Essential 8 medium containing 3 μM retinoic acid (re). Immediately after transduction, the well plates were then centrifuged at 1900 rpm for 60 minutes. On day 2, the medium was replaced with Essential 8 medium containing 3 μM retinoic acid. On day 6, the medium was replaced with EpiLife medium (Life Technologies) supplemented with BMP4 (50 ng / ml, Miltenyi Biotec) and EGF (5 ng / ml, Miltenyi Biotec). The medium was changed every two days throughout the experiment.

[0336] Immunofluorescence analysis showed expression of the keratinocyte marker pankeratin at 21 days of transdifferentiation (FIG. 17).

[0337] Example 14 - From pluripotent stem cells to keratinocytes Transcription factors: TFAP2A, MYC, SOX9, NFKB1. (Mogrify also predicted factors TP63 and NFKBIA, but these were not used.)

[0338] Transdifferentiation strategy: Human induced pluripotent stem cells (32F donor) were plated onto Matrigel-coated (BD Falcon) well plates at 5k cells / cm in Essential 8 medium (Life Technologies) 24 h prior to viral transduction of transcription factors. 2 The cells were seeded with 1000 μg of ...

[0339] Immunofluorescence analysis shows the expression of the keratinocyte marker keratin 14 (KRT14) at day 21 of transdifferentiation (FIG. 18A).

[0340] qPCR analysis shows the expression levels of the keratinocyte-associated genes keratin 14 and keratin 1 at day 21 of transdifferentiation (Figures 18B and 18C).

[0341] Example 15 Several attempts have been made to generate representative cell landscapes, focusing on one or two cell types and based on path-integral quasi-potentials, mechanistic modeling, or stochastic landscapes. We hypothesized that comparing the combined all-to-all TF network differences determined by Mogrify with transcriptional profiles would enable the formation of a 3D landscape representative of human cell types (Figure 7). The landscape positions molecularly similar cell types close to each other in the xy plane and adjusts their height (z direction) according to how likely a cell type is to be a good starting cell source (see details in the online Materials and Methods). Interestingly, we observed that distinct stem cells are positioned at the highest positions. This may suggest that the transcriptional networks of cells at the highest points in the landscape are controlled by fewer TFs, and that as cells differentiate (in valleys), they require more TFs to fine-tune their transcriptional networks.

Claims

1. 1. A method for determining transcription factors necessary to convert a source cell into a cell having at least one characteristic of a target cell type, comprising: - determining the differential expression of genes in said source and target cell types; - determining a network score for each transcription factor in each of the source and target cell types across at least one network based on the differential gene expression, the network containing information on interactions that affect gene expression, wherein the network score for each transcription factor is calculated by performing a weighted sum of gene scores across at least one sub-network centered on the transcription factor; - ranking the transcription factors based on a combination of network score and differential gene expression information, thereby identifying a set of transcription factors for conversion of a source cell to a cell having at least one characteristic of a target cell type.

2. 10. The method of claim 1, comprising determining a gene score for each differentially expressed gene in the source cell type and the target cell type.

3. 3. The method of claim 1 or 2, wherein the gene score is a combination of the log fold change of the differential expression and an adjusted P-value.

4. 4. The method of claim 3, wherein the gene score is a combination of the log fold change and adjusted P-value using Equation 1 below: (In the formula, L S X is the log fold change of gene x in sample s, and P S X is the adjusted p-value of gene x in sample s).

5. The method according to any one of claims 1 to 4, wherein the gene scores are calculated using a tree-based method, preferably against a background, or using Bayesian clustering.

6. The method according to any one of claims 1 to 5, wherein the network comprises information on protein-DNA interactions, protein-protein interactions and / or protein-RNA interactions.

7. 7. The method of claim 6, wherein the network comprises a Search Tool for Retrieval of Interacting Genes / Proteins (STRING) database.

8. The method of claim 6 , wherein the network comprises information on interactions between transcription factors and regulatory regions of genes.

9. 9. The method of claim 8, wherein the regulatory region is a promoter region of a gene and / or the network comprises interactions derived from Motif Activity Response Analysis (MARA).

10. The method of any one of claims 1 to 9, further comprising removing transcriptionally redundant TFs from the ranked list from each cell type.

11. 11. The method of any one of claims 1 to 10, wherein determining differential expression of genes in the source cell type and the target cell type comprises determining gene scores using differential expression analysis of expression data, wherein the expression data comprises expression data obtained by CAGE analysis of gene expression (CAGE).

12. A method described in any one of claims 1 to 11, wherein determining the differential expression of genes in the source cell type and the target cell type comprises determining gene scores using differential expression analysis of expression data, and the expression data comprises expression data obtained from the FANTOM5 dataset.

13. Network score (N) for transcription factor x in cell type S S X,n 13. The method of any one of claims 1 to 12, wherein the n-th signal is calculated over the network n using the following equation 2: where r∈Vx is each gene (r) in the set of nodes (Vx) that make up the local subnetwork of transcription factor x; V x is the set of nodes that make up the local subnetwork of transcription factor x, Lr,n is the number r of steps away from x in network n, G s r is the gene score for gene r, and Or,n is the parental degree of r in network n; Here, the parent gene of gene r is upstream of gene r from transcription factor x in the local subnetwork downstream of gene r.

14. identifying a set of transcription factors for transforming the source cell type into a cell exhibiting at least one characteristic of the target cell type includes comparing the ranked list of the target cell type with the set of identified genes expressed in the source cell type; The method of any one of claims 1 to 13, wherein a transcription factor from the target cell type list is subsequently removed from the ranked list if it is already expressed in the source cell type.

15. 15. The method of any one of claims 1 to 14, wherein determining a network score for each transcription factor in each of the source and target cell types based on differential gene expression across at least one network comprises determining a first network score for each transcription factor across the first network and a second network score for each transcription factor across a second network, wherein the first network comprises protein-DNA interactions between transcription factors with known binding sites within promoter regions of genes, and the second network comprises protein-protein, protein-DNA, and protein-RNA interactions.

16. 16. The method according to any one of claims 1 to 15, - the source cells are selected from the group consisting of skin fibroblasts, epidermal keratinocytes, embryonic stem cells, induced pluripotent stem cells (iPSCs), mesenchymal stem cells, monocytes or cardiac fibroblasts; and / or - the target cell type is selected from the group consisting of chondrocytes, hair follicles, CD4+ T cells, CD8+ T cells, NK-cells, haemopoietic stem cells (HSC), adipose mesenchymal stem cells (MSC), bone marrow mesenchymal stem cells (MSC), oligodendrocytes, oligodendrocyte precursors, skeletal muscle cells, smooth muscle cells, fetal cardiomyocytes, endothelial cells, astrocytes, keratinocytes and epithelial cells.

17. The method of any of claims 1 to 16, wherein at least one characteristic of the target cell is upregulation of any one or more target cell markers and / or a change in cell morphology.

18. 18. The method of any one of claims 1 to 17, further comprising: - Increasing the amount of one or more transcription factors in the source cells that have been determined to be necessary for the conversion of the source cell type to the target cell type. A method comprising:

19. A computer system comprising a processor and a computer readable medium storing instructions which, when executed by the processor, cause the processor to carry out the method of any of claims 1 to 18.

20. 1. A method for identifying an agent useful for promoting conversion of a source cell type into a cell exhibiting at least one characteristic of said target cell type, comprising: - determining one or more transcription factors required for the conversion of a source cell type into a cell exhibiting at least one characteristic of a target cell type, using the method according to claims 1 to 18; - screening one or more candidate agents for their ability to increase the amount of said one or more transcription factors necessary for conversion of a source cell type into a cell exhibiting at least one characteristic of a target cell type; Including, The method, wherein the agent that increases the amount of the one or more transcription factors is an agent useful for promoting the conversion of a source cell type into a cell that exhibits at least one characteristic of the target cell type.

Citation Information

Patent Citations

  • Transcription factor searching method for cell conversion

    JP2013074857A